Speaker adaptation techniques can be classified as intra-lingual or cross-lingual depending on whether or not the source model and the target speaker employ the same language. Most of the work in this field has been focused on the first case, while the second one has been less explored. In this paper we address the cross-lingual paradigm in the framework of a HMM-based speech synthesis system by further developing a formerly proposed approach. This method is able to clone a given speaker into a different language by combining the linguistic structure and the acoustic characteristics of two HTS models. In this work, we discuss the extension of the adaptation procedure to some other source model parameters that were kept unmodified in the initial version, and compare the performance of both versions by means of subjective and objective tests. These results are also contrasted with those obtained by a KLD-based technique proposed in the literature for a similar purpose. While no significant preference for any of the versions of our method is observed, our approach clearly outperforms the KLD-based technique. (C) 2018 Elsevier Ltd. All rights reserved.
Depression is a common mental disorder that is usually addressed by outpatient treatments that favour patients’ inclusion in the society. This raises the need for tools to remotely monitor the emotional state of the patients, which can be carried out via telephone or the Internet using speech processing approaches. However, these strategies lead to privacy concerns caused by the transmission of the patients’ speech and its subsequent storage in servers. The use of speech de-identification to protect the privacy of these patients seems straightforward, but the influence of this procedure in the manifestation of the disease in the patients’ speech has not been addressed yet. Hence, this study evaluates the performance of an automatic depression level estimation system when dealing with original and de-identified speech, in order to analyse the influence of the de-identification procedure in the detection of depression. Two de-identification approaches based on voice transformation via frequency warping and amplitude scaling are assessed, which can be applied to any speaker without additional training. Experiments carried out in the framework of the audio/visual emotion challenge 2014 show that the proposed de-identification approaches achieve promising de-identification results at the expense of a slight degradation of depression detection.
Speaker de-identification approaches must accomplish three main goals: universality, naturalness and reversibility. The main drawback of the traditional approach to speaker de-identification using voice conversion techniques is its lack of universality, since a parallel corpus between the input and target speakers is necessary to train the conversion parameters. It is possible to make use of a synthetic target to overcome this issue, but this harms the naturalness of the resulting de-identified speech. Hence, a technique is proposed in this paper in which a pool of pre-trained transformations between a set of speakers is used as follows: given a new user to de-identify, its most similar speaker in this set of speakers is chosen as the source speaker, and the speaker that is the most dissimilar to the source speaker is chosen as the target speaker. Speaker similarity is measured using the i-vector paradigm, which is usually employed as an objective measure of speaker de-identification performance, leading to a system with high de-identification accuracy. The transformation method is based on frequency warping and amplitude scaling, in order to obtain natural sounding speech while masking the identity of the speaker. In addition, compared to other voice conversion approaches, the proposed method is easily reversible. Experiments were conducted on Albayzin database, and performance was evaluated in terms of objective and subjective measures. These results showed a high success when de-identifying speech, as well as a great naturalness of the transformed voices. In addition, when making the transformation parameters available to a trusted holder, it is possible to invert the de-identification procedure, hence recovering the original speaker identity. The computational cost of the proposed approach is small, making it possible to produce de-identified speech in real-time with a high level of naturalness.
This paper presents a new method for cross-lingual speaker adaptation in the framework of HMM-based speech synthesis. Taking two HTS voice models as input, one for the desired language and another for the aimed speaker identity, it yields a third model that produces speech in the target language while sounding like the target speaker. The method operates at segmental level (spectral information and average fundamental frequency) and does not require any phonetic or linguistic information. Perceptual evaluation experiments show that, when the input models are good enough, the resulting synthetic voice is perceived as similar to the target speaker with no important quality degradation.
Postfiltering is a well known technique that helps increasing the quality of coded speech, enhancing speech intelligibility or alleviating the oversmoothing effect of statistical speech processing methods. This letter presents a new formulation of the radial cepstral postfiltering method that enables the application of different postfiltering factors to low and high frequencies. The transition between bands will be smooth and controllable through an adjustable cut-off frequency. The proposed algorithm can be implemented by means of a simple multiplicative matrix of which an analytical expression is derived. The new method provides a flexible framework to tackle issues related to the quality and intelligibility of synthetic speech.
This paper describes the implementation of the Aholab entry for the Singing Synthesis Challenge: Fill-in the Gap. Our approach in this work makes use of an HTS based Text-to-Speech (TTS) synthesizer for Basque to generate the singing voice. The prosody related parameters provided by the TTS system for a spoken version of the score are modified to adapt them to the requirements of the music score concerning syllables duration and tone, while the spectral parameters are basically maintained. The paper describes the processing details developed to improve the quality of the output signal: the syllable timing, the generation of the intonation with vibrato and the manipulation of the model states. In this entry, the lyrics have been freely translated into Basque and the rhythm has been adapted to a Basque traditional rhythm.
The main drawback of speaker de-identification approaches using voice conversion techniques is the need for parallel corpora to train transformation functions between the source and target speakers. In this paper, a voice conversion approach that does not require training any parameters is proposed: it consists in manually defining frequency warping (FW) based transformations by using piecewise linear approximations. An analysis of the de-identification capabilities of the proposed approach using FW only or combined with FW modification and spectral amplitude scaling (AS) was performed. Experimental results show that, using the manually defined transformations using only FW, it is not possible to obtain de-identified natural sounding speech. Nevertheless, when modifying the FW, both de-identification accuracy and naturalness increase to a great extent. A slight improvement in de-identification was also obtained when applying spectral amplitude scaling.
In silent speech interfaces a mapping is established between biosignals captured by sensors and acoustic characteristics of speech. Recent works have shown the feasibility of a silent interface based on permanent magnet-articulography (PMA). This paper studies the performance of four different mapping methods based on Gaussian mixture models (GMMs), typical from the voice conversion field, when applied to PMA-to-spectrum conversion. The results show the superiority of methods based on maximum likelihood parameter generation (MLPG), especially when the parameters of the mapping function are trained by minimizing the generation error. Informal listening tests reveal that the resulting speech is moderately intelligible for the database under study.
In a previous work we presented a method to combine the acoustic characteristics of a speech synthesis model with the linguistic characteristics of another one. This paper presents a more extensive evaluation of the method when applied to cross-lingual adaptation. A large number of voices from a database in Spanish are adapted to Basque, Catalan, English and Galician. Using a state-of-the-art speaker identification system, we show that the proposed method captures the identity of the target speakers almost as well as standard intra-lingual adaptation techniques.
EuskaraArtikulu honetan Aholabek garatutako Ahots Sintetikoaren Detektorea (Synthetic Speech Detector, SSD) deskribatzen da, eta bere erabilera Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015) nazioarteko norgehiagokan. GMMtan oinarritutako detektorea da, non ereduak Relative Phase Shift (RPS) fasearen informaziorako transformazioa erabiliz osatzen diren. Estrategia desberdinak ebaluatzen dira: eraso espezifikoen ereduak sortu norgehiagokaren antolatzaileek emandako informazioaren bidez, edo mehatxu-seinaleak sortzeko erabili omen den bokoderraren ereduak sortu, aurreko lanetan garatutako informazioa erabiliz. Ebaluazioaren emaitzetan ikusten denez, eraso espezifikoen ereduak ondo aritzen dira eraso ezagunekin, baina ez ezezagunekin. Bokoderren ereduak erabiltzen direnean, informazioa ez da guztiz erabiltzen eta ereduen moldaketa landu behar da. EnglishThis paper introduces the Synthetic Speech Detection system developed by Aholab for the Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015). The detector is a classifier based on Gaussian Mixture Models that are created using the Relative Phase Shift (RPS) transformation for the phase information. Different strategies have been evaluated: modeling the specific attacks using the information provided by the ASVspoof 2015 organizers, and modeling the vocoders possibly used in the spoofing signals, using data from previous works. The evaluation results show that attack specific models work for known attacks but they do not cope with the unknown attacks correctly. When using vocoder models build with other databases, the results suggest that the followed strategy do not take advantage of the available data and thus model adaptation should be explored.
Speaker adaptation techniques use a small amount of data to modify Hidden Markov Model (HMM) based speech synthesis systems to mimic a target voice. These techniques can be used to provide personalized systems to people who suffer some speech impairment and allow them to communicate in a more natural way. Although the adaptation techniques don't require a big quantity of data, the recording process can be tedious if the user has speaking problems. To improve the acceptance of these systems an important factor is to be able to obtain acceptable results with minimal amount of recordings. In this work we explore the performance of an adaptation method based on Frequency Warping which uses only vocalic segments according to the amount of available training data.
This paper presents a system capable of de-identifying speech signals in order to hide and protect the identity of the speaker. It applies a relatively simple yet effective transformation of the pitch and the frequency axis of the spectral envelope thanks to a flexible wideband harmonic model. Moreover, it inserts the parameters of the transformation in the signal by means of watermarking techniques, thus enabling re-identification. Our experiments show that for adequate modification factors its performance is satisfactory in terms of quality, de-identification degree and naturalness. The limitations due to the signal processing framework are discussed as well.
This paper describes the characteristics and structure of a Basque singing voice database of bertsolaritza. Bertsolaritza is a popular singing style from Basque Country sung exclusively in Basque that is improvised and a capella. The database is designed to be used in statistical singing voice synthesis for bertsolaritza style. Starting from the recordings and transcriptions of numerous singers, diarization and phoneme alignment experiments have been made to extract the singing voice from the recordings and create phoneme alignments. This labelling processes have been performed applying standard speech processing techniques and the results prove that these techniques can be used in this specific singing style.
New voice conversion functions based on bilinear frequency warping and constrained amplitude scaling.Good overall conversion performance using more intuitive and informative parameters.Useful as an analysis tool to visualize spectral differences between different voices or styles. Voice conversion functions based on Gaussian mixture models and parametric speech signal representations are opaque in the sense that it is not straightforward to interpret the physical meaning of the conversion parameters. Following the line of recent works based on the frequency warping plus amplitude scaling paradigm, in this article we show that voice conversion functions can be designed according to physically meaningful constraints in such manner that they become highly informative. The resulting voice conversion method can be used to visualize the differences between source and target voices or styles in terms of formant location in frequency, spectral tilt and amplitude in a number of spectral bands.
The Auracle project aimed at investigating how sound alters the gaze behavior of people watching moving images, using low-cost opensource systems and copyleft databases to conduct experiments. We created a database of audiovisual content comprising: several fragments of movies with stereo audio released under Creative Commons licenses, shorts with multitrack audio shot by participants, a comic book augmented with sound (events and music). We set up a low-cost experimental system for gaze and head tracking synchronized with audiovisual stimuli using the Social Signal interpretation (SSI) framework. We ran an eye tracking experiment on the comic book augmented with sound with 25 participants. We visualized the resulting recordings using a tool overlaying heatmaps and eye positions/saccades on the stimuli: CARPE. The heatmaps qualitatively observed don’t show a significant influence of sound on eye gaze. We proceeded with another pre-existing database of audiovisual stimuli plus related gaze tracking to perform audiovisual content analysis in order to find correlations between audiovisual and gaze features. We visualized this exploratory analysis by importing CARPE heatmap videos and audiovisual/gaze features resampled as audio files into a multitrack visualization tool originally aimed at documenting digital performances: Rekall. We also improved a webcam-only eye tracking system, CVC Eye Tracker, by porting half of its processing stages on the GPU, a promising work to create applications relying on gaze interaction.
Speaker adaptation techniques allow hidden Markov model (HMM) based speech synthesis systems to mimic a target voice of which a few samples are available. However, usual adaptation approaches are not applicable when the target voice is dysarthric, i.e. the target speaker has an impairment which pre-vents the correct pronunciation of some phonemes. As a first step towards giving personalized synthetic voices to these particular speakers, this paper explores the possibility of adapting the whole statistical voice model using frequency warping (FW) based transformations trained exclusively with vowels. Perceptual evaluations performed for healthy voices show that the proposed method achieves reasonable results even when the adaptation data exhibit medium/low recording quality.
This paper introduces the Synthetic Speech Detection system developed by Aholab for the Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof 2015). The detector is a classifier based on Gaussian Mixture Models that are created using the Relative Phase Shift (RPS) transformation for the phase information. Different strategies have been evaluated: modeling the specific attacks using the information provided by the ASVspoof 2015 organizers, and modeling the vocoders possibly used in the spoofing signals, using data from previous works. The evaluation results show that attack specific models work for known attacks but they do not cope with the unknown attacks correctly. When using vocoder models build with other databases, the results suggest that the followed strategy do not take advantage of the available data and thus model adaptation should be explored. Index Terms: synthetic speech detection, phase information, anti-spoofing
In the field of speaker verification (SV) it is nowadays feasible and relatively easy to create a synthetic voice to deceive a speech driven biometric access system. This paper presents a synthetic speech detector that can be connected at the front-end or at the back-end of a standard SV system, and that will protect it from spoofing attacks coming from state-of-the-art statistical Text to Speech (TTS) systems. The system described is a Gaussian Mixture Model (GMM) based binary classifier that uses natural and copy-synthesized signals obtained from the Wall Street Journal database to train the system models. Three different state-of-the-art vocoders are chosen and modeled using two sets of acoustic parameters: 1) relative phase shift and 2) canonical Mel Frequency Cepstral Coefficients (MFCC) parameters, as baseline. The vocoder dependency of the system and multivocoder modeling features are thoroughly studied. Additional phase-aware vocoders are also tested. Several experiments are carried out, showing that the phase-based parameters perform better and are able to cope with new unknown attacks. The final evaluations, testing synthetic TTS signals obtained from the Blizzard challenge, validate our proposal.
This article explores the potential of the harmonics plus noise model of speech in the development of a high-quality vocoder applicable in statistical frameworks, particularly in modern speech synthesizers. It presents an extensive explanation of all the different alternatives considered during the design of the HNM-based vocoder, together with the corresponding objective and subjective experiments, and a careful description of its implementation details. Three aspects of the analysis have been investigated: refinement of the pitch estimation using quasi-harmonic analysis, study and comparison of several spectral envelope analysis procedures, and strategies to analyze and model the maximum voiced frequency. The performance of the resulting vocoder is shown to be similar to that of state-of-the-art vocoders in synthesis tasks.