Electro-laryngeal (EL) speech is characterized by constant pitch, limited prosody, and mechanical noise, reducing naturalness and intelligibility. We propose a lightweight adaptation of the state-of-the-art StreamVC framework to this setting by removing pitch and energy modules and combining self-supervised pretraining with supervised fine-tuning on parallel EL and healthy (HE) speech data, guided by perceptual and intelligibility losses. Objective and subjective evaluations across different loss configurations confirm their influence: the best model variant, based on WavLM features and human-feedback predictions (+WavLM+HF), drastically reduces character error rate (CER) of EL inputs, raises naturalness mean opinion score (nMOS) from 1.1 to 3.3, and consistently narrows the gap to HE ground-truth speech in all evaluated metrics. These findings demonstrate the feasibility of adapting lightweight voice conversion architectures to EL voice rehabilitation while also identifying prosody generation and intelligibility improvements as the main remaining bottlenecks.
Pathological speech, caused by dysphonia or produced via electro-larynx devices, often suffers from poor intelligibility and unnatural prosody. In this paper, we investigate the potential of four state-of-the-art voice conversion models: FreeVC, QuickVC, LLVC, and XVC for restoring healthy-sounding speech. All models are fine-tuned on Austrian-German datasets and evaluated using objective and subjective metrics. Results show substantial gains in intelligibility, naturalness, and perceived vocal health. QuickVC, FreeVC, and XVC perform similarly and achieve the highest preference scores, exceeding unprocessed pathological speech by up to 200%. These findings highlight the potential to improve communication for individuals with voice disorders and motivate further development of efficient, high-quality conversion systems.
Medical image segmentation is crucial for clinical applications, but challenges persist due to noise and variability. In particular, accurate glottis segmentation from high-speed videos is vital for voice research and diagnostics. Manual searching for failed segmentations is labor-intensive, prompting interest in automated methods. This paper proposes the first deep learning approach for detecting faulty glottis segmentations. For this purpose, faulty segmentations are generated by applying both a poorly performing neural network and perturbation procedures to three public datasets. Heavy data augmentations are added to the input until the neural network’s performance decreases to the desired mean intersection over union (IoU). Likewise, the perturbation procedure involves a series of image transformations to the original ground truth segmentations in a randomized manner. These data are then used to train a ResNet18 neural network with custom loss functions to predict the IoU scores of faulty segmentations. This value is then thresholded with a fixed IoU of 0.6 for classification, thereby achieving 88.27% classification accuracy with 91.54% specificity. Experimental results demonstrate the effectiveness of the presented approach. Contributions include: (i) a knowledge-driven perturbation procedure, (ii) a deep learning framework for scoring and detecting faulty glottis segmentations, and (iii) an evaluation of custom loss functions.
Vocal fry is a voice quality that occurs in a healthy voice, but it can also be a sign of a voice disorder. In this study, we investigated the relationship between the parameters of voice production, a dedicated psychoacoustic feature, and the perceptual aspects of vocal fry. Two perceptual experiments were carried out to determine whether the fundamental frequency, the open quotient, and the glottal area pulse skewness affect the perception of vocal fry in synthetic vowels. Thirteen listeners participated in the perceptual experiments to assess the following attributes: binary fry (yes/no) and impulsiveness, tonality, and naturalness (7-point Likert scales). The results suggest that the perception of vocal fry is mainly triggered by a low fundamental frequency, but the open quotient also plays a role, with narrower glottal area pulses slightly increasing the probability of perceived fry. Perceived tonality is inversely related to perceived impulsiveness. Internal reference standards of listeners appear to have fixed elements but may also be affected by anchoring and the short-term (i.e., within-vowel) context of the stimuli. In addition, the prominence of the peaks observed in the loudness curve over time appears to be related to graduations of fry.
PURPOSE:Laryngeal high-speed videoendoscopy (LHSV) has been recognized as a highly valuable modality for the scientific investigations of vocal fold (VF) vibrations. In contrast to stroboscopic imaging, LHSV enables visualizing aperiodic VF vibrations. However, the technique is less well established in the clinical care of disordered voices, partly because the properties of aperiodic vibration patterns are not yet described comprehensively. To address this, a computer model for simulation of VF vibration patterns observed in a variety of different phonation types is proposed. METHOD:A previously published kinematic model of mucosal wave phenomena is generalized to be capable of left-right asymmetry and to simulate endoscopic videos instead of only kymograms of VF vibrations at single sagittal positions. The most influential control parameters are the glottal halfwidths, the oscillation frequencies, the amplitudes, and the phase delays. RESULTS:The presented videos demonstrate zipper-like vibration, pressed voice, voice onset, constant and time-varying left-right and anterior-posterior phase differences, as well as left-right frequency differences of the VF vibration. Video frames, videokymograms, phonovibrograms, glottal area waveforms, and waveforms of VF contact area relating to electroglottograms are shown, as well as selected kinematic parameters. CONCLUSION:The presented videos demonstrate the ability to produce vibration patterns that are similar to those typically seen in endoscopic videos obtained from vocally healthy and dysphonic speakers. SUPPLEMENTAL MATERIAL:https://doi.org/10.23641/asha.20151833.
The decision of whether to perform cochlear implantation is crucial because implantation cannot be reversed without harm. The aim of the study was to compare model-predicted time–place representations of auditory nerve (AN) firing rates for normal hearing and impaired hearing with a view towards personalized indication of cochlear implantation. AN firing rates of 1024 virtual subjects with a wide variety of different types and degrees of hearing impairment were predicted. A normal hearing reference was compared to four hearing prosthesis options, which were unaided hearing, sole acoustic amplification, sole electrical stimulation, and a combination of the latter two. The comparisons and the fitting of the prostheses were based on a ‘loss of action potentials’ (LAP) score. Single-parameter threshold analysis suggested that cochlear implantation is indicated when more than approximately two-thirds of the inner hair cells (IHCs) are damaged. Second, cochlear implantation is also indicated when more than an average of approximately 12 synapses per IHC are damaged due to cochlear synaptopathy (CS). Cochlear gain loss (CGL) appeared to shift these thresholds only slightly. Finally, a support vector machine predicted the indication of a cochlear implantation from hearing loss parameters with a 10-fold cross-validated accuracy of 99.2%.
The characterization of voice quality is important for the diagnosis of a voice disorder. Vocal fry is a voice quality which is traditionally characterized by a low frequency and a long closed phase of the glottis. However, we also observed amplitude modulated vocal fry glottal area waveforms (GAWs) without long closed phases (positive group) which we modelled using an analysis-by-synthesis approach. Natural and synthetic GAWs are modelled. The negative group consists of euphonic, i.e., normophonic GAWs. The analysis-by-synthesis approach fits two modelled GAWs for each of the input GAW. One modelled GAW is modulated to replicate the amplitude and frequency modulations of the input GAW and the other modelled GAW is unmodulated. The modelling errors of the two modelled GAWs are determined to classify the GAWs into the positive and the negative groups using a simple support vector machine (SVM) classifier with a linear kernel. The modelling errors of all vocal fry GAWs obtained using the modulating model are smaller than the modelling errors obtained using the unmodulated model. Using the two modelling errors as predictors for classification, no false positives or false negatives are obtained. To further distinguish the subtypes of amplitude modulated vocal fry GAWs, the entropy of the modulator’s power spectral density and the modulator-to-carrier frequency ratio are obtained.
Diplophonia is a type of disordered voice in which two simultaneous pitches are perceived. Most commonly in diplophonic voices, the vocal folds are divided into two parts that vibrate at different frequencies. The glottal area is the projected area of the space between the vocal folds. The glottal area in time is referred to as the glottal area waveform (GAW). The GAW is modeled for diplophonic voice by superimposing two partial GAWs (pGAWs) that are trains of single-peak pulses with different pulse frequencies, i.e., fundamental frequencies ($f_o$s). In current kinematic models of diplophonic vocal fold vibration, the pGAWs are assumed to be quasiperiodic. This assumption is mitigated here by modulating pulse-to-pulse cycle length and amplitude. Both random and deterministic modulations are considered. Deterministic modulations depend on the difference of the pGAWs' instantaneous phases. Model GAWs are fitted to input GAWs using an analysis-by-synthesis approach which we refer to as `modulated pulse trains decomposition' (MPD). MPD is shown to be applicable to diplophonic as well as to nondiplophonic types of dysphonia, which include multi-pulse patterns, random timing behaviours, and chaos. It is mostly robust against modulations but degraded by large random modulations. MPD is compared to a deep autoencoder neural network, and the WaveGlow neural network. In terms of time-domain fitting errors, MPD outperforms the other two approaches unless random modulations are large. MPD outperforms the best of the other two approaches by up to approximately 5 dB. For large random modulations, the deep autoencoder network achieves the smallest fitting errors. In terms of magnitude spectrum fitting errors, WaveGlow is superior except for natural input GAWs containing only nondiplophonic types of dysphonia. Also pulse timing errors are shown to be advantageous for MPD.
The quality and timbre of disordered voices heavily rely on the vibration properties of the vocal folds. We discuss the representation of sagittal phase differences in vocal fold oscillations through a numerical biomechanical model involving lumped elements as well as distributed elements, i.e., delay lines. A dynamic glottal source model is proposed in which the fold displacement along the vertical and the sagittal dimensions is modelled using delay lines. In contrast to other models, with which the reproduction of sagittal phase differences is impossible (e.g., in two-mass models) or not easy to control (e.g., in 3D 16-mass and multi-mass models in general), the one proposed here provides direct control over the amount of phase delay between folds’ oscillations at the posterior and anterior part of the glottis, i.e., the sagittal axis, and at the superior and inferior part of the glottis, i.e., the vertical axis, while keeping the dynamic model simple and computationally efficient. The model is assessed by addressing the reproduction of oscillatory patterns observed in high-speed videoendoscopic data, in which sagittal phase differences are observed. Also, timing asymmetry parameters observed in hemi glottal area waveforms (GAWs) are used for fitting.
In this paper we present the extraction of kinematic vocal fold parameters from videokymographic images and the visual estimation of phase differences between the left and right vocal folds. We used a model of vocal fold vibrations with kinematic parameters to generate synthetic kymograms. To extract kinematic model parameters from the images, we propose a fitting procedure for error minimization. We used the "Structural Dissimilarity Index Measure" (DSSIM) and the "Cross Uncorrelation" (CUC) as error measures. The minimization procedure was used to evaluate 55 clinical kymograms and two sets of 50 synthetic kymograms each. The two sets of synthetic kymograms were generated using the probability density functions (PDFs) of the model parameters, which were obtained by fitting the clinical kymograms using the two error measures. The relative parameter estimation errors range up to 15% for the most important parameters, indicating acceptable performance of the fitting procedure. Additionally, the phase difference was assessed by three observers with integer ratings ranging from 0 (negligible) to 3 (large phase difference). The Fleiss' Kappa of 0.46 and its standard error 0.03 indicated a moderate agreement among the three observers. ROC analysis was carried out for estimating thresholds for predicting the ratings. Moderate to good prediction accuracies (0.69-0.91) for the clinical corpus and very good prediction accuracy (0.92-1.00) for the synthetic corpora were observed.
L’objectif est l’etude des causes des disperiodicites des voix du type 1 qui sont pseudo-periodiques et monophoniques. Un modele qui explique quantitativement les perturbations des durees de cycles glottiques fait appel aux fluctuations de la tension du muscle vocal. Or, ces fluctuations n’expliquent pas l’enrouement qui peut faire suite a une charge vocale ou une laryngite legere, par exemple. C’est pourquoi, nous discutons plusieurs modeles qui montrent qu’une redistribution des amplitudes vibratoires entre le corps et la couverture du pli module les perturbations qui trouvent leur origine au niveau du muscle vocal. Des simulations a l’aide d’un modele corps-couverture suggerent ainsi que les perturbations des durees des cycles glottiques augmentent avec une redistribution des amplitudes vibratoires de la couverture vers le muscle suite a une redistribution des masses vibrantes du muscle vers la couverture.
Breathiness is a voice quality type that may be a sign of a voice disorder. It involves the auditory perception of additive noise caused by turbulent trans-glottal airflow. Ten normal and ten breathy phonations recorded during high-speed videolaryngoscopy are analyzed. A method for extracting additive noise from audio recordings in the presence of modulation noise is presented. It generates quasi-unit pulse trains at a rate equal to the vocal frequency. The cycle shape is obtained by cross-correlating the pulse train with the audio signal. The cycle shape is then Fourier transformed and input to a Fourier synthesizer. Other inputs are the estimates of the instantaneous phase and the amplitude modulation function. The timing and height of the pulses are modulated so as to minimize the modeling error, which is the difference between the recorded audio signal and the output of the synthesizer. The error is assumed to approximate the additive noise in the voice. Differences of energy levels and spectral slope with respect to voice quality (breathy/normal) are reported. No differences with regard to cycle phase (open/quasi-closed) are reported.
We discuss the representation of anterior-posterior (A-P) phase differences in vocal cord oscillations through a numerical biomechanical model involving lumped elements as well as distributed elements, i.e., delay lines. A dynamic glottal source model is illustrated in which the fold displacement along the vertical and the longitudinal dimensions is explicitly modeled by numerical waveguide components representing the propagation on the fold cover tissue. In contrast to other models of the same class, in which the reproduction of longitudinal phase differences are intrinsically impossible (e.g., in two-mass models) or not easy to control explicitely (e.g., in 3D 16-mass and multi-mass models in general), the one proposed here provides direct control over the amount of phase delay between folds oscillations at the posterior and anterior side of the glottis, while keeping the dynamic model simple and computationally efficient. The model is assessed by addressing the reproduction of typical oscillatory patterns observed in high-speed videoendoscopic data, in which A-P phase differences are observed. Experimental results are provided which demonstrate the ability of the approach to effectively reproduce different oscillatory patterns of the vocal folds.
Background and objectives: The description of production kinematics of dysphonic voices plays an important role in the clinical care of voice disorders. However, high-speed videolaryngoscopy is not routinely used in clinical practice, partly because there is a lack of diagnostic markers that may be obtained from high-speed videos automatically. Aim of the study is to propose and test a procedure that automatically detects extra pulses, which may occur in voiced source signals of pathological voices in addition to cyclic pulses. Material and methods: Glottal area waveforms (GAW) are synthesized and used to test a detector for extra pulses. Regarding synthesis, for each GAW a cyclic pulse train is mixed with an extra pulse train, and additive noise. The cyclic pulse trains are varied across GAWs in terms of fundamental frequency, pulse shape, and modulation noise, i.e., jitter and shimmer. The extra pulse trains are varied across GAWs in terms of the height of the extra pulses, and their rates of occurrence. The energy level of the additive noise is also varied. Regarding detection, first, the fundamental frequency is estimated jointly with the cyclic pulse train waveform, second, the modulation noise is estimated, and finally the extra pulse train waveform is estimated. Two versions of the detector are compared, i.e., one that parameterizes the shapes of the cyclic pulses, and one that uses unparameterized pulse shape estimates. Two corpora are used for testing, i.e., one with 100 GAWs containing random extra pulses, and one with 25 GAWs containing extra pulses in the closed phases of each glottal phase representing subharmonic voices. Results and discussion: With pulse shape parameterization (PSP) a maximum mean accuracy of 88.3% is achieved when detecting random extra pulses. Without PSP, the maximum mean accuracy reduces to 82.9%. Detection performance decreases if the energy level of additive noise is higher than -25 dB with respect to the energy of the cyclic pulse train, and if the irregularity strength exceeds 0.1. For bicyclic, i.e., subharmonic voices, the approach fails without PSP, whereas with PSP, a mean sensitivity of 87.4% is achieved for subharmonic voices. Conclusion: A synthesizer for GAWs containing extra pulses, and a detector for extra pulses are proposed. With PSP, favorable detector performance is observed for not too high levels of additive noise and irregularity strengths. In signals with high noise levels, the detector without PSP outperforms the other one. Detection of extra pulses fails if irregularity strength is large. For subharmonic voices PSP must be used. (C) 2019 The Authors. Published by Elsevier Ltd.
Perturbations of the strict periodicity of the glottal vibrations are relevant features of the voice quality of normophonic and dysphonic speakers. Vocal perturbations in healthy speakers are assigned different names according to the range of the typical perturbation frequencies. The objective of the presentation is to model jitter and flutter, which are in the > 20Hz and 10Hz range respectively, via a simulation of the fluctuations of the tension of the thyro-arytenoid muscle and compare simulated perturbations to jitter and flutter observed in vowels sustained by normophonic speakers.
OBJECTIVES:Diplophonia is a common symptom of voice disorder that is in need of objectification. We investigated whether diplophonia can be detected from audio recordings of text readings by means of dedicated audio signal processing, ie, a descendant of a formerly published "Diplophonia Diagram." STUDY DESIGN:Diagnostic study. METHODS:Forty subjects were included who had been clinically rated in the past as diplophonic. For each subject, the audio signal of the German standard text "Der Nordwind und die Sonne" was recorded. First, subject groups regarding the frequency of occurrence of diplophonic episodes were established via manual labeling of audio recordings. Reference boundaries of diplophonic time intervals and the boundaries of voiced time intervals were manually obtained. Each time interval was labeled as diplophonic or nondiplophonic, as well as voiced or unvoiced. The diplophonia rate was defined as the total duration of diplophonation among the total duration of voiced phonation. Based on the diplophonia rate obtained from manual annotations, subjects were distinguished who were (1) frequently diplophonic, (2) unfrequently diplophonic, and (3) nondiplophonic during the reading of the standard text. Second, the grouping was predicted automatically via audio signal processing, and the performance of automatic prediction was evaluated. The audio recordings were analyzed with a purpose-built audio signal processor that estimated the diplophonia rate automatically. Two cut-off threshold classifiers were trained to detect automatically (1) frequently diplophonic, and (2) nondiplophonic subjects. In addition, multinomial logistic regression was performed to enable automatic 3-way classification. RESULTS:Among all subjects, 14 were frequently diplophonic during the reading of the text, 14 were unfrequently diplophonic, and the remaining 12 were nondiplophonic. In automated detection of frequently diplophonic subjects, a sensitivity of 71% and a specificity of 88% were obtained. The sensitivity and specificity regarding automated detection of nondiplophonic subjects were 68% and 92%. In 3-way classification, 62.5% of the subjects were classified into the correct group. CONCLUSIONS:Only two-thirds of the subjects who had been labeled as diplophonic on the base of auditory impression during clinical anamnesis diplophonated during the reading of a standard text. This demonstrates that the ecological validity of audio recordings of standard text readings is limited. Subject groups regarding the frequency of occurrence of diplophonic episodes were established and audio signal processing enabled automated classification. The observed performance of automated classification was promising and may be relevant to future clinical and scientific work. Possible applications include objective clinical voice assessment for diagnostic purposes and feedback based training of clinical raters.
The recently published study of Gaskill et al constitutes a remarkable contribution to the understanding of time-variance of dysphonic voice quality, because time-variance is often neglected in the description of voice quality. This circumstance may cause problems in clinical voice quality evaluation. We acknowledge the authors' contributions to clarify the different semantic levels of voice quality description, which are the speech level, the segmental level, the signal frame level, and the level of cycles. Although we perceive the consideration of time-variant voice quality as a major strength of the study by Gaskill et al, we are looking forward to interesting discussions regarding the interpretation of the cepstral peak prominence (CPP) as an acoustic correlate to perceptual voice quality, which has a long tradition. 1 Noll A. Cepstrum pitch determination. J Acoust Soc Am. 1967; 41: 293-309 Crossref PubMed Scopus (505) Google Scholar , 2 Hillenbrand J. Cleveland R.A. Erickson R.L. Acoustic correlates of breathy vocal quality. J Speech Hear Res. 1994; 37: 769-778 Crossref PubMed Scopus (377) Google Scholar The viewpoints that we would like to add to this discussion are outlined hereafter. Response to Aichinger and Kubin Re: Letter to the Editor “Acoustic and Perceptual Classification of Within-Sample Normal, Intermittently Dysphonic, and Consistently Dysphonic Voice Types”Journal of VoiceVol. 32Issue 3PreviewThanks to Drs Aichinger and Kubin for their consideration of our recent article “Acoustic and Perceptual Classification of Within-sample Normal, Intermittently Dysphonic, and Consistently Dysphonic Voice Types” and for their interesting comments. Our study was an initial attempt at using measures of the cepstral peak prominence (CPP) distribution to identify intermittently dysphonic samples that can be difficult to accurately classify both perceptually and acoustically. In addition to examination of the moments of the CPP distribution (mean, standard deviation, skewness, and kurtosis), we also considered measures for the detection of extreme CPP values used in statistical process control and measures of the relative percentage of dysphonia duration in our analyses. Full-Text PDF
Diplophonia is a type of pathological voice in which two fundamental frequencies (f o ) are present simultaneously. Specialized audio analyzers that can handle up to two f o s in diplophonic voices are in their infancy. We propose the tracking of up to two f o s in diplophonic voices by audio waveform modeling (AWM), which involves obtaining candidates by repetitive execution of the Viterbi algorithm, followed by waveform Fourier synthesis, and heuristic candidate selection with majority voting. Our approach is evaluated with reference f o -tracks obtained from laryngeal highspeed videos of 29 sustained phonations and compared to state-of-the-art tracking algorithms for multiple f o s. An accurate and a fast variant of our algorithm are tested. The median error rate of the accurate variant is 6.52%, whereas the most accurate benchmark achieves 11.11%. The fast variant is more than twice as fast as the fastest relevant benchmark, and the median error rate is 9.52%. Furthermore, illustrative results of connected speech analysis are reported. Our approach may help to improve detection and analysis of diplophonia in clinical research and practice, as well as to advance synthesis of disordered voices.
Jean Schoentgen合作论文数National Fund for Scientific Research, Belgium23