This article is intended as an introductory tutorial for technically inclined clinicians, vocologists and voice pedagogues who want to understand the principles and potentials of voice mapping. Voice mapping has its origins in the Voice Range Profile, or phonetogram, but it is less concerned with the extremes of the voice range, and more with what happens within a relevant range of the voice. It is a voice instrumentation paradigm that is intended to improve the evidential value of voice measurements. It exposes and automatically accounts for the strong co-variation that most voice metrics exhibit with fundamental frequency and sound level. Very many data points are automatically collected in a short time, and their means are mapped by colour onto maps. This results in a robust representation of voice status and function. While individual voices are very different, a voice map’s appearance is reproducible within individuals. Comparing maps across interventions gives rich information, even on subtle changes in a voice. Moreover, by statistically clustering multiple metrics, phonation types can be identified and mapped automatically, thereby enhancing clinical relevance, and facilitating a deeper understanding of voice data.
This study investigates voice mapping as an evaluation framework for text-to-speech (TTS) synthesis quality. The study analyzes six TTS models, including historical and recent ones. The metrics are crest factor, spectrum balance, and cepstral peak prominence (CPPs). We investigated 6 influential TTS models: Merlin, Tacotron 2, Transformer TTS, FastSpeech 2, Glow-TTS, and VITS. The results demonstrate that voice range serves as a primary indicator of model capability, with VITS showing the largest range among tested models. Glow-TTS exhibited superior performance in soft phonation, indicated by higher spectrum balance, despite limited voice range. The results showed that the CPPs values between 7-8 dB indicate natural voice quality, while with CPPs exceeding 10 dB, the speech tends to sound robotic. These findings underscore the need for voice mapping to evaluate vocal effort, and capture how TTS systems handle voice dynamic and expressiveness.
The electroglottographic (EGG) signal offers a non-invasive approach to analyze phonation. It is known, if not obvious, that the onset of vocal fold contacting has a substantial effect on how the vocal folds vibrate and on the quality of the voice. Given that the presence or absence of vocal fold contacting has major consequences also for the interpretation of acoustic metrics, it is compelling to consider the possibility of predicting EGG signals directly from the microphone speech signal. This retrospective study presents a neural network model for EGG signal estimation utilizing a WaveNet architecture augmented with a self-attention mechanism. The model was trained on an existing dataset that comprehensively recorded participants' full voice range. The proposed model effectively captures the temporal dynamics and morphological characteristics of normophonic EGG waveforms, achieving outputs that closely resemble the ground truth in terms of EGG waveshape and extracted EGG metrics. For evaluation, voice mapping was used to display the distribution similarities of extracted metrics from predicted and ground truth EGG waveforms. The model exhibits proficiency in accurately estimating EGG signals in areas of stable and contacting voicing but displays reduced accuracy in transitional and breathy phonatory conditions.
We investigated the causal basis of abrupt frequency jumps in a unique database of New World monkey vocalizations. We used a combination of acoustic and electroglottographic recordings in vivo, excised larynx investigations of vocal fold dynamics, and computational modelling. We particularly attended to the contribution of the vocal membranes: thin upward extensions of the vocal folds found in most primates but absent in humans. In three of the six investigated species, we observed two distinct modes of vocal fold vibration. The first, involving vocal fold vibration alone, produced low-frequency oscillations, and is analogous to that underlying human phonation. The second, incorporating the vocal membranes, resulted in much higher-frequency oscillation. Abrupt fundamental frequency shifts were observed in all three datasets. While these data are reminiscent of the rapid transitions in frequency observed in certain human singing styles (e.g. yodelling), the frequency jumps are considerably larger in the nonhuman primates studied. Our data suggest that peripheral modifications of vocal anatomy provide an important source of variability and complexity in the vocal repertoires of nonhuman primates. We further propose that the call repertoire is crucially related to a species' ability to vocalize with different laryngeal mechanisms, analogous to human vocal registers. This article is part of the theme issue 'Nonlinear phenomena in vertebrate vocalizations: mechanisms and communicative functions'.
Purpose: Literature suggests a dependency of the acoustic metrics, smoothed cepstral peak prominence (CPPS) and harmonics-to-noise ratio (HNR), on human voice loudness and fundamental frequency ( F 0). Even though this has been explained with different oscillatory patterns of the vocal folds, so far, it has not been specifically investigated. In the present work, the influence of three elicitation levels, calibrated sound pressure level (SPL), F 0 and vowel on the electroglottographic (EGG) and time-differentiated EGG (dEGG) metrics hybrid open quotient (OQ), dEGG OQ and peak dEGG, as well as on the acoustic metrics CPPS and HNR, was examined, and their suitability for voice assessment was evaluated. Method: In a retrospective study, 29 women with a mean age of 25 years (± 8.9, range: 18–53) diagnosed with structural vocal fold pathologies were examined before and after voice therapy or phonosurgery. Both acoustic and EGG signals were recorded simultaneously during the phonation of the sustained vowels /ɑ/, /i/, and /u/ at three elicited levels of loudness (soft/comfortable/loud) and unconstrained F 0 conditions. Results: A linear mixed-model analysis showed a significant effect of elicitation effort levels on peak dEGG, HNR, and CPPS (all p < .01). Calibrated SPL significantly influenced HNR and CPPS (both p < .01). Furthermore, F 0 had a significant effect on peak dEGG and CPPS ( p < .0001). All metrics showed significant changes with regard to vowel (all p < .05). However, the treatment had no effect on the examined metrics, regardless of the treatment type (surgery vs. voice therapy). Conclusions: The value of the investigated metrics for voice assessment purposes when sampled without sufficient control of SPL and F 0 is limited, in that they are significantly influenced by the phonatory context, be it speech or elicited sustained vowels. Future studies should explore the diagnostic value of new data collation approaches such as voice mapping, which take SPL and F 0 effects into account.
ObjectivesThis study aims to explore the effects of thyroidectomy—a surgical intervention involving the removal of the thyroid gland—on voice quality, as represented by acoustic and electroglottographic measures. Given the thyroid gland's proximity to the inferior and superior laryngeal nerves, thyroidectomy carries a potential risk of affecting vocal function. While earlier studies have documented effects on the voice range, few studies have looked at voice quality after thyroidectomy. Since voice quality effects could manifest in many ways, that a priori are unknown, we wish to apply an exploratory approach that collects many data points from several metrics.MethodsA voice-mapping analysis paradigm was applied retrospectively on a corpus of spoken and sung sentences produced by patients who had thyroid surgery. Voice quality changes were assessed objectively for 57 patients prior to surgery and 2months after surgery, by making comparative voice maps, pre- and post-intervention, of six acoustic and electroglottographic (EGG) metrics.ResultsAfter thyroidectomy, statistically significant changes consistent with a worsening of voice quality were observed in most metrics. For all individual metrics, however, the effect sizes were too small to be clinically relevant. Statistical clustering of the metrics helped to clarify the nature of these changes. While partial thyroidectomy demonstrated greater uniformity than did total thyroidectomy, the type of perioperative damage had no discernible impact on voice quality.ConclusionsChanges in voice quality after thyroidectomy were related mostly to increased phonatory instability in both the acoustic and EGG metrics. Clustered voice metrics exhibited a higher correlation to voice complaints than did individual voice metrics.
The human voice is notoriously variable, and conventional measurement paradigms are weak in terms of providing evidence for effects of treatment and/or training of voices. New methods are needed that can take into account the variability of metrics and types of phonation across the voice range. The “voice map” is a generalization of the Voice Range Profile (a.k.a. the phonetogram), with the potential to be used in many ways, for teaching, training, therapy and research. FonaDyn is intended as a proof-of concept workbench for education and research on phonation, and for exploring and validating the analysis paradigm of voice-mapping. Version 3.1 of the FonaDyn system adds many new functions, including listening from maps; displaying multiple maps and difference maps to track effects of voice interventions; smoothing/interpolation of voice maps; clustering not only of EGG shapes but also of acoustic and EGG metrics into phonation types; extended multichannel acquisition; 24-bit recording with optional max 140 dB SPL; a built-in SPL calibration and signal diagnostics tool; EGG noise suppression; more Matlab integration; script control; the acoustic metrics Spectrum Balance, Cepstral Peak Prominence and Harmonic Richness Factor (of the EGG); and better window layout control. Stability and usability are further improved. Apple M-series processors are now supported natively.
We can make many different sounds with our voices, communicating not only with what is said, but also in which context it is said, and who is saying it. This large variability makes the voice a rich channel for communication, but it also presents us with challenges when we try to assess the status of a voice using quantitative measurements, rather than by listening. When voices run into trouble, even more variability can be expected. Voice production is usually described as three processes in sequence: respiration – breathing; phonation – the vibration of the vocal folds; and articulation – changing the shape of the vocal tract, which modifies the sound into vowels and consonants. For brevity, let’s look only at some aspects of phonation that can be expected to be clinically relevant.
In voice analysis, the electroglottographic (EGG) signal has long been recognized as a useful complement to the acoustic signal, but only when the vocal folds are actually contacting, such that this signal has an appreciable amplitude. However, phonation can also occur without the vocal folds contacting, as in breathy voice, in which case the EGG amplitude is low, but not zero. It is of great interest to identify the transition from non-contacting to contacting, because this will substantially change the nature of the vocal fold oscillations; however, that transition is not in itself audible. The magnitude of the cycle-normalized peak derivative of the EGG signal is a convenient indicator of vocal fold contacting, but no current EGG hardware has a sufficient signal-to-noise ratio of the derivative. We show how the textbook techniques of spectral thresholding and static notch filtering are straightforward to implement, can run in real time, and can mitigate several noise problems in EGG hardware. This can be useful to researchers in vocology.
Recent investigations on music performances have shown the relevance of singers' body motion for pedagogical as well as performance purposes. However, little is known about how the perception of voice-matching or task complexity affects choristers' body motion during ensemble singing. This study focussed on the body motion of choral singers who perform in duo along with a pre-recorded tune presented over a loudspeaker. Specifically, we examined the effects of the perception of voice-matching, operationalized in terms of sound spectral envelope, and task complexity on choristers' body motion. Fifteen singers with advanced choral experience first manipulated the spectral components of a pre-recorded short tune composed for the study, by choosing the settings they felt most and least together with. Then, they performed the tune in unison (i.e., singing the same melody simultaneously) and in canon (i.e., singing the same melody but at a temporal delay) with the chosen filter settings. Motion data of the choristers' upper body and audio of the repeated performances were collected and analyzed. Results show that the settings perceived as least together relate to extreme differences between the spectral components of the sound. The singers' wrists and torso motion was more periodic, their upper body posture was more open, and their bodies were more distant from the music stand when singing in unison than in canon. These findings suggest that unison singing promotes an expressive-periodic motion of the upper body.
Choir singers report anecdotally that two voices can be perceived as a good match for each other, or not. Could “matching voices” be explained by the spectrum envelopes? Thirteen singers sang a duo in unison or canon with an adjacent prerecorded reference singer, in a moderately reverberant room. Singers controlled the stimulus timbre, using variable filters in medium and high frequency bands. They were asked to adjust the filters, while singing, for “best” and “worst” perceived matching. The singers then performed the song again, but with the filters automatically set to their chosen (dis-)preferences. The Self-to-Other ratio as a function of frequency [SOR(f)] at the ipsilateral ear of the participant was estimated from multiple microphone signals to predict separately the long-time average spectra of Self and Other. Most participants rated the sound with extreme filter settings at ±15 dB as the “worst” match, while “best” matches were fairly evenly distributed. Some but not all participants preferred the spectra to be complementary. However, at low frequencies, SOR(f) was about +10 dB, and very irregular but rarely negative at medium and high frequencies; so how an adjacent singer can be heard at all will require further investigation.
The human voice production mechanism implements a superbly rich communication channel that at once tells us what, who, how, and much more [...]
Individual acoustic and other physical metrics of vocal status have long struggled to prove their worth as clinical evidence. While combinations of metrics or “features” are now being intensely explored using data analytics methods, there is a risk that explainability and insight will suffer. The voice mapping paradigm discards the temporal dimension of vocal productions and uses fundamental frequency (fo) and sound pressure level (SPL) as independent control variables to implement a dense grid of measurement points over a relevant voice range. Such mapping visualizes how most physical voice metrics are greatly affected by fo and SPL, and more so individually than has been generally recognized. It is demonstrated that if fo and SPL are not controlled for during task elicitation, repeated measurements will generate “elicitation noise”, which can easily be large enough to obscure the effect of an intervention. It is observed that, although a given metric’s dependencies on fo and SPL often are complex and/or non-linear, they tend to be systematic and reproducible in any given individual. Once such personal trends are accounted for, ordinary voice metrics can be used to assess vocal status. The momentary value of any given metric needs to be interpreted in the context of the individual’s voice range, and voice mapping makes this possible. Examples are given of how voice mapping can be used to quantify voice variability, to eliminate elicitation noise, to improve the reproducibility and representativeness of already established metrics of the voice, and to assess reliably even subtle effects of interventions. Understanding variability at this level of detail will shed more light on the interdependent mechanisms of voice production, and facilitate progress toward more reliable objective assessments of voices across therapy or training.
Abstract A typical performance situation of a vocal ensemble or choir consists of a group of singers in a room with listeners. The choir singers on stage interact while they sing, since they also hear the sound of the neighboring singer and react accordingly. From a physical point of view, the choir singers can be regarded as sound sources. The properties of the room influence the sound and the listeners perceive the sound event as a sound receiver. Furthermore, the processes in the choir can also be described acoustically, which affects the overall performance. The room influences the timbre of the sound on their way to the audience, the receiver. Reflection, absorption, diffraction, or refraction influence the timbre in the room. The sound in a performance space can be distinguished between a near field very close to the singer and a far field. The distance at which the far field can be assumed is strongly dependent on the acoustics of the room. Especially for singers within a choir, the differentiation between those sound fields is important for hearing oneself and the other singers. The position of the singers, their directivity, and the seating position of the listener in the audience will have an influence on listener perception. Furthermore, this chapter gives background information on intonation and synchronization aspects, which are most relevant for any vocal ensemble situation. Using this knowledge, intuitive behavior and performance practice can be explained and new adaptations can be suggested for singing in vocal ensembles.
For voice analysis, much work has been undertaken with a multitude of acoustic and electroglottographic metrics. However, few of these have proven to be robustly correlated with physical and physiological phenomena. In particular, all metrics are affected by the fundamental frequency and sound level, making voice assessment sensitive to the recording protocol. It was investigated whether combinations of metrics, acquired over voice maps rather than with individual sustained vowels, can offer a more functional and comprehensive interpretation. For this descriptive, retrospective study, 13 men, 13 women, and 22 children were instructed to phonate on /a/ over their full voice range. Six acoustic and EGG signal features were obtained for every phonatory cycle. An unsupervised voice classification model created feature clusters, which were then displayed on voice maps. It was found that the feature clusters may be readily interpreted in terms of phonation types. For example, the typical intense voice has a high peak EGG derivative, a relatively high contact quotient, low EGG cycle-rate entropy, and a high cepstral peak prominence in the voice signal, all represented by one cluster centroid that is mapped to a given color. In a transition region between the non-contacting and contacting of the vocal folds, the combination of metrics shows a low contact quotient and relatively high entropy, which can be mapped to a different color. Based on this data set, male phonation types could be clustered into up to six categories and female and child types into four. Combining acoustic and EGG metrics resolved more categories than either kind on their own. The inter- and intra-participant distributional features are discussed.
In this article we design an experimental setup to detect disturbances in voice recordings, such as additive noise, clipping, infrasound and random muting. The datasets are generated by introducing degradations into clean recordings. We test five different classification algorithms in both single- and multi-label settings: kernel substitution based support vector machine, convolutional neural network, long short-term memory (LSTM), and a hidden Markov model using either Gaussian mixture models or generative models in its state distribution. The LSTM achieved good results in both tests, most notably in the multi-label case where the average balanced accuracy was 82.7% on one dataset.
ObjectivesThe purpose of this study was to assess the outcome following continuous tactile biofeedback of voice sound level administered, with a portable voice accumulator to individuals with Parkinson's disease (PD).MethodNine out of 16 participants with PD completed a 4-week intervention program where biofeedback of voice sound level was administered with the portable voice accumulator VoxLog during speech in daily life. The feedback, a tactile vibration signal from the device, was activated when the wearer used a voice sound level below an individually predetermined threshold level, reminding the wearer to increase voice sound level during speech. Voice use was registered in daily life with the VoxLog during the intervention period as well as during one baseline week, one follow-up week post intervention and 1 week 3 months post intervention. Self-to-other ratio (SOR), which is the difference between voice sound level and environmental noise, was studied in multiple noise ranges.ResultsA significant increase in SOR across all noise ranges of 2.28 dB (SD: 0.55) was seen for participants with scores above the cut-off for normal function (>26 points) on the cognitive screening test Montreal Cognitive Assessment (MoCA) (n = 5). No significant increase was seen for the group of participants with MoCA scores below 26 (n = 4). Forty-four percent ended their participation early, all which scored below 26 on MoCA (n = 7).ConclusionsBiofeedback administered in daily life regarding voice level may help individuals with PD to increase their voice sound level in relation to environmental noise in daily life, but only for a limited subset. Only participants with normal cognitive function as screened by MoCA improved their voice sound level in relation to environmental noise.
Purpose The purpose of this study is to identify the extent to which various measurements of contacting parameters differ between children and adults during habitual range and overlap vocal frequency/intensity, using voice map–based assessment of noninvasive electroglottography (EGG). Method EGG voice maps were analyzed from 26 adults (22–45 years) and 22 children (4–8 years) during connected speech and vowel /a/ over the habitual range and the overlap vocal frequency/intensity from the voice range profile task on the vowel /a/. Mean and standard deviations of contact quotient by integration, normalized contacting speed, quotient of speed by integration, and cycle-rate sample entropy were obtained. Group differences were evaluated using the linear mixed model analysis for the habitual range connected speech and the vowel, whereas analysis of covariance was conducted for the overlap vocal frequency/intensity from the voice range profile task. Presence of a “knee” on the EGG wave shape was determined by visual inspection of the presence of convexity along the decontacting slope of the EGG pulse and the presence of the second derivative zero-crossing. Results The contact quotient by integration, normalized contacting speed, quotient of speed by integration, and cycle-rate sample entropy were significantly different in children compared to (a) adult males for habitual range and (b) adult males and adult females for the overlap vocal frequency/intensity. None of the children had a “knee” on the decontacting slope of the EGG slope. Conclusion EGG parameters of contact quotient by integration, normalized contacting speed, quotient of speed by integration, cycle-rate sample entropy, and absence of a “knee” on the decontacting slope characterize the wave shape differences between children and adults, whereas the normalized contacting speed, quotient of speed by integration, cycle-rate sample entropy, and presence of a “knee” on the downward pulse slope characterize the wave shape differences between adult males and adult females. Supplemental Material https://doi.org/10.23641/asha.15057345
Objective: Effects of exercises using a tool that promotes a semi-occluded artificially elongated vocal tract with real-time visual feedback of airflow - the flow ball - were tested using voice maps of EGG time-domain metrics. Methods: Ten classically trained singers (5 males and 5 females) were asked to sing messa di voce exercises on eight scale tones, performed in three consecutive conditions: baseline ('before'), flow ball phonation ('during'), and again without the flow ball ('after'). These conditions were repeated eight times in a row: one scale tone at a time, on an ascending whole tone scale. Audio and electroglottographic signals were recorded using a Laryngograph microprocessor. Vocal fold contacting was assessed using three time-domain metrics of the EGG waveform, using FonaDyn. The quotient of contact by integration, Q(ci), the normalized peak derivative, Q(Delta), and the index of contacting I-c, were quantified and compared between 'before' and 'after' conditions. Results: Effects of flow ball exercises depended on singers' habitual phonatory behaviours and on the position in the voice range. As computed over the entire range of the task, Q(ci), was reduced by about 2% in five of ten singers. Q(Delta) was 2-6% lower in six of the singers, and 3-4% higher only in the two bass-baritones. (I)c decreased by almost 4% in all singers. Conclusion: Overall, vocal adduction was reduced and a gentler vocal fold collision was observed for the 'after' conditions. Significance: Flow ball exercises may contribute to the modification of phonatory behaviours of vocal pressedness. (C) 2020 Elsevier Ltd. All rights reserved.