VoiceSauce (Shue et al., 2011) is an acoustic signal analysis software widely used for analyzing voice quality for various purposes (e.g., linguistic phonation acoustics: Keating et al., 2023; articulatory correlates: Wu & Zhang, 2023; speech disorder: Asiaee et al., 2022; singing: Meireles, 2016). One of the voice quality measurements provided by VoiceSauce is the amplitude difference between the first and second harmonics (H1–H2), a measurement widely used as a proxy for the degree of constriction of the vocal folds. The built-in formant correction formula (Iseli et al., 2007) in VoiceSauce enables users to remove the amplifying effect of vowel formants on harmonic energy, and consequently compare the harmonic energy across different vowel qualities. However, we found that the formant correction formula is unsuitable for correcting harmonics for formants much lower than the harmonic frequency. Such a correction will result in a large boost in the corrected harmonic energy. In this tutorial, we will demonstrate the rationale and problems of the formant correction formula used in VoiceSauce, how to modify the formant correction formula in VoiceSauce, and how to add additional harmonic parameters to VoiceSauce.
Many languages use phonation types for phonemic or allophonic distinctions. This study examines the acoustic structure of the phonetic space for vowel phonations across languages. Our sample of eleven languages includes languages with contrastive modal, breathy, creaky, lax, tense, harsh, and/or pharyngealized phonations, and languages with allophonic nonmodal phonation on particular tones. In compiling and analyzing this sample we address related issues such as contrast vs. allophony, phonetic similarity across languages, and understanding complex contrasts of several multidimensional phonetic categories via data reduction. Based on extensive acoustic analysis, all of the languages’ phonations were mapped into a single phonetic space, which exhibits dispersion (languages with more categories use more of the space). The space is largely two-dimensional, with dimensions that can be interpreted phonetically (e.g. dimension 2 is like a traditional breathy-to-creaky continuum) and also can be related back to the acoustic measures that structure them, thus indicating which acoustic measures are most important across languages.*
“Creaky voice” is a term that covers multiple kinds of voicing, and there is no single defining acoustic property shared by all subtypes of creaky voice. Here we explore the distinct characteristics of each subtype. We identify three main properties of creaky voice: low f0, irregular f0, and constricted glottis (as shown by electroglottography). Prototypical creaky voice has all of these properties; other subtypes are characterized by different subsets of properties. We will describe, with reference to previous literature, multiply-pulsed creak (a special case of irregular f0, along with low f0), open-glottis creak (low and often irregular f0, but unconstricted glottis), vocal fry (low but regular f0, constricted glottis), and creak with such irregular pulsing that no f0 can be recovered. Building on our previous work [Keating et al. 2015 Proc. ICPhS], we show how various acoustic measures pattern for each subtype. Results from parametric speech synthesis provide support for our acoustic observations.
The IPA currently does not specify how to represent prenasalization, preglottalization or preaspiration. We first review some current transcription practices, and phonetic and phonological literature bearing on the unitary status of prenasalized, preglottalized and preaspirated segments. We then propose that the IPA adopt superscript diacritics placed before a base symbol for these three phenomena. We also suggest how the current IPA Diacritics chart can be modified to allow these diacritics to be fit within the chart.
A speech production experiment with electroglottography investigated how voicing is affected by consonants of differing degrees of constriction. Measures of glottal contact [closed quotient (CQ)] and strength of voicing [strength of excitation (SoE)] were used in conditional inference tree analyses. Broadly, the results show that as the degree of constriction increases, both CQ and SoE values decrease, indicating breathier and weaker voicing. Similar changes in voicing quality are observed throughout the course of the production of a given segment. Implications of these results for a greater understanding of source-tract interactions and for the phonological notion of sonority are discussed.
We describe a new speech corpus designed to sample variability in speaking within individual speakers and across a large number of speakers. The public version of the database comprises audio recordings of 201 speakers performing 12 brief speech tasks over three recording sessions. Most of the tasks are unscripted, and include a phone call and pet-directed speech. The recordings have been orthographically transcribed, and dictionary broad transcriptions have been forcealigned. The database can be downloaded for free.
Little is known about the nature or extent of everyday variability in voice quality. This paper describes a series of principal component analyses to explore within- and between-talker acoustic variation and the extent to which they conform to expectations derived from current models of voice perception. Based on studies of faces and cognitive models of speaker recognition, the authors hypothesized that a few measures would be important across speakers, but that much of within-speaker variability would be idiosyncratic. Analyses used multiple sentence productions from 50 female and 50 male speakers of English, recorded over three days. Twenty-six acoustic variables from a psychoacoustic model of voice quality were measured every 5 ms on vowels and approximants. Across speakers the balance between higher harmonic amplitudes and inharmonic energy in the voice accounted for the most variance (females = 20%, males = 22%). Formant frequencies and their variability accounted for an additional 12% of variance across speakers. Remaining variance appeared largely idiosyncratic, suggesting that the speaker-specific voice space is different for different people. Results further showed that voice spaces for individuals and for the population of talkers have very similar acoustic structures. Implications for prototype models of voice perception and recognition are discussed.
Speakers of North American English are known to use a variety of tap/flap articulations depending on phonetic context (Derrick and Gick, 2011); it is also known that NAE taps/flaps are sometimes associated with a greatly lowered F4 frequency (Warner and Tucker, 2017). It has been less clear whether only certain articulatory variants show this acoustic effect. Since retroflex stops are also associated with lowered F4 (Blumstein and Stevens, 1975), we predict that flap retroflexion is associated with lowered F4. To test this prediction, synchronized ultrasound and audio recordings were made of words containing /t, d/ in a variety of contexts known to give rise to tap/flap variants. Based on visual inspection of ultrasound videos, these were coded as one of four articulatory variants (low tap, high tap, up flap, down flap: Derrick and Gick, 2011); formant frequencies were extracted from the audio at several timepoints relative to the tap/flap. Preliminary results from one speaker support the hypothesis: high taps and down flaps (variants with initial retroflexion) show an F4 drop into the consonant, while high taps and up flaps (variants with final retroflexion) show an F4 rise out of the consonant.
Little is known about human and machine speaker discrimination ability when utterances are very short and the speaking style is variable. This study compares text-independent speaker discrimination ability of humans and machines based on utterances shorter than 2 s in two different speaking styles (read sentences and speech directed towards pets, characterized by exaggerated prosody). Recordings of 50 female speakers drawn from the UCLA Speaker Variability Database were used as stimuli. Performance of 65 human listeners was compared to i-vector-based automatic speaker verification systems using mel-frequency cepstral coefficients, voice quality features, which were inspired by a psychoacoustic model of voice perception, or their combination by score-level fusion. Humans always outperformed machines, except in the case of style-mismatched pairs from perceptually-marked speakers. Speaker representations by humans and machines were compared using multi-dimensional scaling (MDS). Canonical correlation analysis showed a weak correlation between machine and human MDS spaces. Multiple regression showed that means of voice quality features could represent the most important human MDS dimension well, but not the dimensions from machines. These results suggest that speaker representations by humans and machines are different, and machine performance might be improved by better understanding how different acoustic features relate to perceived speaker identity.
Due to within-speaker variability in phonetic content and/or speaking style, the performance of automatic speaker verification (ASV) systems degrades especially when the enrollment and test utterances are short. This study examines how different types of variability influence performance of ASV systems. Speech samples (< 2 sec) from the UCLA Speaker Variability Database containing 5 different read sentences by 200 speakers were used to study content variability. Other samples (about 5 sec) that contained speech directed towards pets, characterized by exaggerated prosody. were used to analyze style variability. Using the i-vector/PLDA framework, the ASV system error rate with MFCCs had a relative increase of at least 265% and 730% in content-mismatched and style-mismatched trials, respectively. A set of features that represents voice quality (F0, F1, F2, F3, H1-H2, H2-H4, H4-H2k, Al, A2, A3, and CPP) was also used. Using score fusion with MFCCs, all conditions saw decreases in error rates. In addition, using the NIST SRE10 database, score fusion provided relative improvements of 11.78% for 5-second utterances, 12.41% for 10-second utterances, and a small improvement for long utterances (about 5 min). These results suggest that voice quality features can improve short-utterance text-independent ASV system performance.
Little is known about how to characterize normal variability in voice quality within and across utterances from normal speakers. Our previous study of female voices suggested that only a few acoustic parameters consistently distinguish speakers, with most of the work being done by idiosyncratic subsets of parameters. The present study extends this research to samples of 50 men’s voices. The men read 5 sentences twice on 3 days—30 sentences per speaker. The VoiceSauce analysis program estimated means and standard deviations for many acoustic parameters (including F0, harmonic amplitude differences, harmonic-to-noise ratios, and formant frequencies) for the vowels and approximant consonants in each sentence. Multidimensional scaling and linear discriminant analysis were used to examine the acoustic characteristics of the overall voice space, and to measure how well each speaker’s set of 30 sentences could be acoustically distinguished from all other speakers’ sentences. Additional analyses of small subsets of voices compared the importance of global versus local details in discriminating voices acoustically. Results will be compared to those for female speakers, and implications for recognition by listening will be discussed. [Work supported by NSF and NIH.]
Despite recent breakthroughs in automatic speaker recognition (ASpR), system performance still degrades when utterances are short and/or when within-speaker variability is large. This study used short test utterances (2-3sec) to investigate the effect of within-speaker variability on state-of-the-art ASpR system performance. A subset of a newly-developed UCLA database is used, which contains multiple speech tasks per speaker. The short utterances combined with a speaking-style mismatch between read sentences and spontaneous affective speech degraded system performance, for 25 female speakers, by 36%. Because humans are more robust to utterance length or withinspeaker variability, understanding human perception might benefit ASpR systems. Perception experiments were conducted with recorded read sentences from 3 female speakers, and a model is proposed to predict the perceptual dissimilarity between tokens. Results showed that a set of voice quality features including F0, F1, F2, F3, H1*-H2*, H2*-H4*, H4*-H2k*, H2k*-H5k, and CPP provides information that complements MFCCs. By fusing the feature set with MFCCs, human response prediction RMS error was .12, which represents a 12% relative error reduction compared to using MFCCs alone. In ASpR experiments with short utterances from 50 speakers, the voice quality feature set decreased the error rate by 11% when fused with MFCCs.
Recent advances in real-time magnetic resonance imaging (rtMRI) of the upper airway for acquiring speech production data provide unparalleled views of the dynamics of a speaker's vocal tract at very high frame rates (83 frames per second and even higher). This paper introduces an effort to collect and make available on-line rtMRI data corresponding to a large subset of the sounds of the world's languages as encoded in the International Phonetic Alphabet, with supplementary English words and phonetically-balanced texts, produced by four prominent phoneticians, using the latest rtMRI technology. The technique images oral as well as laryngeal articulator movements in the production of each sound category. This resource is envisioned as a teaching tool in pronunciation training, second language acquisition, and speech therapy.
Little is known about how to characterize normal variability in voice quality within and across utterances from normal speakers. Given a standard set of acoustic measures of voice, how similar are samples of 50 women’s voices? Fifty women, all native speakers of English, read 5 sentences twice on 3 days—30 sentences per speaker. The VoiceSauce analysis program estimated many acoustic parameters for the vowels and approximant consonants in each sentence, including F0, harmonic amplitude differences, harmonic-to-noise ratios, formant frequencies. Each sentence was then characterized by the mean and standard deviation of each measure. Linear discriminant analysis tested how well each speaker’s set of 30 sentences could be acoustically distinguished from all other speakers’ sentences. Initial work testing just 3 speakers from this sample found that the speakers could be completely discriminated (classified) by these measures, and largely discriminated by just 2 of them. Such a simple result is not expected for the larger sample of speakers. We will present results concerning how successfully speakers can be discriminated, how well different numbers of discriminant functions do, and which acoustic measures do the most work. Implications for recognition by listening will be discussed. [Work supported by NSF and NIH.]
While almost all of Ken Stevens’s research has been influential in linguistic phonetics—especially his work on the quantal nature of speech and on enhancement theory—this presentation will focus more on aspects not covered by others in this session. A noteworthy aspect of Ken’s career is that he frequently collaborated with linguists, notably in research on phonetic features and their structure. In work with Blumstein and with Halle, he provided acoustic correlates of place of articulation, laryngeal, and nasalization features. Much of this work is summarized in his 1980 paper in JASA, “Acoustic correlates of some phonetic categories.” In work with Keyser, he suggested an overall organization of features to define major classes of sounds. This work was published in linguistics journals, e.g., their 1994 paper in Phonology, “Feature geometry and the vocal tract.” Ken’s importance to linguistics, as an engineer interested in linguistic sound systems and eager to work with phoneticians and phonologists, cannot be overestimated, and is a legacy continued by many of his students.
Increasing evidence suggests that voices are best thought of as complex auditory patterns, and that listeners perceive and remember voices with reference to a “prototype” or “average” for that talker. Little is known about how, and how much, individual talkers vary their voice quality across situations that arise in every-day speaking, so the nature and extent of variability underlying these abstract averages, and thus the nature of the averages themselves, is unclear. The theoretical relationship between acoustic similarity and confusability in the context of a prototype model also remains unclear. In this preliminary study, 9 tokens of the vowel /a/ were recorded from 5 females on three dates. Measures of F0, spectral slope, HNR, and formant frequencies and their variability were gathered for all voice samples and acoustic distances between talkers were calculated under the assumption that all acoustic variables were equally important perceptually. Perceptual confusability was assessed in a same/different task, and predictions under the equal perceptual weight assumption were tested. Discussion will focus on how much variability is required before a voice sample no longer sounds like the originating talker, and on how the perceptual importance of each acoustical variable varies across talkers and acoustic contexts. [Work supported by NSF and NIH.]
Little is known about intraspeaker changes in voice across changing speaking situations in everyday life. In this study, we examined acoustic variations between and within 5 talkers and their effect on the likelihood that voice samples would not be identified as coming from the same talker. Talkers were drawn from a large database recorded to capture everyday variations in vocal characteristics. Nine samples of /a/, recorded on three different days, were examined for each talker. Acoustic characteristics were estimated using VoiceSauce and analysis-by-synthesis, and listeners judged whether pairs of voices came from the same or two different talkers. Results indicate that interspeaker variability in voice quality exceeds intraspeaker variability, but differences are smaller than expected. As predicted by models that treat voice quality as an auditory pattern, the acoustic attributes associated with incorrect “different speaker” responses varied from talker to talker, depending on the particular characteristics of the voice in question.