The vowel inventories of Kamu and Larrakia each consist of peripheral /i/-/u/-/e/-/a/-/o/ and another non-low vowel that has not been phonetically characterized. Descriptions of similar six-vowel systems in Western Top End languages offer inconsistent accounts of the non-peripheral vowel, but no instrumental analysis has yet been conducted. Acoustic and phonological analysis reveals that the sixth vowel is best characterized as /u/ in Kamu and /i/ in Larrakia. Height and backness are the critical phonetic properties that differentiate the non-peripheral vowel from the other vowel phonemes in each language. These results enhance our understanding of the structure and properties of higher cardinality vowel systems in Australian languages, and the typology of six-vowel systems more generally.
A minority of Australian languages contrast stops at the same place of articulation – oppositions described in a variety of ways: e.g. fortis vs lenis, geminate vs singleton – but not well understood due to a lack of instrumental data. To shed more light on these contrasts, acoustic phonetic properties of stops were analyzed in medial pre-vocalic environments in Kamu, Larrakia, Warlmanpa, and Warumungu. In all four languages, stops at the same place of articulation are differentiated primarily by duration, and also degree of voicing estimated using Harmonic Ratios. There is a degree of correlation between voicing and duration, but duration is found to be the property that most consistently differentiates stop modes. These data suggest that the opposition in these languages is best characterised as a length contrast, and that these stop oppositions are realized with complex interactions between duration and voicing.
Research on American English (AmE) has found that words containing a diphthong or tense vowel and word-final liquid (file, fire) are sesquisyllabic: adjudged by some speakers as over one syllable (Lavoie & Cohn 1999, Tilsen & Cohn 2016, Popescu & Chitoran 2022). Previous analyses predict that sesquisyllabicity is sensitive to the sequence of articulations in the rhyme and their potential for overlap. This study investigates this prediction with speakers of Australian English (AusE) (n = 30), a non-rhotic variety, in comparison with AmE (n = 45). In a production task, acoustic recordings were made of participants reading words aloud. In a syllable count judgment (SCJ) task, participants judged words presented on a screen as 1, 1.5 or 2 syllables in length. Trials consisted of words with five rhyme types (e.g. buy, hide, pine, file, fire). All possible vowel contrasts were investigated. Mean SCJ was calculated for each rhyme type and production data were annotated for acoustic measurements (e.g. duration). Participants differ in percepts of syllabicity, but overall, in both AusE and AmE, higher SCJs are associated with contexts where vocalic articulations in the rhyme are more distinct and allow for less overlap, including diphthongs and sequences of dissimilar vowel and vocalic liquid articulations.
The shape of the human vocal tract plays a critical role in shaping the acoustic output of speech, yet capturing and modeling its complex 3-D dynamics remains a challenge. We present a novel acoustic evaluation framework for vocal tract morphing based on both volumetric and real-time Magnetic Resonance Imaging (MRI) data. Our approach uses Large Deformation Diffeomorphic Metric Mapping (LDDMM) applied to volumetric mesh data, constrained by dynamic midsagittal motion captured via real-time MRI. Building on previous work that separately used volumetric MRI for static postures and real-time MRI for dynamic imaging, we now relate 3-D shape deformation directly to speech output. Data were collected from an adult female speaker of Australian English, including real-time midsagittal MRI videos with synchronized audio, and high-resolution 3-D scans of sustained articulatory postures. For each intermediate mesh representing vocal tract deformation, we compute acoustic transfer functions and synthesize speech signals. These are temporally aligned with the recorded speech and compared via extracted formant frequencies (F1–F3). By quantifying formant deviations over time, we assess the acoustic plausibility of deformations guided by real-time MRI. To our knowledge, this is the first study to directly validate a diffeomorphic vocal tract morphing pipeline using measured speech acoustics.
Goals of lateral approximant production are imperfectly understood, partly because of the limitations of most existing data, restricted to the midsagittal plane. To provide more complete information about the configuration of the vocal tract for laterals, /l/-production in three vowel contexts by three Australian English speakers was examined for the first time using a combination of real-time and volumetric Magnetic Resonance Imaging (MRI). Laterals produced by all speakers were characterised by bilateral parasagittal airflow, with some asymmetries in side channel geometries. Central occlusion of the oral airway varied in location, timing, and duration across speakers and vowel contexts, but was consistently associated with reduction in overall acoustic intensity relative to context vowels. These data provide further insights into the complex relationships between articulatory, coarticulatory and acoustic properties of lateral approximants, and their realisation in Australian English.
We present a novel framework for analyzing dynamic vocal tract deformations by integrating volumetric Magnetic Resonance Imaging (MRI) data and real-time MRI (rtMRI) boundary constraints within an iterative Large Deformation Diffeomorphic Metric Mapping (LDDMM) framework. More precisely, we apply LDDMM to morph volumetric vocal tract shapes using rtMRI boundary constraints that enable a smooth and anatomically plausible articulatory transformation. We demonstrate the method and discuss the issues involved using a vowel-consonant-vowel sequence. We show the influence of varying the number of rtMRI images on the resulting articulatory transformation.
Japanese learners of American English (AE) can face difficulties producing mid and low AE vowels /æ/, /^/, and /ɔ/, which tend to be perceptually assimilated into native vowel categories. An interactive multimedia visualization of the acoustic-phonetic vowel space (JSpace) was provided to assist Japanese learners of English as a foreign language (EFL), allowing students practice finding and auditing AE vowel sounds in an acoustic space calibrated with reference Japanese vowels. In each training session, students explored a set of seven AE vowel sounds: /æ/-/^/-/ɔ/-/ɪ/-/ɛ/-/u/-/U/. Clicking in the acoustic space generated a vowel sound synthesized using the corresponding first and second formant frequencies. Participants were 28 first-year Japanese undergraduates at a private university in Tokyo who received four weeks of targeted training via the JSpace interactive visualization, during which training their formant frequency settings for the seven AE vowel sounds were tabulated for analysis. From the first to the fourth week, there was a substantial reduction in the spread of first-formant frequencies for just three of the seven AE vowels: /ɪ/, /æ/, /^/. Performance on a forced-choice, minimal-pairs discrimination task was improved from pre-training to post-training assessments.
English rhotics are realized with rich allophony across speakers, contexts and varieties, but Australian English /o/ has not previously been examined in detail. Rhotic approximants produced in three vowel contexts by four speakers of Australian English were captured using real-time and volumetric structural magnetic resonance imaging. /o/ was articulated with bunched tongue postures by two speakers and more apical configurations by two speakers, but all rhotics were characterized by three coordinated gestures: tongue tip, tongue body and labial constrictions. These data shed new light on the complex goals of production of rhotic approximants beyond the midsagittal plane, and their realization and extent of variation in Australian English.
Real-time magnetic resonance image (rtMRI) data of the upper airway provides a rich source of information about vocal tract shaping that can inform phonemic analysis and classification. We describe a multimodal phonemic classifier that combines articulatory data with speech audio features to improve performance. A deep network model processes rtMRI video data using ResNet18 and speech audio using a custom CNN and then combines the two data streams using a Transformer layer to fully explore the correlation of the two streams towards better vowel-consonant-vowel classification via the Transformer's multi-head self-attention mechanism. The classification accuracy of both the unimodal and multimodal models show substantial improvement on previous work (> 38%). The addition of audio features improves classification accuracy in the multimodal model by 7% compared with the unimodal model using articulatory data. We analyze the model and discuss the phonetic implications.
3D analysis of the vocal tract using dynamic MRI remains a technically difficult challenge. Various approaches have been explored such as using parametic models of the vocal tract (Yehia et al., 1997); integrating data across parallel slices of 2D dynamic data (Zhu et al., 2012); applying stack-of-spiral MRI sampling with 3D constrained reconstruction (Zhao et al., 2020); and combining static 3D and dynamic 2D data (Douros et al., 2019). In this work, we follow a similar approach to Douros et al. and explore the relationship between 2D real-time midsagittal images and 3D volumetric scans of the vocal tract. The real-time MRI midsagittal images are recorded during vowel-consonant-vowel vocal tasks, while the 3D volumetric scans are recorded during sustained vowels. We use large deformation diffeomorphic metric mapping as the foundation for this modeling work. We explore techniques to use constraints provided by the real-time MRI midsagittal images to enable smooth deformations of the 3D volumetric data. We focus on the feasibility of the methods and report on the types of constraints explored and the resulting deformations of the 3D volumetric data.
Dynamic phonetic properties of Australian English vowels have been well described in acoustic (Cox, 2006; Elvin et al., 2016) and articulographic studies (Ratko et al., 2023), but the global configuration of the vocal tract during Australian English vowel production has not previously been examined. Midsagittal configuration of the upper airway during production of Australian English vowels was tracked using real-time magnetic resonance imaging (Kennerley et al., 2022) on a 3T scanner (Siemens Magnetom Prisma) using a 64-channel head/neck receiver array coil. Speech audio was simultaneously recorded in-scanner at 16 kHz using a ceramic noise-canceling microphone (Opto-acoustics FOMRI-III). 3D configuration of the vocal tract during sustained productions of the same monophthongs was captured using volumetric imaging of the upper airway at a resolution of 2mm × 2 mm × 2 mm. These data offer new details on configuration of the entire vocal tract and the relationship between articulatory and acoustic targets during production of Australian English vowels.
There are three phonological hypotheses on the Kaytetye segmental inventory. Hypothesis 1 proposes 30 segments: four monophthongs, one diphthong and 25 consonants. Hypothesis 2 proposes 54 segments: two monophthongs and 52 consonants. Hypothesis 3 proposes 55 segments: three monophthongs and 52 consonants. The choice between these three hypotheses has significant implications for models of phonological contrast, phonotactic organization, syllable structure and partial reduplication processes in Kaytetye. We evaluate the three hypotheses against evidence from these domains and find that Hypothesis 1 is the best supported phonological analysis. Companion analysis of the phonetic distribution and functional load of medial Kaytetye monophthong tokens was conducted by phonetically-trained transcribers, and compared with groupings of vowels obtained through unsupervised classification of first and second formant values using finite Gaussian mixture models. Both transcriber-perceived and machine-learnt categorizations agree that none of the four monophthongs are marginal, nor can their qualities be attributed to phonological context effects. These data demonstrate the importance of both phonological and phonetic evidence in evaluating the structure and properties of vowel systems in under-described languages.
Two-formant representation of the universal vowel space is a foundational model in phonetics (Fant & Risberg 1963; Traunmüller & Lacerda 1983), but tools for rapid synthesis of vowel sounds corresponding to points in the vowel space are limited. Stand-alone applications such as the Formant Synthesizer Demo (Beskow, 2000) provide this functionality, and rich tools for formant synthesis are available in Praat (Boersma and Weenink, 2023), Matlab (Rabiner et al., 2023) and other platforms, but these tools are not universally accessible, while web-based speech synthesis tools (e.g., Thapen, 2017) do not typically allow systematic manipulation or quantification of key acoustic parameters. VSpace is a browser-based formant synthesizer developed using the Tone.js Web Audio framework. A universal vowel space is represented as a trapezoid scaled to typical male, female or child speaker formant ranges. Clicking on any point generates a vowel sound corresponding to the first and second formant frequencies selected, synthesized with typical F3 and F4 frequencies excited by a selectable source signal. IPA symbols corresponding to typical formant frequencies for language-specific vowel phonemes may be superimposed on the vowel space. Key synthesis parameters are configurable through the web interface to allow exploration of the acoustic consequences for vowel perception.
Australian English (AusE) pre-/l/ /ʉ:/ is retracted when coarticulated with coda /l/, leading to vowel change through acoustic contrast reduction between /ʉ:-ʊ/ (pool-pull). Younger speakers show smaller contrast in /ʉ:l- ʊl/ targets than older speakers. As vowel trajectories are less well understood, we tested changes in the F2 trajectory. 200 tokens of /ʉ:, ʊ/ in the /hVd, pVl/ contexts (who’d-hood, pool-pull), produced by eight younger (ages = 20–29) and nine older (ages = 54–80) female AusE speakers, were extracted from the audio corpus AusTalk. Formant trajectories in pre-/d/ vowels and /Vl/ rimes were extracted automatically, corrected manually, and time-normalized (0–1). F2 trajectory was fitted with a Generalized Additive Mixed Model using fixed factors Vowel, Coda, and Age (treatment-coded, comparing /ʉ:/ to /ʊ/, /l/ to /d/, younger to older) with Speaker as random intercept with smooths for Vowel and Coda. Both age groups maintained significant vowel contrast in the pre-/d/ context in the entire vowel trajectory. In the /l/-context, older speakers showed a significant F2 difference throughout the rime trajectory, while younger speakers did not. Durational differences may be maintained, as duration contrast was not analysed due to time normalization. Our findings are consistent with an AusE pool-pull merger.
Recent empirical studies have highlighted the large degree of analytic flexibility in data analysis that can lead to substantially different conclusions based on the same data set. Thus, researchers have expressed their concerns that these researcher degrees of freedom might facilitate bias and can lead to claims that do not stand the test of time. Even greater flexibility is to be expected in fields in which the primary data lend themselves to a variety of possible operationalizations. The multidimensional, temporally extended nature of speech constitutes an ideal testing ground for assessing the variability in analytic approaches, which derives not only from aspects of statistical modeling but also from decisions regarding the quantification of the measured behavior. In this study, we gave the same speech-production data set to 46 teams of researchers and asked them to answer the same research question, resulting in substantial variability in reported effect sizes and their interpretation. Using Bayesian meta-analytic tools, we further found little to no evidence that the observed variability can be explained by analysts' prior beliefs, expertise, or the perceived quality of their analyses. In light of this idiosyncratic variability, we recommend that researchers more transparently share details of their analysis, strengthen the link between theoretical construct and quantitative system, and calibrate their (un)certainty in their conclusions.
Acoustic studies have shown that in Australian English (AusE), vowel length contrasts are realised through temporal, spectral and dynamic characteristics. However, relatively little is known about the articulatory differences between long and short vowels in this variety. This study investigates the articulatory properties of three long–short vowel pairs in AusE: /iː–ɪ/ beat – bit, /ɐː–ɐ/ cart – cut and /oː–ɔ/ port – pot, using electromagnetic articulography. Our findings show that short vowel gestures had shorter durations and more centralised articulatory targets than their long equivalents. Short vowel gestures also had proportionately shorter periods of articulatory stability and proportionately longer articulatory transitions to following consonants than long vowels. Long–short vowel pairs varied in the relationship between their acoustic duration and the similarity of their articulatory targets: /iː–ɪ/ had more similar acoustic durations and less similar articulatory targets, while /ɐː–ɐ/ were distinguished by greater differences in acoustic duration and more similar articulatory targets. These data suggest that the articulation of vowel length contrasts in AusE may be realised through a complex interaction of temporal, spatial and dynamic kinematic cues.
Lateral vocalisation is assumed to arise from changes in coronal articulation but is typically characterised perceptually without linking the vocalised percept to a coronal articulation. Therefore, we examined how listeners' perception of coda /l/ as vocalised relates to coronal closure. Perceptual stimuli were acquired by recording laterals produced by six speakers of Australian English using electromagnetic articulography (EMA). Tongue tip closure was monitored for each lateral in the EMA data. Increased incidence of incomplete coronal closure was found in coda /l/ relative to onset /l/. Having verified that the dataset included /l/ tokens produced with incomplete coronal closure-a primary articulatory cue of vocalised /l/-we conducted a perception study in which four highly experienced auditors rated each coda /l/ token from vocalised (3) to non-vocalised (0). An ordinal mixed model showed that increased tongue tip (TT) aperture and delay correlated with vocalised percept, but auditors ratings were characterised by a lack of inter-rater reliability. While the correlation between increased TT aperture, delay, and vocalised percept shows that there is some reliability in auditory classification, variation between auditors suggests that listeners may be sensitive to different sets of cues associated with lateral vocalisation that are not yet entirely understood.
In Australian English rimes, coarticulation between coda /l/ and its preceding vowel has the potential to attenuate cues that contribute to phonological vowel contrast. Therefore, vowel-/l/ coarticulation may increase ambiguity between prelateral vowels. We used a vowel identification task to test the effect of vowel-/l/ coarticulation on vowel disambiguation in perception. Listeners categorized vowels in /hVd/ and /hVl/ contexts. Results showed reduced accuracy of vowels before coda /l/ compared to coda /d/, showing that coda /l/ increases vowel disambiguation difficulty. In particular, reduced perceptual contrast was found for the rime pairs /ʉːl-ʊl, æɔl-æl/ and /əʉl-ɔl/ (e.g., 'fool-full, howl-Hal, dole-doll'). A second experiment tested the effect of reduced perceptual contrast on word recognition. Listeners identified minimal pairs contrasting key vowel pairs in the /CVl/ and /CVd/ contexts. Reduced accuracy and increased response time in /l/ contexts shows that coda /l/ hinders listeners’ ability to identify vowels. The implications of reduced perceptual vowel contrast for compensation for coarticulation and sound change are discussed.
Research on the temporal dynamics of /l/ production has focused primarily on mid-sagittal tongue movements. This study reports how known variations in the timing of mid-sagittal gestures are related to para-sagittal dynamics in /l/ formation in Australian English (AusE), using three-dimensional electromagnetic articulography (3D EMA). The articulatory analyses show (1) consistent with past work, the temporal lag between tongue tip and tongue body gestures identified in the mid-sagittal plane changes across different syllable positions and vowel contexts; (2) the lateral channel is largely formed by tilting the tongue to the left/right side of the oral cavity as opposed to curving the tongue within the coronal plane; and, (3) the timing of lateral channel formation relative to the tongue body gesture is consistent across syllable positions and vowel contexts, even as the temporal lag between tongue tip and tongue body gestures varies. This last result is particularly informative with respect to theoretical hypotheses regarding gestural control for /l/s, as it suggests that lateral channel formation is actively controlled as opposed to resulting as a passive consequence of tongue stretching. These results are interpreted as evidence that the formation of the lateral channel is a primary articulatory goal of /l/ production in AusE.