In the Chengdu Dialect of Mandarin, the /(V)an/ rime words have been described to have undergone a nasal loss process in the last decades. However, no acoustical or physiological evidence has been provided so far. In this study, we investigate this sound change process by directly looking at the velum gesture in the target segments from 4 Chengdu speakers. By means of real-time Magnetic Resonance Imaging (rt-MRI), the velum opening signal was captured along with synchronized and noise suppressed audio. The maximum degree of velum opening was compared between tautosyllabic and heterosyllabic VN sequences for different vowels (N = /n/). Nasal consonant loss was most evident for tautosyllabic /(V)an/ rime words. This sound change, together with the observed diachronic vowel raising in /(V)an/ rimes, is compatible with other research showing a preference for low vowel raising before nasal consonants. This phonetically motivated oral vowel, which is a consequence of nasal coda loss and vowel raising, would form a new phonological contrast in this dialect e.g., from /pa, pan/ to /pa, p./.
It is widely accepted that vowels are longer before voiced consonants than before voiceless ones in many languages (such as English). However, neither voiced-voiceless stop contrasts nor long-short vowel contrasts exist in Mandarin Chinese. This study investigates whether the Chinese learners exhibit a difference in vowel duration as a function of the stop voicing in their L2 English production and perception. The production measurements in the Chinese L2 English TIMIT database demonstrated a similar effect as that of the native speakers. In the perceptual experiment with 56 Chinese participants, the results showed that therewas a general tendency for the subjects to choose voiced stops when the preceding vowels were lengthened, but there was no significant difference between the intermediateand the advanced-level learners. However, the Chinese learners displayed different perceptual patterns in different vowels and stops. The results have some implications for L2 English speech learning.
The study compared the oral stops produced by the Chinese learners of English with those of the American native speakers. We employed the original English TIMIT, the global Chinese TIMIT, and the L2 English TIMIT by Chinese speakers to represent the target language, source language and interlanguage. Because of the quantity and diversity of these databases, this study only selected part of the speech in which the texts were read by most American speakers and Chinese speakers for analysis. Regarding the unbalanced occurrences of stop releases, the mixed effects model was used for the statistics of the release duration comparison between the native and L2 speakers. The results showed that the Chinese speakers produced signif-icantly longer aspirated stops than the American speakers did. A further investigation indicated that the Chinese speakers, un-like the American natives, released the final stop consonant and sometimes with an extra vowel at the end. The prolonged word final stops and inserted schwas become distinguished prosodic features of Chinese speakers’ English interlanguage. Different features of stops and new allophonic rules in English prove to be difficult for Chinese learners in English acquisition. The find-ings can present some pedagogical implications in L2 speech learning.
Although the TIMIT acoustic-phonetic dataset ([1], [2]) was created three decades ago. it remains in wide use, with more than 20000 Google Scholar references, and more than 1000 since 2017. Despite TIMIT's antiquity and relatively small size, inspection of these references shows that it is still used in many research areas: speech recognition, speaker recognition, speech synthesis, speech coding, speech enhancement, voice activity detection, speech perception, overlap detection and source separation, diagnosis of speech and language disorders, and linguistic phonetics, among others. Nevertheless, comparable datasets are not available even for other widely-studied languages, much less for under documented languages and varieties. Therefore, we have developed a method for creating TIMIT-like datasets in new languages with modest effort and cost, and we have applied this method in standard Thai, standard Mandarin Chinese, English from Chinese L2 learners, the Guanzhong dialect of Mandarin Chinese, and the Ga language of West Africa. Other collections are planned or underway. The resulting datasets will be published through the LDC, along with instructions and open-source tools for replicating this method in other languages, covering the steps of sentence selection and assignment to speakers, speaker recruiting and recording, proof-listening, and forced alignment.
This paper describes an effort to build a TIMIT-like corpus in Standard Chinese, which is part of our "Global TIMIT" project. Three steps are involved and detailed in the paper: selection of sentences; speaker recruitment and recording; and phonetic segmentation. The corpus consists of 6000 sentences read by 50 speakers (25 females and 25 males). Phonetic segmentation obtained from forced alignment is provided, which has 93.2% agreement (of phone boundaries) within 20 ms compared to manual segmentation on 50 randomly selected sentences. Statistics on the number of tokens and mean duration of phones and tones in the corpus are also reported. Males have shorter phones/tones but more and longer utterance internal silences than females, demonstrating that males in this dataset speak faster but pause more frequently and longer.