Cantonese speakers struggle with Mandarin sibilants. This study explores how phonotactics, acoustic similarity, and language experience influence acquisition. Mandarin sibilant perception and production were tested via identification and accent rating tasks. Participants included high- and low-experience Cantonese speakers (n=38 each, gender-balanced) and a Mandarin-native control group. Perceptual data showed that the low-experience group had difficulty distinguishing the two alveolar sibilants, exhibiting poorer identification of [s] than [ʂ] in non-rounded vowel contexts—a pattern that reversed in rounded vowel environments (e.g., /uan/). The high-experience group demonstrated overall better and more stable performance, though they showed a similar trend of weaker [s] identification as the low-experience group. In production, the low-experience group generally struggled with articulating [ʂ], while the high-experience group faced particular difficulty producing [ʂ] before rounded vowels. These results indicate that phonotactic constraints and acoustic similarity jointly shape non-native sibilant acquisition. The palatal [ɕ], aided by favorable phonotactics and acoustic overlap, showed better identification. In contrast, [ʂ]’s lower acoustic similarity aided perception but hindered production, especially for less experienced learners. Notably, increased language experience exerted a generally facilitative effect, markedly improving mastery of perceptually dissimilar non-native sounds.
The present study examined the influence of changes in speakers’ fundamental frequency (fo) and vocal tract resonance (VTR) on speech recognition in different types of noise by non-native listeners. The goal was to identify whether the fo-VTR relationship has a similar effect on non-native listeners as it does on native listeners. Twenty-six adults who were native Mandarin speakers learning English as a second language were presented with English Hearing-in-Noise Test (HINT) sentences in four voice conditions with the original male speaker's fo doubled and/or VTR scaled up by a factor of 1.2: (1) low fo low VTR (LfoLVTR, the original recordings); (2) low fo high VTR (LfoHVTR); (3) high fo high VTR (HfoHVTR), and (4) high fo low VTR (HfoLVTR). The stimuli were presented in speech-shaped noise (SSN) and four-talker babble (FTB) at signal-to-noise ratios of −3, 0, +3 dB. The results showed that the non-native listeners performed more poorly with fo-VTR mismatched voices than with fo-VTR matched voices and the negative influence of mismatched voice features was mainly manifested in the HfoLVTR condition. Compared to SSN, FTB had a greater adverse impact on the non-native listeners’ recognition accuracy. Further, the performance difference between matched and mismatched conditions showed distinct patterns across SSN and FTB.
This study examined the effects of change in the talker’s sex-related acoustic properties [fundamental frequency (F0) and vocal tract resonance (VTR)] on speech recognition in noise. The stimuli were HINT sentences with the original male talker’s F0 and VTR being manipulated (doubling F0 and/or scaling up VTR by a factor of 1.2) into four conditions: low F0 low VTR (LF0LVTR, the original recordings), low F0 high VTR (LF0HVTR), high F0 high VTR (HF0HVTR), and high F0 low VTR (HF0LVTR). Randomly selected sentences from each condition were presented to 193 adults for a gender rating task on a 7-point scale. Then, all sentences were mixed with speech-shaped noise at signal-to-noise ratios of −10, −5, 0, and +5 dB, and presented to 42 normal-hearing adults for recognition. The two conditions with matched F0 and VTR (HF0HVTR and LF0LVTR) were perceived as male or female voices and showed no significant differences in recognition accuracy and estimated speech reception thresholds. However, the mismatched conditions HF0LVTR and LF0HVTR showed reduced recognition performance and significantly higher SRTs than the matched conditions. In general, voices with matched F0 and VTR yield equivalent speech recognition in noise, whereas voices with mismatched F0 and VTR may reduce intelligibility in noise.
The purpose of the present study was to examine the influence of visual cues in audiovisual perception of interrupted speech by nonnative English listeners and to identify the role of working memory, long-term memory retrieval, and vocabulary knowledge in audiovisual perception by nonnative listeners. The participants included 31 Mandarin-speaking English learners between 19 and 41 years of age. The perceptual stimuli were noise-filled periodically interrupted AzBio and QuickSIN sentences with or without visual cues that showed a male speaker uttering the sentences. In addition to sentence recognition, the listeners completed a semantic fluency task, verbal (operation span) and visuospatial (symmetry span) working memory tasks, and two vocabulary knowledge tests (Vocabulary Level Test and Lexical Test for Advanced Learners of English). The results revealed significantly better speech recognition in the audio-visual condition than the audio-only condition, but the magnitude of visual benefit was substantially attenuated for sentences that had limited semantic context. The listeners’ vocabulary size in English played a key role in the restoration of missing speech information and audiovisual integration in the perception of interrupted speech. Meanwhile, the listeners’ verbal working memory capacity played an important role in audiovisual integration especially for the difficult stimuli with limited semantic context.
This study examined how talker accentedness affects the recognition of noise-vocoded speech by native English listeners and how contextual information interplays with talker accentedness during this process. The listeners included 20 native English-speaking, normal-hearing adults aged between 19 and 23 years old. The stimuli were English Hearing in Noise Test (HINT) and Revised Speech Perception in Noise (R-SPIN) sentences produced by four native Mandarin talkers (two males and two females) who learned English as a second language. Two talkers (one in each sex) had a mild foreign accent and the other two had a moderate foreign accent. A six-channel noise vocoder was used to process the stimulus sentences. The vocoder-processed and unprocessed sentences were presented to the listeners. The results revealed that talkers’ foreign accents introduced additional detrimental effects besides spectral degradation and that the negative effect was exacerbated as the foreign accent became stronger. While the contextual information provided a beneficial role in recognizing mildly accented vocoded speech, the magnitude of contextual benefit decreased as the talkers’ accentedness increased. These findings revealed the joint influence of talker variability and sentence context on the perception of degraded speech.
The purpose of the study was to examine the acoustic features of sibilant fricatives and affricates produced by prelingually deafened Mandarin-speaking children with cochlear implants (CIs) in comparison to their age-matched normal-hearing (NH) peers. The speakers included 21 children with NH aged between 3.25 and 10 years old and 35 children with CIs aged between 3.77 and 15 years old who were assigned into chronological-age-matched and hearing-age-matched subgroups. All speakers were recorded producing Mandarin words containing nine sibilant fricatives and affricates (/s, ɕ, ʂ, ts, tsʰ, tɕ, tɕʰ, tʂ, tʂʰ/) located at the word-initial position. Acoustic analysis was conducted to examine consonant duration, normalized amplitude, rise time, and spectral peak. The results revealed that the CI children, regardless of whether chronological-age-matched or hearing-age-matched, approximated the NH peers in the features of duration, amplitude, and rise time. However, the spectral peaks of the alveolar and alveolopalatal sounds in the CI children were significantly lower than in the NH children. The lower spectral peaks of the alveolar and alveolopalatal sounds resulted in less distinctive place contrast with the retroflex sounds in the CI children than in the NH peers, which might partially account for the lower intelligibility of high-frequency consonants in children with CIs.
There are abundant acoustic-phonetic cues in speech signals for listeners to encode talker identity. However, speech signals in the real world are always less optimal due to various adverse listening sources. Vocoded speech is one type of simplified signal that has less spectral and/or temporal information in comparison to normal speech. The purpose of this study is to examine whether and how listeners’ judgment of talker accent is affected by noise and tone vocoding. Twelve Mandarin-accented English speakers with varying degree of accentedness and two native English speakers were recorded reading “The Rainbow Passage.” The recorded speech samples from each talker were segmented into small sections that were randomly selected for noise- or tone-excited vocoder processing into 1, 2, 4, 8, and 16 channels. The vocoded and unprocessed speech samples were randomly presented to a group of normal-hearing, monolingual English listeners for accent rating. The listeners judged the degree of talker accent on a 9-point Likert scale with “1” representing no accent and “9” representing extremely strong accent. The data are still in the process of being collected and analyzed. Results and implications of the present study will be discussed.
This study examined accent rating of speech samples collected from 12 Mandarin-accented English talkers and two native English talkers. The speech samples were processed with noise- and tone-vocoders at 1, 2, 4, 8, and 16 channels. The accentedness of the vocoded and unprocessed signals was judged by 53 native English listeners on a 9-point scale. The foreign-accented talkers were judged as having a less strong accent in the vocoded conditions than in the unprocessed condition. The native talkers and foreign-accented talkers with varying degrees of accentedness demonstrated different patterns of accent rating changes as a function of the number of channels.
Purpose: This study assessed the intelligibility of obstruent consonants in prelingually deafened Mandarin-speaking children with cochlear implants (CIs). Method: Twenty-two Mandarin-speaking children with normal hearing (NH) aged 3.25–10.0 years and 35 Mandarin-speaking children with CIs aged 3.77–15.0 years were recruited to produce a list of Mandarin words composed of 17 word-initial obstruent consonants in different vowel contexts. The children with CIs were assigned to chronological age–matched (CA) and hearing age–matched (HA) subgroups with reference to the NH controls. One hundred naïve NH adult listeners were recruited for a consonant identification task that consisted of a total of 2,663 stimulus tokens through an online research platform. For each child speaker, the consonant productions were judged by seven to 12 different adult listeners. An average percentage of consonants correct was calculated across all listeners for each consonant. Results: The CI children in both the CA and HA subgroups showed lower intelligibility in their consonant productions than the NH controls. Among the 17 obstruents, both CI subgroups showed higher intelligibility for stops, but they demonstrated major problems with the sibilant fricatives and affricates and showed a different confusion pattern from the NH controls on these sibilants. Of the three places (alveolar, alveolopalatal, and retroflex) in Mandarin sibilants, both CI subgroups showed the lowest intelligibility and the greatest difficulties with alveolar sounds. For the NH children, there was a significant positive relationship between overall consonant intelligibility and chronological age. For the children with CIs, the best fit regression model revealed significant effects of chronological age and age at implantation, with their quadratic terms included. Conclusions: Mandarin-speaking children with CIs experience major challenges in the three-way place contrasts of sibilant sounds in consonant production. Chronological age and the combined effect of CI-related time variables play important roles in the development of obstruent consonants in the CI children.
Objectives: The purpose of the present study was to investigate the pitch accuracy of vocal singing in children with severe to profound hearing loss who use bilateral cochlear implants (CIs) or bimodal devices [CI at one ear and hearing aid (HA) at the other] in comparison to similarly-aged children with normal-hearing (NH). Design: The participants included four groups: (1) 26 children with NH, (2) 13 children with bimodal devices, (3) 31 children with bilateral CIs that were implanted sequentially, and (4) 10 children with bilateral CIs that were implanted simultaneously. All participants were aged between 7 and 11 years old. Each participant was recorded singing a self-chosen song that was familiar to him or her. The fundamental frequencies (F0) of individual sung notes were extracted and normalized to facilitate cross-subject comparisons. Pitch accuracy was quantified using four pitch-based metrics calculated with reference to the target music notes: mean note deviation, contour direction, mean interval deviation, and F0 variance ratio. A one-way ANOVA was used to compare listener-group difference on each pitch metric. A principal component analysis showed that the mean note deviation best accounted for pitch accuracy in vocal singing. A regression analysis examined potential predictors of CI children's singing proficiency using mean note deviation as the dependent variable and demographic and audiological factors as independent variables. Results: The results revealed significantly poorer performance on all four pitch-based metrics in the three groups of children with CIs in comparison to children with NH. No significant differences were found among the three CI groups. Among the children with CIs, variability in the vocal singing proficiency was large. Within the group of 13 bimodal users, the mean note deviation was significantly correlated with their unaided pure-tone average thresholds (r = 0.582, p = 0.037). The regression analysis for all children with CIs, however, revealed no significant demographic or audiological predictor for their vocal singing performance. Conclusion: Vocal singing performance in children with bilateral CIs or bimodal devices is not significantly different from each other on a group level. Compared to children with NH, the pediatric bimodal and bilateral CI users, in general, demonstrated significant deficits in vocal singing ability. Demographic and audiological factors, known from previous studies to be associated with good speech and language development in prelingually-deafened children with CIs, were not associated with singing accuracy for these children.
The purpose of this study was to examine the impact of spectral degradation on speech processing in non-native listeners. The participants included 27 native English (L1) listeners and 43 native Mandarin listeners who learned English as a second language (L2). The speech stimuli included 12 English vowels embedded in a /hVd/ context, 20 English consonants embedded in a /Ca/ context, and HINT, CUNY, and R-SPIN sentences. All stimuli were processed using 2-, 4-, 6-, 8-, and 12-channel noise vocoders. The results showed that compared to the L1 lis-teners, the L2 listeners demonstrated less improvement in phoneme recognition with increasing number of channels, which was associated with the phoneme confusions due to the impact of their native language. Both consonant and vowel recognition made significant contributions to sentence recognition in the L2 listeners. In addition, the L2 listeners were less effective than the L1 listeners in applying contextual information and lin-guistic knowledge to sentence recognition. However, the facilitating role of contextual cues in sentence recog-nition was consistently present in the L2 listeners but they required more spectral information to maximize the contextual benefit in comparison to the L1 listeners. The overall perceptual performance of the L2 listeners was positively correlated with and predicted by the length of residence in the U.S.
The purpose of this study was to assess the accuracy of consonant production of children with cochlear implants (CIs) judged by naïve adult listeners. A total of 57 Mandarin-speaking children (22 with normal hearing and 35 with CIs) were recruited to produce a list of Mandarin words composed of 17 word-initial obstruent consonants in three different vowel contexts. A total number of 2628 tokens were generated and were divided into 10 subsets. One hundred Mandarin-speaking naïve adult listeners were recruited to identify the consonant productions through Gorilla, the online research platform. Each listener was randomly assigned to one subset. For each child speaker, the consonant productions were judged by 7–12 adult listeners and an average accuracy rate was calculated across all listeners for each consonant. The results revealed that the children with CIs showed lower accuracies and different confusion patterns on their consonant productions than the normal hearing controls. In particular, they demonstrated higher accuracy for stops but had major problems with the fricatives and affricates involved in the alveolar—alveolopalatal—retroflex postalveolar three-way sibilant contrast. Of the three places of the sibilant contrast, they showed the greatest difficulties for the alveolar sounds.
Purpose The purpose of this study was to characterize the acoustic profile and to evaluate the intelligibility of vowel productions in prelingually deafened, Mandarin-speaking children with cochlear implants (CIs). Method Twenty-five children with CIs and 20 age-matched children with normal hearing (NH) were recorded producing a list of Mandarin disyllabic and trisyllabic words containing 20 Mandarin vowels [a, i, u, y, ɤ, ɿ, ʅ, ai, ei, ia, ie, ye, ua, uo, au, ou, iau, iou, uai, uei] located in the first consonant–vowel syllable. The children with CIs were all prelingually deafened and received unilateral implantation before 7 years of age with an average length of CI use of 4.54 years. In the acoustic analysis, the first two formants (F1 and F2) were extracted at seven equidistant time locations for the tested vowels. The durational and spectral features were compared between the CI and NH groups. In the vowel intelligibility task, the extracted vowel portions in both NH and CI children were presented to six Mandarin-speaking, NH adult listeners for identification. Results The acoustic analysis revealed that the children with CIs deviated from the NH controls in the acoustic features for both single vowels and compound vowels. The acoustic deviations were reflected in longer duration, more scattered vowel categories, smaller vowel space area, and distinct formant trajectories in the children with CIs in comparison to NH controls. The vowel intelligibility results showed that the recognition accuracy of the vowels produced by the children with CIs was significantly lower than that of the NH children. The confusion pattern of vowel recognition in the children with CIs generally followed that in the NH children. Conclusion Our data suggested that the prelingually deafened children with CIs, with a relatively long duration of CI experience, still showed measurable acoustic deviations and lower intelligibility in vowel productions in comparison to the NH children.
ObjectiveThis study was aimed at examining the effects of an adaptive non-linear frequency compression algorithm implemented in hearing aids (i.e., SoundRecover2, or SR2) at different parameter settings and auditory acclimatization on speech and sound-quality perception in native Mandarin-speaking adult listeners with sensorineural hearing loss.DesignData consisted of participants’ unaided and aided hearing thresholds, Mandarin consonant and vowel recognition in quiet, and sentence recognition in noise, as well as sound-quality ratings through five sessions in a 12-week period with three SR2 settings (i.e., SR2 off, SR2 default, and SR2 strong).Study SampleTwenty-nine native Mandarin-speaking adults aged 37–76 years old with symmetric sloping moderate-to-profound sensorineural hearing loss were recruited. They were all fitted bilaterally with Phonak Naida V90-SP BTE hearing aids with hard ear-molds.ResultsThe participants demonstrated a significant improvement of aided hearing in detecting high frequency sounds at 8 kHz. For consonant recognition and overall sound-quality rating, the participants performed significantly better with the SR2 default setting than the other two settings. No significant differences were found in vowel and sentence recognition among the three SR2 settings. Test session was a significant factor that contributed to the participants’ performance in all speech and sound-quality perception tests. Specifically, the participants benefited from a longer duration of hearing aid use.ConclusionFindings from this study suggested possible perceptual benefit from the adaptive non-linear frequency compression algorithm for native Mandarin-speaking adults with moderate-to-profound hearing loss. Periods of acclimatization should be taken for better performance in novel technologies in hearing aids.
The role of working memory (WM) and long-term lexical-semantic memory (LTM) in the perception of interrupted speech with and without visual cues, was studied in 29 native English speakers. Perceptual stimuli were periodically interrupted sentences filled with speech noise. The memory measures included an LTM semantic fluency task, verbal WM, and visuo-spatial WM tasks. Whereas perceptual performance in the audio-only condition demonstrated a significant positive association with listeners' semantic fluency, perception in audio-video mode did not. These results imply that when listening to distorted speech without visual cues, listeners rely on lexical-semantic retrieval from LTM to restore missing speech information.
Abstract This study examined the development of vowel categories in young Mandarin -English bilingual children. The participants included 35 children aged between 3 and 4 years old (15 Mandarin-English bilinguals, six English monolinguals, and 14 Mandarin monolinguals). The bilingual children were divided into two groups: one group had a shorter duration (<1 year) of intensive immersion in English (Bi-low group) and one group had a longer duration (>1 year) of intensive immersion in English (Bi-high group). The participants were recorded producing one list of Mandarin words containing the vowels /a, i, u, y, ɤ/ and/or one list of English words containing the vowels /i, ɪ, e, ɛ, æ, u, ʊ, o, ɑ, ʌ/. Formant frequency values were extracted at five equidistant time locations (the 20–35–50–65–80% point) over the course of vowel duration. Cross-language and within-language comparisons were conducted on the midpoint formant values and formant trajectories. The results showed that children in the Bi-low group produced their English vowels into clusters and showed positional deviations from the monolingual targets. However, they maintained the phonetic features of their native vowel sounds well and mainly used an assimilatory process to organize the vowel systems. Children in the Bi-high group separated their English vowels well. They used both assimilatory and dissimilatory processes to construct and refine the two vowel systems. These bilingual children approximated monolingual English children to a better extent than the children in the Bi-low group. However, when compared to the monolingual peers, they demonstrated observable deviations in both L1 and L2.
BACKGROUND:Mandarin Chinese has a rich repertoire of high-frequency speech sounds. This may pose a remarkable challenge to hearing-impaired listeners who speak Mandarin Chinese because of their high-frequency sloping hearing loss. An adaptive nonlinear frequency compression (adaptive NLFC) algorithm has been implemented in contemporary hearing aids to alleviate the problem.PURPOSE:The present study examined the performance of speech perception and sound-quality rating in Mandarin-speaking hearing-impaired listeners using hearing aids fitted with adaptive NLFC (i.e., SoundRecover2 or SR2) at different parameter settings.RESEARCH DESIGN:Hearing-impaired listeners' phoneme detection thresholds, speech reception thresholds, and sound-quality ratings were collected with various SR2 settings.STUDY SAMPLE:The participants included 15 Mandarin-speaking adults aged 32 to 84 years old who had symmetric sloping severe-to-profound sensorineural hearing loss.INTERVENTION:The participants were fitted bilaterally with Phonak Naida V90-SP hearing aids.DATA COLLECTION AND ANALYSIS:The outcome measures included phoneme detection threshold using the Mandarin Phonak Phoneme Perception test, speech reception threshold using the Mandarin hearing in noise test (M-HINT), and sound-quality ratings on human speech in quiet and noise, bird chirps, and music in quiet. For each test, five experimental settings were applied and compared: SR2-off, SR2-weak, SR2-default, SR2-strong 1, and SR2-strong 2.RESULTS:The results showed that listeners performed significantly better with SR2-strong 1 and SR2-strong 2 settings than with SR2-off or SR2-weak settings for speech reception threshold and phoneme detection threshold. However, no significant improvement was observed in sound-quality ratings among different settings.CONCLUSIONS:These preliminary findings suggested that the adaptive NLFC algorithm provides perceptual benefit to Mandarin-speaking people with severe-to-profound hearing loss.
Objective:The purpose of the present study was to examine the effects of NLFC fitting in hearing aids and auditory acclimatisation on speech perception and sound-quality rating in hearing-impaired, native Mandarin-speaking adult listeners. Design:Mandarin consonant, vowel and tone recognition were tested in quiet and sentence recognition in noise (speech-shaped noise at a +5 dB signal-to-noise ratio) with NLFC-on and NLFC-off. Sound-quality ratings were collected on a 0-10 scale at each test session. A generalised linear model and correlational analyses were performed. Study sample:Thirty native Mandarin-speaking adults with moderate-to-severe sensorineural hearing loss were recruited. Results:The hearing-impaired listeners showed significantly higher accuracy with NLFC-on than with NLFC-off for consonant and sentence recognition and the recognition performance improved with both NLFC-on and off as a function of increased length of use. The satisfaction score of sound-quality ratings for different types of sounds significantly increased with NLFC-on than with NLFC-off. The speech recognition results showed moderate to strong correlation with the unaided hearing thresholds. Conclusion:For native Mandarin-speaking listeners with hearing loss, the NLFC technology provided modest but significant improvement in Mandarin fricative and sentence recognition. Subjectively, the naturalness and overall preference of sound-quality satisfaction judgement also improved with NLFC.
The purpose of the present study was two-fold: (1) to examine whether Mandarin-speaking children with CIs showed distinctive durational and amplitude features for the four lexical tones in their tone production; (2) to compare the duration and amplitude patterns of Mandarin lexical tones in monosyllables produced in citation form between CI children and age-matched normal-hearing (NH) children. The participants included 14 prelingually deafened Mandarin-speaking children with CIs and 14 NH children, all aged between 2.9 and 8.3 years old. Each participant produced five CV syllables (fa, fu, pi, xu, ke) in four tones through a tone drill activity. The vowel duration and rms amplitude values at nine equidistant time locations over the vowel duration were obtained. The results revealed that the CI children can produce distinctive duration and amplitude features for the four lexical tones. Their durational pattern and amplitude contours were highly similar to the NH children on tone 1, 2, and 4 but differed from the NH children on tone 3. In addition, NH children showed positively correlated amplitude contour and F0 contour but the CI children demonstrated inconsistent amplitude and F0 contours for tone 3. This finding suggested that the amplitude contour of a tone and the F0 contour of the same tone may not always be highly correlated.
Word-initial stops in Mandarin and English show a distinctive phonological categorization but a similar phonetic realization along the VOT (Voice Onset Time) continuum. Previous research reported that native Mandarin adults produce measurably longer long-lag VOTs than native English adults. The present study examined whether and how the difference between Mandarin and English VOTs is manifested in monolingual children and Mandarin–English bilingual children. The participants included 15 five- to six-year-old sequential bilingual children, 24 corresponding monolingual children (15 Mandarin, 9 English), and 22 monolingual adults (12 Mandarin, 10 English). The bilingual children were divided into two groups (Bi-low and Bi-high) based on the amount of experience in English. Each participant was recorded producing 18 Mandarin words and/or 18 English words containing six stops in each language. The VOT values were measured from the beginning of stop burst to the onset of the voicing. The results showed that the language difference in VOT in the monolingual children was manifested in a pattern similar to the monolingual adults. However, Mandarin and English VOTs showed less separable distributions in the two groups of bilingual children. Further analysis suggested that both groups of bilingual children tended to separate Mandarin and English short-lag VOTs but only the Bi-low children showed different long-lag VOTs between the two languages. These results suggested that due to the bilingual effects and L1–L2 (first language – second language) interactions, even though the bilingual children tried to separate the two VOT systems, they implemented the separation in a different manner than the monolingual speakers.
Robert Allen Fox合作论文数Speech Perception and Acoustics Labs|Ohio State University|University of Wisconsin-Madison7