
Previous studies on sentence production have shown that the accessibility of a referent in the speaker's mind influences the syntactic structure selected during utterance planning. The present study investigated how the effect of referent accessibility on structural choice during sentence production is influenced by language constraints, such as free vs. fixed word order, and by language status as L1 or L2. Hindi, a flexible word-order language with rich case-marking, allows constituent reordering without changing grammatical roles, whereas English, a fixed word-order language, requires a passive construction to place patient referent in the sentence initial position. In an event-description task, referent accessibility was manipulated through visual cueing (agent-cued vs. patient-cued) and referent position (agent-left vs. patient-left), while Hindi-English bilinguals produced sentences in Hindi (L1) and English (L2). In Hindi, visual cueing of the patient resulted in object-initial sentences through scrambling, whereas in English, no significant effect of visual cueing on structural choice was observed. In contrast, referent position exerted a robust influence on sentence structure in both languages. Left-positioned patients elicited more object-initial sentences in Hindi and more passive voice constructions in English. Eye-movement patterns reflected both independent and interactive effects of visual cueing and referent position in the two languages. These findings suggest that the influence of referent accessibility on structural choices is constrained by language flexibility and by the L1/L2 status of language in bilingual speakers.
This study investigates changes in English vowel perception and production among 22 advanced-proficiency second-language (L2) learners from diverse first-language (L1) backgrounds over 18 months at a large U.S. university, with a particular focus on the influence of the L2 English-speaking environment and the impact of the COVID-19 pandemic. The study employed Bayesian mixed-effects models to analyze participants' vowel identification accuracy using a perception oddity task and vowel production (intelligibility) accuracy using listener-based ratings at 1, 3, 6, and 18 months after arrival. The sudden shift to online instruction due to the COVID-19 pandemic, which occurred between Months 6 and 18, and the subsequent return of some participants to their home countries added an additional layer of complexity to the investigation. The findings revealed that the participants demonstrated high initial accuracy in identifying L2 vowels, with improvement observed for most contrasts over the first 6 months. Vowel intelligibility, in contrast, did not improve much, with only one contrast showing increased accuracy at the last data point. Individual trajectories showed variability in both perception and production, emphasizing the unique change trajectories of advanced-proficiency learners.
This research empirically tests the claim that the sonority sequencing principle (SSP), which dictates that syllables rise in sonority from the edges towards the nucleus, is gradient. This was tested on two languages that vary in their adherence to the SSP, that is, German, which does not allow SSP violations outside of sibilant-initial clusters, and Russian, a language that has more SSP violations. We performed corpus-based lexical statistics to test whether the type frequency (Study 1) and probability of attestedness (Study 2) of word-initial consonant clusters are predicted by sonority difference, that is, the degree by which the consonants in the clusters differed in sonority. In Study 1, we found that in both German and Russian, the higher the sonority difference, the higher the type frequency of a consonant cluster. However, this generalisation only held for non-sibilant-initial clusters. In Study 2, we found that the higher the sonority difference, the higher the probability of a cluster being attested in German and Russian. This finding was also observed for sibilant-initial clusters. Overall, these results provide empirical evidence that the SSP is gradient and can be quantitatively observed even within languages that contain SSP violations. Moreover, our findings highlight the need to consider the special status of sibilant-initial clusters in SSP analyses.
Languages differ in how they construct motion events. These crosslinguistic differences, referred to as framing typology, affect how speakers of different languages conceptualize motion events. Differences in event conceptualization may also transfer to the L2 (conceptual transfer), as shown from learners’ grammatical choices and their gestures. However, it is unclear whether L1-based conceptualization patterns can also extend to the acoustic domain of L2 speech production. Therefore, this study examined read speech from two spoken corpora of L2 English speakers whose L1s were of different framing typologies. English production from Korean and Turkish L1 speakers (V-framed) was compared to that of L1 German and L1 English speakers (S-framed). Various acoustic measurements of duration, pitch and intensity were extracted and analyzed. Results show no reliable acoustic differences between the two framing typologies, but rather show differences between native and non-native English speakers. Post hoc acoustic analyses further show how native and non-native English speakers differ. Our results suggest that speech production may be one modality where conceptual transfer of framing typology is not easily observable and that differences in L2 speech production can be best explained as resulting from articulatory differences between native and non-native speakers of the target language.
This article investigates whether lexical frequency drives stress assignment in Greek, a morphology-dependent stress system. We examine the stress distributions of trisyllabic nouns from seven inflection classes across three lexical databases (GreekLex 2, the HNC Golden Corpus, and HelexKids 2.0) and compare them with the stress patterns produced by adults and elementary school children (Grades 3 and 6) in a pseudo-noun elicitation task. Using hypothesis testing for proportions, frequentist and Bayesian generalized linear mixed-effects models, Bayes Factor analyses, and Monte Carlo simulations, we show that speakers' stress choices do not replicate lexical frequency distributions. Adults show systematic inflection-class conditioning: their stress patterns broadly follow the lexicon's dominant stress pattern per inflection class but depart from its distributional frequencies. Children exhibit a strong, frequency-insensitive preference for penultimate stress that persists across inflection classes, with only traces of suffix-specific stress preferences emerging in Grade 6. We argue that lexical frequency operates within grammatically imposed boundaries: the phonological default and inflection-class conditioning together determine the extent to which distributional patterns in the lexicon shape stress assignment.
This study examined how dialectal background and individual production patterns influence nasality perception in American English, focusing on Midland and Inland North listeners. Forty-one adults from the two dialect regions completed nasometric testing to index oral-nasal balance characteristics (nasalance). Nasality perception was assessed using direct magnitude estimation (DME) of phrase-level synthetic stimuli, representing a continuum of synthesized velopharyngeal (VP) port sizes (0-0.20 cm2, in 0.04 cm2 steps). Results showed no significant between-dialect differences in nasalance. In contrast, DME ratings increased systematically with synthesized VP port size and differed by dialect, with Inland North listeners assigning higher ratings than Midland listeners, particularly at larger synthesized VP port sizes, yielding a significant dialect-by-synthesized VP port size interaction. Correlation analyses further examined production-perception relationships and revealed a significant negative association between nasalance and DME ratings for the 0.08 cm2 stimulus among Midland listeners, whereas no such association emerged among Inland North listeners. These findings indicate that perceptual judgments of nasality may vary even when production measures appear comparable. One possible explanation is that dialect contact may reduce production differences while perceptual representations remain relatively stable. The production-perception relationship observed exclusively for the 0.08 cm2 stimulus among Midland listeners further suggests that such links may be most evident when external stimulus contrast is minimal and listeners' production and perceptual norms remain aligned. More broadly, these findings support the view that nasality perception is shaped by experience-based listener factors, consistent with exemplar and normalization accounts of speech perception.
This study investigates how German learners of French acquire liaison and proposes a formal account within the Gradient Harmonic Grammar (GHG) framework. A production experiment with 65 secondary school learners (aged 10-17 years) identified several distinct realisation types (including glottal stop insertion, /l/-insertion, and emerging target-like resyllabification) and revealed systematic effects of lexical and positional frequency. Building on these findings, a revised GHG model is developed that integrates corpus-based activation values for Gradient Symbolic Representations, allowing item-specific variation to be modelled formally. Adjustments to the original analysis by Smolensky and Goldrick include the incorporation of further constraints and frequency-shaped activation values for both liaison and nonliaison consonants. The resulting model successfully captures both variable learner patterns and target-like liaison. The study demonstrates that combining usage-based insights with a gradient approach to underlying representations provides a powerful framework for understanding the phonological mechanisms involved in L2 liaison acquisition.
While bilingual studies have extensively explored crosslinguistic transfer across various domains, prosody remains relatively underexplored. This study addresses the gap by contributing experimental data on prosodic transfer at lexical and pragmatic levels in heritage speakers of Turkish in Germany. Given typological differences between Turkish and German prosody, this population is ideal for examining whether the heritage language (Turkish) is influenced by the majority language (German) in processing prosodic information across linguistic levels. In a recent study, Zora et al. examined how majority language speakers of Turkish interpret prosodic cues, specifically lexical stress and prosodic focus, when assessing sentence acceptability. Forty participants rated sentences containing prosodic violations, revealing that lexical stress violations elicited lower acceptability than prosodic focus violations. This suggests that lexical-level prosody, rather than pragmatic-level prosody, plays a more central role in spoken comprehension for majority language speakers of Turkish. Building on these findings, this study applies the same paradigm to 40 heritage Turkish speakers in Germany. Results show the reverse pattern, such that prosodic focus violations were rated less acceptable than lexical stress violations. Typologically, this aligns with Germanic languages, where prosodic focus holds greater weight in interpretation under the same experimental paradigm. The results further revealed that proficiency modulated sensitivity to prosodic information, such that high-proficiency speakers of Turkish judged both stress and focus violations as incorrect, whereas low-proficiency speakers judged only focus violations as incorrect. These findings highlight the dynamic nature of prosodic transfer, potentially from German to Turkish, and the significant role of language dominance and proficiency in shaping bilingual prosodic processing.
For successful processing of the speech signal, listeners need to perform at least two prosodic tasks: allocating prominences and grouping into smaller units. The separation of these two functions has been a fundamental assumption in the modeling of prosody production and perception across languages. The current study investigates how listeners allocate prominences and demarcate (group) them at the word level in Papuan Malay, an under-researched language of Papua, Indonesia. To this end, a tapping and a grouping task were carried out using acoustically manipulated sequences of strong and weak syllables. It was tested how duration, intensity, spectral tilt and vowel quality affect listeners' prominence and grouping responses. The results show that, taking into account the variation among participants and items, all cues facilitate prominence allocation, in particular if they are strong enough. Spectral tilt was the main cue found to affect grouping, although duration seemed to play a double role for both prominence and grouping. The outcomes are discussed not only for how they improve our understanding of Papuan Malay prosody, but also for their contribution to the separation of perceptual functions in prosodic theory.
This study explores prosodically driven perceptual adjustment of a non-contrastive coarticulatory feature. Specifically, it examines how prosodic prominence and information structure influence the perception of anticipatory vowel nasalization in CVN in American English as a cue to an upcoming nasal consonant. Using a two-alternative forced-choice task, we tested whether the same acoustic degree of nasalization is interpreted differently depending on prosodic and information-structural contexts, and whether these effects are consistent across native English speakers and Korean learners of English. Listeners identified target words (Bob or bomb) from a nasalization continuum embedded in carrier sentences ("(No,) Riley wrote Bob/bomb slowly"), where the coda consonant was masked by noise. Prosodic prominence was manipulated by adjusting pitch and amplitude of surrounding words while keeping the target constant; and information structure was manipulated by including or omitting No, signaling contrastive focus. For a given stimulus, listeners gave more nasal responses (bomb) when the target was prosodically prominent or preceded by No, implying that they compensate for reduced vowel nasality in these contexts. The fact that No influenced perception independent of prosody suggests that information structure can modulate perceptual compensation beyond acoustic-prosodic cues, reflecting a top-down effect of discourse meaning. Finally, although Korean listeners gave fewer nasal responses overall-possibly due to reliance on acoustic detail or perceptual hypercorrection-both groups showed similar sensitivity to prosodic and information-structural cues. These findings taken together point toward a shared perceptual mechanism in both native and non-native perception that integrates fine-grained phonetic detail with prosodic and information-structural contexts.
Previous research on imitation in L1 has shown that explicit instructions generally generate more robust convergence with the model talker than implicit instructions. However, this regularity has not gained sufficient empirical attention in L2 speech imitation. To address this gap, we examined English vowel duration variability as a cue to following voiced versus voiceless consonants (vowel clipping) in 28 Czech and 28 Polish L2 learners of English. Neither Czech nor Polish uses vowel clipping as a voicing cue in the L1. Participants completed three tasks: (1) baseline, (2) imitation, and (3) post-test. Our results confirmed convergence towards the model stimuli during imitation, but this effect diminished in the post-test. The critical manipulation involved the imitation task: half of the participants were explicitly instructed to imitate the model voice as accurately as possible (explicit imitation), whereas the other half performed a simple oral identification task (implicit imitation). Contrary to predictions based on L1 imitation studies, explicit instructions did not increase the magnitude of convergence in L2 learners in our study. Moreover, the results showed that the Czech and Polish participants exhibited comparable degrees of convergence with the native English model, thus replicating the null effect of instruction type across two language groups. We discuss the current results by offering two lines of reasoning as to why explicit instructions did not, in our study or more generally, lead to more imitation in L2 speech.
Echoing recent calls for exploring skill-specific individual-difference factors, this two-part study investigated the predictive effect of pronunciation anxiety, enjoyment, and motivation on Chinese university English-as-a-foreign-language (EFL) learners' global second-language (L2) speech learning. Study 1 cross-sectionally investigated the predictive effects of emotional and motivational factors on 82 participants' comprehensibility and accentedness attainment based on expert raters' evaluation of speech samples elicited from a long-turn monologic task. Statistical results found that comprehensibility was unrelated to participants' emotions and motivation, while pronunciation anxiety negatively predicted accentedness. Study 2 explored 65 participants' speech development following an 8-week explicit pronunciation instruction. Results revealed that participants developed both accentedness and comprehensibility after pronunciation instruction, regardless of their different emotional and motivational traits. Theoretical insights into revising speech acquisition models and pedagogical implications for pronunciation teaching were discussed.
Second language (L2) learners often face a number of obstacles in acquiring target-like speech production patterns. This study examines L1 Japanese speakers' production of L2 English /l/ to investigate how they implement position-dependent allophonic variation while overcoming mismatches between Japanese (L1) and English (L2) in (1) the number of liquid phonemes, (2) their allophonic distributions, and (3) syllable structure. Acoustic and articulatory (ultrasound) analyses of word-initial and word-final English /l/ produced by thirteen L1 Japanese-L2 English speakers and nine L1 English speakers show that L1 Japanese-L2 English speakers produce target-like lateral allophony acoustically, although their lateral production is overall clearer (i.e., less velarised) than that of L1 English speakers. By contrast, L1 Japanese-L2 English speakers' articulation is non-target-like, characterised by different tongue shape dimensions and the absence of coronal-dorsal timing patterns found in L1 English speakers' production. Overall, the results highlight the complex nature of L2 speech production, in which L2 speakers utilise a wider range of phonetic cues to overcome learning challenges than is commonly assumed.
This study investigates how adult learners perceive segmental length in L2 Italian. We tested 104 learners from 5 L1 backgrounds that differ in their phonological treatment of quantity: Finnish (vowel and consonant length), Czech and Slovak (vowel length), German (restricted vowel length), and Spanish (no phonemic length). These groups were compared with 34 native speakers of Italian, a language with distinctive consonant length (e.g., papa 'Pope' vs. pappa 'porridge'). Using an AX discrimination task with 45 pseudowords, we examined the influence of L1, stress position, consonant type, and proficiency. The results reveal that L1 strongly predicts perceptual sensitivity to consonantal length contrasts: Finnish learners showed the highest sensitivity, followed by Slovak, Czech, German, and Spanish learners. Discrimination of vowel-length contrasts was weaker across all groups and showed more variation, likely because in Italian, vowel duration does not cue phoneme identity. Stress position also played a significant role: discrimination of consonantal quantity contrasts was most robust in post- and pre-stressed syllables and weakest in unstressed contexts. Consonant type further influenced performance, with quantity being easier to discriminate in sonorants than in obstruents, although L1-specific effects emerged. Proficiency did not emerge as a uniform global predictor, but descriptive patterns indicated higher performance among advanced learners than among beginner-level and intermediate-level learners. Overall, the multifactorial analysis highlights the role of L1 in shaping learners' perception of L2 length, alongside prosodic effects, segmental properties, and proficiency. The study also underscores the importance of perceptual training in acquiring new phonological contrasts.
Extracting talker identity from speech signals is a core perceptual function, yet the mechanisms underlying second Language (L2) identity processing remain unclear. Grounded in the DIVA model and Source-Filter Theory, this study investigated the coupling between perception and production in Tibetan-Mandarin bilinguals. Participants performed a delayed imitation task and a talker-identity discrimination task. Imitation performance was quantified within a multidimensional acoustic space defined by fundamental frequency (F0), harmonics-to-noise ratio (HNR), and formant dispersion (FD). Results indicated that learners achieved significant acoustic convergence toward the L2 model speaker, which persisted as episodic traces across short-term temporal delays. In the discrimination task, sensitivity improved with acoustic distance but plateaued between medium and large distances, while a significant negative response bias in the near condition revealed a tendency toward perceptual assimilation. Crucially, regression and machine-learning analyses revealed that only FD distance was significantly associated with discrimination sensitivity. Unlike source-related cues such as F0 that fluctuate with context, FD reflects relatively invariant vocal-tract structures. These findings suggest that the formation of L2 talker-identity representations involves a functional anatomical alignment with the target speaker through sensorimotor inverse mapping. By locking onto structural invariants like FD, learners can overcome within-person variability to form detailed episodic identity representations. This study extends the scope of auditory targets in speech production models from segmental to indexical levels.
This study examines (a) the prosodic cues to stress and focus in Ukrainian speakers in their first language (L1 Ukrainian) and in a second language (L2 German), as compared to German speakers and (b) the influence of F0 during online word recognition in L2 German. Analyses of productions of trisyllabic target words differing in stress (Experiment 1) show that stressed vowels are consistently longer and produced with slightly higher intensity than unstressed vowels, with little additional modulation by focus (focused, unfocused) and hardly any group effects (L1 Ukrainian, L2 German, L1 German). The groups differed strongly in F0 though. While L1 Ukrainian speakers produced the initial syllable with high-pitched and stressed syllables with a falling/low F0 contour (suggesting head-edge prominence marking), L1 German speakers produced focus with a rising pitch accent. There were two subgroups of learners: a small group that produced the conditions as in L1 Ukrainian (Type 1) and a larger group that was similar to L1 German (Type 2). In perception (Experiment 2), target words were fixated more when the first syllable was high-pitched (in line with L1 Ukrainian productions). Similar to Germans, Type 2 learners temporarily treated high-pitched syllables as stressed, leading to increased competitor fixations when the pitch peak preceded the stressed syllable (early-peak accent, H+L*, compared to medial-peak accents, L+H*). These findings demonstrate individual differences in the L2 production of stress and focus, which are closely linked to cue use during L2 word recognition.
This study examined the read speech of vowel length contrasts produced by Cantonese-English-German trilinguals, comparing their performance with that of Mandarin-English-German trilinguals, Cantonese-English bilinguals, native English speakers, and native German speakers. Acoustic and statistical analyses of vowel quality and duration across the first language (L1), second language (L2), and third language (L3) yielded several key findings. First, Cantonese-speaking trilinguals were more nativelike than the Mandarin-speaking trilinguals in both the L2 and L3, indicating that the L1 exerts a sustained influence across the multilingual system. Second, Cantonese-English-German trilinguals differed from Cantonese-English bilinguals in the L2 but not the L1, suggesting that reverse transfer from the L3 more strongly affects the L2 than the L1. Finally, individuals who produced larger quality contrasts in L2 vowel length distinctions also demonstrated greater quality contrasts in comparable L3 vowels, whereas those who produced larger duration contrasts in the L1 exhibited reduced duration contrasts in analogous L3 vowels, indicating that L1-L3 and L2-L3 bidirectional interactions emerge among phonetically similar vowels. The study highlights the dynamicity of the multilingual phonological system.
This study reports an experiment conducted to examine the contribution of non-native speech timing to the perception of foreign accent. Native English listeners rated utterances produced in English by speakers whose first language was English or Saudi Arabic, for the degree of perceived foreign accent. The utterances were acoustically modified to significantly reduce segmental and intonational information available to the listeners. The listeners were able to distinguish the native and non-native speaker groups in the acoustically degraded utterances. This suggests that the listeners were able to make use of temporal cues, in the absence of segmental and intonational information, to rate the utterances for the degree of foreign accent. To further investigate this, three temporal measures (articulation rate, durational ratio of unstressed to stressed vowels, utterance-final vowel lengthening) were calculated for each utterance to examine their contribution to the overall perception of foreign accent. Among these measures, articulation rate and, to a lesser extent, the durational ratio of unstressed to stressed vowels played a role in cueing the listeners' perception of foreign accent. However, while the impact of articulation rate on listener ratings varied by speaker group, higher values of the ratio of unstressed to stressed vowel duration, reflecting lower degrees of vowel reduction, consistently predicted foreign accent ratings.