
Abstract This article explores the visual representation of second-language (L2) pronunciation, a crucial but underexamined area in instructed Second Language Acquisition. Building on the Noticing Hypothesis, the study argues that visual aids provide essential ‘graphic facilitation’ that helps learners anchor transient acoustic input into more permanent representations. To address the current lack of systematic classification, the article proposes a taxonomy that categorizes representation into four broad types: orthographic, notational, salience-based, and diagrammatic. Furthermore, an evaluative framework is introduced to assist researchers and practitioners in selecting techniques based on four core dimensions: accessibility, informativeness, cognitive load, and learnability. By moving beyond descriptive taxonomy to offer a structured evaluative matrix, the article seeks to optimize instructional practices and provide a theoretical foundation for future empirical research on the effectiveness of diverse visual interventions for L2 phonological development.
Abstract English and Spanish differ in syllable formation. One difference is that certain vocoid sequences in English are produced in hiatus, whereas the same vocoid sequences in Spanish are produced either as a diphthong or as a hiatus, subject to phonological conditioning. This difference may contribute to crosslinguistic interference in first language (L1) English speakers producing vocoid sequences in second language (L2) Spanish. This interference may increase in cognates (e.g., piano). This paper analyses diphthong/hiatus production in L2 Spanish, to explore acoustic contrast of diphthong/hiatus and test whether production is modulated by L2 proficiency and cognate status. L2 learners and L1 speakers read Spanish words containing the sequence /ia/. Duration and formant trajectories were analyzed. L1 Spanish speakers produced the expected diphthong/hiatus contrast, but L2 learners did not. Learners’ production was modulated by proficiency, starting with a diphthong bias, instead of the expected hiatus bias. There were no strong cognate effects.
Abstract Filler particles (FPs) are used as indicators of fluency in second language (L2) speech research, yet their diverse phonetic exponents and their multifunctionality pose significant methodological challenges. This study critically examines the operationalization and measurement of FPs within the domain of utterance fluency. In order to better capture the multifaceted role of FPs, the domain of filler fluency is proposed alongside established categories such as speed, breakdown, and repair fluency. This study employs both static and dynamic approaches to critically evaluate FP measures in spontaneous, task-based L1–L2 dialogues involving 24 speakers. Static analyses show how the definition of FPs and how they are annotated affects the outcome of FP measurements. The dynamic, time-series methods explored here incorporate section-wise analyses, cumulative FP frequency trajectories, and sliding window techniques to capture intra-individual variability in FP use, revealing patterns that static averages often overlook.
Abstract This study investigates how listeners integrate dynamic cues to perceive the Mandarin medial glide /j/ over time, and how this process is modulated by speech rate and second-language (L2) background. We tested 30 adults in three groups: Mandarin native speakers, intermediate Japanese learners, and beginning Japanese learners. A female Beijing Mandarin talker produced disyllabic carrier phrases of the form mao + target syllable, containing the four target syllables miao, mian, mia, and yao , at three speech rates (fast, mid, slow). For each token, the first syllable mao was kept intact, while the second syllable was progressively revealed across ten gates spanning the interval from the onset of the second syllable to the early steady-state portion of the vowel. On each trial, participants heard a two-syllable gated sequence and rated how much the second syllable sounded as if it began with an /i/-like quality on a five-point scale. Cumulative-link and linear mixed-effects models revealed robust effects of gate, speech rate, syllable structure, and group. Ratings increased steadily with gate level, were highest in slow speech, and followed a stable gradient of MNS > JP_INT > JP_BEG; CGVX syllables, especially miao , were easier and faster than the GVX syllable yao . The results highlight time-based cue integration as a locus of L2 difficulty, consistent with dynamic cue-integration accounts of speech perception and contemporary models of L2 phonology.
This scoping review maps 25 years of empirical research on L2 English pronunciation (1996-2020), covering 463 studies published in 35 prominent journals across second language acquisition, second language learning and teaching, and phonetics and phonology. Using Arksey and O'Malley's framework, it traces developments in participants' L1 profiles, phonological features, and the treatment of the construct of intelligibility. Four major trends emerge: a gradual diversification of participant profiles, a move away from native-speaker benchmarks, increased attention to suprasegmental features, and a rising - yet still unstable - presence of intelligibility as a research focus. The analysis highlights both changes over time and persistent divides between domains of research. The review calls for more diverse samples, clearer definitions of intelligibility, and stronger cross-disciplinary dialogue. Such developments could help consolidate theoretical foundations, improve comparability across studies, and advance both research and pedagogy in L2 pronunciation.
Filler particles (FPs) are used as indicators of fluency in second language (L2) speech research, yet their diverse phonetic exponents and their multifunctionality pose significant methodological challenges. This study critically examines the operationalization and measurement of FPs within the domain of utterance fluency. In order to better capture the multifaceted role of FPs, the domain of filler fluency is proposed alongside established categories such as speed, breakdown, and repair fluency. This study employs both static and dynamic approaches to critically evaluate FP measures in spontaneous, task-based L1-L2 dialogues involving 24 speakers. Static analyses show how the definition of FPs and how they are annotated affects the outcome of FP measurements. The dynamic, time-series methods explored here incorporate section-wise analyses, cumulative FP frequency trajectories, and sliding window techniques to capture intra-individual variability in FP use, revealing patterns that static averages often overlook.
L2 pronunciation research has mainly focused on linguistic dimensions of speech, such as comprehensibility, intelligibility and accentedness. However, successful L2 communication also involves social dimensions, like acceptability or pleasantness. Moreover, L2 pronunciation research has mostly examined L2 English. This study addresses these gaps by examining how Finnish and Danish L1 speakers evaluate L2 speech comprehensibility, acceptability (operationalized as language switching and suitability for occupation), and pleasantness in their respective languages, as well as relationships between these dimensions.In Finnish, acceptability items were intercorrelated, and more strongly correlated to pleasantness than to comprehensibility. In the Danish data the two acceptability items were not intercorrelated. A statistically significant difference also emerged for acceptability in terms of language switching: Finnish L1 raters scored acceptability lower than Danish L1 raters. These results suggest that acceptability ratings are not only linguistic judgments, but also reflections of social inclusion, which may present differently in different contexts.
This study investigates how listeners integrate dynamic cues to perceive the Mandarin medial glide /j/ over time, and how this process is modulated by speech rate and second-language (L2) background. We tested 30 adults in three groups: Mandarin native speakers, intermediate Japanese learners, and beginning Japanese learners. A female Beijing Mandarin talker produced disyllabic carrier phrases of the form mao + target syllable, containing the four target syllables miao, mian, mia, and yao, at three speech rates (fast, mid, slow). For each token, the first syllable mao was kept intact, while the second syllable was progressively revealed across ten gates spanning the interval from the onset of the second syllable to the early steady-state portion of the vowel. On each trial, participants heard a two-syllable gated sequence and rated how much the second syllable sounded as if it began with an /i/-like quality on a five-point scale. Cumulative-link and linear mixed-effects models revealed robust effects of gate, speech rate, syllable structure, and group. Ratings increased steadily with gate level, were highest in slow speech, and followed a stable gradient of MNS > JP_INT > JP_BEG; CGVX syllables, especially miao, were easier and faster than the GVX syllable yao. The results highlight time-based cue integration as a locus of L-2 difficulty, consistent with dynamic cue-integration accounts of speech perception and contemporary models of L-2 phonology.
This article reviews previous literature on the concepts of fluency and prosody to examine how these two concepts interact in both foreign language (L2) learners' productions and native speaker (L1) perceptions of L2 speech. First, it presents a comprehensive overview of prosodic features of speech that play a role in L2 production (utterance fluency). Next, it explores the relationship between perceived fluency and other perception measures known to be influenced by prosodic accuracy, such as accentedness, comprehensibility, and intelligibility. Finally, it examines the influence of embodied visual information (e.g., hand, arm, and head gestures or facial expressions) on the production and perception of prosody and fluency in the L2. As such, it contributes to our understanding of the interplay between fluency and prosodic accuracy in both spoken and multimodal communication, and informs L2 learners and teachers on the potential benefits of embodiment for L2 prosody and fluency.
Second-language (L2) instructors often use hand gestures to teach pronunciation, yet empirical benefits vary in size and scope. Focusing on L2 Mandarin tone production, we explored whether physically exaggerating and emotionally emphasizing tone gestures enhanced pronunciation. Participants imitated videos of a native speaker producing tones in three conditions: Speech Alone (S), Speech + Gesture (SG), and Speech + Gesture + Enthusiasm (SGE). In S, the native speaker spoke tones without arm movements or enthusiastic facial expressions. In SG and SGE, exaggerated hand gestures followed tone contours, and in SGE, the speaker also produced enthusiastic facial expressions. While gesture and enthusiastic expressions yielded only modest pronunciation benefits-slightly higher F for Tone 1 and greater Tone 3 lengthening toward native values-self-ratings of motivation, enjoyment, preference, and helpfulness were substantially higher for SG and SGE than S, suggesting that gesture and enthusiastic expressions in L2 may influence affective experience more than correct pronunciation.
Speech fluency is often assessed using articulation rate and pause frequency. However, not all pauses hinder fluency: when placed strategically, they structure discourse and enhance comprehensibility. To better characterize speaker fluency, it is crucial to consider where pauses occur. Traditional approaches rely on categorical syntactic boundaries (e.g., clauses or phrases), but inadequately capture syntactic complexity. We propose a continuous measure of pause placement based on syntactic distance between adjacent words. Using spontaneous English speech from Japanese learners and native speakers, we show that syntactic distance robustly predicts both pause location and duration across proficiency levels. We compare its contribution to proficiency classification against baseline and categorical models. The syntactic distance model outperforms all others, explaining 87% of variance (versus 65% for baseline and 76% for clause/phrase models), with strongest model fit and lowest prediction error. This measure provides a robust and meaningful predictor of L2 speech fluency.
This study investigates the relationship between speech rhythm, utterance fluency, and native listeners' perceptions of comprehensibility and fluency in second language (L2) English. Eighty-two advanced Spanish-Catalan bilingual learners of English completed a spontaneous speaking task, from which temporal rhythm measures (including durational variability metrics and vowel reduction measures) and fluency measures (speed, breakdown, and repair) were extracted. Speech samples were rated for comprehensibility and fluency by L1 English listeners. Correlational and regression analyses revealed that while rhythm and fluency measures did not consistently relate to one another, both independently contributed to listeners' ratings. In addition, vowel reduction ratios were stronger predictors of ratings than traditional durational variability metrics. These findings suggest that rhythm and fluency constitute partially independent constructs that both shape global speaking proficiency perceptions. Pedagogically, the results highlight rhythm as a promising instructional target for enhancing learners' fluency and comprehensibility.