This study reports an experiment conducted to examine the contribution of non-native speech timing to the perception of foreign accent. Native English listeners rated utterances produced in English by speakers whose first language was English or Saudi Arabic, for the degree of perceived foreign accent. The utterances were acoustically modified to significantly reduce segmental and intonational information available to the listeners. The listeners were able to distinguish the native and non-native speaker groups in the acoustically degraded utterances. This suggests that the listeners were able to make use of temporal cues, in the absence of segmental and intonational information, to rate the utterances for the degree of foreign accent. To further investigate this, three temporal measures (articulation rate, durational ratio of unstressed to stressed vowels, utterance-final vowel lengthening) were calculated for each utterance to examine their contribution to the overall perception of foreign accent. Among these measures, articulation rate and, to a lesser extent, the durational ratio of unstressed to stressed vowels played a role in cueing the listeners' perception of foreign accent. However, while the impact of articulation rate on listener ratings varied by speaker group, higher values of the ratio of unstressed to stressed vowel duration, reflecting lower degrees of vowel reduction, consistently predicted foreign accent ratings.
Many adults learn languages with written forms that differ from their first language(s). Empirical research has demonstrated the influential role of written input on developing L2 phonology. However, existing studies are limited by (1) focusing on learning languages that share the same orthographic script, predominantly the Latin alphabet, (2) small sample sizes, and (3) limited consideration of L2 proficiency. This study investigated the influence of Arabic and English written input when lexically encoding the difficult /f-v/ phonological contrast for L1 Arabic-speaking learners of L2 English. A word learning study was completed by 114 L1 Arabic speakers, with varying English proficiency, and 117 L1 English-speaking controls. Mixed-effects modeling of L1 Arabic accuracy revealed an inhibitory effect of any written input when learning words differing by the difficult contrast. Performance improved with increasing L2 proficiency; however, the inhibitory effect of written input for words differing by /f-v/ persisted into high levels of L2 proficiency.
Stress placement in English loanwords into Mirpur Pahari (MP) is used to explore whether a usage-based approach can inform the incorporation of external factors (e.g., exposure to a donor language; here, English) into formal phonological analysis of loanword adaptation alongside internal factors (e.g., phonology of the recipient language; here, MP). Stress placement in English loanwords into MP shows across- and within-speaker variation between conformity to MP stress rules (formalized in classical Optimality Theory; OT) and retention of stress on the syllable that is stressed in English; this is a challenge for most theories of loanword adaptation. In our hybrid approach, variable adaptation patterns in individual speakers' loanword realizations in production data from twelve MP speakers in the UK are correlated with degree of exposure to English, operationalized as vocabulary size, and a unified formal account is sketched through usage-based weighting of constraints in Stochastic OT.
Rhythm metrics can detect second language development of target-like speech rhythm but interpretation of the results from metrics in learners’ speech is problematic because the mapping of metrics to underpinning phonological features is indirect. We investigate speech rhythm in first language (L1) Arabic / second language (L2) English, which differ in key properties contributing to the percept of rhythm: unstressed vowel reduction and syllable structure. Our production data are interpreted using additional measures, of stressed and unstressed vowels and of consonant cluster realization, alongside standard rhythm metrics; this combination facilitates disambiguation of competing interpretations of the metric results. The findings confirm the importance of using multiple rhythm metrics to study L2 speech rhythm and demonstrate how simple additional measures can guide interpretation of their results. In this study the metrics results showed that the speech produced by the L2 speakers, regardless of their length of residence in the UK, exhibited lower vocalic durational variability than the speech produced by the native Arabic and English speakers. However, closer inspection of the degree of vowel reduction by the native and nonnative groups confirms that no single metric captures the complex nature of the observed L2 rhythm patterns. Future L2 studies are advised not to draw firm conclusions about the degree of vowel reduction and consonant cluster realization in L2 speech based solely on the results of the rhythm metrics.
In 2022 we planned speech data collection with speakers of Syrian and Jordanian dialects to inform an updated Syrian Arabic dialectology in response to sustained displacement of millions of Syrians. The pandemic imposed remote data collection, but an internet-based approach also facilitated recruitment with this highly distributed speech community. Their vulnerable situation brings barriers, however, since most prospective participants have limited internet data and rarely use email. We collected self-recorded short audio files in which participants read scripted materials and described pictures. Three platforms were tested: Gorilla, Phonic, and Awesome Voice Recorder (AVR, smartphone app). Gorilla/Phonic offer stimulus presentation advantages, and so were piloted thoroughly, but the audio quality obtained was not suitable for phonetic analysis. AVR yields full spectrum .wav files but requires participants to submit files by email, so we recruited local fieldworkers, or Public Involvement Coordinators (PICs), to support participants with recording and file submission. Through surveys and interviews, we asked PICs and 20% of participants about their experience of working with us. The results confirm PIC fieldworker involvement was crucial to the success of the project, which generated high quality audio data suitable for phonetic analysis from 134 speakers within three months (Almbark, Hellmuth, & Brown, forthcoming).
Recent empirical studies have highlighted the large degree of analytic flexibility in data analysis that can lead to substantially different conclusions based on the same data set. Thus, researchers have expressed their concerns that these researcher degrees of freedom might facilitate bias and can lead to claims that do not stand the test of time. Even greater flexibility is to be expected in fields in which the primary data lend themselves to a variety of possible operationalizations. The multidimensional, temporally extended nature of speech constitutes an ideal testing ground for assessing the variability in analytic approaches, which derives not only from aspects of statistical modeling but also from decisions regarding the quantification of the measured behavior. In this study, we gave the same speech-production data set to 46 teams of researchers and asked them to answer the same research question, resulting in substantial variability in reported effect sizes and their interpretation. Using Bayesian meta-analytic tools, we further found little to no evidence that the observed variability can be explained by analysts' prior beliefs, expertise, or the perceived quality of their analyses. In light of this idiosyncratic variability, we recommend that researchers more transparently share details of their analysis, strengthen the link between theoretical construct and quantitative system, and calibrate their (un)certainty in their conclusions.
In 2022 we planned speech data collection with speakers of Syrian and Jordanian dialects to inform an updated Syrian Arabic dialectology in response to sustained displacement of millions of Syrians. The pandemic imposed remote data collection, but an internet-based approach also facilitated recruitment with this highly distributed speech community. Their vulnerable situation brings barriers, however, since most prospective participants have limited internet data and rarely use email. We collected self-recorded short audio files in which participants read scripted materials and described pictures. Three platforms were tested: Gorilla, Phonic and Awesome Voice Recorder (AVR, smartphone app). Gorilla/Phonic offer stimulus presentation advantages, so were piloted thoroughly, but the audio quality obtained was not suitable for phonetic analysis, ruling out their use in the main study. AVR yields full spectrum wav files but requires participants to submit files by email, so we recruited local fieldworkers to support participants with recording and file submission. We asked fieldworkers and participants about their experience of working with us, through surveys and interviews. The results confirm fieldworker involvement was crucial to the success of the project which generated high quality audio data, suitable for phonetic analysis, from 134 speakers within three months (Almbark, Hellmuth, & Brown, forthcoming).
Multicultural London English (MLE; Kerswill and Torgersen 2008; Cheshire et al., 2011) arose in working class areas of London around 30 years ago through intensive, multiethnic social contact (Kerswill and Torgersen, 2021). The variety’s highly systematic phonological system has been seen as displacing traditional London vernacular varieties including Cockney, which have moved further East (Fox, 2015). MLE and Standard Southern British English are two strands among many, interwoven with features of the earlier London vernacular and features of non-MLE varieties (Sharma and Sankaran, 2011). Using preliminary data from a new project, Generations of London English, we present an acoustic analysis that shows both the stability and emerging maturity of MLE as well as strands of continuity from the past. The GOAT vowel shows a diverse range of variants indexing social class, ethnicity, age, and gender. By contrast, all Londoners participate in GOOSE-fronting, which is thus primarily an index of age, and women lead FOOT-fronting, a city-wide change showing less ethnic and social class sensitivity. A generalized use of labels such as MLE risks mischaracterizing the sociophonetic reality. We argue for analysis of individual phonetic profiles as unique intersections of historically layered features, diffusing differently through a heterogeneous population.
In 2022 we planned speech data collection with speakers of Syrian and Jordanian dialects to inform an updated Syrian Arabic dialectology in response to sustained displacement of millions of Syrians. The pandemic imposed remote data collection, but an internet-based approach also facilitated recruitment with this highly distributed speech community. Their vulnerable situation brings barriers, however, since most prospective participants have limited internet data and rarely use email. We collected self-recorded short audio files in which participants read scripted materials and described pictures. Three platforms were tested: Gorilla, Phonic and Awesome Voice Recorder (AVR, smartphone app). Gorilla/Phonic offer stimulus presentation advantages, so were piloted thoroughly, but the audio quality obtained was not suitable for phonetic analysis, ruling out their use in the main study. AVR yields full spectrum wav files but requires participants to submit files by email, so we recruited local fieldworkers to support participants with recording and file submission. We asked fieldworkers and participants about their experience of working with us, through surveys and interviews. The results confirm fieldworker involvement was crucial to the success of the project which generated high quality audio data, suitable for phonetic analysis, from 134 speakers within three months (Almbark, Hellmuth, & Brown, forthcoming).
Diglossia in Arabic differs from bilingualism in functional differentiation and mode of acquisition of the two registers used by all speakers raised in an Arabic-speaking environment. The ‘low’ (L) regional spoken dialect is acquired naturally and used in daily life, but the ‘high’ (H) variety, Modern Standard Arabic, is learned and used in formal settings. Register variation between the two ends of this H–L continuum is ubiquitous in everyday interaction, such that authors have proposed distinct intermediate register levels, despite evidence of mixing of H and L features, within and between utterances, at all linguistic levels. The role of sentence prosody in register variation in Arabic is uninvestigated to date. The present study examines three variables (F0 variation, intonational choices and post-lexical utterance-final laryngealization) in 400+ turns at talk produced by one speaker of San’ani Arabic in a 20 min sociolinguistic interview, coded for register on three levels: formal (fusħa), ‘middle’ (wusṭaː) and dialect (ʕaːmijja). The results reveal a picture of key shared features across all register levels, alongside distinct properties which serve to differentiate the registers at each end of the continuum, at least some of which appear to be under the speaker’s control.
Dialect variation spans different linguistic levels of analysis. Two examples include the typical phonetic realisations produced and the typical range of intonational choices made by individuals belonging to a given dialect group. Taking the modeling principles of a specific automatic accent recognition system, the work here characterises and observes the variation that exists within these two levels of analysis among eight Arabic dialects. Using a method that has previously shown promising performance on English accent varieties, we first model the segmental level of analysis from recordings of Arabic speakers to capture the variation in the phonetic realisations of the vowels and consonants. In doing so, we show how powerful this model can be in distinguishing between Arabic dialects. This paper then shows how this modeling approach can be adapted to instead characterise prosodic variation among these same dialects from the same speech recordings. This allows us to inspect the relative power of the segmental and prosodic levels of analysis in separating the Arabic dialects. This work opens up the possibility of using these modeling frameworks to study the extent and nature of phonetic and prosodic variation across speech corpora.
Post-focal compression (PFC) of F0 is a known cue to focus in English and Beijing Mandarin (BM), but PFC is neither present in Taiwan Mandarin (TM) production nor interpreted as a cue to focus in perception [1].Studies of variation in L2 English production by BM and TM learners of English confirm transfer of some L1 patterns into their L2 English [2].This paper explores for the first time how BM and TM listeners' interpret PFC in English.It also seeks to clarify L2 listeners' interpretation of PFC in contexts where discourse-new postfocal material carries a post-focal prominence in English [3].Following [4] we presented L1 BM, TM and English listeners with a series of written discourse contexts and two prosodically congruous or incongruous audio responses in a betweenparticipants design.One set of listeners in each language group judged SVO English stimuli produced in either all-new context (NN) or with initial narrow focus followed by discourse-given post-focal material (FG).Another set of listeners in each group judged the same all-new (NN) stimuli against recordings with initial narrow focus followed by discourse-new post-focal material (FN).Results indicate differential interpretation of onfocus and post-focal prosody matching a L1 perceptual transfer hypothesis.
This study seeks to explore the impact of the input modality, language exposure, context, and gender on the production patterns of two non-native sounds, /tʃ/ and /v/, in Saudi Arabic.A production task was conducted to test 67 Saudi speakers in three conditions: aural-only (auditory inputs), written-only (orthographic inputs), and aural-written (auditory-orthographic inputs).Language exposure had a main effect on the production of the two sounds.Context was a major factor influencing the production accuracy of /v/ but not /tʃ/; /v/ was more likely to be devoiced in wordfinal position.The written input resulted in a decrease in the production accuracy for /tʃ/but not /v/, suggesting that the effect of the input type varies for different nonnative sounds.
The variety of English spoken in the city of York, UK, is of sociolinguistic interest due to ‘recycling’ of traditional dialectal forms such as Definite Article Reduction (‘ to t’pub ’) and Past-Reference come (‘ I come home late last night ’) by younger (typically male) speakers; in apparent time studies based on the York English Corpus (YEC), middle-aged speakers (aged 50-70) used these forms less than older speakers (>70), so the patterns had previously appeared to be falling out of use. In this paper we first argue for the existence of a distinctive ‘Yorkshire rise-fall’ nuclear contour, which is sufficiently different in form and distribution from rise-fall contours reported for other varieties of British English that it can be characterized as a traditional (prosodic) feature of Yorkshire dialects. We then explore whether the observed patterns of variation in lexical-grammatical variables are mirrored in variation and change in use of this distinctive Yorkshire rise-fall nuclear contour, in apparent time, via qualitative analysis of data from the YEC.
Presently there is no consensus regarding the interpretation and analysis of the stress system of Moroccan Arabic. This paper tests whether the acoustic realisation of syllables support one widely adopted interpretation of lexical stress, according to which stress is either penultimate or final depending on syllable weight. The experiment reports on word-initial syllables that differ in presumed stress status. Target words were embedded in a carrier sentence within a scripted mock dialogue to ensure that the measurements reflect lexical stress rather than phrase-level prominence. Results from all four acoustic parameters tested (f0, duration, Centre of Gravity and vowel quality) showed that there were no differences as a function of presumed stress status, thus failing to support an interpretation according to which stressed syllables are acoustically differentiated. We consider the results in relation to previous claims and observations, and conclude that the absence of acoustic correlates of presumed stress is compatible with the view that Moroccan Arabic lacks lexical stress.