Vowels have long been known to vary beyond static representations in formant space, and instead vary with respect to time-dependent dynamic acoustic information, including duration and spectral change across their timecourse. At the same time, it remains unclear both the extent to which a vowel can vary in its dynamic properties across many dialects of a language, and how vowel dynamicity contributes to the successful disambiguation of dialects. This paper presents an analysis of the dynamic acoustic realization of seven vowels across 30 dialects of North American and British Isles Englishes, represented in 1.3 million tokens from almost 5000 speakers. We find that not only do dialects dramatically differ in the dynamic realization of vowels, but also that these dynamic spectral properties of vowels are almost always the most informative set of acoustic dimensions for distinguishing between dialects. These findings thus provide the first large-scale evidence in support of the importance of time-dependent acoustic information in the cross-dialectal realization of vowels for English.
This study presents a real-time acoustic-dynamic analysis of & lstrok;-vocalisation in Polish newsreel speech broadcast between 1944 and 1994. Using a 150,000-token dataset, we examined three phonetic variants-[l], [L] (reduced velarised lateral), and [w]-in postconsonantal onset (C/V), intervocalic onset (V/V), and postvocalic coda (V/C) positions, across three recording periods. We measured F2-F1 trajectories as a proxy for acoustic darkness, modelled using Generalised Additive Mixed Models (GAMMs) as a function of time, adjacent phonetic context, lexical stress, sequence duration, and lexical frequency. Results confirm a diachronic shift from the initial [l, L, w] system to a reduced [(L), w] system, with the change progressing fastest in V/C position and latest in C/V position, with V/V position showing signs of an earlier merger of phonetic variants than in C/V position. Across all positions, darker variants were favoured adjacent to /(sic), u/, while the influence of adjacent consonants proved more Polish-specific, with labial and velar consonants not yielding darker lateral realisations, contrary to findings from prior English-language studies. Our findings demonstrate that, in the lenition of [l] at its very late stages through to completion, the sound change continues to be shaped by the position of the segment within the syllable, but without being driven by the coarticulatory factors that typically initiate earlier stages, namely the influence of the adjacent segment. (c) 2026 The Author(s). Published by Elsevier Ltd.
This study examines how articulation rate is structured across a large and heterogeneous set of English varieties with respect to both speech production constraints and social factors. Through Bayesian distributional regression modelling of articulation rate mean and variance across 33 varieties of British and North American English drawn from 16 spontaneous speech corpora collected as part of the SPeech Across Dialects of English (SPADE) project, we observe that the within-speaker utterance length effect on articulation rate (whereby longer utterances are produced at a faster rate and with less between-utterance variance) is large and broadly consistent across varieties. Specifically, the effect on mean rate exhibits modest variability across varieties, and between-utterance variance shows high cross-variety consistency. Small but consistent effects of speaker gender and age on articulation rate are observed across all varieties in both size and direction, and neither effect can be attributed to differences in the length of utterances, indicating that the source of these social differences lies outside of any systematic variability in utterance length. These findings across 33 hetereogeneous varieties show that the constraints imposed by utterance length on articulation rate are largely invariant across the regional and stylistic diversity of naturally-occurring English speech, while social factors exert consistent, independent influences.
While automatic tools for speech annotation are now commonplace within phonetic research pipelines, many tasks require substantial manual correction or training sets to perform accurately. Simultaneously, large speech models such as wav2vec2 have been shown to perform well at speech classification tasks, raising the question of how these models may be applied to phonetic annotation tasks. We introduce wav2VOT: a tool for the automatic estimation of voice onset time, closure duration, and burst realisation using wav2vec2. We demonstrate that wav2VOT performs comparably with current approaches on unseen datasets, and can estimate with high accuracy with fine-tuning. Analysis of wav2VOT predictions demonstrate high fidelity across stop voicing and place of articulation. These results demonstrate that large speech models are capable of producing accurate annotations, and further motivate exploration of large speech models as tools in phonetic research pipelines.
The Syrian refugee crisis is among the largest globally. We are developing a social robot tailored to the needs of displaced Syrians hosted in Scotland. As part of a mixed-methods study to understand the needs of this population and the possible use cases of the robot, we conducted two focus groups with Syrian refugees and asylum seekers residing in Glasgow. Using thematic analysis, we identified these participants’ unmet needs and existing gaps in access to services. Participants observed an Arabic-speaking social robot, and together we explored its potential as a solution to help navigate bureaucratic processes and access services. The participants expressed curiosity and enthusiasm about the robot. As they shared experiences of homelessness and displacement, they also highlighted bureaucracy and the English language as key barriers to accessing services. This study identifies key design requirements for developing a multilingual support robot for refugees and asylum seekers.
In daily life, we interact with each other using the social, regional, and ethnic communication styles typical of our local communities. Successful communication further rests on our ability to seamlessly adjust to our interlocutors following the norms and expectations of our local social setting as well as conversational context and goals. However, despite significant advances in speech technology, most artificial speech systems-particularly, most social robots-still use a single, "standard", non-local communication style for all users, social settings and interaction goals. Recent research has shown that when they interact with digital agents, humans transfer and adapt their sociolinguistic behaviours, including communication bias. Despite this, the barriers set up by this inherent communication bias have never been systematically studied for HRI; and the potential benefits to user engagement from socially inclusive, diverse communication styles have not been explored. We argue that social robotics researchers should also consider sociolinguistic factors constraining human interaction. To explore the implications, we describe two hypothetical robots designed to support the local communication style of two regions of the United Kingdom, and we consider the potential sociolinguistic impact each robot might have on its conversational partners and the wider society.
International students in UK universities often struggle with interactions in English, particularly on their first days in the country. We have developed a multilingual support robot tailored to their needs. To evaluate the performance of the robot, 60 international students asked the robot, in either English or their native language (Modern Standard Arabic or Mandarin Chinese), for support on topics including campus directions, local tax exemption, financial aid, and official documents. Overall, users preferred to use their native language when interacting with the support robot, and using their native language in the robot interaction also had a positive effect on their perception of the interaction itself.
Social robots are increasingly being deployed in public spaces, where they face not only technological difficulties and unexpected user utterances, but also objections from stakeholders who may not be comfortable with introducing a robot into those spaces. We describe our difficulties with deploying a social robot in two different public settings: 1) Student services center; 2) Refugees and asylum seekers drop-in service. Although this is a failure report, in each use case we eventually managed to earn the trust of the staff and form a relationship with them, allowing us to deploy our robot and conduct our studies.
Conversational User Interfaces (CUIs), including chatbots, virtual agents and social robots, are increasingly shaping how we communicate, seek support and access services. Yet, as these systems grow more sophisticated, concerns about bias and fairness in their design and deployment have become increasingly urgent. We propose a multidimensional approach to bias and fairness in CUIs that spans four interconnected themes: conceptual grounding, verbal communication, multimodal expression and interactional dynamics. Rather than framing bias merely as a technical flaw, we argue that it should be understood as a relational, interactional and design-based phenomenon. Accordingly, in this workshop, we aim to foster critical discussion around how CUIs encode social norms, perpetuate or mitigate exclusion, and shape perceptions of fairness through their language, embodiment and behaviour. By bringing together researchers, designers and policymakers, the workshop will explore pathways towards more equitable and transparent CUIs. The goal is to promote a relational understanding of fairness, one that centres user experience and social context, to guide future work in conversational AI.
Modern phonetic research regularly makes use of automatic tools for the annotation of speech data, however few tools exist for the annotation of many variable phonetic phenomena. At the same time, pre-trained self-supervised models, such as wav2vec2.0, have been shown to perform well at speech classification tasks and latently encode fine-grained phonetic information. We demonstrate that wav2vec2.0 models can be trained to automatically classify stop burst presence with high accuracy in both English and Japanese, robust across both finely-curated and unprepared speech corpora. Patterns of variability in stop realisation are replicated with the automatic annotations, and closely follow those of manual annotations. These results demonstrate the potential of pre-trained speech models as tools for the automatic annotation and processing of speech corpus data, enabling researchers to `scale-up' the scope of phonetic research with relative ease.
Voice disorders, or dysphonia, in children impact communication, social interactions, and quality of life, emphasizing the need for effective assessment tools with accurate reference norms. Acoustic measures taken during sound prolongation are widely used to evaluate voice quality, but variability in children’s performance and limited norms from children from diverse backgrounds pose challenges for clinicians. This study investigated voice quality and variability in sound prolongation tasks among 5–12-year-old school children, contributing to the development of acoustic reference data. Method: A total of 275 primary school-aged children in Scotland participated, producing sustained phonations of [a], [s], and [z] to evaluate respiratory and phonatory performance. Durations and acoustic measures, including jitter, shimmer, Harmonics-to-Noise Ratio (HNR), Cepstral Peak Prominence (CPP), and s/z ratio, were analyzed to capture variability in performance. Results: Analysis indicated significant age-related increases in sound prolongation durations, with older children (7–12 years) outperforming younger children (5–6 years), reflecting enhanced respiratory capacity and vocal fold control. While jitter, shimmer, and HNR did not differ significantly across age groups, CPP values were higher in older children, indicating improved vocal stability and harmonic richness. Median s/z ratios also showed significant age-related changes, highlighting developmental changes in phonatory and respiratory coordination. Notably, children exhibited longer average sound prolongation durations than previously reported norms, with considerable variability in performance. No significant sex differences were found, except for the s/z ratio, where females had higher values. These findings contribute and advance the growing body of reference data for assessing voice quality in children and emphasize the importance of factors such as age and sex in large, diverse samples. The study highlights the need to account for developmental variability and robust, comprehensive methodologies to contextualize voice quality issues in children.
Objectives: This study investigates phonetic backward transfer in first-generation multilingual Indians in Glasgow ('Glaswasians'). According to the Revised Speech Learning Model (SLM-r) bilinguals' languages interact and influence each other in a shared phonetic space which can over time lead to an assimilation or dissimilation of sound categories. This prediction is applied to explore whether and how the sound systems of Glaswasians' native languages, Hindi and Indian English (IE), are influenced by the host variety Glasgow English (GE) after migrating to Glasgow. It also examines whether GE will affect IE more than Hindi due to linguistic proximity. Methodology: Two speaker groups were recruited. Both groups are native multilingual speakers of Hindi and Indian English, but differ in their language environment. The 'Glaswasian' group live in Glasgow, United Kingdom, and are constantly exposed to the ambient dialect of Glasgow English. The 'Indian' group live in India, where Hindi and Indian English are unaffected by contact with Glasgow English. Speakers were recorded reading sentence lists in English and Hindi containing three types of sounds: /b dg/, /l/, /u/. Data analysis: The analysis looked for effects of Group (Glaswasian/Indian) and Language (Hindi/English) on multiple acoustic properties - pre-voicing and burst intensity for /b d g/; F1, F2, F3 for /u/; F2-F1 difference for l/. Acoustic measures were taken in Praat; repeated measures analyses of variance (ANOVAs) and t-tests were conducted in R. Findings: Results were mixed for Glaswasians. Assimilation emerged in /u/ in Indian English (IE) and Hindi for F2, F3, and also for F2-F1 difference in /l/ in IE; partial assimilation appeared in both languages for /b d g/; and dissimilation appeared in Hindi for /l/ and /u/. However, IE did not show more influence from GE than Hindi, suggesting that linguistic proximity may not necessarily modulate the nature of interaction between sound systems. Originality: This paper demonstrates backward transfer across languages and dialects. It illustrates the complexity of backward transfer and examines the impact of a majority language (English) onto a minority language (Hindi) in a novel group. In addition, it compares a bilingual experimental group with a bilingual 'control' group, which is rarely done. Significance: This study contributes to the knowledge of cross-linguistic transfer and phonetic variation and change while highlighting the various ways in which backward transfer can be manifested.
This study evaluates both automated transcription (WhisperX) and forced alignment (MFA) in developing a semi-automated pipeline for obtaining acoustic vowel measures from field recordings from 275 children speaking a non-standard, English dialect, Scottish English. As expected, manual correction of speech transcriptions before forced alignment improves the quality of acoustic vowel measures with respect to manually-annotated data, though speech style and recording environment present some challenges for both tools. Adaptation of the MFA pre-trained english us arpa acoustic model towards the children's speech also improves the quality of acoustic measures, though greater improvement was not found by increasing training sample size.
Phonetic theories of sound change posit that coarticulatory factors systematically motivate the fine phonetic variation which promotes sound change time (Ohala 1981; Harrington and Schiel 2017). This chapter presents the first direct empirical evidence of how coarticulatory factors effectively control the progression of a sound change as it plays out in a community over real-time. Specifically, we show how the acoustic quality of word-initial /l/ in spontaneous Glaswegian vernacular speech changed across four decades, and in particular, we find that the change towards acoustic darkening of the lateral is both propelled, and resisted, entirely consistently with the kinds of predictions we would make from synchronic phonetic observation (Recasens and Espinosa 2005; Simonet 2015), namely that the darkening of the lateral takes place in acoustically "darker" preceding and following phonological contexts, and before acoustically "lighter" contexts, which resist darkening over time. An additional analysis of formant trajectories, using GAMM modelling, illustrates the diachronic impact of coarticulatory context on the dynamic acoustic quality of initial /l/. It also reveals how women led in the acoustic darkening of initial /l/ in Glasgow, underscoring the interaction of phonetic and social factors in the propagation of this real-time change.
Speech rate has been shown to vary across social categories such as gender, age, and dialect, while also being conditioned by properties of speech planning. The effect of utterance length, where speech rate is faster and less variable for longer utterances, has also been shown to reduce the role of social factors once it has been accounted for, leaving unclear the relationship between social factors and speech production in conditioning speech rate. Through modelling of speech rate across 13 English speech corpora, it is found that utterance length has the largest effect on speech rate, though this effect itself varies little across corpora and speakers. While age and gender also modulate speech rate, their effects are much smaller in magnitude. These findings suggest utterance length effects may be conditioned by articulatory and perceptual constraints, and that social influences on speech rate should be interpreted in the broader context of how speech rate variation is structured.
We are developing a social robot to work alongside human support workers who help new arrivals in a country to navigate the necessary bureaucratic processes in that country. The ultimate goal is to develop a robot that can support refugees and asylum seekers in the UK. As a first step, we are targeting a less vulnerable population with similar support needs: international students in the University of Glasgow. As the target users are in a new country and may be in a state of stress when they seek support, forcing them to communicate in a foreign language will only fuel their anxiety, so a crucial aspect of the robot design is that it should speak the users' native language if at all possible. We provide a technical description of the robot hardware and software, and describe the user study that will shortly be carried out. At the end, we explain how we are engaging with refugee support organisations to extend the robot into one that can also support refugees and asylum seekers.
Deploying a social robot in the real world means that it must interact with speakers from diverse backgrounds, who in turn are likely to show substantial accent and dialect variation. Linguistic variation in social context has been well studied in human-human interaction; however, the influence of these factors on human interactions with digital agents, especially embodied agents such as robots, has received less attention. Here we present an ongoing project where the goal is to develop a social robot that is suitable for deployment in ethnically-diverse areas with distinctive regional accents. To help in developing this robot, we carried out an online survey of Scottish adults to understand their expectations for conversational interaction with a robot. The results confirm that social factors constraining accent and dialect are likely to be significant issues for human-robot interaction in this context, and so must be taken into account in the design of the system at all levels.