This study investigates how human listeners perceive and locate Northern English accents, with a focus on the linguistic features that attract attention during accent recognition. Although sociolinguistic research often centers on specific phonetic variables, it is unclear whether these align with the cues non-linguists naturally notice when identifying regional varieties. To address this, we used a real-time linguistic attention method to examine which features listeners attended to as they attempted to identify the origin of speakers from five Northern English cities: Leeds, Liverpool, Manchester, Newcastle, and Sheffield. Crucially, listeners were not directed to focus on any particular features. Data from 98 participants revealed substantial variation in recognition accuracy. While Liverpool and Newcastle voices were frequently identified correctly, accents from Leeds, Manchester, and Sheffield proved more difficult to place. Participants consistently attended to salient features such as the bath and strut vowels, along with more locally specific features like fricated /k/ in Liverpool and glottalized /t/ in Newcastle. Accents with multiple distinct cues were more reliably identified, suggesting that a cluster of reinforcing features enhances perceptual success. The findings show that listener attention is guided by both cognitive and social salience, and that listeners rely on a broader and more socially grounded set of features than those often prioritized in computational models of accent classification. By revealing which features listeners attend to spontaneously, this study contributes to understanding the cognitive processes underpinning regional accent perception and offers new insights into the interplay between linguistic variation, salience, and social indexing.
This article introduces Salient Language in Context (SLIC), a web application designed for collecting real-time attention data from listeners in response to audio samples. SLIC enables researchers to capture timestamped attention points as listeners engage with spoken language, allowing for an innovative approach to studying real-time perception of speech. The application facilitates the collection of both quantitative and qualitative data, providing researchers with not only precise reaction timing but also listener justifications for their responses. SLIC addresses long-standing challenges in sociolinguistics and perceptual dialectology by enabling researchers to examine which linguistic features capture attention, when they do so, and how listeners interpret them in context. The technical aspects of SLIC are discussed in the article and include its secure online infrastructure and integration with scripts written for Praat, which allow for efficient data visualization and interpretation. A case study using SLIC demonstrates its effectiveness in examining listener perceptions of segmental features in Liverpool English, revealing distinct patterns of attention and commentary strategies. By providing a robust method for investigating real-time language perception, SLIC has a broad range of potential applications.
Previous research has provided strong evidence that speech patterns can help to distinguish between people with early stage neurodegenerative disorders (ND) and healthy controls. This study examined speech patterns in responses to questions asked by an intelligent virtual agent (IVA): a talking head on a computer which asks pre-recorded questions. The study investigated whether measures of response length, speech rate and pausing in responses to questions asked by an IVA help to distinguish between healthy control participants and people diagnosed with Mild Cognitive Impairment (MCI) or Alzheimer's disease (AD). The study also considered whether those measures can further help to distinguish between people with MCI, people with AD, and healthy control participants (HC). There were 38 people with ND (31 people with MCI, 7 people with AD) and 26 HC. All interactions took place in English. People with MCI spoke fewer words compared to HC, and people with AD and people with MCI spoke for less time than HC. People with AD spoke at a slower rate than people with MCI and HC. There were significant differences across all three groups for the proportion of time spent pausing and the average pause duration: silent pauses make up the greatest proportion of responses from people with AD, who also have the longest average silent pause duration, followed by people with MCI then HC. Therefore, the study demonstrates the potential of an IVA as a method for collecting data showing patterns which can help to distinguish between diagnostic groups.
The diagnosis of Mild Cognitive Impairment (MCI) characterises patients at risk of dementia and may provide an opportunity for disease-modifying interventions. Identifying persons with MCI (PwMCI) from adults of a similar age without cognitive complaints is a significant challenge. The main aims of this study were to determine whether generic speech differences were evident between PwMCI and healthy controls (HC), whether such differences were identifiable in responses to recent or remote memory questions, and to determine which speech variables showed the clearest between-group differences. This study analysed recordings of 8 PwMCI (5 females, 3 males) and 14 HC of a similar age (8 females, 6 males). Participants were recorded interacting with an intelligent virtual agent: a computer-generated talking head on a computer screen which asks pre-recorded questions when prompted by the interviewee through pressing the next key on a computer keyboard. Responses to recent and remote memory questions were analysed. Mann-Whitney U tests were used to test for statistically significant differences between PwMCI and HC on each of 12 speech variables, relating to temporal characteristics, number of words produced and pitch. It was found that compared to HC, PwMCI produce speech for less time and in shorter chunks, they pause more often and for longer, take longer to begin speaking and produce fewer words in their answers. It was also found that the PwMCI and HC were more alike when responding to remote memory questions than when responding to recent memory questions. These findings show great promise and suggest that detailed speech analysis can make an important contribution to diagnostic and stratification systems in patients with memory complaints.
A body of research has shown that there are linguistic differences in the way people with epilepsy talk about their seizures when compared to those with non-epileptic seizures. We extend this line of research by presenting the results of a phonetic analysis comparing speech samples from people with a confirmed diagnosis of epilepsy (7 patients), to those with a confirmed diagnosis of non-epileptic seizures (8 patients). Variables considered include features of pitch, intensity, duration and pausing in their responses to questions from a neurologist during medical history-taking. We find only limited evidence of differences between the two diagnostic groups (epilepsy vs. non-epileptic seizures). We discuss possible reasons for this lack of evidence.
This study investigates prototypically "turn-final" pitch features (fall-to-low) at points of possible turn-completion where the same speaker continues. It is shown that points of possible turn-completion accompanied by fall-to-low and followed by same-speaker continuation only rarely engender incoming talk. It is shown that such points are frequently accompanied by nonpitch talk-projecting phonetic features and that the presence of these features may constrain the nature of any incoming talk. The results of the study should serve as caution to researchers with regard to an overemphasis on intonation when describing and analyzing talk-in-interaction. Data are from audio recordings of American English telephone calls.
Visual representations of acoustic data are becoming more common in Conversation Analysis (CA) and Interactional Linguistics (IL) research. This article provides a survey of visual representations of acoustic data in the journal Research on Language and Social Interaction. Shortcomings in their preparation and use are identified and discussed. Comparisons are made with visual representations prepared by expert phoneticians. Suggestions are made as to how visual representations could be prepared and used more effectively to support CA/IL researchers' claims and to allow readers to independently verify them.
An auspicious but unexplored environment for studying phonetic variation in naturalistic interaction is where two or more participants say the same thing at the same time. Working with a core dataset built from the multimodal Augmented Multi-party Interaction corpus, the principles of conversation analysis were followed to analyze the sequential organization of the talk and to explain the phonetic variation observed. Acoustic divergence and equivalence between simultaneous responses are described. Phonetic features discussed include duration and timing, pitch, loudness, and phonation type. The interactional factors that explain the acoustic divergences are established through turn-by-turn analysis and consideration of gaze direction and other visible features. It is argued that any research on phonetic variation in naturalistic talk that disregards the local organization of interaction will always be incomplete.
Can very young children deploy laughter interactionally? Using data from video recordings of 52 interactions between six mothers and their young children, this article examines one particular kind of sequence in which interactionally ordered child laughter occurs. In that sequence, the young child commits some kind of potential transgression (e.g., breaking wind, standing on objects on the floor, or playing in a proscribed location). The child's mother then draws attention to the potential transgression in some way (e.g., by admonishing the child, requesting a change to the child's behavior, issuing a particular kind of child-directed gaze), thus treating the child's action as constituting a transgression. At some point following the potential transgression, the child laughs. What is shown is that even young children can fit their laughter to the ongoing interactional sequence. It is argued that the child's laughter provides for a display of affiliation from the mother.
The analysis of language use in real-world contexts poses particular methodological challenges. We codify responses to these challenges as a series of methodological imperatives. To demonstrate the relevance of these imperatives to clinical investigation, we present analyses of single episodes of interaction where one participant has a speech and/or language impairment: atypical prosody, echolalia and dysarthria. We demonstrate there is considerable heuristic and analytic value in taking this approach to analysing the organization of interaction involving individuals with a speech and/or language impairment.
Investigations into the management of turn-taking have typically focussed on pitch and other prosodic phenomena, particularly pitch-accents. Here, non-pitch phonetic features and their role in turn-taking are described. Through sustained phonetic and interactional analysis of a naturally occurring, 12-minute long telephone call between two adult speakers of British English, sets of talk-projecting and turn-projecting features are identified. Talk-projecting features include the avoidance of durational lengthening, articulatory anticipation, continuation of voicing, the production of talk in maximally close proximity to a preceding point of possible turn-completion, and the reduction of consonants and vowels. Turn-projecting features include the converse of each of the talk-projecting features, and two other distinct features: release of plosives at the point of possible turn-completion, and the production of audible outbreaths. We show that features of articulatory and phonatory quality and duration are relevant factors in the design and treatment of talk as talk- or turn-projective.
eprints@whiterose.ac.uk https://eprints.whiterose.ac.uk/ Reuse Unless indicated otherwise, fulltext items are protected by copyright with all rights reserved. The copyright exception in section 29 of the Copyright, Designs and Patents Act 1988 allows the making of a single copy solely for the purpose of non-commercial research or private study within the limits of fair dealing. The publisher or other rights-holder may allow further reproduction and re-use of this version refer to the White Rose Research Online record for this item. Where records identify the publisher as the copyright holder, users can verify any specific terms of use on the publisher’s website.
The empirical focus of this paper is a conversational turn-taking phenomenon in which conjunctions produced immediately after a point of possible syntactic and pragmatic completion are treated by co-participants as points of possible completion and transition relevance. The data for this study are audio-video recordings of 5 unscripted face-to-face interactions involving native speakers of US English, yielding 28 'trail-off' conjunctions. Detailed sequential analysis of talk is combined with analysis of visible features (including gaze, posture, gesture and involvement with material objects) and technical phonetic analysis. A range of phonetic and visible features are shown to regularly co-occur in the production of 'trail-off' conjunctions. These features distinguish them from other conjunctions followed by the cessation of talk.
eprints@whiterose.ac.uk https://eprints.whiterose.ac.uk/ Reuse Unless indicated otherwise, fulltext items are protected by copyright with all rights reserved. The copyright exception in section 29 of the Copyright, Designs and Patents Act 1988 allows the making of a single copy solely for the purpose of non-commercial research or private study within the limits of fair dealing. The publisher or other rights-holder may allow further reproduction and re-use of this version refer to the White Rose Research Online record for this item. Where records identify the publisher as the copyright holder, users can verify any specific terms of use on the publisher’s website.
There is a need to get to grips with the phonetic design of talk in its totality and without a separation of prosodic and non-prosodic aspects. Features of duration, phonation and articulation are all shown to be systematic features of rush-throughs, and bound up with the turn-holding function of the practice. Data are drawn from audio and video recordings made in a range of interactional settings, all involving speakers of English from the UK or the US. The paper concludes by reviewing some of the reasons why this holistic approach is desirable, namely: empirical findings, the parametric nature of speech, and a commitment to a mode of enquiry which takes seriously observable details of all kinds.
Linguists, and other analysts of discourse, regularly make appeal to affectual states in determining the meaning of utterances. We examine two kinds of sequence that occur in everyday conversation. The first involves one participant making an explicit lexical formulation of a co-participant's, affecutal state (e. g., 'you sound happy' 'don't sound so depressed'). The second involves responses to 'positive informings' and 'negative informings'. Through consideration of sequential organization, participant orientation, and phonetic detail, we suggest that the attribution of analytic categories affect is problematic. We argue that phonetic characteristics which might be thought to he associated with affect may better be accounted for with reference to the management of particular sequential-interactional tasks. The findings that stance does not inhere in any single turn at talk or any single linguistic aspect leads us to suggest that future investigations into stance and affect will need to pay attention simultaneously, to matters of both linguistic-phonetic and sequential organization.
In everyday English conversation, talk can be produced such that it is simultaneously a grammatical ending of what precedes it, and a beginning of what follows (e.g. "that's what I'd like to have is a fresh one"). A range of features of phonetic design (including pitch, loudness, duration, and articulatory characteristics) are shown to be deployed in systematic ways in order to handle the dual tasks of avoiding the signaling of transition relevance at the end of the pivot, and marking out the fittedness of the pivot to both what precedes and what follows. Turns built with pivots are found to be most often engaged in assessing, enquiring, or reporting, though their more general application as a practice for the continuation of a turn past a point of possible syntactic and pragmatic completion is emphasized.
Repetition poses certain problems for pragmatics, as evidenced by Sperber and Wilson's claim that "the effects of repetition on utterance interpretation are by no means constant". This is particularly apposite when we examine repetitions produced in naturally occurring talk. As part of an ongoing study of how phonetics relates to the dynamic evolution of meaning within the sequential organisation of talk-in-interaction, we present a detailed phonetic and pragmatic analysis of a particular kind of self-repetition.The practice of repetition we are concerned with exhibits a range of forms: "have another go tomorrow... have another go tomorrow", "it might do... it might do", "it's a shame.... it's a shame". The approach we adopt emphasises the necessity of exploring participants' displayed understandings of pragmatic inferences and attempts not to prejudge the relevance of phonetic (prosodic) parameters. The analysis reveals that speakers draw on a range of phonetic features, including tempo and loudness as well as pitch, in designing these repetitions. The pragmatic function of repetitions designed in this way is to close sequences of talk.Our findings raise a number of theoretical and methodological issues surrounding the prosodypragmatics interface and participants' understanding of naturally occurring discourse. (c) 2006 Elsevier B.V. All rights reserved.
We describe and exemplify a methodology for providing an integratedaccount of the communicative function of parametric phonetic detail and its rela-tionshipwith interactional organization. We exemplify our analytic approach bydocumenting two different phonetic designs of stand-alone ‘so’ in a corpus ofrecorded American English telephone conversations. These two designs - whichencompass particular loudness, pitch and laryngeal characteristics - correlate withdifferent communicative functions and have different consequences for the inter-actional-sequential organization of the talk. We argue that if phonology is to betruly concerned with function and linguistic contrast, we need to induce thosefunctions and domains of contrast from a thoroughgoing phonetic and sequentialanalysis of talk-in-interaction.
This report is based on phonetic and interactional analysis of a collection of increments drawn from audio recordings of British and North American talk-in-interaction.An increment is a grammatically fitted continuation of a turn at talk following the reaching of a point of possible syntactic, pragmatic, and prosodic completion.Parametric phonetic analysis reveals that a range of phonetic parameters (including pitch, loudness, rate of articulation, and articulatory characteristics) mark out an increment as a continuation of its host.Interactional analysis reveals that increments deal with a range of interactional exigencies including, but not limited to, possible problems of understanding and alignment arising from the host turn.