Successful speech communication in multi-talker scenarios requires a skillful combination of sustained attention and rapid attention switching. While the neurophysiology literature offers detailed insights into the neural underpinnings of sustained attention, there remains considerable uncertainty on how attention switching takes place. In this study, using EEG recordings from normal-hearing adults in an immersive multi-talker environment, we measured the neural encoding of two competing speech streams amid background babble. Participants were cued to switch attention between streams every 15-30 s. Neural tracking was assessed via Temporal Response Functions (TRF), confirming reliable decoding of attentional focus. Our results indicate asymmetric disengagement and engagement processes during attention switches, where the neural tracking of the new target stream emerges before disengaging from the previous target, revealing a transient simultaneous encoding of two speech streams. That transition was closely mirrored by a reduction in EEG alpha power, informing on the cognitive effort during different phases of the attention switch. We then isolated cortical activity reflecting lexical prediction mechanisms to determine how lexical context is updated after an attention switch, comparing four context-accumulation strategies that were constructed using Large Language Models. Our findings elucidate both the temporal and contextual mechanisms underlying auditory attention shifts, pointing to the possibility that listeners carry out a reset in lexical context after switching attention. By focusing on dynamic attentional reallocation, this study offers insights into the brain's capacity for flexible speech processing in complex listening environments.
Understanding the neural basis of speech communication is essential for uncovering how sounds are translated into meaning, how that changes with development, ageing and speech-related deficits, as well as contributing to brain-computer interfaces research. While traditional neurophysiological studies have relied on simplified, controlled paradigms, recent advances have shifted the field towards more ecologically valid approaches. Here, we describe the evolving landscape of experimental designs in speech neurophysiology, from discrete to continuous stimuli and from socially isolated listening to dynamic, multiagent communication. Realistic paradigms in that space challenge conventional methods, offering richer insights into neural encoding, functional brain mapping and neural entrainment. At the same time, they introduce significant analytical and technical complexities, particularly when incorporating social interaction. By synthesising findings across studies, we highlight how these ecologically valid speech paradigms have been contributing to refining theories of language processing and open new avenues for research. In doing so, this review critically evaluates of whether the move towards realism in speech neurophysiology represents a technological trend or a transformative leap in understanding the neural underpinnings of speech communication.
Encoding models enable measurement of how our brains represent sensory inputs using electro-and magneto-encephalography (MEEG). Evaluating how closely encoding models reflect the underlying brain functions is a crucial premise for model interpretation and hypothesis testing. However, the ground-truth neural activity is unknown, preventing model evaluation with respect to the target neural signal. Existing evaluation metrics must therefore relate model's predictions to noisy MEEG measurements, where most variance is stimulus-unrelated. Here, I introduce an evaluation framework where model predictions are compared to a ground-truth approximation, obtained by aligning MEEG signals with predictions using canonical correlation analysis and via participant averaging. The resulting metric (CPA-PA) yields single-participant evaluations outperforming conventional scores by 300-1000
Abstract Speech comprehension involves the inference of abstract information from continuous acoustic signals. Prior work suggests that electrophysiological activity is synchronized with abstract linguistic structures (phrases and sentences) during the processing of isochronous syllable sequences. It is yet unclear whether this prior evidence generalizes to natural speech comprehension, which requires the flexible processing of continuous speech, where syllables and other types of linguistic units are anisochronous. Our magnetoencephalography experiment investigated neural synchronization to acoustic (syllables) and abstract units (phrases and sentences) using continuous speech ranging from artificial isochronous to more natural anisochronous. We find that neural synchronization to phrases and sentences, but not syllables, is resilient to naturalistic anisochrony. This suggests that linguistic structure processing reflects endogenous inferences that are fundamentally distinct from the exogenous processing of syllables driven by speech acoustics. Lateralization and linear regression results extend this functional dissociation as hemispheric asymmetry: stimulus-independent leftward lateralization for linguistic structure processing but stimulus-driven rightward lateralization (or bilaterality) for both syllable and acoustic processing. Our findings provide a more realistic characterization of the flexible neural mechanisms supporting the efficient comprehension of natural speech.
Abstract The first year of life is considered a sensitive period for the acquisition of phonetic categories, a hallmark of successful native language specialization. The extent to which this process depends on early auditory experience and intrinsic biological constraints remains unresolved. We measured neural encoding of continuous natural speech in hearing children (HC) and cochlear implant (CI) users with congenital or acquired deafness, contrasting children with and without access to auditory input in the first year of life. Speech encoding was present across all groups, but its specificity depended on early input: auditory phonetic features were encoded only in children exposed to speech within the first year, whereas visually discriminable phonetic features were encoded regardless of auditory deprivation. These findings show that early sensory input gates phonetic attunement; this constraint is not limited to or grounded in audition but instead reveals a sensitive period that is modality-flexible in mechanism and experience-dependent in expression.
Human speech is inherently social. Yet our understanding of the neural substrates underlying continuous speech perception relies largely on neural responses to monologues, leaving substantial uncertainty about how social interactions shape the neural encoding of speech. Here, we bridge this gap by studying how EEG responses to speech change when the input includes a social element. Specifically, we compared the neural encoding of synthesised undirected monologues, directed monologues, and dialogues in Experiment 1. In Experiment 2, we extended this by using podcasts, addressing the additional challenges of real speech dialogue, such as dysfluencies. Using temporal response function analyses, we show that the presence of a social component strengthens envelope tracking — despite identical acoustic properties — indicating heightened listener engagement. Neural responses to synthesised speech showed a strong correlation with those for real speech podcasts, with a stronger alignment emerging for more socially-relevant speech material. In addition, we demonstrate that robust neural indices of sound and lexical-level processing can be derived using real podcast recordings despite the presence of dysfluencies. Finally, we discuss the importance of dysfluency in social speech experiments, presenting a simulation quantifying the potential impact of dysfluency on lexical level analyses. Together, these findings advance our understanding of continuous speech neurophysiology by highlighting the impact of social elements in shaping auditory neural processing in a controlled manner and providing a framework for future investigation and analysis of social speech listening and speech interaction. Significance Statement Human speech is rarely produced or processed in a social vacuum. Yet, our understanding of continuous speech neurophysiology mostly comes from experiments involving speech monologues. This study reveals how social context modulates the neural encoding of speech. We directly contrast neural signals recorded when participants listened to monologues and dialogues, using controlled material from speech synthesis and real podcast recordings. We found that the social element amplifies the neural encoding of speech features, reflecting greater engagement. We also show strong correlation between synthetic and real podcast neural responses, scaling with social relevance. Finally, we demonstrate that lexical processing can be measured robustly even amid natural dysfluencies. These insights advance our understanding of speech neurophysiology, informing future research on social speech. ### Competing Interest Statement The authors have declared no competing interest. Taighde Éireann Research Ireland, 18/CRT/6224, 13/RC/2106_P2
Humans seamlessly process multi-voice music into a coherent perceptual whole. Yet the neural strategies supporting this experience remain unclear. One fundamental component of this process is the formation of melody, a core structural element of music. Previous work on monophonic listening has provided strong evidence for the neurophysiological basis of melody processing, for example indicating predictive processing as a foundational mechanism underlying melody encoding. However, considerable uncertainty remains about how melodies are formed during polyphonic music listening, as existing theories (e.g., divided attention, figure–ground model, stream integration) fail to unify the full range of empirical findings. Here, we combined behavioral measures with non-invasive electroencephalography (EEG) to probe spontaneous attentional bias and melodic expectation while participants listened to two-voice classical excerpts. Our uninstructed listening paradigm eliminated a major experimental constraint, creating a more ecologically valid setting. We found that attention bias was significantly influenced by both the high-voice superiority effect and intrinsic melodic statistics. We then employed transformer-based models to generate next-note expectation profiles and test competing theories of polyphonic perception. Drawing on our findings, we propose a weighted-integration framework in which attentional bias dynamically calibrates the degree of integration of the competing streams. In doing so, the proposed framework reconciles previous divergent accounts by showing that, even under free-listening conditions, melodies emerge through an attention-guided statistical integration mechanism.
Speakers accommodate their speech to meet the needs of their listeners, producing different speech registers. One such register is L2 Accommodation (L2A), which is the way native speakers address non-native listeners, typically characterized by features such as slow speech rate and phonetic exaggeration. Here, we investigated how register impacts the cortical encoding of speech at different levels of language integration. Specifically, we tested the hypothesis that enhanced comprehension of L2A compared with Native Directed Speech (NDS) involves more than just a slower speech rate, influencing speech processing from acoustic to semantic levels. Electroencephalography (EEG) signals were recorded from Spanish native listeners, who were learning English (L2 learners), and English native listeners (L1 listeners) as they were presented with audio-stories. Speech was presented in English in three different speech registers: L2A, NDS, and a control register (Slow-NDS) which is a slowed down version of NDS. We measured the cortical encoding of acoustic, phonological, and semantic information with a multivariate temporal response function analysis (TRF) on the EEG signals. We found that L2A promoted L2 learners' cortical encoding at all the levels of speech and language processing considered. First, L2A led to a more pronounced encoding of the speech envelope. Second, phonological encoding was more refined when listening to L2A, with phoneme perception getting closer to that of L1 listeners. Finally, L2A also enhanced the TRF-N400, a neural signature of semantic integration. Conversely, L2A impacted acoustic but not linguistic speech encoding in L1 listeners. In contrast, slow-NDS altered the cortical encoding of sound acoustics in L1 listeners but did not impact semantic or phonological encoding. Taken together, these results support our hypothesis that L2A accommodates speech processing in L2 listeners beyond what can be achieved by simply speaking slowly, impacting the cortical encoding of sound and language at different abstraction levels. In turn, this study provides objective metrics that are sensitive to the impact of register on the hierarchical encoding of speech, which could be extended to other registers and cohorts.
Enhancing speech perception in everyday noisy acoustic environments remains an outstanding challenge for hearing aids. Speech separation technology is improving rapidly, but hearing devices cannot fully exploit this advance without knowing which sound sources the user wants to hear. Even with high-quality source separation, the hearing aid must know which speech streams to enhance and which to suppress. Advances in EEG-based decoding of auditory attention raise the potential of neurosteering, in which a hearing instrument selectively enhances the sound sources that a hearing-impaired listener is focusing their attention on. Here, we present and discuss a real-time brain-computer interface system that combines a stimulus-response model based on canonical correlation analysis for real-time EEG attention decoding, coupled with a multi-microphone hardware platform enabling low-latency real-time speech separation through spatial beamforming. We provide an overview of the system and its various components, discuss prospects and limitations of the technology, and illustrate its application with case studies of listeners steering acoustic feedback of competing speech streams via real-time attention decoding. A software implementation code of the system is publicly available for further research and explorations.
Auditory neurophysiology has increasingly focused on naturalistic auditory tasks, for example involving continuous speech and music listening. Probing such neural signatures involves relating the sensory input with the corresponding neural signal, for example by utilizing system identification methodologies. The resulting neural indices have advanced our understanding of foundational neural mechanisms, such as auditory attention and prediction, as well as shedding light on language development, impairments impacting auditory communication, and ageing. One key challenge is to determine whether results from a given experiment generalize or are specific to the experimental design. Here, we re-analyze data from a variety of studies on continuous sound listening to test the sensitivity of neural signals to a key experimental design choice: the segmentation of the sound stimulus. We measured the neural tracking of the sound envelope, finding a reduced tracking through the first 20 seconds of a sound segment. The effect was measured in correspondence with the onset of a segment but was unrelated to the duration of the experiment. This phenomenon was consistent across datasets, regardless of whether the stimulus was speech or music, and independent of attention (target vs. masker speech), intelligibility (speech vs. time-reversed speech), and modality (speech vs. audio-visual speech). This result is akin to slow neural habituation effect or short-term habituation (STH) seen in traditional discrete-stimuli event-related (ERP) studies.Clinical Significance—Our findings indicate that the segmentation choice for experiments involving continuous amplitude-modulated sounds affects the envelope tracking measures. This is a very important issue to consider as it can alter the interpretation of existing and future results. Substantial changes in the neural measurements due to that design choice would be undesirable or, at least, important to consider when comparing data from different experiments.
Objective. Speech comprehension involves detecting words and interpreting their meaning according to the preceding semantic context. This process is thought to be underpinned by a predictive neural system that uses that context to anticipate upcoming words. However, previous studies relied on evaluation metrics designed for continuous univariate sound features, overlooking the discrete and sparse nature of word-level features. This mismatch has limited effect sizes and hampered progress in understanding lexical prediction mechanisms in ecologically-valid experiments.Approach. We investigate these limitations by analyzing both simulated and actual electroencephalography (EEG) signals recorded during a speech comprehension task. We then introduce two novel assessment metrics tailored to capture the neural encoding of lexical surprise, improving upon traditional evaluation approaches.Main results. The proposed metrics demonstrated effect-sizes over 140% larger than those achieved with the conventional temporal response function (TRF) evaluation. These improvements were consistent across both simulated and real EEG datasets.Significance. Our findings substantially advance methods for evaluating lexical prediction in neural data, enabling more precise measurements and deeper insights into how the brain builds predictive representations during speech comprehension. These contributions open new avenues for research into predictive coding mechanisms in naturalistic language processing.
Successful speech communication in multi-talker scenarios requires a skilful combination of sustained attention and rapid attention switching. While the neurophysiology literature offers detailed insights into the neural underpinnings of sustained attention, there remains considerable uncertainty on how attention switching takes place. In this study, using EEG recordings from normal-hearing adults in an immersive multi-talker environment, we measured the neural encoding of two competing speech streams amid background babble. Participants were cued to switch attention between streams every 15–30 seconds. Neural tracking was assessed via Temporal Response Functions (TRF), confirming reliable decoding of attentional focus. Our results indicate asymmetric disengagement and engagement processes during attention switches, where the neural tracking of the new target stream emerges before disengaging from the previous target, revealing a transient simultaneous encoding of two speech streams. That transition was closely mirrored by a reduction in EEG alpha power, informing on the cognitive effort during different phases of the attention switch. We then isolated cortical activity reflecting lexical prediction mechanisms to determine how lexical context is updated after an attention switch, comparing four numerical hypotheses that were constructed using Large Language Models. Our findings elucidate both the temporal and contextual mechanisms underlying auditory attention shifts, pointing to the possibility that listeners carry out a reset in lexical context after switching attention. By focusing on dynamic attentional reallocation, this study offers insights into the brain’s capacity for flexible speech processing in complex listening environments. ### Competing Interest Statement The authors have declared no competing interest. William Demant Fonden, https://ror.org/02x3xhs38, 21-0628, 22-0552 Taighde Éireann – Research Ireland, 18/CRT/6223
Cortical signals have been shown to track acoustic and linguistic properties of continuous speech. This phenomenon has been measured in both children and adults, reflecting speech understanding by adults as well as cognitive functions such as attention and prediction. Furthermore, atypical low-frequency cortical tracking of speech is found in children with phonological difficulties (developmental dyslexia). Accordingly, low-frequency cortical signals may play a critical role in language acquisition. A recent investigation with infants Attaheri et al., 2022 (1) probed cortical tracking mechanisms at the ages of 4, 7 and 11 months as participants listened to sung speech. Results from temporal response function (TRF), phase-amplitude coupling (PAC) and dynamic theta-delta power (PSD) analyses indicated speech envelope tracking and stimulus-related power (PSD) for delta and theta neural signals. Furthermore, delta- and theta-driven PAC was found at all ages, with theta phases displaying stronger PAC with high-frequency amplitudes than delta. The present study tests whether these previous findings replicate in the second half of the full cohort of infants (N = 122) who were participating in this longitudinal study (first half: N=61, (1); second half: N=61). In addition to demonstrating good replication, we investigate whether cortical tracking in the first year of life predicts later language acquisition for the full cohort (122 infants recruited, 113 retained) using both infant-led and parent-estimated measures and multivariate and univariate analyses. Increased delta cortical tracking in the univariate analyses, increased ~2Hz PSD power and stronger theta-gamma PAC in both multivariate and univariate analyses were related to better language outcomes using both infant-led and parent-estimated measures. By contrast, increased ~4Hz PSD power in the multi-variate analyses, increased delta-beta PAC and a higher theta/delta power ratio in the multi-variate analyses were related to worse language outcomes. The data are interpreted within a "Temporal Sampling" framework for developmental language trajectories.
Purpose: Developmental language disorder (DLD) is typically characterized by a core grammatical impairment, yet spoken word production can also exhibit atypical phonology, for example the omission of unstressed syllables. This may indicate a role for the sensory/neural processing of syllable stress patterns in the aetiology of DLD. Here we explore this connection by investigating the accuracy of the production of multisyllabic words from a speech rhythm/syllable stress perspective.Method: We adapted a computerized speech copying task originally designed for children with dyslexia, in which participants copy aloud familiar targets like “alligator”. Fifty-seven children with and without DLD were tested with a 30-item battery comprising equal numbers of 2-syllable, 3-syllable and 4-syllable words. The children with DLD (N=20) were compared with both age-matched typically-developing control children (AMC, N=21) and younger typically-developing control children with less mature language skills (younger language controls, YLC, N=16). Similarity of the child’s productions to the target in terms of the speech amplitude envelope (AE) and pitch contour was computed using two similarity metrics, correlation and mutual information. Both the speech AE and the pitch contour contain important information about stress patterns and intonational information.Results: Children with DLD were significantly less accurate than AMC at copying the multi-syllabic targets, for both the AE and the pitch contour. The opportunity to repeat the targets had no impact on performance, for any group. Word length effects were similar across groups. Conclusion: The spoken production of multisyllabic words by children with DLD is atypical regarding both the AE and the pitch contour. This is consistent with a theoretical explanation of DLD based on impaired sensory/neural processing of speech rhythm.
Neurophysiology research has demonstrated that it is possible and valuable to investigate sensory processing in scenarios involving continuous sensory streams, such as speech and music. Over the past 10 years or so, novel analytic frameworks combined with the growing participation in data sharing has led to a surge of publicly available datasets involving continuous sensory experiments. However, open science efforts in this domain of research remain scattered, lacking a cohesive set of guidelines. This paper presents an end-to-end open science framework for the storage, analysis, sharing, and re-analysis of neural data recorded during continuous sensory experiments. We propose a data structure that builds on existing custom structures (Continuous-event Neural Data or CND), providing precise naming conventions and data types, as well as a workflow for storing and loading data in the general-purpose BIDS structure. The framework has been designed to interface with existing EEG/MEG analysis toolboxes, such as Eelbrain, NAPLib, MNE, and mTRF-Toolbox. We present guidelines by taking both the user view (rapidly re-analyse existing data) and the experimenter view (store, analyse, and share), making the process straightforward and accessible. Additionally, we introduce a web-based data browser that enables the effortless replication of published results and data re-analysis.
Purpose: Developmental language disorder (DLD) is a multifaceted disorder. Recently, interest has grown in prosodic aspects of DLD, but most investigations of possible prosodic causes focus on speech perception tasks. Here, we focus on speech production from a speech amplitude envelope (AE) perspective. Perceptual studies have indicated a role for difficulties in AE processing in DLD related to sensory/neural processing of prosody. We explore possible matching AE difficulties in production. Method: Fifty-seven children with and without DLD completed a computerized imitation task, copying aloud 30 familiar targets such as “alligator.” Children with DLD ( n = 20) were compared with typically developing children (age-matched controls [AMC], n = 21) and younger language controls (YLC, n = 16). Similarity of the child's productions to the target in terms of the continuous AE and pitch contour was computed using two similarity metrics, correlation, and mutual information. Both the speech AE and the pitch contour contain important information about stress patterning and intonational information over time. Results: Children with DLD showed significantly reduced imitation for both the AE and pitch contour metrics compared to AMC children. The opportunity to repeat the targets had no impact on performance for any group. Word length effects were similar across groups. Conclusions: The spoken production of multisyllabic words by children with DLD is atypical regarding both the AE and the pitch contour. This is consistent with a theoretical explanation of DLD based on impaired sensory/neural processing of low-frequency (slow) amplitude and frequency modulations, as predicted by the temporal sampling theory. Supplemental Material: https://doi.org/10.23641/asha.27165690
Slow cortical oscillations play a crucial role in processing the speech amplitude envelope, which is perceived atypically by children with developmental dyslexia. Here we use electroencephalography (EEG) recorded during natural speech listening to identify neural processing patterns involving slow oscillations that may characterize children with dyslexia. In a story listening paradigm, we find that atypical power dynamics and phase-amplitude coupling between delta and theta oscillations characterize dyslexic versus other child control groups (typically-developing controls, other language disorder controls). We further isolate EEG common spatial patterns (CSP) during speech listening across delta and theta oscillations that identify dyslexic children. A linear classifier using four delta-band CSP variables predicted dyslexia status (0.77 AUC). Crucially, these spatial patterns also identified children with dyslexia when applied to EEG measured during a rhythmic syllable processing task. This transfer effect (i.e., the ability to use neural features derived from a story listening task as input features to a classifier based on a rhythmic syllable task) is consistent with a core developmental deficit in neural processing of speech rhythm. The findings are suggestive of distinct atypical neurocognitive speech encoding mechanisms underlying dyslexia, which could be targeted by novel interventions.
Background : IDyOM (Information Dynamics of Music) is the statistical model of music the most used in the community of neuroscience of music. It has been shown to allow for significant correlations with EEG (Marion, 2021), ECoG (Di Liberto, 2020) and fMRI (Cheung, 2019) recordings of human music listening. The language used for IDyOM-Lisp- is not very familiar to the neuroscience community and makes this model hard to use and more importantly to modify. New method : IDyOMpy is anew Python re-implementation and extension of IDyOM. This new model allows for computing the information content and entropy for each melody note after training on a corpus of melodies. In addition to those features, two new features are presented: probability estimation of silences and enculturation modeling. Results : We first describe the mathematical details of the implementation. We extensively compare the two models and show that they generate very similar outputs. We also support the validity of IDyOMpy by using its output to replicate previous EEG and behavioral results that relied on the original Lisp version (Gold, 2019; Di Liberto, 2020; Marion, 2021). Finally, it reproduced the computation of cultural distances between two different datasets as described in previous studies (Pearce, 2018). Comparison with existing methods and Conclusions : Our model replicates the previous behaviors of IDyOM in a modern and easy-to-use language-Python. In addition, more features are presented. We deeply think this new version will be of great use to the community of neuroscience of music.
Hearing impairment alters the sound input received by the human auditory system, reducing speech comprehension in noisy multi-talker auditory scenes. Despite such difficulties, neural signals were shown to encode the attended speech envelope more reliably than the envelope of ignored sounds, reflecting the intention of listeners with hearing impairment (HI). This result raises an important question: What speech-processing stage could reflect the difficulty in attentional selection, if not envelope tracking? Here, we use scalp electroencephalography (EEG) to test the hypothesis that the neural encoding of phonological information (i.e., phonetic boundaries and phonological categories) is affected by HI. In a cocktail-party scenario, such phonological difficulty might be reflected in an overrepresentation of phonological information for both attended and ignored speech sounds, with detrimental effects on the ability to effectively focus on the speaker of interest. To investigate this question, we carried out a re-analysis of an existing dataset where EEG signals were recorded as participants with HI, fitted with hearing aids, attended to one speaker (target) while ignoring a competing speaker (masker) and spatialised multi-talker background noise. Multivariate temporal response function (TRF) analyses indicated a stronger phonological information encoding for target than masker speech streams. Follow-up analyses aimed at disentangling the encoding of phonological categories and phonetic boundaries (phoneme onsets) revealed that neural signals encoded the phoneme onsets for both target and masker streams, in contrast with previously published findings with normal hearing (NH) participants and in line with our hypothesis that speech comprehension difficulties emerge due to a robust phonological encoding of both target and masker. Finally, the neural encoding of phoneme-onsets was stronger for the masker speech, pointing to a possible neural basis for the higher distractibility experienced by individuals with HI.