Large language models (LLMs) capture long-range contextual structure in natural language and have recently been shown to align with the human brain's contextualized linguistic encoding. This makes them a promising computational probe for studying how context-dependent linguistic information is represented during natural speech perception. Speech perception often occurs in multi-talker environments, where attention must dynamically select among competing streams, yet how contextual information from attended and unattended speech is neurally encoded remains underexplored. Here, we investigate how auditory attention modulates neural tracking of context-dependent linguistic representations using electrocorticography (ECoG) and stereoelectroencephalography (sEEG) recordings from three epilepsy patients engaged in a two-conversation "cocktail party" paradigm. To model neural responses to attended and unattended speech streams, we used contextual word embeddings generated by large language models. We find that LLM-derived features reliably predict brain activity for the attended stream and that contextual information from the unattended stream also contributes to neural prediction. Importantly, these contributions extend beyond low-level acoustic features and shallow syntactic information, and depend on the surrounding linguistic context. Moreover, neural tracking of the unattended stream reflects shorter-range contextual integration than that of the attended stream. Together, these findings indicate that neural responses to speech reflect context-dependent linguistic representations from multiple concurrent speech streams, with attention modulating the depth and timescale of contextual integration. Our results highlight the utility of LLMs for probing higher-level linguistic representations in complex, naturalistic listening environments.
Abstract Objective In individuals with drug-resistant epilepsy, accurately identifying the brain regions where seizures originate is a critical prerequisite to guide surgical treatment and achieve seizure freedom. To accomplish this, intracranial EEG is considered the gold standard, providing the spatiotemporal high-resolution data necessary to pinpoint epileptogenic activity. However, this precision is achieved through an invasive procedure with significant patient burden, which is fundamentally limited by the electrode placement and spatial coverage. Methods In this study, we investigated the potential utility of preoperative resting-state fMRI to non-invasively map alterations in brain dynamics at the whole brain level. Region-wise brain dynamics were quantified with complementary measures of local autocorrelation decay rates. We assessed the capacity of these derived features to effectively identify intracranial EEG confirmed seizure onset zones in 18 individuals with drug-resistant medial temporal lobe epilepsy. Overall, the study cohort contained 3867 implanted electrodes of which 159 classified as seizure onset zones by two independent board-certified epileptologists. Results Overall, our findings reveal more constrained temporal dynamics for brain regions associated with seizure onsets compared to non-seizure onset zones. Individual-level prediction showed a performance better than chance in 15 of the 18 patients. The overall predictive performance across all patients yielded a median AUC of 0.81, a median true positive rate of 0.75, and a median true negative rate of 0.83. Furthermore, in a subset of 13 patients, those with negative seizure outcomes showed higher probabilities of seizure onset zone predictions outside the resection area compared to those with good outcomes. Significance Overall, our findings suggest that altered temporal dynamics derived from preoperative resting-state fMRI represent a promising non-invasive approach for delineating epileptogenic tissue, potentially informing intervention strategies and guiding electrode placement.
How does the human brain detect and respond to disruptions in breathing? While animal studies have advanced our understanding of respiratory control, breathing distress in humans remains difficult to treat. It often arises not only from pulmonary lesions or brainstem dysfunction but also from how higher brain regions interpret breathing signals shaped by emotion and experience. We recorded intracranial cortical activity in neurosurgical patients during an interoceptive task involving transient breathing challenges. Conscious detection of these disruptions was predicted by early responses in the anterior insula, which routed signals to orbitofrontal and premotor cortices for appraisal and compensation. These cortical regions preferentially encoded inspiratory effort or airflow, revealing signal-specific processing that echoes functional segregation in brainstem centers. The present findings identify a dynamic insular-frontal circuit for sensing and adapting to respiratory challenges, offering insight into the neural basis of breathing awareness and its disruption in disease.
Millions of people worldwide are living with movement and sensory impairments owing to spinal cord injury, stroke and other neurological conditions. Here we report a double neural bypass (DNB), a hybrid neuroprosthetic system designed to restore both immediate and lasting gains in movement and sensation after a severe, complete spinal cord injury. The DNB links an intracortical brain-computer interface with targeted and patterned neuromodulation of the spinal cord and cortex. This allows brain signals associated with movement intention to directly control the movement of the user's own hand in real time while also promoting long-term sensorimotor recovery-even after the system is turned off. The DNB system uses recurrent artificial neural networks and reinforcement learning for fine grasp control, together with patterned spinal cord stimulation and activity-informed intracortical microstimulation ('cortical mirroring') to promote neuroplasticity and durable recovery of function. In a participant with chronic C4 sensory/C5 motor complete tetraplegia, this hybrid approach enabled recovery of functional abilities including self-feeding and manipulation of delicate objects, while also producing significant and persistent improvements in elbow flexion and wrist tactile sensation. These findings demonstrate the potential of combining a sensorimotor neuroprosthesis with targeted brain and spinal neuromodulation to restore clinically relevant function in severe paralysis.
Electrical brain stimulation (EBS) has been studied as a tool to improve speech perception. Available EBS methods compromise between precision and power and rarely consider temporal dynamics. Noninvasive EBS delivers currents to wide areas but may lack regional specificity and is susceptible to current shunting through the skin. We introduce a novel minimally-invasive electrical brain stimulation approach using clinically implanted epicranial electrodes and cranial bolts, circumventing the current shunting limitations of traditional transcranial stimulation. Two types of contacts were used in five epilepsy patients (one female) over six experimental sessions undergoing intracranial monitoring: epicranial electrodes (served as reference electrodes for clinical recordings) and cranial bolts (holding fixtures for depth electrodes). Participants monaurally listened to pseudo-randomly presented sentences from matrix sentence speech-in-noise task in 3 conditions: no-stimulation, 50 ms, or 200 ms stimulation lag relative to sentence onset. EBS parameters were 100 Hz, biphasic, amplitude-balanced, square-wave pulses where the amplitude (1-6 mA) was modulated by the ongoing speech envelope. Accuracy in the task improved in 5 of 6 sessions while not reaching significance across sessions possibly limited due to the limited number of sessions and trials (p=0.087 and 0.068 for 50 ms and 200 ms delay, linear mixed-effects model). In 5/6 sessions, there was increased accuracy with 50 ms delay, and in 3 of those, there was further improvement with 200 ms delay. We successfully demonstrated the dynamic EBS patterns delivered through minimally-invasive electrodes during speech perception with 83% responder rate for improvement in speech perception with short latency (50 ms) and 50% with longer latency (200 ms) stimulation.
Understanding speech in noisy environments is difficult for many people, and current hearing aids often fail because they amplify all sounds rather than the talker of interest. Auditory attention decoding (AAD) offers a potential solution by using the listener's brain signals to identify and enhance the attended speaker, but it has been unclear whether this can provide real-time perceptual benefits. Here we used high-resolution intracranial electroencephalography in patients undergoing neurosurgical procedures to implement a closed-loop system that achieves the decoding fidelity necessary to dynamically amplify the attended talker. Across multiple experiments, the system improved speech intelligibility, reduced listening effort and was consistently preferred by subjects. It also tracked both instructed and self-initiated attention shifts. By providing direct evidence that a real-time, brain-controlled hearing system can enhance perception, this work establishes a key performance benchmark for future auditory brain-computer interfaces and advances AAD from a theoretical concept to a validated solution for personalized assistive hearing.
We sample visual scenes with short gaze fixations separated by saccades. While low-level integration is known, semantic integration of foveal vision across multiple fixations remains unclear. We hypothesized that the brain responds to changes in semantic information from one fixation to the next, and therefore postulated a neural signal associated with semantic novelty for each saccade. Novelty was measured using a deep network on foveal vision. Novelty modulated frontal and occipital fixation-related potentials in human EEG during natural viewing of full-length movies (3.4x106 saccades). Intracranial recordings in humans (9.0x104 saccades) and non-human primates (3.3x104 saccades) revealed broadband high-frequency activity modulations in ventromedial visual and frontal brain areas. This modulation was stronger for movies than static images, and frontal modulation preceded occipital modulation, suggesting top-down effects. This modulation of fixation-related activity with novelty suggests that foveal representations are integrated across saccades to construct scene representations during natural viewing in primates.
Music listening is one of the most compelling and rewarding activities humans engage in spontaneously. But what exactly catches people's attention when listening to music remains unclear. Musicologists have argued that tension-release dynamics in music constitute crucial features of the listening experience. They arise from the intertwining of different low-level and high-level features and sit at the core of music enjoyment by engaging listeners dynamically. This study aims to characterize the relationship between tension dynamics and engagement during naturalistic music listening. Using canonical correlation analysis, we decoded the music envelope from EEG and ECoG responses and found that musical tension patterns, as reported by listeners, were predictive of fluctuations in the coupling between the music and the neural response. Importantly, tension dynamics were significantly correlated with neural measures of envelope tracking even after controlling for loudness and musical expectations, confirming the specific and crucial role of musical tension in engaging listeners. These results shed new light on how musical structure gives rise to internal response modulations that, in turn, dynamically reflect musical engagement. This interplay may underlie the pervasive and emotionally rewarding nature of music. ### Competing Interest Statement The authors have declared no competing interest.
Abstract Transient high-frequency activity (HFA) in local field potentials exhibits substantial variability across trials and behavioral states, yet its functional role in sensory representation remains poorly understood. Here, we tested whether transient HFA events carry stimulus-related information through their temporal, spectral, and phase-dependent properties. Using intracranial recordings from 21 human participants performing a visual localizer task, we extracted transient HFA events and quantified their features using information-theoretic and decoding analyses. Individual event features carried only modest information about stimulus identity, and response magnitude and morphology-related properties contributed minimally to decoding performance. Instead, decoding was dominated by temporal alignment and low-frequency phase. Critically, representing transient events as distributed low-frequency phase configurations across electrodes substantially improved cross-trial decoding performance, whereas disrupting distributed phase structure eliminated this decoding advantage. These findings indicate that stimulus-related structure in transient HFA does not primarily arise from isolated local event properties, but instead emerges through distributed, phase-dependent network dynamics. More broadly, the results provide a framework for understanding how transient high-frequency neural activity contributes to sensory representations across distributed cortical networks.
Restoring sensorimotor function in individuals with severe paralysis remains an important challenge in the field of medicine. Here we show, for the first time, an individual with severe upper-limb sensorimotor impairment can interact with real-world objects through another individual using an implanted brain-computer interface (iBCI) and wireless muscle activation, while also engaging in cooperative rehabilitation. This ‘human avatar’ paradigm enabled the iBCI study participant to generate motor intent signals, decoded in real time, to control the hand of other participants wirelessly through high resolution neuromuscular electrical stimulation. Working with an able-bodied participant, the iBCI participant achieved volitional control of the other participant’s hand to discriminate objects of different compliance with 64% accuracy (p < 0.001) under blindfolded conditions. Discrimination was enabled by an S1 encoding paradigm combining multi-electrode recruitment and graded amplitude, which yielded perceptual accuracies exceeding 90% during calibration. Extending this paradigm, the iBCI participant worked with another participant with upper limb impairment and collaborated in a cooperative task: the iBCI participant’s decoded motor commands drove spinal cord and neuromuscular stimulation in the other participant’s arm, enabling bottle pouring with a 94% success rate (p < 0.0001). Qualitative interviews revealed strong subjective satisfaction, highlighting the potential of this form of cooperative rehabilitation. These findings show that cortically interfaced human avatars can restore both motor control and discriminative sensation across individuals, introducing a new framework for cooperative iBCI-based rehabilitation and perhaps even remote human interaction in the future.
Although recent work has made headway in understanding the neural temporospatial dynamics of conscious perception, much of that work has focused on visual paradigms. To determine whether there are shared mechanisms for perceptual consciousness across sensory modalities, here we test within the auditory domain. Participants completed an auditory threshold task while undergoing intracranial electroencephalography. Recordings from >2,800 grey matter electrodes were analyzed for broadband gamma power (a range which reflects local neural activity). For perceived trials, we find nearly simultaneous activity in early auditory regions, the right caudal middle frontal gyrus, and the non-auditory thalamus; followed by a wave of activity that sweeps through auditory association regions into parietal and frontal cortices. For not perceived trials, significant activity is restricted to early auditory regions. These findings show the cortical and subcortical networks involved in auditory perception are similar to those observed with vision, suggesting shared mechanisms for conscious perception.
Auditory foundation models, including auditory large language models (LLMs), process all sound inputs equally, independent of listener perception. However, human auditory perception is inherently selective: listeners focus on specific speakers while ignoring others in complex auditory scenes. Existing models do not incorporate this selectivity, limiting their ability to generate perception-aligned responses. To address this, we introduce intention-informed auditory scene understanding (II-ASU) and present Auditory Attention-Driven LLM (AAD-LLM), a prototype system that integrates brain signals to infer listener attention. AAD-LLM extends an auditory LLM by incorporating intracranial electroencephalography (iEEG) recordings to decode which speaker a listener is attending to and refine responses accordingly. The model first predicts the attended speaker from neural activity, then conditions response generation on this inferred attentional state. We evaluate AAD-LLM on speaker description, speech transcription and extraction, and question answering in multitalker scenarios, with both objective and subjective ratings showing improved alignment with listener intention. By taking a first step toward intention-aware auditory AI, this work explores a new paradigm where listener perception informs machine listening, paving the way for future listener-centered auditory systems. Demo available.
We present a method for spatially resolving the electric field potential throughout the entire volume of the human brain from electroencephalography (EEG) data. The method is not a variation of the well-known 'source reconstruction' methods, but rather a direct solution to the EEG inverse problem based on our recently developed model for brain waves that demonstrates the inadequacy of the standard 'quasi-static approximation' that has fostered the belief that such a reconstruction is not physically possible. The method retains the high temporal/frequency resolution of EEG yet has spatial resolution comparable to (or better than) functional MRI (fMRI), without its significant inherent limitations. The method is validated using simultaneous EEG/fMRI data in healthy subjects, intracranial EEG data in epilepsy patients, comparison with numerical simulations, and a direct comparison with standard state-of-the-art EEG analysis in a well-established attention paradigm. The method is then demonstrated on a very large cohort of subjects performing a standard gambling task designed to activate the brain's 'reward circuit'. The technique uses the output from standard extant EEG systems and thus has potential for immediate benefit to a broad range of important basic scientific and clinical questions concerning brain electrical activity. By offering an inexpensive and portable alternative to fMRI, it provides a realistic methodology to efficiently promote the democratization of medicine.
Sensory stimulation of the brain reverberates in its recurrent neural networks. However, current computational models of brain activity do not separate immediate sensory responses from this intrinsic dynamic. We apply a vector-autoregressive model with external input (VARX), combining the concepts of 'functional connectivity' and 'encoding models', to intracranial recordings in humans. This model captures the extrinsic effect of the stimulus and separates that from the intrinsic effect of the recurrent brain dynamic. We find that the intrinsic dynamic enhances and prolongs the neural responses to scene cuts, eye movements, and sounds. Failing to account for these extrinsic inputs leads to spurious recurrent connections that govern the intrinsic dynamic. We also find that the recurrent connectivity during rest is reduced during movie watching. The model shows that an external stimulus can reduce intrinsic noise. It also shows that sensory areas have mostly outward, whereas higher-order brain areas have mostly incoming connections. We conclude that the response to an external audiovisual stimulus can largely be attributed to the intrinsic dynamic of the brain, already observed during rest.
The human brain's ability to transform acoustic speech signals into rich linguistic representations has inspired advancements in automatic speech recognition (ASR) systems. While ASR systems now achieve human-level performance under controlled conditions, prior research on their parallels with the brain has been limited by the use of biologically implausible models, narrow feature sets, and comparisons that primarily emphasize predictability of brain activity without fully exploring shared underlying representations. Additionally, studies comparing the brain to text-based language models overlook the acoustic stages of speech processing, an essential part in transforming sound to meaning. Leveraging high-resolution intracranial recordings and a recurrent ASR model, this study bridges these gaps by uncovering a striking correspondence in the hierarchical encoding of linguistic features, from low-level acoustic signals to high-level semantic processing. Specifically, we demonstrate that neural activity in distinct regions of the auditory cortex aligns with representations in corresponding layers of the ASR model and, crucially, that both systems encode similar features at each stage of processing-from acoustic to phonetic, lexical, and semantic information. These findings suggest that both systems, despite their distinct architectures, converge on similar strategies for language processing, providing insight in the optimal computational principles underlying linguistic representation and the shared constraints shaping human and artificial speech processing.
Music perception requires integrating individual notes into their broader musical context, yet how musical expertise shapes the neural encoding of this information across the auditory hierarchy remains unclear. Here, we address this by using the hierarchical representations from a generative music transformer model to predict human brain activity. We recorded scalp electroencephalography (EEG) from expert musicians and non-musicians, as well as intracranial EEG (iEEG) from six neurological patients, during listening to classical piano pieces. We found that deeper layers of the transformer, which represent more disentangled and contextual musical features, were more predictive of neural responses in both groups. However, this neural correspondence was significantly enhanced by musical expertise: for non-musicians, prediction accuracy plateaued across the model's final layers, whereas for musicians it continued to increase. This enhanced encoding in experts was also strongly lateralized to the left hemisphere. Finally, the iEEG recordings revealed an anatomical gradient for this function, where neural sites progressively farther from the primary auditory cortex encoded musical context more strongly. Our study reveals how musical training refines the hierarchical neural processing of music and provides a neuro-computational account of this remarkable cognitive skill.
The discrete events of our narrative experience are organized by the neural substrate that underlies episodic memory. This narrative process is segmented into distinct units by event boundaries, which facilitate a replay process that acts to consolidate each event into a narrative memory. High-frequency oscillations (HFOs) may synchronize neural activity during these processes. We use intracranial recordings from participants viewing and freely recalling a continuous, audiovisual stimulus. We find that hippocampal HFOs increase following event boundaries and hippocampal-cortical coincident HFOs (co-HFOs) occur in cortical regions that underlie event segmentation (inferior parietal, precuneus, lateral occipital, and inferior frontal cortices). Event-specific co-HFO patterns that occur during event viewing reoccur following event boundaries for the subsequent three events and during recall. This is consistent with models that support replay as a mechanism for memory consolidation. Therefore, HFOs may coordinate activity across brain regions that facilitate event segmentation, encode memory of discrete events, and bind representations to assemble memory of a coherent, continuous experience.
Decoding continuous language from neural signals remains a significant challenge in the intersection of neuroscience and artificial intelligence. We introduce Neuro2Semantic, a novel framework that reconstructs the semantic content of perceived speech from intracranial EEG (iEEG) recordings. Our approach consists of two phases: first, an LSTM-based adapter aligns neural signals with pre-trained text embeddings; second, a corrector module generates continuous, natural text directly from these aligned embeddings. This flexible method overcomes the limitations of previous decoding approaches and enables unconstrained text generation. Neuro2Semantic achieves strong performance with as little as 30 minutes of neural data, outperforming a recent state-of-the-art method in low-data settings. These results highlight the potential for practical applications in brain-computer interfaces and neural decoding technologies.
Listeners effortlessly extract multi-dimensional auditory objects, such as a localized talker, from complex acoustic scenes. However, the neural mechanisms that enable simultaneous encoding and linking of distinct sound features-such as a talker's voice and location-are not fully understood. Using invasive intracranial recordings in seven neurosurgical patients (four male, three female), we investigated how the human auditory cortex processes and integrates these features during naturalistic multi-talker scenes and how attentional mechanisms modulate such feature integration. We found that cortical sites exhibit a continuum of feature sensitivity, ranging from single-feature-sensitive sites (responsive primarily to voice spectral features or to location features) to dual-feature-sensitive sites (responsive to both features). At the population level, neural response patterns from both single- and dual-feature-sensitive sites jointly encoded the attended talker's voice and location. Notably, single-feature-sensitive sites encoded their primary feature with greater precision but also represented coarse information about the secondary feature. Sites selectively tracking a single, attended speech stream concurrently encoded both voice and location features, demonstrating a link between selective attention and feature integration. Additionally, attention selectively enhanced temporal coherence between voice- and location-sensitive sites, suggesting that temporal synchronization serves as a mechanism for linking these features. Our findings highlight two complementary neural mechanisms-joint population coding and temporal coherence-that enable the integration of voice and location features in the auditory cortex. These results provide new insights into the distributed, multi-dimensional nature of auditory object formation during active listening in complex environments.
Humans live in an environment that contains rich auditory stimuli, which must be processed efficiently. The entrainment of neural oscillations to acoustic inputs may support the processing of simple and complex sounds. However, the characteristics of this entrainment process have been shown to be inconsistent across species and experimental paradigms. It is imperative to establish whether neural activity in response to speech is a result of combination of simple evoked responses or of entrainment of neural oscillations in human participants. In this study, 12 participants with intracranial electrodes listened to natural speech and neural entrainment as evidenced by oscillatory activity persisting beyond the evoked responses was assessed. Neural activity was recorded from 165 contacts in Heschl’s gyrus and superior temporal gyrus. First, acoustic edges in the speech envelope induced coherence between speech and auditory cortex activity. Further, entrainment in the theta-alpha band outlasted the acoustic stimulation. This activity exceeded what could be expected from a simple evoked response. These findings suggest that speech has the potential to entrain neural oscillations in the human auditory cortex.
José R. Herrero合作论文数Departament d'Arquitectura de Computadors (DAC);Universitat Polit鑓nica de Catalunya5