Artificial neural networks are widely used in modeling sentence processing but often exhibit overconfident point estimates in their output behavior, even in the presence of ambiguous or conflicting linguistic cues. This limitation is illustrated by reversal anomalies—sentences with unexpected role reversals inducing a conflict between syntactic and semantic information, and has been observed in established models such as the Sentence Gestalt (SG) model. To address this, we introduce a Bayesian formulation of the SG model by applying an extension of the ensemble Kalman filter for Bayesian inference at the level of model parameters. Framing sentence comprehension as a Bayesian inverse problem allows us to characterize posterior predictive uncertainty in the model’s output representations, rather than relying on point estimates. Through numerical experiments and comparisons with a standard maximum likelihood-trained SG model, we show that the Bayesian approach yields systematically less overconfident output activations under cue conflict, reflecting increased uncertainty when processing linguistic ambiguities.
The N400 component of ERPs is modulated by how predictable a word is, but predictability is usually quantified with lexical cloze —the probability that readers supply that exact word in offline sentence completion tasks. This form-based metric is at odds with decades of evidence that the N400 is primarily sensitive to meaning . Here, we asked whether a measure of semantic feature predictability can better account for N400 amplitude modulation. We reanalysed two independent EEG datasets (N = 26 and N = 334), computing lexical and semantic cloze for each critical word. Across both datasets, semantic cloze emerged as a better predictor of the N400 data than lexical cloze. Using the same materials, we then compared semantic and lexical cloze with probabilities from four large language models (GPT-2, GPT-2.7b, RoBERTa, ALBERT). None of the LLM-derived predictors outperformed semantic cloze. Our findings support the view that the N400 primarily reflects semantic—not exact-word—processing. Methodologically, we argue that replacing lexical cloze with semantic cloze can substantially increase the explanatory power of N400 studies, and caution against substituting human norms with raw LLM probabilities. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft (DFG, German Re-search Foundation) – Project ID 317633480 (SFB1287)
Artificial neural networks (ANNs) are widely used in modeling sentence processing but often exhibit deterministic behavior, contrasting with human sentence comprehension, which manages uncertainty during ambiguous or unexpected inputs. This is exemplified by reversal anomalies-sentences with unexpected role reversals that challenge syntax and semantics-highlighting the limitations of traditional ANN models, such as the Sentence Gestalt (SG) Model. To address these limitations, we propose a Bayesian framework for sentence comprehension, applying an extension of the ensemble Kalman filter (EnKF) for Bayesian inference to quantify uncertainty. By framing language comprehension as a Bayesian inverse problem, this approach enhances the SG model's ability to reflect human sentence processing with respect to the representation of uncertainty. Numerical experiments and comparisons with maximum likelihood estimation (MLE) demonstrate that Bayesian methods improve uncertainty representation, enabling the model to better approximate human cognitive processing when dealing with linguistic ambiguities.
Prediction error, both at the level of sentence meaning and at the level of the next presented word, has been shown to successfully account for N400 amplitudes. Here we address the question of whether people differ in the representational level at which they implicitly predict upcoming language. To this end, we compute a measure of prediction error at the level of sentence meaning (magnitude of change in hidden layer activation, termed semantic update, in a neural network model of sentence comprehension, the Sentence Gestalt model) and a measure of prediction error at the level of the next presented word (surprisal from a next word prediction language model). When using both measures to predict N400 amplitudes during the reading of naturalistic texts, results showed that both measures significantly accounted for N400 amplitudes even when the other measure was controlled for. Most important for current purposes, both effects were significantly negatively correlated such that people with a reversed or weak surprisal effect showed the strongest influence of semantic update on N400 amplitudes, and random-effects model comparison showed that individuals differ in whether their N400 amplitudes are driven by semantic update only, by surprisal only, or by both, and that the most common model in the population was either semantic update or the combined model but clearly not the pure surprisal model. The current approach of combining large-scale models implementing different theoretical accounts with advanced model comparison techniques enables fine-grained investigations into the computational processes underlying N400 amplitudes, including interindividual differences. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, RA 2715/2-1 (Emmy Noether grant), 318763901 (SFB 1294, project B09)
Recent research has shown that the internal dynamics of an artificial neural network model of sentence comprehension displayed a similar pattern to the amplitude of the N400 in several conditions known to modulate this event-related potential. These results led Rabovsky et al. (2018) to suggest that the N400 might reflect change in an implicit predictive representation of meaning corresponding to semantic prediction error. This explanation stands as an alternative to the hypothesis that the N400 reflects lexical prediction error as estimated by word surprisal (Frank et al., 2015). In the present study, we directly model the amplitude of the N400 elicited during naturalistic sentence processing by using as predictor the update of the distributed representation of sentence meaning generated by a sentence gestalt model (McClelland et al., 1989) trained on a large-scale text corpus. This enables a quantitative prediction of N400 amplitudes based on a cognitively motivated model, as well as quantitative comparison of this model to alternative models of the N400. Specifically, we compare the update measure from the sentence gestalt model to surprisal estimated by a comparable language model trained on next-word prediction. Our results suggest that both sentence gestalt update and surprisal predict aspects of N400 amplitudes. Thus, we argue that N400 amplitudes might reflect two distinct but probably closely related sub-processes that contribute to the processing of a sentence.
The N400 component of the event related brain potential is widely used to investigate language and meaning processing. However, despite much research the component’s functional basis remains actively debated. Recent work showed that the update of the predictive representation of sentence meaning (semantic update, or SU) generated by the Sentence Gestalt model (McClelland, St. John, & Taraban, 1989) consistently displayed a similar pattern to the N400 amplitude in a series of conditions known to modulate this event-related potential. These results led Rabovsky, Hansen, and McClelland (2018) to suggest that the N400 might reflect change in a probabilistic representation of meaning corresponding to an implicit semantic prediction error. However, a limitation of this work is that the model was trained on a small artificial training corpus and thus could not be presented with the same naturalistic stimuli presented in empirical experiments. In the present study, we overcome this limitation and directly model the amplitude of the N400 elicited during naturalistic sentence processing by using as predictor the SU generated by a Sentence Gestalt model trained on a large corpus of texts. The results reported in this paper corroborate the hypothesis that the N400 component reflects the change in a probabilistic representation of meaning after every word presentation. Further analyses demonstrate that the SU of the Sentence Gestalt model and the amplitude of the N400 are influenced similarly by the stochastic and positional properties of the linguistic input.
Finding the structure of a sentence-the way its words hold together to convey meaning-is a fundamental step in language comprehension. Several brain regions, including the left inferior frontal gyrus, the left posterior superior temporal gyrus, and the left anterior temporal pole, are supposed to support this operation. The exact role of these areas is nonetheless still debated. In this paper we investigate the hypothesis that different brain regions could be sensitive to different kinds of syntactic computations. We compare the fit of phrase-structure and dependency structure descriptors to activity in brain areas using fMRI. Our results show a division between areas with regard to the type of structure computed, with the left anterior temporal pole and left inferior frontal gyrus favouring dependency structures and left posterior superior temporal gyrus favouring phrase structures.
The meaning of a word depends on its lexical semantics and on the context in which it is embedded. At the basis of this lays the distinction between lexical retrieval and integration, two basic operations supporting language comprehension. In this paper, we investigate how lexical retrieval and integration are implemented in the brain by comparing MEG activity to word representations generated by computational language models. We test both non-contextualized embeddings, representing words independently from their context, and contextualized embeddings, which instead integrate contextual information in their representations. Using representational similarity analysis over cortical regions and over time, we observed that brain activity in the left anterior temporal pole and inferior frontal regions shows higher similarity with contextualized word embeddings compared to non-contextualized embeddings, between 300 and 500 ms after word presentation. On the other hand, non-contextualized word embeddings show higher similarity with brain activity in the left lateral and anterior temporal lobe at earlier latencies – areas and latencies related to lexical retrieval. Our results highlight how lexical retrieval and context integration can be tracked in the brain using word embeddings obtained with computational models. These results also suggest that the distinction between lexical retrieval and integration might be framed in terms of context-independent and contextualized representations.
Neural decoding of speech and language refers to the extraction of information regarding the stimulus and the mental state of subjects from recordings of their brain activity while performing linguistic tasks. Recent years have seen significant progress in the decoding of speech from cortical activity. This study instead focuses on decoding linguistic information. We present a deep parallel temporal convolutional neural network (1DCNN) trained on part-of-speech (PoS) classification from magnetoencephalography (MEG) data collected during natural language reading. The network is trained on data from 15 human subjects separately, and yields above-chance accuracies on test data for all of them. The level of PoS was targeted because it offers a clean linguistic benchmark level that represents syntactic information and abstracts away from semantic or conceptual representations.
Backward saccades during reading have been hypothesized to be involved in structural reanalysis, or to be related to the level of text difficulty. We test the hypothesis that backward saccades are involved in online syntactic analysis. If this is the case we expect that saccades will coincide, at least partially, with the edges of the relations computed by a dependency parser. In order to test this, we analyzed a large eye-tracking dataset collected while 102 participants read three short narrative texts. Our results show a relation between backward saccades and the syntactic structure of sentences.
We present the Narrative Brain Dataset, an fMRI dataset that was collected during spoken presentation of short excerpts of three stories in Dutch. Together with the brain imaging data, the dataset contains the written versions of the stimulation texts. The texts are accompanied with stochastic (perplexity and entropy) and semantic computational linguistic measures. The richness and unconstrained nature of the data allows the study of language processing in the brain in a more naturalistic setting than is common for fMRI studies. We hope that by making NBD available we serve the double purpose of providing useful neural data to researchers interested in natural language processing in the brain and to further stimulate data sharing in the field of neuroscience of language.
Language comprehension involves the simultaneous processing of information at the phonological, syntactic, and lexical level. We track these three distinct streams of information in the brain by using stochastic measures derived from computational language models to detect neural correlates of phoneme, part-of-speech, and word processing in an fMRI experiment. Probabilistic language models have proven to be useful tools for studying how language is processed as a sequence of symbols unfolding in time. Conditional probabilities between sequences of words are at the basis of probabilistic measures such as surprisal and perplexity which have been successfully used as predictors of several behavioural and neural correlates of sentence processing. Here we computed perplexity from sequences of words and their parts of speech, and their phonemic transcriptions. Brain activity time-locked to each word is regressed on the three model-derived measures. We observe that the brain keeps track of the statistical structure of lexical, syntactic and phonological information in distinct areas.
Following earlier work in multimodal distributional semantics, we present the first results of our efforts to build a perceptually grounded semantic model. Rather than using images, our models are built on sound data collected from freesound.org. We compare three models: one bag-of-words model based on user-provided tags, a model based on audio features, using a ‘bag-of-audio-words’ approach and a model that combines the two. Our results show that the models are able to capture semantic relatedness, with the tag-based model scoring higher than the sound-based model and the combined model. However, capturing semantic relatedness is biased towards language-based models. Future work will focus on improving the sound-based model, finding ways to combine linguistic and acoustic information, and creating more reliable evaluation data.
We present preliminary results in the domain of sound labeling and sound representation. Our work is based on data from the Freesound database, which contains thousands of sounds complete with tags and descriptions, under a Creative Commons license. We want to investigate how people represent and categorize different sounds, and how language reflects this categorization. Moreover, following recent developments in multimodal distributional semantics (Bruni et al. 2012), we want to assess whether acoustic information can improve the semantic representation of lexemes. We have built two different distributional models on the basis of a subset of the Freesound database, containing all sounds that were manually classified as SoundFX (e.g. footsteps, opening and closing doors, animal sounds). The first model is based on tag co-occurrence. On the basis of this model, we created a network of tags that we partitioned using cluster analysis. The clustering intuitively seems to correspond with different types of scenes. We imagine that this partitioning is a first step towards linking particular sounds with relevant frames in FrameNet. The second model is built using a bag-of-auditory-words approach. In order to assess the goodness of the semantic representations, the two models are compared to human judgment scores from the WordSim353 and MEN database.
Embodiment theory predicts that mental imagery of object words recruits neural circuits involved in object perception. The degree of visual imagery present in routine thought and how it is encoded in the brain is largely unknown. We test whether fMRI activity patterns elicited by participants reading objects' names include embodied visual-object representations, and whether we can decode the representations using novel computational image-based semantic models. We first apply the image models in conjunction with text-based semantic models to test predictions of visual-specificity of semantic representations in different brain regions. Representational similarity analysis confirms that fMRI structure within ventral-temporal and lateral-occipital regions correlates most strongly with the image models and conversely text models correlate better with posterior-parietal/lateral-temporal/inferior-frontal regions. We use an unsupervised decoding algorithm that exploits commonalities in representational similarity structure found within both image model and brain data sets to classify embodied visual representations with high accuracy (8/10) and then extend it to exploit model combinations to robustly decode different brain regions in parallel. By capturing latent visual-semantic structure our models provide a route into analyzing neural representations derived from past perceptual experience rather than stimulus-driven brain activity. Our results also verify the benefit of combining multimodal data to model human-like semantic representations.
Simulated Annealing Optimality Theory (SA-OT) is a recent update of Optimality Theory, adding a model of performance to a theory of linguistic competence. Our aim is to show how SA-OT can be a useful paradigm for language change simulations. Performance "errors" are considered to be one of the causes of variation and change. We have chosen to model the evolution of sentential negation (SN). The descriptive background adopts Jespersen's Cycle, according to which the evolution of sentential negation follows three main stages (1. pre-verbal, 2. discontinuous, and 3. post-verbal). Therefore, we advance a novel model for SN, based on SA-OT. It reproduces the three pure and the two observed mixed stages, whereas it correctly predicts the lack of an intermediate stage between 3 and 1. The success of the approach corroborates the computational, performance-based approach to the data. Finally, we employ the iterated learning paradigm to reproduce historical changes in a "simulated corpus study". This enterprise turns out to be more difficult than one would naively believe.