Artificial neural networks are widely used in modeling sentence processing but often exhibit overconfident point estimates in their output behavior, even in the presence of ambiguous or conflicting linguistic cues. This limitation is illustrated by reversal anomalies—sentences with unexpected role reversals inducing a conflict between syntactic and semantic information, and has been observed in established models such as the Sentence Gestalt (SG) model. To address this, we introduce a Bayesian formulation of the SG model by applying an extension of the ensemble Kalman filter for Bayesian inference at the level of model parameters. Framing sentence comprehension as a Bayesian inverse problem allows us to characterize posterior predictive uncertainty in the model’s output representations, rather than relying on point estimates. Through numerical experiments and comparisons with a standard maximum likelihood-trained SG model, we show that the Bayesian approach yields systematically less overconfident output activations under cue conflict, reflecting increased uncertainty when processing linguistic ambiguities.
Understanding why readers misinterpret role-reversed sentences like the waitress that the customer served (versus the customer that the waitress served) is important in understanding how readers process who does what to whom in sentences. In such sentence pairs, we infer that readers experience temporary illusions of plausibility from the absence of the usual N400 effect between the plausible and implausible verb served. A recent finding is that an N400 effect is observed if presentation of the verb is delayed. We first use ERPs to show that time alone rather than new information is sufficient to produce this delay effect. Our meta-analysis of published studies demonstrates a small but consistent effect of similar delays. Next, we distinguish between two accounts of the illusion addressing the delay effect: In one, the illusion results from conflict between syntactic and semantic cues (Sentence Gestalt (SG) model; Rabovsky et al., 2018), in the other, from the slow application of thematic roles (e.g., Liao et al., 2022). Using speeded lexical decisions, we demonstrate an immediate influence of thematic roles on verb processing, favouring the SG model. The role reversal effect on lexical decision times but not the N400 further supports the SG model because the model predicts a dissociation between preactivation of specific words (reflected in lexical decisions) versus sentence meaning (reflected in the N400). Together, the findings suggest that the rapid pace of reading can give rise to temporary illusions by preventing readers from resolving cue conflict rather than by omitting processing steps.
The ERP components P3 and P600 have been proposed to reflect phasic activity of the locus coeruleus norepinephrine (LC/NE) system in response to deviant and task-relevant stimuli across cognitive domains. However, causal evidence for this link remains limited. Here, we used continuous transcutaneous auricular vagus nerve stimulation (taVNS), a noninvasive method proposed to modulate LC/NE activity, to test whether these components are indeed sensitive to NE manipulation. Forty participants completed both an active visual oddball task and a sentence processing task, including both syntactic and semantic violations, while receiving continuous taVNS at the cymba conchae in one session and sham stimulation at the earlobe in another session. We observed robust P3 and P600 effects. Crucially, however, taVNS had no effect on P3 or P600 amplitude. The physiological NE markers, salivary alpha-amylase level and baseline pupil size, were also unaffected by the stimulation, suggesting that the taVNS protocol and/or task may not have been sufficient to successfully engage the LC/NE system. Beyond the stimulation, however, exploratory analyses revealed correlations between the syntactic P600 and both the P3 and salivary alpha-amylase levels, supporting the idea that the P600 might be related to both the P3 and NE. Overall, our findings do not allow for theoretical implications concerning a potential causal link between the two components and NE but highlight the need for more standardized taVNS protocols.
Unexpected words within a context elicit large N400 brain potentials. However, sometimes the N400 at an unexpected word is small when stereotypical agent and patient roles are reversed, such as at “arrested” in “the cop that the thief arrested.” In a study of 74 native German speakers, we demonstrate evidence that readers can avoid this so-called N400 semantic illusion if the verb is delayed with neutral information such as “that evening,” but are less able to do so if the delay contains cues that could further strengthen the canonical interpretation, such as “with handcuffs.” In doing so, we provide a conceptual replication of a relatively new finding and extend previous research by showing that the semantic content of the delay is important. Moreover, we demonstrate evidence that the effect of only the neutral delay increases as the experiment progresses. We propose an interpretation of these findings with reference to the Sentence Gestalt model [Rabovsky, M., Hansen, S. S., & McClelland, J. L. Modelling the N400 brain potential as change in a probabilistic representation of meaning. Nature Human Behaviour, 2, 693, 2018], which accounts for the initial illusion as resulting from uncertainty and an erroneous interpretation based on a strong semantic attractor. Two additional, novel contributions of the work are a demonstration that the illusion can be elicited in German, despite its explicit subject–object case marking, and an exploration of illusion effect among individual readers.
The N400 component of ERPs is modulated by how predictable a word is, but predictability is usually quantified with lexical cloze —the probability that readers supply that exact word in offline sentence completion tasks. This form-based metric is at odds with decades of evidence that the N400 is primarily sensitive to meaning . Here, we asked whether a measure of semantic feature predictability can better account for N400 amplitude modulation. We reanalysed two independent EEG datasets (N = 26 and N = 334), computing lexical and semantic cloze for each critical word. Across both datasets, semantic cloze emerged as a better predictor of the N400 data than lexical cloze. Using the same materials, we then compared semantic and lexical cloze with probabilities from four large language models (GPT-2, GPT-2.7b, RoBERTa, ALBERT). None of the LLM-derived predictors outperformed semantic cloze. Our findings support the view that the N400 primarily reflects semantic—not exact-word—processing. Methodologically, we argue that replacing lexical cloze with semantic cloze can substantially increase the explanatory power of N400 studies, and caution against substituting human norms with raw LLM probabilities. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft (DFG, German Re-search Foundation) – Project ID 317633480 (SFB1287)
The ERP components P3 and P600 have been proposed to reflect phasic activity of the locus coeruleus norepinephrine (LC/NE) system in response to deviant and task-relevant stimuli across cognitive domains. Yet, causal evidence for this link remains limited. Here, we used continuous transcutaneous auricular vagus nerve stimulation (taVNS), a non-invasive method proposed to modulate LC/NE activity, to test whether these components are indeed sensitive to noradrenergic manipulation. Forty participants completed both an active visual oddball task and a sentence processing task including both syntactic and semantic violations, while receiving continuous taVNS at the cymba conchae in one session and sham stimulation at the earlobe in another session. We observed robust P3 and P600 effects. Crucially though, taVNS had no effect on P3 or P600 amplitude. The physiological NE markers salivary alpha amylase level and baseline pupil size were also unaffected by the stimulation, suggesting that the taVNS protocol and/or task may not have been sufficient to successfully engage the norepinephrine system. Beyond the stimulation, however, exploratory analyses revealed correlations between the syntactic P600 and both the P3 and salivary alpha amylase levels, supporting the idea that the P600 might be related to both the P3 and NE. Overall, our findings do not allow for theoretical implications concerning a potential causal link between the two components and NE but highlight the need for more standardized taVNS protocols. ### Competing Interest Statement The authors have declared no competing interest. Research Focus Cognitive Sciences of the University of Potsdam Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, 317633480, SFB1287
The P600 ERP component is elicited by a wide range of anomalies and ambiguities during sentence comprehension and remains important for neurocognitive models of language processing. It has been proposed that the P600 is a more domain-general component, signaling phasic norepinephrine release from the locus coeruleus in response to salient stimuli that require attention and behavioral adaptation. Because such norepinephrine release promotes explicit memory formation, we here investigated whether the P600 during sentence reading (encoding) is thus predictive of such explicit memory formation using a subsequent old/new word recognition task. Indeed, the P600 amplitude during our encoding task was related to behavioral recognition effects in the memory task on a trial-by-trial basis, although only for one type of violation. Recognition performance was better for semantically, but not syntactically, violated words that had previously elicited a larger P600. However, the P600 to both types of violations during encoding was positively related to a more subtle, neural marker of recognition, namely, the amplitude of the recollection ERP component in response to old words. In summary, we find that the P600 predicts later recognition memory both on the behavioral and neural level. Such explicit memory effects further link the late positivity to norepinephrine activity, suggesting a more domain-general nature of the component. The connection between the P600 and later recognition indicates that the neurocognitive processes that deal with salient and anomalous aspects in the linguistic input in the moment will also be involved in keeping this event available for later recognition.
Artificial neural networks (ANNs) are widely used in modeling sentence processing but often exhibit deterministic behavior, contrasting with human sentence comprehension, which manages uncertainty during ambiguous or unexpected inputs. This is exemplified by reversal anomalies-sentences with unexpected role reversals that challenge syntax and semantics-highlighting the limitations of traditional ANN models, such as the Sentence Gestalt (SG) Model. To address these limitations, we propose a Bayesian framework for sentence comprehension, applying an extension of the ensemble Kalman filter (EnKF) for Bayesian inference to quantify uncertainty. By framing language comprehension as a Bayesian inverse problem, this approach enhances the SG model's ability to reflect human sentence processing with respect to the representation of uncertainty. Numerical experiments and comparisons with maximum likelihood estimation (MLE) demonstrate that Bayesian methods improve uncertainty representation, enabling the model to better approximate human cognitive processing when dealing with linguistic ambiguities.
The brain’s remarkable ability to extract patterns from sequences of events has been demonstrated across cognitive domains and is a central assumption of predictive processing theories. While predictions shape language processing at the level of meaning, little is known about the underlying learning mechanism. Here, we investigated how continuous statistical inference in a semantic sequence influences the neural response. 60 participants were presented with a semantic oddball-like roving paradigm, consisting of sequences of nouns from different semantic categories. Unknown to the participants, the overall sequence contained an additional manipulation of transition probability between categories. Two Bayesian sequential learner models that captured different aspects of probabilistic learning were used to derive theoretical surprise levels for each trial and investigate online probabilistic semantic learning. The N400 ERP component was primarily modulated by increased probability with repeated exposure to the categories throughout the experiment, which essentially represents repetition suppression. This N400 repetition suppression likely prevented sizeable influences of more complex predictions such as those based on transition probability, as any incoming information was already continuously active in semantic memory. In contrast, the P600 was associated with semantic surprise in a transition probability model over recent observations, possibly indicating a working memory update in response to violations of these conditional dependencies. The results support probabilistic predictive processing of semantic information and demonstrate that continuous update of distinct statistics differentially influences language related ERPs.### Competing Interest StatementThe authors have declared no competing interest.
During language comprehension, anomalies and ambiguities in the input typically elicit the P600 event-related potential component. Although traditionally interpreted as a specific signal of combinatorial operations in sentence processing, the component has alternatively been proposed to be a variant of the oddball-sensitive, domain-general P3 component. In particular, both components might reflect phasic norepinephrine release from the locus coeruleus (LC/NE) to motivationally significant stimuli. In this preregistered study, we tested this hypothesis by relating both components to the task-evoked pupillary response, a putative biomarker of LC/NE activity. 36 participants completed a sentence comprehension task (containing 25% morphosyntactic violations) and a non-linguistic oddball task (containing 20% oddballs), while the EEG and pupil size were co-registered. Our results showed that the task-evoked pupillary response and the ERP amplitudes of both components were similarly affected by both experimental tasks. In the oddball task, there was also a temporally specific relationship between the P3 and the pupillary response beyond the shared oddball effect, thereby further linking the P3 to NE. Because this link was less reliable in the linguistic context, we did not find conclusive evidence for or against a relationship between the P600 and the pupillary response. Still, our findings further stimulate the debate on whether language-related ERPs are indeed specific to linguistic processes or shared across cognitive domains. However, further research is required to verify a potential link between the two ERP positivities and the LC/NE system as the common neural generator.
Prediction error, both at the level of sentence meaning and at the level of the next presented word, has been shown to successfully account for N400 amplitudes. Here we address the question of whether people differ in the representational level at which they implicitly predict upcoming language. To this end, we compute a measure of prediction error at the level of sentence meaning (magnitude of change in hidden layer activation, termed semantic update, in a neural network model of sentence comprehension, the Sentence Gestalt model) and a measure of prediction error at the level of the next presented word (surprisal from a next word prediction language model). When using both measures to predict N400 amplitudes during the reading of naturalistic texts, results showed that both measures significantly accounted for N400 amplitudes even when the other measure was controlled for. Most important for current purposes, both effects were significantly negatively correlated such that people with a reversed or weak surprisal effect showed the strongest influence of semantic update on N400 amplitudes, and random-effects model comparison showed that individuals differ in whether their N400 amplitudes are driven by semantic update only, by surprisal only, or by both, and that the most common model in the population was either semantic update or the combined model but clearly not the pure surprisal model. The current approach of combining large-scale models implementing different theoretical accounts with advanced model comparison techniques enables fine-grained investigations into the computational processes underlying N400 amplitudes, including interindividual differences. ### Competing Interest Statement The authors have declared no competing interest. Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, RA 2715/2-1 (Emmy Noether grant), 318763901 (SFB 1294, project B09)
The present EEG study with 32 healthy participants investigated whether affective knowledge about a person influences the visual awareness of their face, additionally considering the impact of facial appearance. Faces differing in perceived trustworthiness based on appearance were associated with negative or neutral social information and shown as target stimuli in an attentional blink task. As expected, participants showed enhanced awareness of faces associated with negative compared to neutral social information. On the neurophysiological level, this effect was connected to differences in the time range of the early posterior negativity (EPN)—a component associated with enhanced attention and facilitated processing of emotional stimuli. The findings indicate that the social-affective relevance of a face based on emotional knowledge is accessed during a phase of attentional enhancement for conscious perception and can affect prioritization for awareness. In contrast, no clear evidence for influences of facial trustworthiness during the attentional blink was found.
Prediction errors drive implicit learning in language, but the specific mechanisms underlying these effects remain debated. This issue was addressed in an EEG study manipulating the context of a repeated unpredictable word (repetition of the complete sentence or repetition of the word in a new sentence context) and sentence constraint. For the manipulation of sentence constraint, unexpected words were presented either in high-constraint (eliciting a precise prediction) or low-constraint sentences (not eliciting any specific prediction). Repetition-induced reduction of N400 amplitudes and of power in the alpha/beta frequency band was larger for words repeated with their sentence context as compared with words repeated in a new low-constraint context, suggesting that implicit learning happens not only at the level of individual items but additionally improves sentence-based predictions. These processing benefits for repeated sentences did not differ between constraint conditions, suggesting that sentence-based prediction update might be proportional to the amount of unpredicted semantic information, rather than to the precision of the prediction that was violated. In addition, the consequences of high-constraint prediction violations, as reflected in a frontal positivity and increased theta band power, were reduced with repetition. Overall, our findings suggest a powerful and specific adaptation mechanism that allows the language system to quickly adapt its predictions when unexpected semantic information is processed, irrespective of sentence constraint, and to reduce potential costs of strong predictions that were violated.
Recent research has shown that the internal dynamics of an artificial neural network model of sentence comprehension displayed a similar pattern to the amplitude of the N400 in several conditions known to modulate this event-related potential. These results led Rabovsky et al. (2018) to suggest that the N400 might reflect change in an implicit predictive representation of meaning corresponding to semantic prediction error. This explanation stands as an alternative to the hypothesis that the N400 reflects lexical prediction error as estimated by word surprisal (Frank et al., 2015). In the present study, we directly model the amplitude of the N400 elicited during naturalistic sentence processing by using as predictor the update of the distributed representation of sentence meaning generated by a sentence gestalt model (McClelland et al., 1989) trained on a large-scale text corpus. This enables a quantitative prediction of N400 amplitudes based on a cognitively motivated model, as well as quantitative comparison of this model to alternative models of the N400. Specifically, we compare the update measure from the sentence gestalt model to surprisal estimated by a comparable language model trained on next-word prediction. Our results suggest that both sentence gestalt update and surprisal predict aspects of N400 amplitudes. Thus, we argue that N400 amplitudes might reflect two distinct but probably closely related sub-processes that contribute to the processing of a sentence.
The N400 component of the event-related potential (ERP) is the most widely used brain signal in research on semantic processing. It has been discovered now more than 30 years ago, in 1980, as a larger negativity for semantically incongruent sentence continuations such as “I take my coffee with cream and dog” (as compared to congruent continuations such as “sugar”). The N400 has meanwhile been shown to be modulated by a very wide variety of lexical and semantic variables and has taught us a lot about how meaning is processed in language and beyond. This chapter reviews the literature on the N400 component including its relationship to the subsequent P600 component and discusses implications for the neurocognition of semantic processing.