A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind. In this work, we study reading dynamics for ``noisy-channel garden-path'' sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error. We find evidence for targeted regressions -- eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis. We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics.
Digging-in effects, where disambiguation difficulty increases with longer ambiguous regions, have been cited as evidence for self-organized sentence processing, in which structural commitments strengthen over time. In contrast, surprisal theory predicts no such effect unless lengthening genuinely shifts statistical expectations, and neural language models appear to show the opposite pattern. Whether digging-in is a robust real-time phenomenon in human sentence processing – or an artifact of wrap-up processes or methodological confounds – remains unclear. We report two experiments on English NP/Z garden-path sentences using Maze and self-paced reading, comparing human behavior with predictions from an ensemble of large language models. We find no evidence for real-time digging-in effects. Critically, items with sentence-final versus nonfinal disambiguation show qualitatively different patterns: positive digging-in trends appear only sentence-finally, where wrap-up effects confound interpretation. Nonfinal items – the cleaner test of real-time processing – show reverse trends consistent with neural model predictions.
Under surprisal theory, linguistic representations affect processing difficulty only through the bottleneck of surprisal. Our best estimates of surprisal come from large language models, which have no explicit representation of structural ambiguity. While LLM surprisal robustly predicts reading times across languages, it systematically underpredicts difficulty when structural expectations are violated – suggesting that representations of ambiguity are causally implicated in sentence processing. Particle filter models offer an alternative where structural hypotheses are explicitly represented as a finite set of particles. We prove several algorithmic consequences of particle filter models, including the amplification of garden-path effects. Most critically, we demonstrate that resampling, a common practice with these models, inherently produces real-time digging-in effects – where disambiguation difficulty increases with ambiguous region length. Digging-in magnitude scales inversely with particle count: fully parallel models predict no such effect.
Human language processing can be studied through both behavior and brain activity, yet it remains unclear whether these two data types reflect sensitivity to the same information. One influential view holds that both behavioral and neural responses are largely determined by processing effort, often estimated by word surprisal together with the context-independent properties of word frequency and length. At the same time, neural responses have been shown to encode richer aspects of linguistic content, including meaning. Here, we use neural network language models to operationalize these alternatives and systematically compare, within the same analytic computational framework, the predictive power of low-dimensional effort-based predictors and high-dimensional embedding representations that encode contextualized linguistic content, including meaning. Across 8 behavioral datasets and 5 neural datasets (4 fMRI and 1 ERP), we find that processing effort captures substantial variance in both behavioral and neural measures of language processing, in line with much previous work. However, for brain responses---but not for behavioral measures---embedding representations carry substantial predictive power beyond the estimates of processing effort. These results therefore suggest that neural data provide access to rich, high-dimensional dynamics of language comprehension, whereas behavioral data reflect a bottlenecking of these dynamics into a small set of theoretically motivated properties of contextualized linguistic input.
Intercomprehension refers to partial intelligibility of an unfamiliar language (L2) by a speaker of a related language (L1). How is this zero-shot cross-language comprehension possible? In this work, we extend past work on algorithmic models of noisy-channel inference to model intercomprehension in a Bayesian framework. The model uses an LM in L1 only for scoring latent hypotheses about the translations of observed L2 utterances, and a general-purpose noise model to infer a mapping between L2 and L1 words based on either form-based similarity or symbolic rules. We then conduct a human behavioral experiment, eliciting inferences for utterances in Dutch, Italian, and Ukrainian from speakers of English, Spanish, and Russian, respectively. Our full model shows a closer alignment to the distribution of human intercomprehension performance than ablations, and also compares favorably to zero-shot prompting of much larger models. These results provide a cognitively plausible computational model of intercomprehension, and highlight the flexible inferences made by comprehenders under wide uncertainty in real-world cross-language scenarios. We share our code publicly.
Studying early speech development at scale requires automatic tools, yet automatic phoneme recognition, especially for young children, remains largely unsolved. Building on decades of data collection, we curate TinyVox, a corpus of more than half a million phonetically transcribed child vocalizations in English, French, Portuguese, German, and Spanish. We use TinyVox to train BabAR, a cross-linguistic phoneme recognition system for child speech. We find that pretraining the system on multilingual child-centered daylong recordings substantially outperforms alternatives, and that providing 20 seconds of surrounding audio context during fine-tuning further improves performance. Error analyses show that substitutions predominantly fall within the same broad phonetic categories, suggesting suitability for coarse-grained developmental analyses. We validate BabAR by showing that its automatic measures of speech maturity align with developmental estimates from the literature.
A central endeavor in psycholinguistic research has been to determine the processing profile of syntactically ambiguous strings. Previous work investigating syntactic attachment ambiguities has shown that discarding a locally grammatically available, but globally failing, parse is costly. However, little is known about how comprehenders cope with semantic parsing ambiguities. Using the case study of scopally ambiguous definite descriptions such as the rabbit in the big hat, we examine whether comparable penalties arise for non-lexical semantic ambiguities. In a series of reference resolution tasks, we find dispreference for strings that are globally defined but fail to refer under alternative semantic parses, compared to strings where all readings successfully refer to the same individual. Crucially, this effect is only detectable when the alternative failing reading gives rise to a REFERENTIAL GARDEN PATH, where a dynamic constraint evaluation process temporarily settles on a unique referent before eventually failing. We conclude that failing alternative readings cause dispreference for a definite description, but only when the failing interpretation constitutes a red herring.
We investigate how children form early grammatical generalizations using the test case of the English regular plural. While some previous research provides evidence that children apply abstract rules to produce novel plurals well before 24mo., other studies have revealed that children use plural forms inconsistently with familiar and novel nouns, and demonstrate limited or variable receptive plural knowledge through (at least) 36mo. This is at odds with typical trajectories in child language development, where receptive knowledge precedes expressive. However, previous studies have not studied both receptive and expressive knowledge in the same children; additionally, differences in experimental materials across studies limit interpretability. In a cross-sectional design we tested 128 24-36-month-olds on two complementary experimental tasks: a receptive (eyetracking) task to evaluate children's understanding of plurals and an expressive (storybook) task to test their plural production. In the former, children heard sentences directing their gaze to an onscreen plural or singular target. In the latter, they heard a singular object labeled, and a prompt eliciting their plural production. We manipulated both novelty (novel vs. familiar object words, e.g., cats vs. wugs) and phonological form (/s/ vs. /z/ plurals, e.g., cats vs. dogs). We found strong, age-related evidence of expressive knowledge of the plural, but much more limited evidence of receptive knowledge (which was not predicted by age). Performance on the expressive task only predicted performance on the receptive task when additional grammatical cues were included (e.g., "there are two wugs" vs. "can you find the wugs?"). This work highlights the complexity of emerging grammatical generalizations in language acquisition and emphasizes the role of redundant grammatical cues in processing.
We present OneStop Eye Movements, a large-scale corpus of eye movements in reading, in which native (L1) speakers read newswire texts in English and answer reading comprehension questions. OneStop has 152 hours of eye movement recordings from 360 participants for 2.6 million word tokens, more data than all the existing public broad coverage English L1 eye tracking datasets combined. The eye movement data was collected for extensively piloted reading comprehension materials comprising 486 reading comprehension questions and auxiliary text annotations geared towards behavioral analyses of reading comprehension. Importantly, OneStop includes multiple reading regimes: ordinary reading, information seeking, repeated reading of the same text, and reading simplified text. The combination of the unprecedented size, high-quality reading comprehension materials and multiple reading scenarios, aims to enable new research avenues in the study of reading and human language processing. It further aims to facilitate the integration of eye tracking data in Natural Language Processing (NLP), Artificial Intelligence (AI), Human Computer Interaction (HCI) and educational applications.
Sometimes sentences sound acceptable when they are ungrammatical or semantically implausible. In this article, we study "comparative illusion" (CI) sentences where people often rate a sentence like More people have been to Russia than I have to be acceptable while in fact it is semantically anomalous. We provide a potential explanation for this language illusion from the noisy-channel framework. We hypothesize that comprehenders make rational inferences over the perceived sentence by entertaining alternative "close" plausible interpretations, where closeness is determined by possible production errors. In four experiments, (a) we identified a linguistic construction that elicits a salient CI illusion effect, (b) we established a range of plausible interpretations of the CI sentence, and (c) we found that the probability for comprehenders to assign a certain plausible interpretation to the CI sentence is proportional to how likely they think that interpretation is to be produced as the CI sentence during noisy language communication. This work contributes to a growing body of literature supporting rational noisy-channel inference during language comprehension. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
Human spoken language uses a continuous stream of acoustic signals to communicate about continuous features of the world, by using discrete forms — words — that segment the world into categories. Here we investigate how discreteness (the segmentation of a continuous signal space into discrete forms) and systematicity (the consistent alignment of these forms with what they refer to in the world) can emerge under communicative pressure. In an exploratory study, participants were paired with one another and played a game in which they varied the pitch of auditory signals to communicate about a continuous color space, generalizing from a small, shared set of signal-color pairings. The emergent systems exhibited both discreteness and systematicity, but only systematicity robustly predicted successful communication. These findings offer insight into the cognitive strategies that could support the creation and evolution of language, highlighting how pressures for effective communication can shape continuous signal spaces into structured, learnable systems.
Two prominent, yet contrasting, theoretical views are available to characterize the underlying drivers of language evolution: on the one hand, task-specific utility maximization; on the other hand, task-agnostic communicative efficiency. The latter has recently been grounded in an information-theoretic tradeoff between communicative complexity and informativeness, known as the Information Bottleneck (IB) principle. Here, we integrate these two views and propose an information-constrained emergent communication framework that trades off utility, informativeness, and complexity. To train agents within our framework, we develop a method, called Vector-Quantized Variational Information Bottleneck (VQ-VIB), that allows agents to interact using information-constrained discrete communication embedded in a continuous vector space. We test this approach in three domains and show that pressure for informativeness facilitates faster learning and better generalization to novel domains. At the same time, limiting complexity yields better alignment with actual human languages. Lastly, we find that VQ-VIB outperforms previously proposed emergent communication methods; we posit that this is due to the semantically-meaningful communication embedding space that VQ-VIB affords. Overall, our work demonstrates the role of cognitively-motivated optimality principles in inducing aspects of human-like communication among artificial agents.
The ability to build and reason about models of the world is essential for situated language understanding. But evaluating world modeling capabilities in modern AI systems-especially those based on language models-has proven challenging, in large part because of the difficulty of disentangling conceptual knowledge about the world from knowledge of surface co-occurrence statistics. This paper presents Elements of World Knowledge (EWOK), a framework for evaluating language models' understanding of the conceptual knowledge underlying world modeling. EWOK targets specific concepts from multiple knowledge domains known to be important for world modeling in humans, from social interactions (help, deceive) to spatial relations (left, right). Objects, agents, and locations in the items can be flexibly filled in, enabling easy generation of multiple controlled datasets. We then introduce EWOK-CORE-1.0, a dataset of 4,374 items covering 11 world knowledge domains. We evaluate 20 open-weights large language models (1.3B-70B parameters) and compare them with human performance. All tested models perform worse than humans, with results varying drastically across domains. Performance on social interactions and social properties was highest and performance on physical relations and spatial relations was lowest. Overall, this dataset highlights simple cases where even large models struggle and presents rich avenues for targeted research on LLM world modeling capabilities.
During real-time language comprehension, our minds rapidly decode complex meanings from sequences of words. The difficulty of doing so is known to be related to words' contextual predictability, but what cognitive processes do these predictability effects reflect? In one view, predictability effects reflect facilitation due to anticipatory processing of words that are predictable from context. This view predicts a linear effect of predictability on processing demand. In another view, predictability effects reflect the costs of probabilistic inference over sentence interpretations. This view predicts either a logarithmic or a superlogarithmic effect of predictability on processing demand, depending on whether it assumes pressures toward a uniform distribution of information over time. The empirical record is currently mixed. Here we revisit this question at scale: we analyze six reading datasets, estimate next-word probabilities with diverse statistical language models, and model reading times using recent advances in nonlinear regression. Results support a logarithmic effect of word predictability on processing difficulty, which favors probabilistic inference as a key component of human language processing.
We report an experiment eliciting ordering preferences for BINOMIAL EXPRESSIONS (e.g. bread and butter vs. butter and bread) in order to investigate the respective influences of productive and item-specific knowledge in language processing. Binomial ordering preferences reflect both (i) productive constraints involving phonological, semantic, and lexical properties, and (ii) item-specific relative frequencies. Bayesian and exemplar-based computational models of acquisition and use predict influences of both productive and item-specific knowledge on ordering preferences, with item-specific knowledge playing a smaller role the lower the expression's overall frequency. Our results confirm this prediction, but also reveal a role of item-specific knowledge even for binomials with overall frequency less than one in ten million. These findings bring a quantitative perspective to the debate over the roles of productive and item-specific knowledge in language.