Pronouns that precede their antecedents are called cataphors. Upon encountering a cataphor, comprehenders engage in an active, eager search for its antecedent later in the sentence. The processing of cataphoric dependencies has been used to probe comprehenders’ incremental expectations and how grammatical knowledge guides those expectations. Past experiments have investigated whether knowledge of Binding Principles B and C constrain comprehenders’ initial expectations for coreference or whether grammatical knowledge applies at a delay. Although studies suggest that the Binding Principles B and C influence the earliest stages of cataphor resolution, recent findings have called into question whether the main experimental paradigm used to test the time-course of constraint-sensitivity – self-paced reading – has the necessary temporal precision to adjudicate between competing accounts. In this paper we investigate the application of Principle B during cataphor processing using eye-tracking-while-reading, an experimental paradigm with more fine-grained temporal sensitivity. Our results converge with prior findings that Principle B strongly constrains active cataphor resolution. We end by discussing how to model Principle B sensitivity via predictions made at the level of the discourse representation.
Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to high-resource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages.
Surprisal theory posits that the processing difficulty of a word is determined by its predictability in context, offering a potential link between human sentence processing and next-word predictions from language models. While language model (LM) surprisals successfully predict reading times in naturalistic text, they systematically underpredict the magnitude of difficulty observed in controlled studies of syntactic ambiguity, particularly in garden path sentences. This mismatch might arise from differences in the computational constraints between humans and LMs. Here we test one such hypothesis, specifically, that LMs may be able to simultaneously consider a greater number of distinct sentence interpretations at once, compared to humans. Using Recurrent Neural Network Grammars (RNNGs) with word-synchronous beam search, we systematically vary the number of simultaneous parses used to compute word surprisal, and then use these surprisals to predict human reading times. Reducing the number of simultaneous active parses indeed increases the magnitude of predicted garden path effects, but not nearly enough to capture the full magnitude of the effects in humans. This suggests that differences in the number of simultaneous parses available to LMs and humans cannot reconcile LM-based surprisal with human sentence processing.
Real-time sentence comprehension relies on memory resources to establish long-distance syntactic dependencies. There is convincing evidence that formation of many dependencies is mediated by hierarchical constraints. However, it remains an open question how richly hierarchical information is represented and employed in memory processes. Here we ask whether hierarchical relations between noun phrases in a sentence as in x c-commands y (Reinhart, 1976) constrain antecedent retrieval for a local anaphor. We measure the real-time reactivation of c-commanding target antecedents and distractors in processing the Turkish reciprocal birbirleri via three visual world studies. Unlike existing studies that confounded multiple cues, we disentangled c-command from other structural information (clause-mateness, case, subjecthood) and linear order/recency. Experiment 1 compared the availability of c-commanding subjects and non-c-commanding, clause-mate, similarly case-marked distractors; and showed that c-commanding subjects were rapidly distinguished from distractors within the reciprocal window, irrespective of their linear order. Experiment 2 revealed that immediate availability of c-commanding subjects extends to c-commanding indirect objects. Experiment 3 was a pre-registered, high-power replication that yielded similar results. We found limited evidence for interference from distractors, which was not replicated. Overall, we find that hierarchical, item-to-item relations between noun phrases rapidly determine antecedent availability in retrieval beyond other cues. We suggest that hierarchical information may guide access to c-commanding items during retrieval if item representations include hierarchically informed features targeted by retrieval cues; alternatively, hierarchical information may shape the organization of items in memory, with c-commanding items represented in a privileged store allowing direct access during retrieval.
At first glance, the brain's language network appears to be universal, but languages clearly differ. Does the brain adapt to the specific details of individual grammatical systems? Here, we present a magnetoencephalography (MEG) study on case and agreement in Hindi and Nepali. Both languages use split-ergative case systems. However, these systems interact with verb agreement differently-in Hindi, case features conspire to determine which noun phrase (NP) the verb agrees with (subject, object, or neither), but in Nepali the verb always agrees with the subject NP. We found that NPs with different case values elicit different MEG signals around 200-500 and 600-900 ms. In subsequent exploratory analyses, we failed to find a reliable difference in this brain activity between the two languages corresponding to the different relations between case and agreement. However, we identified a portion of the left temporoparietal junction as exhibiting a statistically nonsignificant effect that may warrant further investigation.
Real-time language processing relies on a capacity-limited working memory system to encode and maintain (extra-)linguistic information. Prior research shows that structural information is rapidly used to distinguish targets from distractors in memory access, but the exact mechanisms underpinning this remain unclear. Here we consider two hypotheses. One is that structurally prominent, target representations are actively maintained in a privileged, focal attentional state that allows rapid, direct access. Another possibility is that structural prominence serves as a retrieval cue to distinguish target representations stored in a passive, non-focal state. These hypotheses can be distinguished by determining the relative time course of memory access to targets vs. distractors. To do this, we conducted a secondary analysis of a published dataset that measured the looking behavior during the online resolution of the structurally constrained Turkish local reciprocal birbirleri in a visual world paradigm. Bayesian analyses on gaze proportions and new fixations show evidence for reliable and faster access to targets relative to distractors. We interpret this as evidence for a functional divide within linguistic memory that is shaped by hierarchical syntactic structure, whereby structurally prominent target items are held in a privileged focal state, while distractors that lack this privileged status are accessed via a slower retrieval mechanism.
Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.
In the process of extracting a meaning from a text, our eyes linger much more on some words than others, and we often reread earlier portions of the text. These disruptions to the reading process are particularly common in syntactically ambiguous sentences. What explains the difficulty presented by these sentences? One prominent hypothesis explains it as a special case of the impact of a word's predictability (operationalized via surprisal) on the difficulty of processing the word. This contrasts with theories that attribute these disruptions to errors in the structure-building process. Earlier attempts to address this debate have been inconclusive because of small numbers of participants, coarse measurements of the reading process that are ill-suited to disentangling these competing views, and a limited range of surprisal estimates. Here, we conduct a large-scale study ([Formula: see text]) examining eye movements during the reading of syntactically challenging sentences, using 409 types of surprisal estimates from language models with multiple architectures and training settings. We find a stark dissociation: Early effects of syntactic disambiguation are well-approximated by language model surprisal, but syntactic disambiguation incurs a significant additional cost, reflected in an increase in rereading that is not explained by language model surprisal. We conclude that surprisal can capture routine structure-building, but not the cost of detecting or correcting errors in the structure-building process.
Sentence comprehension relies on encoding linguistic items in memory and accessing them subsequently to form linguistic dependencies. This makes processing susceptible to memory interference. Interference, such as the distortion of memory representations or access to irrelevant memory items, can lead to misinterpretation or grammatical errors. Over the years, research on agreement attraction has debated whether this hallmark of memory interference reflects limits of the retrieval mechanism, or inaccuracy of the encoded representations that retrieval targets. We present some evidence in favor of representational accounts of memory interference. Our findings include partial evidence for three kinds of representational effects: (a) the ungrammaticality illusion, a pattern by which attraction arises without misleading retrieval cues; (b) number errors rather than noun errors in final interpretation; and (c) mitigation of attraction when additional markers of the subject's number are available, which we label feature updating. Together, the findings seem to suggest that feature distortion in the content of memory representations contributes to attraction effects. We propose that models of memory mechanisms that mediate dependency formation should incorporate malleable representations rather than stable ones.
Sentence processing models posit that syntactic processing difficulty contributes to the processing time of each word (e.g., Lewis & Vasishth, 2005). On the other hand, the E-Z Reader 10 model of eye movements in reading (Reichle et al., 2009) proposes that syntactic processing affects eye movements only when integration of an input word fails. We present two eye movement experiments designed to address this tension. Readers were presented with ungrammatical sentences with a singular subject and a plural verb, but with a plural local noun that often induces an illusion of grammaticality, e.g., *The editor of the magazines are reviewing the articles before publication. After each critical sentence was removed from the screen, the reader made a binary acceptability judgment. When readers rejected the ungrammatical critical sentences, there was also pronounced disruption to incremental reading. However, on the majority of trials readers accepted these sentences, and on those trials the eye movement record revealed only a very small increase in reading time compared to corresponding grammatical sentences; critically, this small effect was entirely explained by the difference between conditions in the lexical surprisal of the verb. Thus, as predicted by E-Z Reader 10, an effect of agreement computation above and beyond the effect of surprisal was in evidence only when agreement computation failed, i.e., the sentence was rejected. We discuss implications for models of incremental sentence processing and reading.
Reading seems smooth and effortless, but this appearance is deceiving: In the process of extracting a meaning from a text, our eyes linger much more on some words than others, and we often reread earlier portions of the text. These disruptions are particularly common in syntactically challenging sentences. What explains this complex pattern of eye movements? One prominent hypothesis explains these patterns as a special case of the impact of a word's predictability (its surprisal) on the difficulty of recognizing the word. An alternative hypothesis attributes these disruptions to syntactic structure-building operations that apply after word recognition. Earlier attempts to address this debate have been inconclusive because of small numbers of participants, coarse measurements that are ill-suited to disentangling these competing views, and a limited range of predictability estimates. Here, we conduct a large-scale study (n = 368) examining eye movements during the reading of syntactically challenging sentences, using 407 types of predictability estimates from large language models with multiple architectures and training settings. We find a stark dissociation: Early reading measures are well approximated by language model surprisal, but syntactic disambiguation incurs a significant additional cost, reflected in an increase in rereading that is not explained by surprisal. We further show that when rereading parts of the sentence, readers strategically target the words most useful to amend the structure of the sentence. We conclude that linguistic knowledge guides moment-by-moment reading in two dissociable ways: Forward reading is driven by the word's surprisal, and backward reading reflects syntactic structure building operations.
Syntactic dependency formation in comprehension is subject to retrieval interference that occurs when comprehenders need to activate stored information in memory to form and interpret a linguistic dependency. For example, retrieving a subject phrase to attach it to the verb might result in agreement attraction errors. It remains unclear whether this interference arises as part of routine dependency formation or as part of a repair mechanism that is activated when predictive dependency formation fails (e.g., Wagers et al., 2009). For example, it has been argued that reflexive anaphors resist attraction in comprehension because number/gender features of unpredictable elements are not associated with a strong 'prediction error' signal that might trigger retrieval-based and error-prone repair processes (Parker & Phillips, 2017). We test a version of the "Error-driven Retrieval" hypothesis by examining the interaction between reflexive attraction and the predictability of the anaphor. In two reading time experiments and one offline interpretation experiment, we find that the predictability of a reflexive dependency does not modulate its susceptibility to interference effects in comprehension. We propose that attraction is better captured as part of routine retrieval processes and that the (in)sensitivity of reflexives to structurally irrelevant distractors should be explained through other mechanisms.
Intransitive verbs fall into two different syntactic classes, unergatives and unaccusatives. It has long been argued that verbs describing an agentive action are more likely to appear in an unergative syntax, and those describing a telic event to appear in an unaccusative syntax. However, recent work by Kim et al. (2024) found that human ratings for agentivity and telicity were a poor predictor of the syntactic behavior of intransitives. Here we revisit this question using interpretable dimensions, computed from seed words on opposite poles of the agentive and telic scales. Our findings support the link between unergativity/unaccusativity and agentivity/telicity, and demonstrate that using interpretable dimensions in conjunction with human judgments can offer valuable evidence for semantic properties that are not easily evaluated in rating tasks.
Eye tracking has been a popular methodology used to study the visual, cognitive, and linguistic processes underlying word recognition and sentence parsing during reading for several decades. However, the successful use of eye tracking requires researchers to make deliberate choices about how they apply this technique, and there is wide variability across labs and fields with respect to which choices are “standard.” We aim to provide an easy-to-reference guideline that can help new researchers with their entrée into eye-tracking-while-reading research. Because the standards do – and should – vary from field to field or study to study as is appropriate for the research question, we do not set a rigid recipe for handling eye tracking data, but rather provide a conceptual framework within which researchers can make informed decisions about how to treat their data so that it is most informative for their research question. Therefore, this paper provides a description of eye movements in reading and an overview of psycholinguistic research on the topic, an overview of experiment design considerations, a description of the data processing pipeline and important choice points and implications, an overview of common dependent measures and their calculation, and a summary of resources for data analysis.
As they process complex linguistic input, language comprehenders must maintain a mapping between lexical items (e.g., morphemes) and their syntactic position in the sentence. We propose a model of how these morpheme-position bindings are encoded, maintained, and reaccessed in working memory, based on working memory models such as "serial-order-in-a-box" and its SOB-Complex Span version. Like those models, our model of linguistic working memory derives a range of attested memory interference effects from the process of binding items to positions in working memory. We present simulation results capturing similarity-based interference as well as item distortion effects. Our model provides a unified account of these two major classes of interference effects in sentence processing, attributing both types of effects to an associative memory architecture underpinning linguistic computation.
This paper investigates whether agreement attraction is modulated by distributional properties determining subject-likelihood by examining the degree to which bare nouns and full determiner phrases (DPs) cause agreement attraction effects in Romanian. Romanian represents an ideal testing ground for this, given two distributional constraints making bare nouns less subject-like: Locative Determiner Omission, preventing locative prepositions from taking nouns with definite articles (unless modified by adjectives), and the Naked Noun Constraint, disallowing bare nouns as preverbal subjects. We predicted that bare nouns should be less likely to trigger agreement attraction than overt DPs. We conducted four speeded forced-choice sentence continua-tion tasks on Romanian native speakers to test this prediction. We observe that overt DPs cause significantly more attraction than bare nouns. We suggest that the results are consistent with a cue-based retrieval mecha-nism for forming agreement dependencies, where cues that determine subjecthood are used to reactivate elements in working memory upon processing a verb. These cues can be language specific, and in Romanian, this means that agreement attraction is sensitive to the morphophonological overtness of the determiner.
Prediction has been proposed as an overarching principle that explains human information processing in language and beyond. To what degree can processing difficulty in syntactically complex sentences – one of the major concerns of psycholinguistics – be explained by predictability, as estimated using computational language models, and operationalized as surprisal (negative log probability)? A precise, quantitative test of this question requires a much larger scale data collection effort than has been done in the past. We present the Syntactic Ambiguity Processing Benchmark, a dataset of self-paced reading times from 2000 participants, who read a diverse set of complex English sentences. This dataset makes it possible to measure processing difficulty associated with individual syntactic constructions, and even individual sentences, precisely enough to rigorously test the predictions of computational models of language comprehension. By estimating the function that relates surprisal to reading times from filler items included in the experiment, we find that the predictions of language models with two different architectures sharply diverge from the empirical reading time data, dramatically underpredicting processing difficulty, failing to predict relative difficulty among different syntactic ambiguous constructions, and only partially explaining item-wise variability. These findings suggest that next-word prediction is most likely insufficient on its own to explain human syntactic processing.