A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind. In this work, we study reading dynamics for ``noisy-channel garden-path'' sentences, which temporarily appear well-formed but feature late-appearing violations of expectation that can be resolved not by inferring an alternative syntactic structure, but by inferring the presence of an error. We find evidence for targeted regressions -- eye movements towards regions that are promising loci of possible errors in light of later-arriving information, showing patterns consistent with the posterior inferences of a model of noisy-channel processing with reanalysis. We discuss the implications of these findings for theories of noisy-channel language comprehension and information-theoretic explanations of reading dynamics.
The syntax of human languages has long been argued to be complex and even unlearnable from the input alone. However, the success of large language models (LLMs) has challenged this idea. I argue for a simple view of syntax, where the syntax of a language is just the set of dependency rules, with no phrase structure or transformation rules—constructs central to Chomsky’s transformational grammar. This approach accounts for diverse phenomena in human language processing and explains crosslinguistic word order universals. Moreover, it better explains human data for cases that differentiate these accounts and eliminates the syntax learnability problem. I speculate that LLMs, similar to children, learn the dependency grammar from linguistic patterns, leading to their impressive syntactic competence.
Real-life language comprehension frequently requires non-literal interpretation and inferences about speaker intent. What is the structure of these so-called pragmatic abilities? We applied a dimensionality reduction approach to a large behavioral dataset (776 participants, each completing an 8-hour battery of diverse non-literal comprehension tasks). By examining covariation in performance across tasks, we identified three interpretable components of pragmatic language use: adherence to social conventions, extracting meaning from intonation, and causal reasoning based on world knowledge. Thus, pragmatic language use is relatively low-dimensional cognitively, and its distinct components may a) draw on dissociable neural substrates, b) exhibit distinct developmental trajectories and differential susceptibility to genetic brain disorders, and c) be variably challenging for artificial intelligence systems.
A sentence like "The authors that no critics recommended have ever received acknowledgment for a best-selling novel" is sometimes rated as acceptable even though, strictly speaking, it is ungrammatical because the negative polarity word "ever" is not licensed where it is. This behavioral effect is sometimes called a "negative polarity illusion". Here we propose that the lossy context surprisal theory of Hahn et al. (2022) – whereby people have an imperfect encoding of complex sentences – might explain this effect. We hypothesize that people have poor memory representation of the determiners in the main-clause and embedded-clause subjects and could entertain a determiner exchange that licenses ever. We propose that more similar determiners in those positions would trigger stronger illusion effects. Acceptability judgment tasks with six novel determiner pairs (e.g., "few" and "many", "few" and "most") support our proposal, showing, specifically, that a novel sentence, "Many authors that few critics recommended have ever received acknowledgment for a best-selling novel", triggered a much stronger illusion than the canonical one even without time pressure. These results offer further support for the suggestion that human language processing is imperfect and resource-rational: in face of working memory limitations, humans rationally reconstruct what is most likely from noisy linguistic input to facilitate downstream processing.
Intercomprehension refers to partial intelligibility of an unfamiliar language (L2) by a speaker of a related language (L1). How is this zero-shot cross-language comprehension possible? In this work, we extend past work on algorithmic models of noisy-channel inference to model intercomprehension in a Bayesian framework. The model uses an LM in L1 only for scoring latent hypotheses about the translations of observed L2 utterances, and a general-purpose noise model to infer a mapping between L2 and L1 words based on either form-based similarity or symbolic rules. We then conduct a human behavioral experiment, eliciting inferences for utterances in Dutch, Italian, and Ukrainian from speakers of English, Spanish, and Russian, respectively. Our full model shows a closer alignment to the distribution of human intercomprehension performance than ablations, and also compares favorably to zero-shot prompting of much larger models. These results provide a cognitively plausible computational model of intercomprehension, and highlight the flexible inferences made by comprehenders under wide uncertainty in real-world cross-language scenarios. We share our code publicly.
Conversation is a dynamic, multi-modal activity involving the exchange of complex streams of information like words, prosody, gesture, eye contact, and backchannels. Understanding how these different channels interact in naturalistic scenarios is essential for understanding the mechanisms governing human communication.Past studies suggested that the duration of words is tied to their predictability in context, but it remains unclear whether this relationship is speaker-oriented (e.g. retrieval or production-based) or due to listener-oriented, intelligibility-based pressures (i.e. emphasizing unpredictable words to ease comprehension).This study aims to examine the relationship between predictability and additional streams of speaker and listener behavior, to test how much intelligibility-oriented principles impact conversation.We use the GPT-2 large language model to assess the relationship between surprisal, a measure of unpredictability, and several variables known to play an important role in conversation --- the prosodic features of duration, pitch, and intensity, and the timing of listener backchannels. We perform this analysis on the CANDOR corpus of naturalistic spoken video call conversation between strangers in English.In keeping with previous results using n-gram predictability, we find that GPT-2 surprisal predicts significantly higher values for duration. Moreover, surprisal also predicts maximum pitch and maximum intensity even when controlling for duration. Additionally, listener backchannels were more likely to overlap high-surprisal words compared to low-surprisal words, suggesting that listeners provide verbal feedback and acknowledgement of unpredictable, i.e. informative, words.The results provide additional support for intelligibility-based accounts, which hold that language production is sensitive to a pressure for successful communication, not just speaker-oriented pressures.
Human language use is robust to errors: comprehenders can and do mentally correct utterances that are implausible or anomalous. How are humans able to solve these problems in real time, picking out alternatives from an unbounded space of options using limited cognitive resources? And can language models trained on next-word prediction for typical language be augmented to handle language anomalies in a human-like way? Using a language model as a prior and an error model to encode likelihoods, we use Sequential Monte Carlo with optional rejuvenation to perform incremental and approximate probabilistic inference over intended sentences and production errors. We demonstrate that the model captures previously established patterns in human sentence processing, and that a trade-off between human-like noisy-channel inferences and computational resources falls out of this model. From a psycholinguistic perspective, our results offer a candidate algorithmic model of rational inference in language processing. From an NLP perspective, our results showcase how to elicit human-like noisy-channel inference behavior from a relatively small LLM while controlling the amount of computation available during inference. Our model is implemented in the Gen.jl probabilistic programming language, and our code is available at https://github.com/thomashikaru/noisy_channel_model .
The noisy channel language comprehension proposal posits that comprehenders detect and correct errors when interpreting sentences. This study replicates and extends Zhan et al. (2023), testing the model in Mandarin Chinese with three syntactic alternations: (1) Active-Passive-BA sentences, (2) Double Object (DO)-Initial position Prepositional Object (PO)-Final position PO sentences, and (3) Transitive-Initial position Adverbial Intransitive-Final position Adverbial Intransitive sentences. In each alternation, the first two structures were adopted from Zhan et al. (2023), while the third was introduced in this study. These alternations require different numbers and types of edits to transform implausible sentences into plausible ones. Participants read test items and answer corresponding comprehension questions, which indicate whether they interpret the item literally. The results aligned with Zhan et al. (2023)'s findings, indicating that Mandarin participants were most likely to make inferences for implausible sentences resulting from deleting or inserting a single morpheme, followed by those formed by a noun phrase exchange across a function word, and least likely to make inferences for implausible sentences obtained through a noun phrase exchange across a main verb. The inclusion of novel structures reinforces the robustness of the noisy-channel framework and highlights how language-specific properties influence language comprehension.
Language comprehension relies on integrating the perceived utterance with prior expectations. Previous investigations of expectations about sentence structure (the structural prior) have found that comprehenders often interpret rare constructions nonliterally. However, this work has mostly relied on analytic languages like English, where word order is the main way to indicate syntactic relations in the sentence. This raises the possibility that the structural prior over word order is not a universal part of the sentence processing toolkit, but rather a tool acquired only by speakers of languages where word order has special importance as the main source of syntactic information in the sentence. Moving away from English to make conclusions about more general cognitive strategies (Blasi et al., 2022), we investigate whether the structural prior over word order is a part of language processing more universally using Hindi and Russian, synthetic languages with flexible word order. We conducted two studies in Hindi (Ns = 50, 57, the latter preregistered) and three studies with the same materials, translated, in Russian (Ns = 50, 100, 100, all preregistered), manipulating plausibility and structural frequency. Structural frequency was manipulated by comparing simple clauses with the canonical word order (subject-object-verb in Hindi, subject-verb-object in Russian) to ones with a noncanonical (low frequency) word order (object-subject-verb in Hindi, object-verb-subject in Russian). We found that noncanonical sentences were interpreted nonliterally more often than canonical sentences, even though we used flexible-word-order languages. We conclude that the structural prior over word order is always evaluated in language processing, regardless of language type. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Individuals with "agrammatic" receptive aphasia have long been known to rely on semantic plausibility rather than syntactic cues when interpreting sentences. In contrast to early interpretations of this pattern as indicative of a deficit in syntactic knowledge, a recent proposal views agrammatic comprehension as a case of "noisy-channel" language processing with an increased expectation of noise in the input relative to healthy adults. Here, we investigate the nature of the noise model in aphasia and whether it is adapted to the statistics of the environment. We first replicate findings that a) healthy adults (N = 40) make inferences about the intended meaning of a sentence by weighing the prior probability of an intended sentence against the likelihood of a noise corruption and b) their estimate of the probability of noise increases when there are more errors in the input (manipulated via exposure sentences). We then extend prior findings that adults with chronic post-stroke aphasia (N = 28) and healthy age-matched adults (N = 19) similarly engage in noisy-channel inference during comprehension. We use a hierarchical latent mixture modeling approach to account for the fact that rates of guessing are likely to differ between healthy controls and individuals with aphasia and capture individual differences in the tendency to make inferences. We show that individuals with aphasia are more likely than healthy controls to draw noisy-channel inferences when interpreting semantically implausible sentences, even when group differences in the tendency to guess are accounted for. While healthy adults rapidly adapt their inference rates to an increase in noise in their input, whether individuals with aphasia do the same remains equivocal. Further investigation of comprehension through a noisy-channel lens holds promise for a parsimonious understanding of language processing in aphasia and may suggest potential avenues for treatment.
The factors that affect the acceptability of long-distance extractions have long been debated, with multiple accounts proposed. Liu et al. (2022) proposed a succinct probability-based account of a sub-class of these kinds of materials, wh-questions with long-distance dependencies across sentence-complement verbs (e.g., “What did Mary whine that John bought?”). The explanation that they proposed was that the acceptability of such sentences depends on the probability of the verb-frame of the intermediate verb (e.g., “whine that”). In the current work, we evaluate some potentially simpler probability-based accounts on Liu et al.'s original data set, and show how an alternative (but also probability-based) approach accounts for the data better. We replicate their experiment and conduct the same analysis on the new dataset, finding the same results. Finally, we apply the same analysis to wh-questions with predicate adjectives (e.g., “What was Mary glad that John bought?”), and again find similar results. We conclude that the acceptability of such constructions is higher the more probable the words and constructions that make up the sentence are.
Conversation is a dynamic, multimodal activity involving the exchange of complex streams of information like words, prosody, gesture, eye contact, and backchannels. Understanding how these different channels interact in naturalistic scenarios is essential for understanding the mechanisms governing human communication. Past studies suggested that the duration of words is tied to their predictability in context, but it remains unclear whether this relationship is speaker-oriented (e.g., retrieval or production-based) or due to listener-oriented, intelligibility-based pressures (i.e., emphasizing unpredictable words to ease comprehension). This study aims to examine the relationship between predictability and additional acoustic variables, to test how much intelligibility-oriented principles impact conversation. We use the GPT-2 large language model to assess the relationship between surprisal, a measure of unpredictability, and several variables known to play an important role in conversation-the prosodic features of duration, intensity, and pitch. We perform this analysis on the CANDOR corpus of naturalistic spoken video call conversation between strangers in English. In keeping with previous results using n-gram predictability, we find that GPT-2 surprisal predicts significantly higher values for duration. Moreover, surprisal also predicts maximum pitch and pitch range even when controlling for duration, with mixed evidence for an effect of surprisal on intensity. Additionally, we investigated listener backchannels (short interjections like "yeah" or "mhm") and found that listener backchannels tended to be accompanied and followed by a spike in the surprisal of speakers' words. Finally, we demonstrate a divergence between the effect of context window size on the model fit of surprisal to maximum pitch and to other variables. The results provide additional support for intelligibility-based accounts, which hold that language production is sensitive to a pressure for successful communication, not just speaker-oriented pressures. Our data and analysis code are shared: https://osf.io/sqpn6/?view_only=e4d9e36c68b54863bc781e359463e1fe.
Successful communication requires frequent inferences. In a large-scale individual-differences investigation, we searched for dissociable components in the ability to make such inferences, commonly referred to as pragmatic language ability. In Experiment 1, n=376 participants each completed an 8-hour behavioral battery of 18 diverse pragmatic tasks in English. Controlling for IQ, an exploratory factor analysis revealed three clusters, which can be post-hoc interpreted as corresponding to i) understanding social conventions (critical for phenomena like indirect requests and irony), ii) interpreting emotional and contrastive intonation patterns, and iii) making causal inferences based on world knowledge. This tripartite structure largely replicated a) in a new sample of n=400 participants (Experiment 2), which additionally ensured that the intonation cluster is not an artifact of the auditory presentation modality, and b) when applying Bayesian factor analysis to the entire dataset. This research uncovers important structure in the toolkit underlying human communication and can inform our understanding of pragmatic difficulties in individuals with developmental and acquired brain disorders, and pragmatic successes and failures in neural network language models.
Sometimes sentences sound acceptable when they are ungrammatical or semantically implausible. In this article, we study "comparative illusion" (CI) sentences where people often rate a sentence like More people have been to Russia than I have to be acceptable while in fact it is semantically anomalous. We provide a potential explanation for this language illusion from the noisy-channel framework. We hypothesize that comprehenders make rational inferences over the perceived sentence by entertaining alternative "close" plausible interpretations, where closeness is determined by possible production errors. In four experiments, (a) we identified a linguistic construction that elicits a salient CI illusion effect, (b) we established a range of plausible interpretations of the CI sentence, and (c) we found that the probability for comprehenders to assign a certain plausible interpretation to the CI sentence is proportional to how likely they think that interpretation is to be produced as the CI sentence during noisy language communication. This work contributes to a growing body of literature supporting rational noisy-channel inference during language comprehension. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
What makes a word memorable? An important claim from past work is that words are encoded by their meanings and not their forms. If true, then, following rational analysis, memorable words should uniquely pick out a particular meaning, which means they should have few or no synonyms, and they should be unambiguous. Across two large-scale recognition-memory experiments (2,222 target words and > 600 participants each, plus 3,780 participants for the norming experiments), we found that memory performance is overall high, and some words are consistently remembered better than others. Critically, the most memorable words indeed have a one-to-one relationship with their meanings-with number of synonyms being a stronger contributor than number of meanings-and number of synonyms outperforms other predictors (such as imageability, frequency, or contextual diversity) of memorability that have been proposed in the past. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
How do comprehenders interpret semantically implausible sentences? Previous studies proposed a noisy-channel framework of sentence comprehension, where communication between a speaker and a comprehender happens in a noisy channel. The comprehender rationally adopts an interpretation of a sentence based on how likely the interpretation is (the semantic prior) and how likely is the interpretation corrupted into the perceived sentence because of noise (the likelihood). The theory predicted that comprehenders would be more likely to adopt a literal interpretation of an implausible sentence if their prior of implausible sentences were higher. To test this hypothesis, Gibson et al. manipulated the proportion of implausible test sentences in two sets of experiments, where participants read a number of sentences and answer a comprehension question following each sentence. Although their results supported the hypothesis, the experiment could be confounded (a) by participants' adaptation effect (due to different experiment lengths) and (b) by different participants having different strategies to do the task (due to the between-subject design). In our study, we manipulated the semantic prior and controlled for these potential confounds. We found participants exposed to more implausible sentences were indeed more likely to interpret implausible sentences literally. Our results hence offer additional support for the noisy-channel framework.
We report the results of two acceptability judgment experiments on English materials, which were designed in order to help disentangle predictions of syntactic theories with transformations from nontransformational theories. The materials in these experiments were motivated from examples from Pickering & Barry (1991), who provided intuitive evidence that there is little processing cost for connecting a fronted prepositional phrase to its verb, even if it is the second postverbal argument of a verb in the declarative form. For example, the PP on which connects to the verb put in the sentence This is the saucer on which Mary put the cup into which she poured the milk. If there is a transformation of phrases from declarative structures to interrogative structures (as proposed in Chomsky (1957) and all versions of related theories since), then there is a long-distance connection between the fronted PP and its base position following the NP object, for example, the cup into which she poured the milk, which is not complete until the end of the sentence. In contrast, in a theory without transformations, the PP can be directly associated with its role-assigning verb put when this verb is encountered. If there is cost for processing making dependency connections that is proportional to their distances, then transformational theories predict a large processing cost for this kind of structure, relative to controls. In contrast, nontransformational theories predict no large cost. The results of the two rating experiments consistently supported the predictions of the non-transformational theories relative to those of the transformational theories. We argue that, in line with other current evidence, the nontransformational theories appear to better support the available empirical data.