Language acquisition involves drawing systematic generalizations from messy data. On one hypothesis, this is facilitated by a domain-general bias for children to "regularize" their input, sharpening the statistical distributions in their input towards more systematic extremes. We introduce a general computational framework for modeling a different explanation: on this view, children expect that their data are a noisy realization of a restrictive underlying grammatical system. We implement a learner that evaluates a choice among composite context-free grammars, in which a restricted set of "core" rules, comprising the particular grammatical processes that the learner is currently trying to acquire, operate alongside a less restricted set of "noise" rules, representing other independent processes that have yet to be learned, and conspire to introduce distortions into the data. Our Noisy Grammar Learner partitions its data into portions that serve as evidence for one of the possible core grammars in its hypothesis space, and portions generated by these noise processes. It does so without knowing in advance how much noise occurs or what its properties are. We compare our learner to a common implementation of the general regularization bias approach, and show that both can account for children's behavior in a representative artificial language learning experiment. However, we find that only our approach succeeds on two naturalistic case studies in early syntax acquisition: learning the rules governing canonical word-order and case-marking, given natural language data with "noise" from non-canonical sentence types. We show that our learner succeeds because its architecture allows a natural way to express linguistically-motivated expectations about the character of those rules. This suggests that, in certain domains, successful learning from messy data may be enabled by a hypothesis space comprising restrictive grammatical options.
A central goal of linguistic theory is to find a precise characterization of the notion "possible human language", in the form of a computational device that is capable of describing all and only the languages that can be acquired by a typically developing human child. The success of recent large language models (LLMs) in NLP applications arguably raises the possibility that LLMs might be computational devices that meet this goal. This would only be the case if, in addition to succeeding in learning human languages, LLMs struggle to learn "impossible" human languages. Kallini et al. (2024) conducted experiments aiming to test this by trainingGPT-2 on a variety of synthetic languages, and found that it learns some more successfully than others. They present these asymmetries as support for the idea that LLMs' inductive biases align with what is regarded as "possible" for human languages, but the most significant comparison has a confound that makes this conclusion unwarranted.
In this chapter we outline some computationally-significant dimensions of variation in the syntactic patterns found in natural languages. The relevant patterns are, more specifically, best thought of as particular configurations of elements that are related by syntactic dependencies. We begin by carefully introducing the key notion of a syntactic dependency via some abstract examples in section 2, and showing how it relates to constituency. Against this backdrop, section 3 introduces the idea that many “linguistically interesting” configurations of syntactic dependencies can be described in terms of discontinuous constituency. This notion provides a descriptive framework for the core of the chapter, section 4, where we present four dimensions along which analyses set in this framework can be classified. These dimensions of variation are computationally significant in the sense that they bear on the question of what kinds of computational machinery or formal system is needed to appropriately describe linguistic phenomena; we mention these consequences briefly and provide relevant references, but our focus here is squarely on the critical generalizations themselves rather than the implicated formal system. Finally, in section 5, we briefly describe an alternative to the discontinuous-constituency perspective that we adopt through much of the chapter, which provides a basis for linking the issues we have discussed to a distinct tradition of grammatical frameworks.
This paper proposes a formal model of regular languages enriched with unbounded copying. We augment finite-state machinery with the ability to recognize copied strings by adding an unbounded memory buffer with a restricted form of first-in-first-out storage. The newly introduced computational device, finite-state buffered machines (FS-BMs), characterizes the class of regular languages and languages de-rived from them through a primitive copying operation. We name this language class regular copying languages (RCLs). We prove a pumping lemma and examine the closure properties of this language class. As suggested by previous literature (Gazdar and Pullum 1985, p.278), regular copying languages should approach the correct characteriza-tion of natural language word sets.
AbstractIn this concluding chapter of the handbook, each contributor has written a 500 word mini-essay presenting their view of the future of experimental syntax, from the important theoretical questions on the horizon, to the methodological challenges that experimental syntacticians need to solve to answer those questions. The hope is that this final chapter will serve both as concrete inspiration for future studies in experimental syntax and as a benchmark for measuring the success of the field in the years to come.
A prominent strand of theorizing in linguistics models mean- ing in language by specifying an “interpretation function” which relates morphosyntactic objects (i.e., those representa- tions whose properties are uncovered by research in morphol-ogy and syntax) to elements of non-linguistic experience. Such theorizing has, for the most part, proceeded in relative isola-tion from developments in the other cognitive sciences. A re- cent body of experimental work growing out of this tradition has, however, pressed the question of precisely how linguistic representations relate to other faculties of mind. We present the beginnings of a two-step formal proposal for how to do this, specifying: (i) the co-domain of the linguistic interpretation function as the Language of Thought (LoT); (ii) what this mental language is like, (iii) which expression of this language is the semantic value of a sentence like ‘Most of the dots are yellow’, and (iv) how that LoT expression is interpreted by other cognitive faculties, in ways that produce the choices of verification procedure that have been empirically observed.
Children acquire their language’s canonical word order from data that contains a messy mixture of canonical and non-canonical clause types. We model this as noise-tolerant learning of grammars that deterministically produce a single word order. In simulations on English and French, our model successfully separates signal from the noise introduced by non-canonical clause types, in order to identify that both languages are SVO. No such preference for the target word order emerges from a comparison model which operates with a fully-gradient hypothesis space and an explicit numerical regularization bias. This provides an alternative general mechanism for regularization in various learning domains, whereby ten-dencies to regularize emerge from a learner’s expectation that the data are a noisy realization of a deterministic underlying system.
Abstract This chapter provides an overview of some ways to bring evidence from experimental language-processing work to bear on theories of syntax, where the relevant syntactic theories are formulated as explicit, self-contained formal grammars. The focus is on linking hypotheses that take the form of complexity metrics, at differing levels of abstraction. At one level are information-theoretic metrics such as surprisal and entropy reduction, where a grammar’s role is to define a probability distribution over generated expressions; a key point emphasized here is the way a grammar’s generative mechanisms constrain the shapes of these distributions. At a lower, explicitly algorithmic, level of abstraction are metrics based on the memory load incurred by particular parsing algorithms. The concluding section outlines applications of these ideas to modern syntactic theory.
Natural languages like English connect pronunciations with meanings. Linguistic pronunciations can be described in ways that relate them to our motor system (e.g., to the movement of our lips and tongue). But how do linguistic meanings relate to our nonlinguistic cognitive systems? As a case study, we defend an explicit proposal about the meaning of most by comparing it to the closely related more: whereas more expresses a comparison between two independent subsets, most expresses a subset-superset comparison. Six experiments with adults and children demonstrate that these subtle differences between their meanings influence how participants organize and interrogate their visual world. In otherwise identical situations, changing the word from most to more affects preferences for picture-sentence matching (experiments 1-2), scene creation (experiments 3-4), memory for visual features (experiment 5), and accuracy on speeded truth judgments (experiment 6). These effects support the idea that the meanings of more and most are mental representations that provide detailed instructions to conceptual systems.
This paper contributes to a body of research using parasitic gaps (PGs; Engdahl 1983, Culicover and Postal 2001, a.o.) to explore the nature of movement (Nissenbaum 2000, Legate 2003, Overfelt 2015b, Bondarenko and Davis 2019, Davis 2020, a.o.). In particular, this paper examines the implications of a particular variety of PG for the hypothesis that movement paths often involve several successive-cyclic steps:
Natural languages like English connect pronunciations with meanings. Linguistic pronunciations can be described in ways that relate them to our motor system (e.g., to the movement of our lips and tongue). But how do linguistic meanings relate to our nonlinguistic cognitive systems? As a case study, we defend an explicit proposal about the meaning of most by comparing it to the closely related more : whereas more expresses a comparison between two independent subsets, most expresses a subset–superset comparison. Six experiments with adults and children demonstrate that these subtle differences between their meanings influence how participants organize and interrogate their visual world. In otherwise identical situations, changing the word from most to more affects preferences for picture–sentence matching (experiments 1–2), scene creation (experiments 3–4), memory for visual features (experiment 5), and accuracy on speeded truth judgments (experiment 6). These effects support the idea that the meanings of more and most are mental representations that provide detailed instructions to conceptual systems.
The classification of grammars that became known as the Chomsky hierarchy was an exploration of what kinds of regularities could arise from grammars that had various conditions imposed on their structure. Intersubstitutability is closely related to the way different levels on the Chomsky hierarchy correspond to different kinds of memory. This chapter deals with the general concept of a string-rewriting grammar, which provides the setting in which the Chomsky hierarchy can be formulated. An unrestricted rewriting grammar works with a specified set of nonterminal symbols, and specified set of terminal symbols. From the very outset there were doubts about about whether context-free grammars (CFGs) could form the basis of a theory of natural language syntax. The grammar's rewrite rules correspond to the automaton's transitions. Chomsky argued that even if the generative capacity of CFGs turned out to be sufficient for English, the resulting grammars would be unreasonably complex.
A central constraint on minimalist derivations is cyclicity, as instantiated in the extension condition (Ext), which permits trees to grow only at (or near) the root. A second key idea concerns the notion of derivational state: how much information about the derivational past can the applicability of a certain operation be contingent on? Conditions along the lines of Phase Impenetrability and Shortest Move, for example, can sometimes limit the amount of information that conditions subsequent derivational steps; some versions of such conditions result in a finite bound (Fin). Interestingly, Ext and Fin define points of variation across mildly context-senstive grammar formalisms, systems that have been proposed to characterize the computational structure of syntactic derivations. On the one hand, Minimalist Grammars (MGs; Stabler, 2011) abide by Ext, making use of the standard minimalist operations of Merge and Move, as do Combinatory Categorial Grammars (CCGs; Steedman, 1996), whose derivations involve bottom-up concatenation of lexical items, and the closely related Linear Indexed Grammars (LIGs; Gazdar, 1988). In contrast, the adjoining operation of Tree Adjoining Grammars (TAGs; Joshi and Schabes, 1997) allows trees to grow “in the middle”, contra Ext. Turning to Fin, both MG and TAG operate with bounded derivational state, whereas CCG, with its unboundedly large categories, and LIG, with its stack-valued non-terminals, permit unbounded state. These relationships are summarized in the table in (1).
Aravind Joshi famously hypothesized that natural language syntax was characterized (in part) by mildly context-sensitive generative power. Subsequent work in mathematical linguistics over the past three decades has revealed surprising convergences among a wide variety of grammatical formalisms, all of which can be said to be mildly context-sensitive. But this convergence is not absolute. Not all mildly context-sensitive formalisms can generate exactly the same stringsets (i.e. they are not all weakly equivalent), and even when two formalisms can both generate a certain stringset, there might be differences in the structural descriptions they use to do so. It has generally been difficult to find cases where such differences in structural descriptions can be pinpointed in a way that allows linguistic considerations to be brought to bear on choices between formalisms, but in this paper we present one such case. The empirical pattern of interest involves wh-movement dependencies in languages that do not enforce the wh-island constraint. This pattern draws attention to two related dimensions of variation among formalisms: whether structures grow monotonically from one end to another, and whether structure-building operations are conditioned by only a finite amount of derivational state. From this perspective, we show that one class of formalisms generates the crucial empirical pattern using structures that align with mainstream syntactic analysis, and another class can only generate that same string pattern in a linguistically unnatural way. This is particularly interesting given that (i) the structurally-inadequate formalisms are strictly more powerful than the structurally-adequate ones from the perspective of weak generative capacity, and (ii) the formalism based on derivational operations that appear on the surface to align most closely with the mechanisms adopted in contemporary work in syntactic theory (merge and move) are the formalisms that fail to align with the analyses proposed in that work when the phenomenon is considered in full generality.
This paper has two closely related aims. The main aim is to lay out one specific way in which the derivational aspects of a grammatical theory can contribute to the cognitive claims made by that theory, to demonstrate that it is not only a theory's posited representations that testable cognitive hypotheses derive from. This requires, however, an understanding of grammatical derivations that initially appears somewhat unnatural in the context of modern generative syntax. The second aim is to argue that this impression is misleading: certain accidents of the way our theories developed over the decades have led to a situation that makes it artificially difficult to apply the understanding of derivations that I adopt to modern generative grammar. Comparisons with other derivational formalisms and with earlier generative grammars serve to clarify the question of how derivational systems can, in general, constitute hypotheses about mental phenomena.
Recent psycholinguistic evidence suggests that human parsing of moved elements is ‘active’, and perhaps even ‘hyper-active’: it seems that a leftward-moved object is related to a verbal position rapidly, perhaps even before the transitivity information associated with the verb is available to the listener. This paper presents a formal, sound and complete parser for Minimalist Grammars whose search space contains branching points that we can identify as the locus of the decision to perform this kind of active gap-finding. This brings formal models of parsing into closer contact with recent psycholinguistic theorizing than was previously possible.
This paper makes two related but distinct claims concerning the relationship between islandhood and the clausal ellipsis construction known as stripping. The first claim is that (at least a certain version of) this construction is island insensitive: no unacceptability results from having a correlate inside an island. This claim is supported by evidence from a formal acceptability judgment study. The second claim concerns the question of how to best account for this phenomenon of island- insensitivity in stripping: we claim that this island-insensitivity is best explained via the notion of island-repair, i.e., the ellipsis site involves the structure of island yet the ellipsis operation ameliorates island violations as opposed to the alternatives that have been dubbed evasion approaches. By this we mean that the island-insensitivity cannot be explained by positing a smaller, non-island structure in the ellipsis site; while this approach does of course explain the lack of an island effect, we show that it is incompatible with other facts about the crucial example sentences. If we instead assume that movement out of an island is grammatical if the island is properly contained inside a clausal ellipsis site, then positing a complete island structure inside the ellipsis site can explain all the properties of these crucial examples.
Much recent research in experimental psycholinguistics revolves around the resolution of long-distance dependencies, and the manner in which the human sentence processor “retrieves’” elements from earlier in a sentence that must be related in some way to the material currently being processed. At present there is no obvious way for the issues raised by this research to be framed in terms of an MG parser. Stabler’s 2013 top-down MG parser does not involve any corresponding notion of “retrieval’”: it requires that a phrase’s position in the derivation tree be completely identified before the phrase can be scanned, which means that a filler cannot be scanned without committing to a particular location for its corresponding gap. This chapter attempts to develop a parsing algorithm that is inspired by Stabler, but which allows a sentence-initial filler to be scanned immediately while delaying the choice of corresponding gap position.
Language is a sub-component of human cognition. One important, though often unattained goal for both cognitive scientists and linguists is to explicate how the meanings of words and sentences relate to the more general, non-linguistic, cognitive systems that are used to evaluate whether sentences are true or false. In the present paper, we explore one such relationship: an interface between the linguistic structures referring to individuals and non-individuals (specifically, count-nouns like ‘cows’ and mass-nouns like ‘beef’) and the non-linguistic cognitive systems that quantify and compare number and area. While humans may be flexible in how they use language across contexts, in two experiments using standard psychophysical testing we find that participants evaluate a count-noun sentence via numerical representations and evaluate a corresponding mass-noun sentence via non-numerical representations; consistent with a principled interface between language and cognition for evaluating these terms. This was the case even when the visual display was held constant across conditions and only the noun type was varied, further suggesting an important difference in how area and number, as well as count and mass nouns, are represented. These findings speak to issues concerning the semantics-cognition interface, the mass-count distinction, and the psychophysics of quantity representation.
Kotek et al. (Nat Lang Semant 23: 119–156, 2015) argue on the basis of novel experimental evidence that sentences like ‘Most of the dots are blue’ are ambiguous, i.e. have two distinct truth conditions. Kotek et al. furthermore suggest that when their results are taken together with those of earlier work by Lidz et al. (Nat Lang Semant 19: 227–256, 2011), the overall picture that emerges casts doubt on the conclusions that Lidz et al. drew from their earlier results. We disagree with this characterization of the relationship between the two studies. Our main aim in this reply is to clarify the relationship as we see it. In our view, Kotek et al.’s central claims are simply logically independent of those of Lidz et al.: the former concern which truth condition(s) a certain kind of sentence has, while the latter concern the procedures that speakers choose for the purposes of determining whether a particular truth condition is satisfied in various scenes. The appearance of a conflict between the two studies stems from inattention to the distinction between questions about truth conditions and questions about verification procedures.