
In this era of rapid artificial intelligence (AI) expansion, computational approaches are reshaping methods for language documentation and description. We survey the history of computational methods that have been applied to research in languages with limited digital resources and also present cutting-edge methods, such as large language models (LLMs), that have the potential to benefit documentary and descriptive fieldwork. We highlight how these methods affect data collection and annotation, transcription and phonological analysis, morphosyntactic description, and translation. Linguists, natural language processing engineers, and speech communities must consider how the use of computational methods such as data mining and machine learning should influence ethical best practices in linguistic field methods and how communities can continue to guide the documentation and maintenance of their languages in the age of AI. Looking forward, LLMs and making computational methods broadly usable through user interfaces are likely to emerge as prominent themes in documentary and descriptive research.
We focus on studying large language models (LLMs) and their ability to successfully interpret African American Language (AAL) in ways that do not distort the communicative intent of the speakers or writers. We discuss research that quantifies the gap between how well LLMs interpret AAL and how well they interpret White Mainstream English as well as research identifying the causes of this difference. This difference has an impact both on users of applications developed using LLMs and on application performance when interpreting AAL. Given this impact, researchers have begun developing approaches to mitigate these differences. We close by discussing the role of culture in speaker interactions and the difficulties in encoding culture in LLMs.
While studies of gestures generally focus on their cospeech use, gestures can sometimes be produced on their own, with a dedicated time slot. Such “pure gestures” have yielded three key findings. First, their iconic content can be divided “on the fly” among traditional slots of the inferential typology of language. This finding extends to novel gestures, visual animations, and (iconically modulated) sign language classifiers, which suggests that the inferential typology is mostly derived through productive rather than lexical means. Second, some pure gestures have a grammar reminiscent of some constructions of sign languages (despite the fact that the latter are not gestural systems). While this could suggest that Universal Grammar specifies certain form-to-function mappings, an alternative is that these mappings have deeper cognitive roots. Finally, pure gestures have a dedicated iconic syntax, which resembles that of sign language classifiers. Overall, these results highlight the importance of gestures for theoretical linguistics and the fruitfulness of multimodal investigations that include gestures and signs.
This autobiographical article describes key elements of Manfred Bierwisch's scientific work and life. It was completed only after his death and contains parts that he wrote himself, together with excerpts from other texts and interviews. The article outlines the main features of his linguistic work and in addition sheds light on the difficult political circumstances under which it was accomplished.
Relative clauses can range over degrees, just as they can range over individuals. But this class of degree relatives is rarely studied uniformly across the diverse constructions they help form. As a result, there have been a number of construction-specific proposals to model the semantics of degree relatives: in equatives, amount relatives, and wh -exclamatives. The goals of this article are to review a standard semantics of relativization writ large, to supplement it with standard assumptions from degree semantics, and to explain how the previously strange behavior of a variety of constructions formed from degree relatives comes out as a natural consequence of a handful of straightforward assumptions. Specifically, I argue that degree relatives are just relative clauses that range over degrees, and that these degree readings of relative clauses are available wherever we have ( a ) the appropriate morphology (e.g., a relativizer that can range over degrees), ( b ) a context of utterance that makes salient some informative (monotonic) dimension of measurement, and ( c ) an individual referent that is associated with a single, determinate measure along that dimension in the context of utterance.
Many linguists choose their discipline because of childhood experiences with languages. I had grown up using American Sign Language and had spent my childhood with Deaf people, but I had no models of how to become a linguist working with sign languages. This area of linguistics was just coming into being, and my models were other scientists who were just as new to this field as I was. This is not a complete account of my work over the course of my career, nor of all the mentors and colleagues who have touched my life and influenced my work. Instead, it is a selective account that traces a life of science built during a period of profound cultural changes in the lives of Deaf people and their sign languages.
Equative sentences like Hesperus is Phosphorus are interesting not only because they convey information about identity but also because of the atypical linguistic means they employ to do so. An equative sentence appears to be composed of a pair of referential singular terms (grammatical arguments) and to lack a logical predicate. Yet in the absence of such a predicate, how can equatives constitute well-formed, meaningful sentences? This rarely acknowledged problem of symmetry is assumed to have an obvious predicativist solution, according to which equatives would actually contain some predicative constituent(s). We expose the problem of symmetry, critically review various predicativist solutions, and outline an alternative hypothesis according to which equatives are distinctive both syntactically, by virtue of their symmetrical [DP DP] structure, and semantically, by virtue of their meaning corresponding to an instruction to perform a mental act of identification rather than to a traditional predication (the ascription of a property or relation).
In this review, I describe two research programs in the study of social meaning, both of which feature prominently in current-day sociolinguistics. The first one—social meaning as reasoning—seeks to understand the reasoning processes involved in the use and interpretation of sociolinguistic variants. The second one—social meaning as link between language and power — seeks to understand the ways in which language is related to power and the reproduction of social structure. I argue that these research programs are intimately linked: A proper understanding of the social relations at play in interactions is required for a satisfactory model of social meaning as a cognitive reasoning phenomenon, and, conversely, a proper understanding of how people use and understand language is crucial to understanding its role in the creation and reproduction of social inequalities. I sketch out a model that combines these two perspectives, inspired by work in game-theoretic pragmatics.
Allocutive markers (AMs) (i.e., markers that encode addressee features in syntax) and honorifics (which encode the social relation in syntax) raise questions about the interface between morphosyntax and discourse. On the AM side, the founding literature suggests that AMs are unembeddable. However, recent studies reveal that AMs are freely available in embedded contexts in some languages. This raises questions about where exactly discourse participants are represented in syntactic structures. On the honorific side, unlike traditional phi-features (number, person, gender), honorific features are relational and dynamic and encode the social relations between the speaker, the addressee, and potentially third persons as well. This raises the question of how a discourse-sensitive, relational feature can be formally encoded. Further issues include the morphosyntactic nature of AMs—the honorifics’ interaction within the verbal and nominal domains. We present theoretical and crosslinguistic advances that have been made on these issues as well as suggestions concerning where future research in this area could go.
Common ground is the information that the participants in a conversation treat as background information for the purposes of their interaction. We review two traditions of research on common ground. The formal tradition, consisting mainly of theoretical linguists and philosophers of language, has developed increasingly sophisticated formal models of common ground to generate predictions about an expanding range of empirical phenomena. Meanwhile, the psycholinguistic tradition has focused on a narrower range of phenomena while developing more realistic theories of the psychological mechanisms that allow us to select and represent common ground. After summarizing these two traditions, we consider several reasons why they should be reintegrated, and we argue that the best way to bring them back together would be to adopt a cognitive-pluralist approach, whereby language users have access to a variety of mechanisms for managing background information, which are more or less available and efficient depending on the communicative situation and the kind of information mentally represented as well as the cognitive demands of each mechanism.
This article examines the relationship between natural language and causal cognition and argues that linguistic expressions both reflect and constrain the ways in which humans conceptualize and represent causation. It opens with a survey of foundational philosophical theories of causation, focusing on the tension between metaphysical accounts and judgment-based approaches. The discussion then turns to a range of linguistic phenomena—including causative constructions, conditional sentences, discourse coherence, aspectual interpretation, and argument structure—demonstrating how causal relations are systematically encoded across grammatical domains. Building on insights from linguistics, philosophy, and cognitive science, the article reviews recent developments of a semantic framework for modeling causal knowledge in language. Rather than assuming a uniform mapping between causal relations and their linguistic expressions, the framework accounts for systematic variation in how causality is selected, structured, and communicated. In doing so, it positions natural language as a key source of evidence for understanding the architecture of causal reasoning and its representation in human cognition.
Minimalist research on syntactic dependencies over the past two decades has sought a unified understanding of the behavior of ϕ-agreement (such as subject–verb agreement) and of more traditionally studied long-distance dependencies, such as wh -movement. Both have been attributed to a single abstract operation, Agree. In this review, I discuss proposed constraints on Agree-based dependencies arising from the structures in which Agree operates, from the features and feature structures in terms of which Agree is stated, and from the specifications of participants to Agree (probes and goals) that control fine-grained aspects of how features are transferred (interaction) and when probing halts (satisfaction). Relevant theoretical concepts include cyclic structure building, minimality, phases, and feature geometries.
A key function of language is to enable concept construction, but empirically disentangling the contribution of language from the contribution of other (e.g., sensory) experiences is challenging. Comparing visual knowledge across sighted people, people born blind, and artificial intelligence (AI) trained exclusively on text provides rare insight. Blind people acquire rich visual knowledge, including normative meanings of light emission and visual perception verbs; similarity of colors; and size, shape, and texture of distal objects. Evidence from text-trained AI models suggests that such visual knowledge can be derived from language. Going beyond meanings of single words, blind people also construct causal intuitive theories of color, light, and visual perception, enabling generative inferences about visual phenomena (e.g., inferring the likelihood that a sighted character will see an object at a distance and the likely number of colors for a given artifact). Language enables concept construction from the ground up, without sensory evidence, and is intimately linked to intuitive theories.
In speech, information is conveyed using words, syntactic structures, and prosody to distinguish new information from discourse-given or inferable information and link the meanings of words and phrases to discourse antecedents. This review examines the role of prosody in communicating information structure. The point of departure is seminal work on information structure developed with reference to English, and corresponding work laying the foundation for current approaches to prosody as a phonological phenomenon. Corpus and experimental studies are reviewed for evidence that ( a ) speakers produce prosody in relation to focus and givenness and ( b ) listeners perceive and interpret prosodic cues to information structure meaning. Empirical findings show qualified evidence that listeners attend to prosody in processing and comprehending speech, though production data clearly show a many-to-many correspondence between form and meaning. A proposal that bridges these findings relates phonological and/or phonetic prominence scales to scales over information structure.
Human skills in the form of data annotation and generation have long been the cornerstone of building and testing machine learning algorithms. With the advent of large language models (LLMs) and their unsupervised learning on large quantities of data, the role of human data experts has been called into question. This article examines the changing role of human data annotation and linguistic analysis and provides an overview of past, current, and possible future paths. We discuss the targeted use of annotated data for fine-tuning and for feedback and reward models to supplement and improve LLMs, as well as the use of human linguistic skills for low-resource languages and for creating multimodal data sets, such as those required for human–robot interaction.
Sensory language has fascinated researchers, as it is here that meaning most clearly straddles biology and culture. Since the seminal work on color vocabularies, typologists have attempted to describe and explain worldwide linguistic patterns in this area. One proposal that has captured the attention of many is the idea that there is a hierarchy of the senses that can account for phenomena as diverse as lexicalization patterns, frequency of use, diachronic stability, order of acquisition, and more. We argue, to the contrary, that a universal sensory hierarchy is no longer tenable. Emerging crosslinguistic data do not support a single hierarchy of the senses that is applicable across distinct linguistic and psycholinguistic properties. This does not mean we must abandon sensory language typology altogether. Alternative methods for identifying crosslinguistic regularities—such as semantic maps—show considerable promise. Moreover, while the preoccupation with a sensory hierarchy has led to an overly narrow research focus, moving beyond it opens up new avenues of research in this area. Future research has potential to uncover new patterns of sensory language structure and use across diverse languages, account for their distribution over cultures, and deepen our understanding of how language interfaces with cognition.
This article provides an overview of variability in natural language by looking at the limits of intraindividual, interindividual, and cross-linguistic variability at all levels of grammar. We review evidence for the hypothesis that variability forms an integral part of natural language and often provides valuable insights into speakers’ linguistic competence. We discuss in turn ( a ) different subtypes of variability; ( b ) the difference between systematic linguistic variability, as driven by grammar or processing-related factors, and random noise, including performance errors; ( c ) the phenomenon of hidden variability, where linguistic expressions can differ in their underlying structure despite parallel surface strings; ( d ) the extent to which variability is accounted for in existing grammatical theories; ( e ) variability in experimental quantitative data as a window into the underlying grammatical systems and processing mechanisms; and ( f ) variability in language as an important trigger for diachronic change and language acquisition.
The study of scalar meanings or intensification has focused primarily on morphological means, yet there are many spoken languages where these concepts are expressed systematically by iconic prosody. Languages employ a combination of prosodic cues, including increased duration, raised pitch, special pitch patterns, and special voice quality, to signal scalar increases of property concepts, quantity, exhaustivity, duration, and so forth. In some languages, attitudinal meanings may also be expressed. Various labels have been used to refer to these iconic prosodic processes; below, the term prosodic intensification is used. This crosslinguistic overview looks at prosodic intensification from several angles: its phonetic realization (and orthographic representation), its meanings, its target domains, its iconic properties, and its status within each language's system (grammar or pragmatics?). It is shown that prosodic intensification is common not only in lesser-known languages but also in spoken and/or informal written registers of well-known languages and that this phenomenon is likely underreported. It is suggested that the underreporting of prosodic intensification, as well as researchers’ reluctance to treat its functions as part of grammar, is due to a persistent scholarly bias toward morphosyntactic over prosodic means.
Although research on typing has not exactly been sparse, studying typing within a psycholinguistic framework has not been a common approach. This paper argues in favor of this practice. By reviewing findings on patterns of typing errors and statistical learning in typed production, as well as influences of various factors on typing, including the similarity between the target word and its context, we show that typing has much in common with other modalities of language production and should be viewed as reflecting the general architecture of the language production system. We then discuss some of the contributions of typing research to the action monitoring literature due to the unique position that typing occupies at the interaction of phonological, orthographic, visual, and motor processes. We end by encouraging greater integration of typing research into psycholinguistic frameworks, not simply to confirm the predictions of such theories but to break new frontiers and push for new domains of inquiry.
Protactile language, a tactile language that has emerged within the DeafBlind community in the United States over the past two decades, challenges conventional assumptions about language, modality, and communication. Originating from a grassroots movement to center touch as a valid epistemology, protactile has developed distinct linguistic structures grounded in contact space, reciprocity, and embodied intersubjectivity. This article reviews the emergence and linguistic development of protactile, highlighting key structural innovations and areas of ongoing research. We discuss how protactile reconfigures foundational concepts such as phonology and interactional structure. We also present ethical considerations involved in studying a community-based language—emerging or otherwise—and emphasize the need for research practices grounded in collaboration and accountability. By centering the tactile experience and DeafBlind lived experiences, protactile contributes to a broader understanding of human language, how it functions, and how it emerges within diverse sensory and cultural ecologies.