
Abstract All linguistic research has the potential to reproduce or challenge racial notions. —Linguistic Society of America Statement on Race (2019) The LSA Statement on Race stems from a larger conversation around undertheorized treatment of race and ethnicity in linguistics research and practice. In this commentary, we define racial identity and ethnicity and explain their relevance for linguistic research. We discuss considerations that linguistic researchers should take prior to research, during study design, and following research, and we offer specific recommendations when soliciting or using race and ethnicity data. These recommendations aim to help researchers avoid social harm, ensure ethical compliance and research integrity, and improve descriptive accuracy, especially for undersampled groups, by balancing research transparency with generalizability. We consider issues germane to collecting self-disclosures of ethnicity and racial identity in a range of study types spanning several subfields of linguistics. We give concrete examples of questions that may arise in planning studies in computational and corpus-based linguistics, formal linguistics, experimental linguistics, and qualitative linguistics. We speak to ethical considerations, including the importance of using locally constructed labels, analyst positionality, and respect for communities. Our goals are to provide linguistic researchers with a firmer basis for conceptualizing racial identity and ethnicity particularly as pertains to linguistics, and to supply a guide that aids linguists in reflecting on their own study design, positionality, and responsibility to participants and communities.
The effect of sound change and analogy upon inflectional paradigms has been traditionally described through Sturtevant's Paradox, which states that sound change is regular but generates irregularity, whereas analogy is irregular but generates regularity. While past work has explored trends in sound change and analogy qualitatively, quantitative investigation with large data sets remains underexploited. We tackle this by exploring the effects of sound change and analogy from Latin to French in large etymologically paired inflected lexicons containing the complete paradigms of 310 verbs with 11,593 total forms. We employ a novel method combining the automated application of historical sound changes and entropy-based quantitative analysis to examine separately the effects of sound change and analogy. The results confirm the role of some oft-cited predictors of analogy like token frequency and morphological regularity, but offer no support for others like markedness. Results also confirm the complexifying role of sound change, and the simplifying role of analogy, on aspects of morphological complexity like the number of inflection classes and the amount of allomorphy, but suggest that these forces have no comparable effect on more modern measures of complexity like average conditional entropies between inflected forms.
Inflectional morphology refers to the mapping from grammatical information to surface forms, which are typically realized as morphemes. This mapping often exhibits fusion, where several abstract features are expressed in a single morpheme that cannot be decomposed into meaningful parts. Here, we discuss crosslinguistic generalizations of morphological fusion. We argue that fusion reflects principles of efficient processing, as formalized by the memory-surprisal tradeoff (Hahn, Degen, & Futrell 2021), which is based on information-theoretic models of language processing from psycholinguistics. We first show that the existence of fusion itself can, in some situations, lead communicative codes to be more efficient under our processing model. Particularly, we reveal via simulation that the fusion of highly correlated features is more efficient for processing, whereas agglutination is more efficient when features are less correlated. We next discuss crosslinguistic patterns of fusion in real languages. First, we analyze well-known generalizations about features that are commonly fused across languages (e.g. tense, aspect, and mood), as well as a typological pattern regarding suppletion. In both cases, we find that the universals we study tend to reflect a tendency toward more efficient structure under our model of language processing. Finally, we use paradigm and frequency data from four languages to study informational fusion, a gradable measure of fusion defined in Rathi et al. 2021. We find that informational fusion is higher when features are highly correlated, which suggests that gradable fusion is also influenced by optimization for the memory-surprisal tradeoff.
This article introduces the phenomenon of discontinuous harmony, where the target and trigger of harmony are separated by intervening nonharmonizing words. We present a case study from Gu & eacute;bie (Kru; C & ocirc;te d'Ivoire) in which particle verbs are split via focus movement. Despite appearing at opposite edges of the clause, the verb controls harmony on the particle without affecting the vowels of intervening material. While discontinuous harmony would appear to violate locality, we offer an analysis that involves local harmony followed by syntactic movement that separates the trigger and target. This analysis thus relies on a cyclic interleaving of syntactic and phonological operations, where syntactic information persists through phonological evaluation and is available to later cycles of syntax.
This article introduces a novel linguistics pedagogy resource called the Language Profiles Project (LPP), an open access resource for linguistics instructors and students. The goal of the LPP is to make it easier and more engaging for instructors in North American institutions to incorporate underrepresented languages into undergraduate linguistics courses. More concretely, a language profile combines data sets for use in linguistics courses with contextual information about the language and culture. In this article, I describe the resource, the motivations for creating it, and some ways in which it could be used.
In acquiring a syntax, children must detect evidence for abstract structural dependencies that can be realized in variable ways in the surface forms of sentences. In What did David fix?, learners must identify a nonlocal relation between a fronted object of the verb (what) and the phonologically null 'gap' in canonical direct object position after the verb, where it is thematically interpreted. How do learners identify a nonadjacent dependency between an expression and something that has no overt phonological form? We propose that identifying abstract syntactic dependencies requires statistical inference over both overt linguistic material and unsatisfied grammatical expectations: noticing when a predicted argument for a verb is unexpectedly missing may serve as evidence for the gap of an argument movement dependency. We provide computational support for this hypothesis. We develop a learner that uses predicted but unexpectedly missing objects of verbs to identify possible gaps of object movement, and identifies which surface morphosyntactic properties of sentences are correlated with these possible movement gaps. We find that it is in principle possible for a learner using this mechanism to identify the majority of sentences with object movement in child-directed English, and that prior knowledge of which verbs require objects provides an important guide for identifying which surface distributions characterize object movement. This provides a computational account for why verb argument-structure knowledge developmentally precedes the acquisition of movement in a language like English. More broadly, these findings illustrate how statistical learning and learning from violated expectations can be combined to novel effect in the domain of language acquisition.
In many areas in linguistic study it is difficult to decide where the study of language ends and the study of other aspects of human cognition begins. In this article, we discuss a particularly striking case of this, the use of the signing space (loci) for marking linguistic relations. The use of loci in the nominal and verbal domains has received a wide range of analyses, from those considering loci to be abstract linguistic mechanisms such as semantic indices and syntactic agreement to those considering them to be making use of nonlinguistic mechanisms such as spatial cognition. We defend the view that the use of loci is both fundamentally linguistic (they are modifiers) and fundamentally spatial (they express an association with space), providing possible descriptive content in both the verbal and the nominal domain. This analysis allows for a uniform account of loci use in the two linguistic domains and accounts for an important, yet less noticed, property of loci, which is that their distribution is pragmatically conditioned for the purpose of disambiguation.
In light of the growing number of undergraduates from racially minoritized backgrounds at newly emergent Minority-Serving Institutions and other colleges and universities, linguists have a special responsibility to engage such students, particularly through projects that connect to students' linguistic and cultural backgrounds. This article describes undergraduates' learning experiences in a research collective committed to community-centered collaborative work to advance sociolinguistic justice for the Mexican Indigenous diasporic community in California. The discussion centers the voices of undergraduate team members to demonstrate the benefits of students' learning with respect to the research process, linguistics as a discipline, and understanding of self, family, and community.
The Phonomaton, a public web-facing facility, computes phonological derivations based on a user's underlying representations and rules. The tool allows a formal implementation of phonological analyses using familiar methods and lets students interactively explore the mechanics of feature systems and serial derivations. We demonstrate a number of the program's features and end with a discussion of its implementation in the classroom.
The pitch contours of Mandarin two-character words are generally understood as being shaped by lexical tones on the constituent single-character words, in interaction with articulatory constraints imposed by factors such as speech rate, coarticulation with adjacent tones, segmental makeup, and predictability. This study shows that tonal realization is also partially determined by words' meanings. We first show, on the basis of a corpus of Taiwan Mandarin spontaneous conversations, using a generalized additive regression model and focusing on the rise-fall tonal pattern, that after controlling for effects of speaker and context, word type is a stronger predictor of tonal realization than all of the previously established word-form-related predictors combined. Importantly, the addition of information about meaning in context improves prediction accuracy even further. We then proceed to show, using computational modeling with context-specific word embeddings, that token-specific pitch contours predict word type with 50% accuracy on held-out data, and that context-sensitive, token-specific embeddings can predict the shape of pitch contours with 40% accuracy. These accuracies, which are an order of magnitude above chance level, suggest that the relation between words' pitch contours and their meanings are sufficiently strong to be potentially functional for language users. The theoretical implications of these empirical findings are discussed.
Abstract: This paper responds to commentaries by several authors on our target article, ‘Bringing signed languages into the study of regular sound change’ (Law et al. 2025a). We provide some additional context on the research program that spurred the target article and draw on several themes discussed in both the target article and commentaries, specifically (i) the affordances and effects of modality in (theories of) language change, (ii) iconicity (and indexicality), (iii) variation and irregularity in language change, and (iv) language transmission. We highlight the methodological advances afforded by new technologies to study phonetic variation in signed languages and advocate for increased attention to the systematic and comparative study of phonetic variation in signed languages as a window into processes of phonetic and phonological change in signed languages.
Abstract: In response to the target article by Law, Power, and Quinto-Pozos, I argue that signed and spoken languages share a common core of phonological mental representations, consisting of events (points in time and/or space), features (monadic properties of events), and precedence (a dyadic relation of temporal order between events). In addition to this, signed languages also include dyadic spatial relations between events/points in space and time. Illustrations of the EFPS model for signed phonology are drawn from some simple ASL signs, and an ongoing diachronic change in ASL motorcycle is analyzed using parallel events, partial reduplication, and underspecification.