How do artificial neural networks bind concepts to form complex semantic structures? Here, we propose a simple neural code, whereby the existence and the type of relations between entities are represented by the distance and the direction between their embeddings, respectively. We test this hypothesis in a variety of Large Language Models (LLMs), each input with natural-language descriptions of minimalist tasks from five different domains: arithmetic, visual scenes, family trees, metro maps and social interactions. Results show that the true semantic structures can be linearly recovered with a Polar Probe targeting a subspace of LLMs' layer activations. Second, this code emerges mostly in middle layers and improves with LLM performance. Third, these Polar Probes successfully generalize to new entities and relation types, but degrades with the size of the semantic structure. Finally, the quality of the polar representation correlates with the LLM's ability to answer questions about the semantic structure. Together, these findings suggest that LLMs learn to build complex semantic structures by binding representations with a simple geometrical principle.
We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to span a wide range of phonemic contrasts. Given only 10 hours of speech in a previously unseen language, systems must produce discrete units that are mapped to a predefined phoneme inventory, through either a many-to-one or a one-to-one assignment. The resulting sequences are evaluated for unit quality, recognition and segmentation. We provide four pretrained multilingual HuBERT and SpidR baselines, and show that phonemic information is available enough in current models for derived units to correlate well with phonemes, though with variations across languages.
Are large language models (LLMs) sensitive to the distinction between humanly possible and impossible languages? This question was recently used in a broader debate on whether LLMs and humans share the same innate learning biases. Previous work has answered it in the positive by comparing LLM learning curves on existing language datasets and on "impossible" datasets derived from them via various perturbation functions. Using the same methodology, we examine this claim on a wider set of languages and impossible perturbations. We find that in most cases, GPT-2 learns each language and its impossible counterpart equally easily, in contrast to previous findings. We also apply a more lenient condition by testing whether GPT-2 provides any kind of separation between the whole sets of natural vs. impossible languages, based on cross-linguistic variance in metrics derived from the learning curves. Taken together, these perspectives show that GPT-2 provides no systematic separation between the possible and the impossible.
The waggle dance of bees has given rise to some of the most striking and detailed studies of animal communication. But because of its gradient character, the waggle dance has widely been taken to have properties that are wholly distinct from those of human language. We argue that this is mistaken, and that the waggle dance represents the oldest instantiation of an iconic system also found in human language, notably in sign language. The waggle dance helps bees locate a food source through four properties: (1) food distance is conveyed through the duration of the waggling phase; and (2) food direction is conveyed through the orientation of the waggle run. In addition, (3) while in bees that dance horizontally, the waggle run points towards the food source, in bees that dance vertically the information involves transposition: the angle of the dance relative to 'upwards' is interpreted as the angle of the food direction relative to the sun. Finally, (4) the number of waggle runs increases with food quality. We show that properties 1 and 2 are instantiated in sign language classifier predicates, highly iconic constructions that produce visual animations of the orientation and movement of an entity. Furthermore, classifier movement (property 3) can be interpreted either directly or with 'viewpoint shift', a more flexible version of transposition. Property 4 seems to be instantiated more generally in the pragmatics of human and animal communication, as repetition can convey intensification and/or excitement (e.g. Go, go, go!). We further show experimentally that properties 1-3 are instantiated in some gestures understood by non-signers. Thus the waggle dance is a primitive form of a semantic system also found (through convergent evolution) in human language. It is remarkably ancient, at least 20 million years old according to phylogenetic reconstructions. While the horizontal dance (without transposition) is usually thought to be ancestral, a closer look at extant phylogenies suggests that the vertical dance (with transposition) might be more primitive, and furthermore that pre-adaptations guarantee that transposition might have been available from the start.
We argue that female Diana monkeys (Cercopithecus diana) can form complex calls by combining an A call with other elementary calls. We reject (on both empirical and conceptual grounds) a combination-free analysis based on accidental homophony, and we consider two main analyses: the Acoustic Theory takes the combination to be merely acoustic, whereas the Affixal Theory takes A to function as a suffix. We provide limited arguments for the Affixal Theory, and through comparison with another closely related monkey species, we date these combinations to at least 6 million years ago.
Large language models have shown strong performance on broad-domain knowledge and reasoning benchmarks, but it remains unclear how well language models handle specialized animal-related knowledge under a unified closed-book evaluation protocol. We introduce BAGEL, a benchmark for evaluating animal knowledge expertise in language models. BAGEL is constructed from diverse scientific and reference sources, including bioRxiv, Global Biotic Interactions, Xeno-canto, and Wikipedia, using a combination of curated examples and automatically generated closed-book question-answer pairs. The benchmark covers multiple aspects of animal knowledge, including taxonomy, morphology, habitat, behavior, vocalization, geographic distribution, and species interactions. By focusing on closed-book evaluation, BAGEL measures animal-related knowledge of models without external retrieval at inference time. BAGEL further supports fine-grained analysis across source domains, taxonomic groups, and knowledge categories, enabling a more precise characterization of model strengths and systematic failure modes. Our benchmark provides a new testbed for studying domain-specific knowledge generalization in language models and for improving their reliability in biodiversity-related applications.
According to much of theoretical linguistics, a fair amount of our linguistic knowledge is innate. One of the best-known (and most contested) kinds of evidence for a large innate endowment is the argument from the poverty of the stimulus (APS). An APS obtains when human learners systematically make inductive leaps that are not warranted by the linguistic evidence. A weakness of the APS has been that it is very hard to assess what is warranted by the linguistic evidence. Current artificial neural networks appear to offer a handle on this challenge, and a growing literature has started to explore the potential implications of such models to questions of innateness. We focus on Wilcox, Futrell, and Levy’s (2024) use of several different networks to examine the available evidence as it pertains to wh-movement, including island constraints. WFL conclude that the (presumably linguistically neutral) networks acquire an adequate knowledge of wh-movement, thus undermining an APS in this domain. We examine the evidence further, looking in particular at parasitic gaps and across-the-board movement, and argue that current networks do not succeed in acquiring or even adequately approximating wh-movement from training corpora roughly the size of the linguistic input that children receive. We also show that the performance of one of the models improves considerably when the training data are artificially enriched with instances of parasitic gaps and across-the-board movement. This finding suggests, albeit tentatively, that the networks’ failure when trained on natural, unenriched corpora is due to the insufficient richness of the linguistic input, thus supporting the APS.
State-of-the-art neural networks can be trained to become remarkable solutions to many problems. But while these architectures can express symbolic, perfect solutions, trained models often arrive at approximations instead. We show that the choice of regularization method plays a crucial role: when trained on formal languages with standard regularization ($L_1$, $L_2$, or none), expressive architectures not only fail to converge to correct solutions but are actively pushed away from perfect initializations. In contrast, applying the Minimum Description Length (MDL) principle to balance model complexity with data fit provides a theoretically grounded regularization method. Using MDL, perfect solutions are selected over approximations, independently of the optimization algorithm. We propose that unlike existing regularization techniques, MDL introduces the appropriate inductive bias to effectively counteract overfitting and promote generalization.
Human languages balance communicative informativity with complexity, conveying as much as needed through the simplest means required to do so. Yet, these concepts—informativity and complexity—have been operationalized in various ways, and it remains unclear which definitions best capture empirical linguistic patterns. A particularly successful operationalization is that offered by the Information Bottleneck framework, which suggests a balance between complexity and informativity across domains like color, kinship, and number. However, we show that the notion of complexity employed by this framework has some counterintuitive consequences. Focusing on color terms, we then study to what extent this and other notions of complexity play a role in explaining cross‐linguistic regularity. We propose a method to assess their explanatory contributions; and to probe whether they enter in a joint optimization or in a trade‐off competition. This offers a more general framework to study language change and the forces that shape it, where instead of showing that a given model is compatible with existing data, the data is used to adjudicate between candidate measures.
Multiple members of the tit and chickadee (= Parid) family combine two classes of calls, F and D, in a rigid order FD. In Japanese tits, FD has been argued on the basis of multiple experiments to involve syntax and non-trivial compositionality. How ancient are these call combinations? We show that FD combinations (as well as individual F and D calls) are present in nearly all Parid species, and almost absent in their closest relatives, the Remizidae and Stenostiridae. Using phylogenetic tools and ancestral reconstruction methods, we infer that FD combinations very likely emerged between 11 and 26 million years ago in the eastern Himalayas. This result contributes to evolutionary animal linguistics using a comparative phylogenetic approach to reconstruct the evolution of call combinations.
Most Parid species produce specific, order-constrained mobbing calls. These calls elicit responses from both conspecifics and heterospecifics, with evidence indicating that such responses occur only when the calls are organised in this specific order. One notable exception is the coal tit (Periparus ater), a species that employs similar types of notes, yet does not exhibit clear order constraints within its mobbing sequences. Despite this apparent absence of order constraints, a recent experiment has demonstrated that coal tits may be sensitive to the order of notes in heterospecific calls. Therefore, the relative significance of note order in conspecific and heterospecific communication among coal tits remains unclear. We conducted a playback experiment to examine the effects of note order (natural coal tit order, typical Parid order and reversed order) and species identity (conspecific, familiar heterospecific-the great tit, Parus major, or artificial notes) on coal tit mobbing responses. Our findings indicate that coal tits exhibited a strong response to conspecific calls, regardless of the order of the notes; conversely, they displayed little to no response to heterospecific calls and artificial notes, irrespective of note order. A similar pattern was observed when assessing the general community response. This unexpectedly low response to familiar heterospecific calls may be attributable to a reduced density of great tits in the area we tested: ecological factors, such as community composition, may influence heterospecific mobbing behaviours and the subsequent biological interpretations of playback experiments. This study also underscores the necessity of conducting comparative research on closely related species to evaluate the potential generality of findings, such as strong order constraints recently observed in great tits and Japanese tits.
While recent “animal linguistics” treats call form as arbitrary, various results suggest that some animals use a biological code to understand the calls of unrelated/unfamiliar species. To clarify matters, we distinguish among three degrees of interspecies comprehension. In the first (“Understand thy neighbor”), a species understands the calls of a neighboring species through exposure. In the second (“call convergence”), it understands the calls of an unrelated/unfamiliar species through evolutionary convergence and resemblance to familiar calls. In the third degree (“featural interpretation”), it uses a rule associating a meaning to a specific acoustic feature—hence a new road to (featural) compositionality.
It takes several years for the developing brain of a baby to fully master word repetition-the task of hearing a word and repeating it aloud. Repeating a new word, such as from a new language, can be a challenging task also for adults. Additionally, brain damage, such as from a stroke, may lead to systematic speech errors with specific characteristics dependent on the location of the brain damage. Cognitive sciences suggest a model with various components for the different processing stages involved in word repetition. While some studies have begun to localize the corresponding regions in the brain, the neural mechanisms and how exactly the brain performs word repetition remain largely unknown. We propose to bridge the gap between the cognitive model of word repetition and neural mechanisms in the human brain by modeling the task using deep neural networks. Neural models are fully observable, allowing us to study the detailed mechanisms in their various substructures and make comparisons with human behavior and, ultimately, the brain. Here, we make first steps in this direction by: (1) training a large set of models to simulate the word repetition task; (2) creating a battery of tests to probe the models for known effects from behavioral studies in humans, and (3) simulating brain damage through ablation studies, where we systematically remove neurons from the model, and repeat the behavioral study to examine the resulting speech errors in the "patient" model. Our results show that neural models can mimic several effects known from human research, but might diverge in other aspects, highlighting both the potential and the challenges for future research aimed at developing human-like neural models.
Compositionality is a means of constructing complex objects through the transformation and combination of simpler elements. While it is common to view compositionality as inherently complex, and thus to assume that compositionality is a byproduct of advanced language expertise, we argue otherwise. We propose that, although compositionality produces complex outcomes, the underlying processes are simple and can often be reduced to the general mechanism of function application. Accordingly, we explore the origins of compositionality not only in compositional language but also, and at an earlier stage, in the development of compositional representations and thoughts in young infants. Infants correctly composed simple noun-verb sentences at 14 months, facial expressions with objects at 12 months, and mental physical transformations at 10 months. This offers evidence for function application, the essence of compositionality, in infancy—emerging well before and outside the development of compositional language. A series of three studies provide evidence in infants for function application, the essence of compositionality, emerging before and outside the development of compositional language.
The syntactic structures of sentences can be readily read-out from the activations of large language models (LLMs). However, the “structural probes” that have been developed to reveal this phenomenon are typically evaluated on an indiscriminate set of sentences. Consequently, it remains unclear whether structural and/or statistical factors systematically affect these syntactic representations. To address this issue, we conduct an in-depth analysis of structural probes on three controlled benchmarks. Our results are three-fold. First, structural probes are biased by a superficial property: the closer two words are in a sentence, the more likely structural probes will consider them as syntactically linked. Second, structural probes are challenged by linguistic properties: they poorly represent deep syntactic structures, and get interfered by interacting nouns or ungrammatical verb forms. Third, structural probes do not appear to be affected by the predictability of individual words. Overall, this work sheds light on the current challenges faced by structural probes. Providing a benchmark made of controlled stimuli to better evaluate their performance.
How did the very first meaning components arise in animals? We argue that answers interact in interesting ways with data on current and ancestral animal communication systems. Using standard notions of evolutionary stability in biology, we develop a simple framework to analyze the emergence of three meaning components: individual signals, nontrivial combinations, and pragmatic principles of competition among signals. We show that for elementary signals to arise, they should have null cost, or be understood from the start. While this conclusion dovetails with the traditional idea that signals often originate in cues, i.e., informative byproducts of non-communicative processes, the two scenarios (null cost vs. understanding from the start) can be distinguished in case studies involving ancestral meaning reconstruction. For nontrivial combinations of the form CC’ (such as pyow-hack sequences in putty-nosed monkeys and ABC-D sequences in Japanese tits), we show that their emergence is heavily constrained because they should initially give rise to some miscommunication, as CC’ could also be understood as the (trivial) combination of separate utterances C and C’. Finally, we investigate the evolution of two pragmatic principles that were posited in recent animal linguistics: the Informativity Principle and the Urgency Principle. We argue that both have a clear evolutionary path, especially if they start appearing in production, and then in comprehension. Overall, recent work in animal linguistics can be fruitfully combined with simple principles of evolutionary stability and with ancestral signal reconstruction to address in a precise fashion questions about the very first meaning operations in nature.
We consider the possible role of current large language models (LLMs) in the study of human linguistic cognition. We focus on the use of such models as proxies for theories of cognition that are relatively linguistically-neutral in their representations and learning but differ from current LLMs in key ways. We illustrate this potential use of LLMs as proxies for theories of cognition in the context of two kinds of questions: (a) whether the target theory accounts for the acquisition of a given pattern from a given corpus; and (b) whether the target theory makes a given typologically-attested pattern easier to acquire than another, typologically-unattested pattern. For each of the two questions we show, building on recent literature, how current LLMs can potentially be of help, but we note that at present this help is quite limited.
Originally formalized with symbolic representations, syntactic trees may also be effectively represented in the activations of large language models (LLMs). Indeed, a ''Structural Probe'' can find a subspace of neural activations, where syntactically-related words are relatively close to one-another. However, this syntactic code remains incomplete: the distance between the Structural Probe word embeddings can represent the \emph{existence} but not the type and direction of syntactic relations. Here, we hypothesize that syntactic relations are, in fact, coded by the relative direction between nearby embeddings. To test this hypothesis, we introduce a ''Polar Probe'' trained to read syntactic relations from both the distance and the direction between word embeddings. Our approach reveals three main findings. First, our Polar Probe successfully recovers the type and direction of syntactic relations, and substantially outperforms the Structural Probe by nearly two folds. Second, we confirm that this polar coordinate system exists in a low-dimensional subspace of the intermediate layers of many LLMs and becomes increasingly precise in the latest frontier models. Third, we demonstrate with a new benchmark that similar syntactic relations are coded similarly across the nested levels of syntactic trees. Overall, this work shows that LLMs spontaneously learn a geometry of neural activations that explicitly represents the main symbolic structures of linguistic theory.
Given a three-valued definition of validity, which choice of three-valued truth tables for the connectives can ensure that the resulting logic coincides exactly with classical logic? We give an answer to this question for the five monotonic consequence relations $st$ , $ss$ , $tt$ , $ss\cap tt$ , and $ts$ , when the connectives are negation, conjunction, and disjunction. For $ts$ and $ss\cap tt$ the answer is trivial (no scheme works), and for $ss$ and $tt$ it is straightforward (they are the collapsible schemes, in which the middle value acts like one of the classical values). For $st$ , the schemes in question are the Boolean normal schemes that are either monotonic or collapsible.
Unsupervised on-the-fly back-translation, in conjunction with multilingual pretraining, is the dominant method for unsupervised neural machine translation. Theoretically, however, the method should not work in general. We therefore conduct controlled experiments with artificial languages to determine what properties of languages make back-translation an effective training method, covering lexical, syntactic, and semantic properties. We find, contrary to popular belief, that (i) parallel word frequency distributions, (ii) partially shared vocabulary, and (iii) similar syntactic structure across languages are not sufficient to explain the success of back-translation. We show however that even crude semantic signal (similar lexical fields across languages) does improve alignment of two languages through back-translation. We conjecture that rich semantic dependencies, parallel across languages, are at the root of the success of unsupervised methods based on back-translation. Overall, the success of unsupervised machine translation was far from being analytically guaranteed. Instead, it is another proof that languages of the world share deep similarities, and we hope to show how to identify which of these similarities can serve the development of unsupervised, cross-linguistic tools.