
Abstract Phoneme inventory size is related to speaker population size, but the cause of this relationship has remained a mystery. The correlation was first reported by Hay and Bauer (2007), who considered several possible causes of this relationship, including ascertainment bias, similarity between relatives, loss of phonemes in small, isolated populations, and maintenance of larger phoneme inventories in larger speaker populations. Here, we test these proposed causes of the relationship and show that the correlation between phoneme inventory size and speaker population size is largely driven by similarity between relatives and neighbours in both speaker population size and phoneme inventory size. However, we also show that phoneme inventory size is related to aspects of linguistic isolation—languages confined to islands or with few bordering languages tend to have lower phoneme inventory sizes than their relatives. We also explore the geographic bias of phoneme inventory databases and show that in some cases there is a positive bias towards regions with larger phoneme inventories. However, we find no evidence that less-studied languages tend to have lower reported phoneme inventory sizes. Our analyses demonstrate the importance of taking sources of non-independence and bias in typological data into account when conducting global analyses of linguistic diversity.
Languages have been claimed to globally disfavor ergatively-aligned case marking, possibly reflecting processing preferences. Phylogenetic research on Pama-Nyungan has cast doubt on this, showing a preference for retaining ergativity in the family. In this study, we re-examine the evolution of ergative alignment of case marking in a sample of over 200 languages from eleven language families (Pama-Nyungan among them) in a continuous-time Markov Chain model. Since one language can adopt multiple alignment patterns of case marking in its grammar ('split' alignment), we incorporate a third, split state into the models. In addition, we compare the goodness of fit across various models where clades can show different rates of change, defining evolutionary regimes that can cut across families. The results show that, across all families, the best-fitting model requires two connected regimes: one favoring non-ergative (S=A) and the other favoring split-ergative alignment (S=A and S not equal A). These differences notwithstanding, the stationary probabilities of the model suggest that in the long run, non-ergative alignments are expected to be much more frequent than ergative alignment in case marking.
Abstract Human languages are widely understood to be conventional: for a given meaning, a population of language users expect a certain form to be used. Given the lack of unconventional natural languages with which to compare, there is little direct empirical evidence which explains why linguistic conventionality matters. We carry out a series of emergent language experiments with artificial agents designed to exhibit the effect of conventionality on robustness, learnability, and cognitive efficiency. By experimentally manipulating the interaction dynamics through which agents align their vocabularies, we create conditions that either promote or prevent population-level convergence, resulting in conventional and unconventional emergent languages while holding communicative success constant. We find that conventional languages are indeed more robust, better learnable, and more efficient, when compared with less conventional languages.
Human communication is highly adaptable in the sense that, for instance, new words or combinations of words emerge when there is a communicative pressure. However, another way in which humans can communicate about novel experiences is through meaning extensions of words that are already in their lexicon. One of the key driving factors of such polysemic extension is metaphorical conceptual mappings, which first become usualized, and then conventionalized within a community. In this article, we set out to investigate whether this kind of conceptual mapping can be achieved in an experimental semiotic study and whether the usualization trajectories of concrete and metaphorical meanings differ from each other. We do so by engaging participants in a referential game in which dyads of participants have to communicate about concrete and metaphorical meanings. We find that, over the course of the experiment, participants adapt elements of their kinematics from concrete to metaphorical meanings. We also find that the usualization trajectory is different between the two groups of meanings: both meanings initially undergo reduction in terms of movement path. However, concrete meanings continue to change and become more efficient across all rounds, while metaphorical meanings stabilize earlier and do not change significantly after their initial kinematic changes. In contrast to concrete meanings, we find no significant change in terms of speed and volume for metaphorical meanings. The study provides evidence that both concrete and metaphorical meanings can be expressed through pantomime, but their trajectory of change is different.
This study tested, for the first time on a diverse cross-linguistic dataset, the existence of a negative correlation between a word’s length and its semantic ambiguity. This correlation can be predicted on various grounds: as a way to make communication more efficient, as a mechanical consequence of the greater availability of short wordforms, or as a corollary of Zipf’s laws of abbreviation and meaning, which state that frequent words are shorter and have more meanings. We tested the correlation between word length and number of meanings in a broad range of languages, our biggest analysis studying 633,308 wordforms in 1,952 languages representing 192 families. We operationalise word ambiguity in three ways: as whether or not a word colexifies several meanings, using Lexibank data, as the number of synsets in Wordnet, and using BERT-derived predictions of the degree of ambiguity of a word. For all three measures of ambiguity, we evidence a robust correlation between a word’s brevity and ambiguity. The correlation remains substantial when controlling for word frequency, suggesting that it is more than a by-product of the fact that frequent words tend to be both shorter and more ambiguous.
Linguistic phylogenies are commonly inferred from abstract cognate classifications that encode relationships among lexemes. Although widespread, this practice has well-recognized limitations: it discards the phylogenetic signal contained in segmental word forms; restricts the range of evolutionary questions that can be addressed; and treats cognacy judgments, which are hypotheses, as observed data. We introduce a comparative framework that addresses these limitations by modeling the evolution of aligned cognate word forms directly. Our approach adapts the TKF91 model of molecular evolution, originally developed to account for insertion and deletion events in DNA sequences, to the domain of linguistic data. By operating on segmental strings rather than abstract character codings, the framework enables phylogenetic inference from observable word forms and supports quantitative investigation of sound change. We demonstrate its utility through analyses that illuminate patterns of segmental stability and the evolution of phonological inventories.
We present an experimental communication game study designed to investigate the effects of noise, context, and time pressures on communication systems. This study is inspired by and follows on from an earlier study, designed to model the emergence of grammatical focus using a simple nonlinguistic communication game paradigm in which participants communicated line figures by selecting cells in a grid to send to each other. In our study, the communication task was more challenging, and a confound between noise and effort pressures was removed. We found that signals became reduced over time, though to a lesser extent given noise, which made communication significantly harder. Time pressures reduced the effect of noise. In general, signals became more stable over time, but noise interfered with the stabilization process. Overall, this study supports the findings of earlier work that employed a much simpler communication task and contributes to our understanding of how information structure gets into language.
Very early in life, from a few weeks to 5 months of age, human infants tend to focus on vowels over consonants to process words. In the present review, we discuss recent evidence suggesting that, just like young infants, nonhuman animals also tend to focus on vowels to identify sequences of sounds. This early use of vowels to recognize words might be linked to acoustic properties that are orthogonal to lexical processing. For example, vowels tend to be more salient than consonants, in the sense that vowels tend to be longer, more stable and are produced with more intensity than consonants. Thus, the recent data with nonhuman animals sheds a new light on this early stage of phoneme processing as the result of biological, and evolutionary relevant predispositions.
We argue that, as well as an empirical approach borrowed from experimental psychology, studies of language evolution can also benefit from an explorative, participatory approach. This is based on a reflection on an experimental semiotics study where the process of arriving at an effective experimental design was equally valuable for developing the theory as the final results of the experiment. We suspect that this process is commonplace in many studies, but there is no formal method for documenting or exploiting any insights gained. We present methods from video game design and ethnography as candidates for addressing this gap and suggest they can be used in a hybrid approach that combines an exploratory phase of cyclic iteration with a final, more traditional linear phase. We illustrate these methods with two case studies and argue that a participatory approach can harness the creative power of our participants and help us reveal important aspects of our theories.
It is unknown how rare sound changes appear and spread through the lexicon. The bilabial trills of Malekula are one such example of a rare sound, and the recent assembly of a lexical database for Vanuatu (the Vanuatu Voices database) affords a unique opportunity for a quantitative historical analysis. We built a linguistic phylogeny of Malekula languages, and performed phylogenetic ancestral state reconstruction of fourteen semantic values which exhibit bilabial trills to track their historical evolution. We found a surprising degree of dynamism, with trills spreading gradually through the lexicon and showing frequent losses and reappearances. Our results are consistent with frequent borrowing of trills between neighbouring languages. We suggest that the rapid dynamism and evidence for borrowing are explained by the low functional load and identity attachment of trills respectively.
We present the common task framework approach to testing causal theories about the evolution of language. There are now many theories about how symbolic communication emerged, but less work trying to compare, synthesize and test these theories. We suggest that the first step is to formalize the theories as causal graphs using tools from the field of causal inference. This helps recognize the critical causal links that differentiate theories. The second step is to use methods from lab-based experimental semiotics to specify a 'common task' or 'arena': an experimental environment and a task for individuals to complete. The different theories suggest different designs for this arena, and the experimental results can be used as a measure of the relative success of each theory. In this paper, we provide an example from anthropological theories of the emergence of symbolic communication, suggesting that an effective arena contains an asymmetry of information, division of labour and contextually distal meanings. We run experiments in arenas based on collaborative construction and fire maintenance. The results indicate that the effectiveness of pointing can limit the emergence of symbolic signals, a problem that has previously not been worked into theories. In this way, we hope the common task framework can be used as a method to further develop theories of language evolution.
There are more than 7,000 languages on our planet today, but they are not evenly distributed. There are over 100 languages in Vanuatu, but only one in Sāmoa. Why might this be? This paper explores this question for one particular region: Remote Oceania. Remote Oceania comprises the eastern Solomon Islands (Temotu), Vanuatu, New Caledonia, Fiji, Polynesia and Micronesia. The region features large differences in language richness, with some islands having 20 languages and many having only one. This paper explores one hypothesis as to why this might be: more levels of political complexity reduces language diversification. In this study, we evaluate the strength of this claim by modelling political complexity as a predictor of language richness together with other relevant factors such as time-depth, size of island, rainfall, etc. The results show that political complexity has a significant effect, but that it is not robust. Taking into account phylogenetic non-independence, in particular, reduces the effect, suggesting that there are relevant unaccounted for variables which are phylogenetically structured. The paper discusses further limitations of the study and possible expansions in the research area.
Human spoken language uses a continuous stream of acoustic signals to communicate about continuous features of the world, by using discrete forms — words — that segment the world into categories. Here we investigate how discreteness (the segmentation of a continuous signal space into discrete forms) and systematicity (the consistent alignment of these forms with what they refer to in the world) can emerge under communicative pressure. In an exploratory study, participants were paired with one another and played a game in which they varied the pitch of auditory signals to communicate about a continuous color space, generalizing from a small, shared set of signal-color pairings. The emergent systems exhibited both discreteness and systematicity, but only systematicity robustly predicted successful communication. These findings offer insight into the cognitive strategies that could support the creation and evolution of language, highlighting how pressures for effective communication can shape continuous signal spaces into structured, learnable systems.
The use of phylogenetic methods in linguistics has provided new insights into the structure, age, and spread of language families. Despite increasing recognition of Japonic as one of the world’s primary language families, research on the family’s phylogeny remains limited. This study presents a new reconstruction of Japonic language history based on NichiRyuuLex, a new lexical dataset comprising data from 48 Japanese and 33 Ryukyuan lects for 256 concepts. The study takes a novel approach, combining lexical and phonotactic data to increase precision in the phylogenetic parameter estimates, providing a more informative reconstruction. The analyses presented here confirm previous findings on the age of the family as a whole, estimating a Japanese-Ryukyuan split at around 500 BCE, supporting that the time of diversification coincides with the influx of Bronze Age rice agriculturists during the Yayoi Period. The topology of the mainland Japanese clade uncovered in the analyses unifies two divisions recognised in traditional Japanese dialectology (the East-West division, and the center-periphery division). The topology of the Ryukyuan clade largely followed geographical segmentation—Northern Ryukyuan (Amami; Okinawa) vs. Southern Ryukyuan (Miyako; Macro-Yaeyama). The ancestor of Ryukyuan was dated to around the 9th century, coinciding with the first traces of cereal farming in the Northern Ryukyu Islands. In sum, the study provides greater certainty about the linguistic history and internal structure of mainland Japanese, as well as a more detailed perspective on the history of the Ryukyuan languages.
This study performs primary data collection, transcription, and cognate coding for eight South West Tibetic languages (Lowa, Gyalsumdo, Nubri, Tsum, Yohlmo, Kagate, Jirel, and Sherpa). This includes partial cognate coding, which analyses linguistic relations at the morpheme level. Prior resources and inferences are leveraged to conduct a Bayesian phylogenetic analysis. This helps estimate the extent to which the historical relationships between the languages represent a tree-like structure. We argue that small-scale projects like this are critical to wider attempts to reconstruct the cultural evolutionary history of Sino-Tibetan and other families.
Recent research has provided supportive evidence for the role of humidity in the evolution of tones. However, there remain numerous challenges in delving deeper into the intricate relationship between the tone system and climatic factors: precisely tracking and identifying potentially relevant climate factors at appropriate temporal and spatial scales, while effectively controlling the potential interference caused by geographical proximity and language inheritance. Based on a substantial database of 1,525 language varieties in China and 41 years of monthly climate data, this study has delved into the correlation between multiple climate factors and number of tones, examined the mediating role of voice quality in this process, and further analyzed the interrelationship between climate factors and pitch variations. The findings reveal that climate factors influencing voice quality and the number of tones are diverse, with specific humidity, precipitation, and average temperature playing pivotal roles. After controlling the influence of language inheritance and geographical proximity, the chain of climate -> voice quality -> number of tones remains significant in China. Specifically, people living in a humid and warm environment tend to exhibit better voice quality. Meanwhile, regions with higher specific humidity and precipitation tend to have a richer and more diverse range of tone types. These findings enrich the theoretical framework of the interaction between language and the environment and provide robust empirical support for understanding the natural mechanisms of language evolution.
Unlike studies of the evolutionary relationship between languages, the dialect-level variation within a language has seldom been studied within the framework of a phylogenetic tree, because frequent lexical borrowing muddles the evidence of shared ancestry. The phonological history of Japanese is an exceptional case study where the phenomenon called accentual class merger enables the phylogenetic analysis of dialectal pitch-accent systems in a way that is not subject to borrowing. However, previous studies have lacked statistical analysis and failed to evaluate the relative credence of alternative hypotheses. Here we developed a novel substitution model that describes the mutation of pitch-accent systems driven by accentual class merger and integrated the model into the framework of Bayesian phylogenetic inference with geographical diffusion. Applying the method to data on the pitch-accent variation in modern Japanese dialects and historical documents collected from literature, we reconstructed the evolutionary history and spatial diffusion of pitch-accent systems. Our result supports the monophyly of each of three groups of pitch-accent systems in conventional categorization, namely Tokyo type, Keihan type, and N-kei (N-pattern) type of Kyushu, whereas the monophyly of the Tokyo type has been highly controversial in previous studies. The divergence time of the mainland pitch-accent systems was estimated to be from mid-Kofun to early Heian period. Also, it is suggested that the modern Kyoto dialect did not inherit its accent patterns from Bumoki but from an unrecorded lineage which survived from the Muromachi period. Analyses on geographical diffusion suggest that the most recent common ancestor (MRCA) of all the taxa and that of Keihan type were located in or around the Kinki region, whereas the MRCA of N-kei type was located in northern to central Kyushu. The geographical location of the MRCA of Tokyo type remains unclear, but the Kinki and Kanto regions are the most plausible candidates.
Can nonhuman animals use the same acoustic signal to transmit different illocutions on different occasions? This communicative capacity is known as vocal functional flexibility and occurs, for example, in speech, when a sentence serves different illocutionary forces or functions on different occasions based on changes to visual and intonational cues. Although common in human speech, there is a lack of clear evidence for this ability in other species. Here, we examined a likely candidate, the Brown-headed Cowbird (Molothrus ater), which is a vocal-learning songbird species that develops a repertoire of structurally distinct song types. Most of this species' songs are directed towards conspecific males and females less than a meter away, making it unusually easy to determine the apparent target of songs, unlike the broadcast songs done by most songbirds. Songs directed to other males have clear aggressive/threatening intent, while those to females involve courtship/sexual intent. Extensive prior work shows that male cowbirds perform the visual display that accompanies singing differently in these two social settings and also modulate the intonation of song types differently. Because of these display and tonal modulations, constancy of song type usage across male- vs female-directed singing would provide evidence of vocal functional flexibility. Herein, we examined 4,828 songs in three captive flocks containing twenty-four males and thirty females during the breeding season. Males did not use their song types randomly and had strongly favored songs and less commonly used ones. Importantly, favored song types and less commonly used ones were the same whether directing courtship song to a female, aggressive song to another male or singing nonsocially with no receiver nearby. Results were consistent within and across the three flocks, providing strong evidence of vocal functional flexibility. These findings indicate that some species may evolve the ability to modulate and exaggerate visual display components and prosody more than vocal presentation per se because a learned phonological system in this and possibly other species is constrained by its vital role as an indicator trait.
This paper presents a scientometric study of the evolution of evolutionary linguistics, a multidisciplinary field that investigates the origin and evolution of language. We apply network science methods to analyse changes in the connections among core concepts discussed in the Causal Hypotheses in Evolutionary Linguistics Database, a searchable database of causal hypotheses in evolutionary linguistics. Our analysis includes a multipartite network of 416 papers, 742 authors, and 1,786 variables such as 'population birth rate' and 'linguistic complexity'. Our findings indicate a significant increase in the size of concept networks from 1886 to 2022, providing an account of the growth and diversification of evolutionary linguistics as a field. We describe eight major clusters of concepts, and characterize the connections within and between clusters. Finally, we identify hypotheses cutting across clusters of concepts that have a high-betweenness centrality, implying that they might have a higher impact on the field if proven right (or wrong). Furthermore, we discuss the role of databases in cultural evolution and scientometrics, emphasizing the value of interdisciplinary connections and the potential for further cross-disciplinary collaboration in the field of Evolutionary Linguistics.