The application of phylogenetic methods to linguistic data requires several critical methodological choices, yet their relative influence on the resulting trees remains poorly understood. This study conducts a systematic sensitivity analysis to evaluate the impact of three key factors: (1) the phylogenetic inference method (neighbor-joining, maximum parsimony, and Bayesian inference), (2) the character coding scheme (binary vs. multistate), and (3) the cognacy (i.e. homology) criterion (root-based vs. stem-based cognacy). Using the Japonic language family as a model, we generated a diverse set of phylogenetic trees and compared them using multiple tree metrics and distance measures between trees. Variance partitioning through distance-based redundancy analysis reveals a clear hierarchy of influence: the choice of inference method is the primary driver of variation between trees, followed by the character coding scheme, while the cognacy criterion exerts a comparatively minor effect. Specifically, we show that the choice of inference method fundamentally shapes tree topology, whereas the character coding scheme impacts branch lengths and evolutionary tempo. While maintaining a consistent and transparent definition of cognacy remains good practice, our results indicate that the choice of one definition over another does not always lead to a significant loss of phylogenetic signal. Therefore, phylolinguistic research should prioritize first and foremost the rigorous selection of inference method.
The family tree remains the main metaphor for describing the evolutionary history and relationships of languages. Languages, like biological species, evolve following processes such as descent with modification and divergence from a common ancestor, which can be modeled using phylogenetic trees. This chapter first gives basic definitions of the concepts involved in the Tree model and phylogenies, and reviews the merits and limitations of the two competing approaches to linguistic diversification: the Tree model and the Wave model. Second, it discusses the sources of non‐tree‐like signal in linguistic data (borrowings, language mixing, parallel innovations, and incomplete lineage sorting) and how to identify these phenomena. Third, it describes the procedures used to infer phylogenies (including tree topology and datings), and how to interpret the results and assess their validity. Finally, it stresses that phylogenetic inference is not only a method for classifying languages, but more importantly a tool with many possible applications related to the evolution of structural features and rates of change, the geographical origin and dispersal routes of ancient human populations, and questions related to the domestication and management of plant and animal species.
Japanese pitch accent is traditionally described as a free accent system allowing up to n + 1 possible patterns for an n-mora word. However, Japanese accentuation is heavily constrained by factors such as word length, prosodic structure, and lexical stratum, leading to the hypothesis that the system fundamentally reduces to only two dominant patterns: unaccented and thematic (antepenultimate). Evaluating this claim requires moving beyond raw counts of empirically attested patterns to quantify the skewness of their distribution. In this study, we introduce information-theoretic metrics, specifically Shannon's conditional entropy and ecological effective numbers of diversity, to measure the uncertainty an diversity of accent patterns in a representative sample of 8,173 Japanese common nouns. Our findings show that while overall pattern richness reaches the theoretical maximum (7 patterns), global effective diversity remains low (3.12 by type frequency, 3.43 by token frequency). Crucially, with knowledge of word length, prosodic structure, and lexical stratum the number of accent patterns is reduced to effectively two (2.1) dominant patterns, providing a precise, quantitative validation of the two-pattern hypothesis. We further demonstrate that prosodic structure is the single most informative factor, reducing accentuation uncertainty by 23%, and that native nouns exhibit greater residual uncertainty than Sino-Japanese and foreign words. Finally, the strong alignment between type and token results confirms that dictionary-based word lists closely mirror natural speech distributions in phonological diversity.
We investigate and compare the evolution of two aspects of culture, languages and weaving technologies, amongst the Kra-Dai (Tai-Kadai) peoples of southwest China and Southeast Asia, using Bayesian Markov-Chain Monte Carlo methods to uncover phylogenies. The results show that languages and looms evolved in related but different ways and bring some new insights into the spread of the Kra-Dai speakers across Southeast Asia. We found that the languages and looms used by Hlai speakers of Hainan are outgroups in both linguistic and loom phylogenies and that the looms used by speakers of closely related languages tend to belong to similar types. However, we also found differences at a deep level both in the details of the evolution of looms and languages and in their overall patterns of change, and we discuss possible reasons for this.
a n l a n g u a g e sThe Ryukyuan (Ry.) languages 1 are spoken in the Ryukyu Islands, a chain of around 50 inhabited islands stretching from southeast of Kyushu to northeast of Taiwan and naturally delimited by the Kuroshio Current.Ryukyuan subdivides into at least five languages, which are not mutually intelligible: Amami (Ama.),Okinawan (Oki.),Miyako (Miy.),Yaeyama (Yae.), and Yonaguni (Yonag., a.k.a.Dunan).Amami and Okinawan belong to Northern Ryukyuan, while Miyako, Yaeyama, and Yonaguni form together the Southern Ryukyuan group (Pellard 2015).All Ryukyuan languages are highly endangered: fluent speakers are usually in their late sixties or older and are bilingual in Japanese (Jp.), while younger generations are monolingual in Japanese.The Ryukyuan languages form a sister branch of Japanese within the Japanese-Ryukyuan family 2 (Pellard 2015).Ryukyuan and Japanese most likely split during the first half of the first millenium ad, and the ancestor of Ryukyuan was then spoken on Kyushu for several centuries before it migrated to the Ryukyus around the 10th century (Pellard 2013a; 2015; 2016a).The following will address, after a brief historiographic overview, those phonological (vowels, consonants, accent/tone), grammatical (verbs, adjectives, case, kakari-musubi), and lexical (pronouns, demonstratives) topics for which the contribution of Ryukyuan is, or has been claimed to be, important for the reconstruction of proto-Japanese-Ryukyuan (pJR), the common ancestor of Japanese and Ryukyuan.2 r y u k y u a n h i s t o r i c a l l i n g u i s t i c s Though the close similarity between Japanese and Ryukyuan (usually Okinawan) had long been repeatedly noticed and their genetic relationship recognized before (Osterkamp 2015), the first historical-comparative study of Ryukyuan is Chamberlain's (1895) comparative grammar of (Tokyo) Japanese and (Shuri) Okinawan.Chamberlain's pioneer work is however mainly of historiographic interest today, as it suffers from the limitations of its pre-Neogrammarian methodology and
Formal and computational linguistics can enhance descriptive linguistics of endangered languages by providing them with precise models and quantitative perspectives. We exemplify the benefits of such an approach with the case of Asama's verb inflectional morphology. We show that a Word-and-Paradigm framework can provide interesting insights and allow for both the identification and the quantification of the sources of uncertainty in the implicative relations within Asama's verb paradigms. We describe Asama's verb morphology by considering whole forms rather than exponents only, and we factor its alternation patterns in two types: segmental alternations and suprasegmental alternations. Measures of Shannon's conditional entropy are then used to estimate the respective contributions of these factors to the complexity of the system. Suprasegmental alternations turn out to be the major source of uncertainty in implicative relations, and vowel length and tone alternations cannot be treated separately but strongly interact. We also show how the principal parts of the inflectional system can be determined with conditional entropy measures of n-ary implicative relations.
Formal and computational linguistics can enhance descriptive linguistics of endangered languages by providing them with precise models and quantitative perspectives. We exemplify the benefits of such an approach with the case of Asama’s verb inflectional morphology. We show that a Word-and-Paradigm framework can provide interesting insights and allow for both the identification and the quantification of the sources of uncertainty in the implicative relations within Asama’s verb paradigms. We describe Asama’s verb morphology by considering whole forms rather than exponents only, and we factor its alternation patterns in two types: segmental alternations and suprasegmental alternations. Measures of Shannon’s conditional entropy are then used to estimate the respective contributions of these factors to the complexity of the system. Suprasegmental alternations turn out to be the major source of uncertainty in implicative relations, and vowel length and tone alternations cannot be treated separately but strongly interact. We also show how the principal parts of the inflectional system can be determined with conditional entropy measures of n -ary implicative relations.
The methods of spatial statistics have been successfully applied to the study of linguistic variation, especially for detecting the existence of spatial patterns in the geographical distribution of linguistic features. However, the use of local indicators of spatial autocorrelation for detecting spatial clusters have been limited to continuous variables, and we propose to apply the new method of Anselin and Li (2019) for categorical variables to linguistic data. We illustrate this method with the case of Japanese rendaku , or sequential voicing, whose dialectal variation is still poorly documented. Focusing on regional differences in the frequency of rendaku, we examined the occurrence of rendaku for four lexemes in 4,921 place names from all Japan. A statistical analysis of local spatial association and an unsupervised density-based cluster analysis revealed the existence of two cluster areas of high rendaku frequency centered around Wakayama and Fukushima-Yamagata prefectures. This suggests that rendaku is more frequent in those dialects, and we recommend that further studies in the dialectal variation of rendaku start by looking at those areas.
Robbeets et al. 1 argue that the dispersal of the so-called “Transeurasian” languages, a highly disputed language superfamily comprising the Turkic, Mongolian, Tungusic, Koreanic, and Japonic language families, was driven by Neolithic farmers in the West Liao River region of China. They adduce evidence from linguistics, archaeology, and genetics to support their claim. An admirable feature of the Robbeets et al.’s paper is that all their datasets can be accessed. However, a closer investigation of all three types of evidence reveals fundamental problems with each of them. Robbeets et al.’s analysis of the linguistic data does not conform to the minimal standards required by traditional scholarship in historical linguistics and contradicts their own stated sound correspondence principles. A reanalysis of the genetic data finds that they do not conclusively support the farming-driven dispersal of Turkic, Mongolian, and Tungusic, nor the two-wave spread of farming to Korea. Their archaeological data contain little phylogenetic signal, and we failed to reproduce the results supporting their core hypotheses about migrations. Given the severe problems we identify in all three parts of the “triangulation” process, we conclude that there is neither conclusive evidence for a Transeurasian language family nor for associating the five different language families with the spread of Neolithic farmers from the West Liao River region.
The native Japanese name of the Buddha hotoke < poto2ke2 has no internal etymology and is likely to be a loanword introduced together with Buddhism. The hypothesis of a link with Korean pwuche < pwuthye ‘Buddha’ and of their ultimate origin as deriving from a Chinese rendering of Sanskrit Buddha makes sense from both a linguistic and historical point of view. Still, the last part of the Japanese and Korean forms has no correspondent in Chinese and has remained unaccounted for hitherto. From the comparison with the pattern ‘Buddha-lord’ for the name of Buddha in several Asian languages, it is hypothesized that the enigmatic final element was originally a word for ‘lord, ruler, king’. This hypothesis is confirmed by the attestation of such a word in toponyms and in nobility titles recorded in ancient Chinese, Korean, and Japanese chronicles.