The family tree remains the main metaphor for describing the evolutionary history and relationships of languages. Languages, like biological species, evolve following processes such as descent with modification and divergence from a common ancestor, which can be modeled using phylogenetic trees. This chapter first gives basic definitions of the concepts involved in the Tree model and phylogenies, and reviews the merits and limitations of the two competing approaches to linguistic diversification: the Tree model and the Wave model. Second, it discusses the sources of non‐tree‐like signal in linguistic data (borrowings, language mixing, parallel innovations, and incomplete lineage sorting) and how to identify these phenomena. Third, it describes the procedures used to infer phylogenies (including tree topology and datings), and how to interpret the results and assess their validity. Finally, it stresses that phylogenetic inference is not only a method for classifying languages, but more importantly a tool with many possible applications related to the evolution of structural features and rates of change, the geographical origin and dispersal routes of ancient human populations, and questions related to the domestication and management of plant and animal species.
Latent space models for network data characterize each node through a vector of latent features whose pairwise similarities define the edge probabilities among the pairs of nodes. Although this formulation has led to successful implementations, the overarching focus has been on directly inferring node embeddings through the latent features, rather than learning the generative process underlying these embeddings. This focus prevents borrowing information across the node features and limits the ability to infer higher-level architectures governing network formation. For example, routinely-studied networks often exhibit multiscale structures informing on nested modular hierarchies among nodes, which could be learned via tree-based representations of dependencies among the latent features. We pursue this direction by bridging latent variable representations of network data with concepts from phylogenetic inference to design a novel latent space model that explicitly characterizes the generative process of the node feature vectors through a branching Brownian motion, with branching structure parametrized by a tree. This tree constitutes the main object of interest and is learned under a Bayesian perspective leveraging priors inherited from phylogenetic literature to infer tree-based modular hierarchies across nodes, which explain heterogeneous multiscale patterns in the network. Identifiability results are derived along with posterior consistency theory. The inference potentials of our model are illustrated in simulations and two real-data applications from criminology and neuroscience, where our formulation learns core structures hidden to state-of-the-art alternatives.
We investigate and compare the evolution of two aspects of culture, languages and weaving technologies, amongst the Kra-Dai (Tai-Kadai) peoples of southwest China and Southeast Asia, using Bayesian Markov-Chain Monte Carlo methods to uncover phylogenies. The results show that languages and looms evolved in related but different ways and bring some new insights into the spread of the Kra-Dai speakers across Southeast Asia. We found that the languages and looms used by Hlai speakers of Hainan are outgroups in both linguistic and loom phylogenies and that the looms used by speakers of closely related languages tend to belong to similar types. However, we also found differences at a deep level both in the details of the evolution of looms and languages and in their overall patterns of change, and we discuss possible reasons for this.
Multiple members of the tit and chickadee (= Parid) family combine two classes of calls, F and D, in a rigid order FD. In Japanese tits, FD has been argued on the basis of multiple experiments to involve syntax and non-trivial compositionality. How ancient are these call combinations? We show that FD combinations (as well as individual F and D calls) are present in nearly all Parid species, and almost absent in their closest relatives, the Remizidae and Stenostiridae. Using phylogenetic tools and ancestral reconstruction methods, we infer that FD combinations very likely emerged between 11 and 26 million years ago in the eastern Himalayas. This result contributes to evolutionary animal linguistics using a comparative phylogenetic approach to reconstruct the evolution of call combinations.
Elite civil servants may move between the public and private sectors throughout their career, a process of interest for the public and social scientists. However, data on career paths are rarely completely available, calling for inference tools that can handle many missing values. We consider public–private paths of elite French civil servants and introduce binary Markov switching models with Bayesian data augmentation. Our procedure combines two complementary data sources: (1) detailed observations of some individual trajectories collected from LinkedIn and (2) less informative ‘traces’ left by all individuals in administrative records, which we model for missing data imputation. This framework allows for varying parameters across individuals and time, yet maintains properties of hidden Markov models, enabling posterior exploration with a tailored sampler. By integrating both sources, we can consider the whole population rather than just a sample, and avoid biases that arise when using only a single source. This allows us to properly test substantive hypotheses on career paths across different public organizations. We notably show that the probability ENA graduates exit the public sector has not increased since 1990, but the probability of return has increased. We identify four clusters of organizations, with distinct patterns of public–private behaviours.
Approximate Bayesian Computation (ABC) methods have become essential tools for performing inference when likelihood functions are intractable or computationally prohibitive. However, their scalability remains a major challenge in hierarchical or high-dimensional models. In this paper, we introduce permABC, a new ABC framework designed for settings with both global and local parameters, where observations are grouped into exchangeable compartments. Building upon the Sequential Monte Carlo ABC (ABC-SMC) framework, permABC exploits the exchangeability of compartments through permutation-based matching, significantly improving computational efficiency. We then develop two further, complementary sequential strategies: Over Sampling, which facilitates early-stage acceptance by temporarily increasing the number of simulated compartments, and Under Matching, which relaxes the acceptance condition by matching only subsets of the data. These techniques allow for robust and scalable inference even in high-dimensional regimes. Through synthetic and real-world experiments – including a hierarchical Susceptible-Infectious-Recover model of the early COVID-19 epidemic across 94 French departments – we demonstrate the practical gains in accuracy and efficiency achieved by our approach.
Sign languages are naturally occurring languages. As such, their emergence and spread reflect the histories of their communities. However, limitations in historical recordkeeping and linguistic documentation have hindered the diachronic analysis of sign languages. In this work, we used computational phylogenetic methods to study family structure among 19 sign languages from deaf communities worldwide. We used phonologically coded lexical data from contemporary languages to infer relatedness and suggest that these methods can help study regular form changes in sign languages. The inferred trees are consistent in key respects with known historical information but challenge certain assumed groupings and surpass analyses made available by traditional methods. Moreover, the phylogenetic inferences are not reducible to geographic distribution but do affirm the importance of geopolitical forces in the histories of human languages.
In some applied scenarios, the availability of complete data is restricted, often due to privacy concerns; only aggregated, robust and inefficient statistics derived from the data are made accessible. These robust statistics are not sufficient, but they demonstrate reduced sensitivity to outliers and offer enhanced data protection due to their higher breakdown point. We consider a parametric framework and propose a method to sample from the posterior distribution of parameters conditioned on various robust and inefficient statistics: specifically, the pairs (median, MAD) or (median, IQR), or a collection of quantiles. Our approach leverages a Gibbs sampler and simulates latent augmented data, which facilitates simulation from the posterior distribution of parameters belonging to specific families of distributions. A by-product of these samples from the joint posterior distribution of parameters and data given the observed statistics is that we can estimate Bayes factors based on observed statistics via bridge sampling. We validate and outline the limitations of the proposed methods through toy examples and an application to real-world income data.
Assuming X is a random vector and A a non-invertible matrix, one sometimes need to perform inference while only having access to samples of Y = AX. The corresponding likelihood is typically intractable. One may still be able to perform exact Bayesian inference using a pseudo-marginal sampler, but this requires an unbiased estimator of the intractable likelihood. We propose saddlepoint Monte Carlo, a method for obtaining an unbiased estimate of the density of Y with very low variance, for any model belonging to an exponential family. Our method relies on importance sampling of the characteristic function, with insights brought by the standard saddlepoint approximation scheme with exponential tilting. We show that saddlepoint Monte Carlo makes it possible to perform exact inference on particularly challenging problems and datasets. We focus on the ecological inference problem, where one observes only aggregates at a fine level. We present in particular a study of the carryover of votes between the two rounds of various French elections, using the finest available data (number of votes for each candidate in about 60,000 polling stations over most of the French territory). We show that existing, popular approximate methods for ecological inference can lead to substantial bias, which saddlepoint Monte Carlo is immune from. We also present original results for the 2024 legislative elections on political centre-to-left and left-to-centre conversion rates when the far-right is present in the second round. Finally, we discuss other exciting applications for saddlepoint Monte Carlo, such as dealing with aggregate data in privacy or inverse problems.
TraitLab is a software package for simulating, fitting and analysing tree-like binary data under a stochastic Dollo model of evolution. The model also allows for rate heterogeneity through catastrophes, evolutionary events where many traits are simultaneously lost while new ones arise, and borrowing, whereby traits transfer laterally between species as well as through ancestral relationships. The core of the package is a Markov chain Monte Carlo (MCMC) sampling algorithm that enables the user to sample from the Bayesian joint posterior distribution for tree topologies, clade and root ages, and the trait loss, catastrophe and borrowing rates for a given data set. Data can be simulated according to the fitted Dollo model or according to a number of generalized models that allow for heterogeneity in the trait loss rate, biases in the data collection process and borrowing of traits between lineages. Coupled pairs of Markov chains can be used to diagnose MCMC mixing and convergence and to debias MCMC estimators. The raw data, MCMC run output, and model fit can be inspected using a number of useful graphical and analytical tools provided within the package or imported into other popular analysis programs. TraitLab is freely available and runs within the Matlab computing environment with its Statistics and Machine Learning toolbox, no other additional toolboxes are required.
We consider the asymptotic properties of Approximate Bayesian Computation (ABC) for the realistic case of summary statistics with heterogeneous rates of convergence. We allow some statistics to converge faster than the ABC tolerance, other statistics to converge slower, and cover the case where some statistics do not converge at all. We give conditions for the ABC posterior to converge, and provide an explicit representation of the shape of the ABC posterior distribution in our general setting; in particular, we show how the shape of the posterior depends on the number of slow statistics. We then quantify the gain brought by the local linear post-processing step.
Cite the source of the dataset as: Sagart L, Jacques G, Lai Y, Ryder RJ, Thouzeau V, Greenhill SJ, List J- M. 2019 Dated language phylogenies shed light on the ancestry of Sino-Tibetan. Proceedings of the National Academy of Sciences, 201817972.
Phylogenetic inference is an intractable statistical problem on a complex space. Markov chain Monte Carlo methods are the primary tool for Bayesian phylogenetic inference, but it is challenging to construct efficient schemes to explore the associated posterior distribution or assess their performance. Existing approaches are unable to diagnose mixing or convergence of Markov schemes jointly across all components of a phylogenetic model. Lagged couplings of Markov chain Monte Carlo algorithms have recently been developed on simpler spaces to diagnose convergence and construct unbiased estimators. We describe a contractive coupling of Markov chains targeting a posterior distribution over a space of phylogenetic trees with branch lengths, scalar parameters and latent variables. We use these couplings to assess mixing and convergence of Markov chains jointly across all components of the phylogenetic model on trees with up to 200 leaves. Samples from our coupled chains may also be used to construct unbiased estimators.
Robbeets et al. 1 argue that the dispersal of the so-called “Transeurasian” languages, a highly disputed language superfamily comprising the Turkic, Mongolian, Tungusic, Koreanic, and Japonic language families, was driven by Neolithic farmers in the West Liao River region of China. They adduce evidence from linguistics, archaeology, and genetics to support their claim. An admirable feature of the Robbeets et al.’s paper is that all their datasets can be accessed. However, a closer investigation of all three types of evidence reveals fundamental problems with each of them. Robbeets et al.’s analysis of the linguistic data does not conform to the minimal standards required by traditional scholarship in historical linguistics and contradicts their own stated sound correspondence principles. A reanalysis of the genetic data finds that they do not conclusively support the farming-driven dispersal of Turkic, Mongolian, and Tungusic, nor the two-wave spread of farming to Korea. Their archaeological data contain little phylogenetic signal, and we failed to reproduce the results supporting their core hypotheses about migrations. Given the severe problems we identify in all three parts of the “triangulation” process, we conclude that there is neither conclusive evidence for a Transeurasian language family nor for associating the five different language families with the spread of Neolithic farmers from the West Liao River region.
We propose a model of the evolution of a matrix along a phylogenetic tree, in which transformations affect either entire rows or columns of the matrix. This represents the change of both lexical and phonological aspects of linguistic data, by allowing for new words to appear and for systematic phonological changes to affect the entire vocabulary. We implement a Sequential Monte Carlo method to sample from the posterior distribution, and infer jointly the phylogeny, model parameters, and latent variables representing cognate births and phonological transformations. We successfully apply this method to synthetic and real data of moderate size.
Functional load (FL) quantifies the contributions by phonological contrasts to distinctions made across the lexicon. Previous research has linked particularly low values of FL to sound change. Here, we broaden the scope of enquiry into FL to its evolution at higher values also. We apply phylogenetic methods to examine the diachronic evolution of FL across 90 languages of the Pama–Nyungan (PN) family of Australia. We find a high degree of phylogenetic signal in FL, indicating that FL values covary closely with genealogical structure across the family. Though phylogenetic signals have been reported for phonological structures, such as phonotactics, their detection in measures of phonological function is novel. We also find a significant, negative correlation between the FL of vowel length and of the following consonant—that is, a time-depth historical trade-off dynamic, which we relate to known allophony in modern PN languages and compensatory sound changes in their past. The findings reveal a historical dynamic, similar to transphonologization, which we characterize as a flow of contrastiveness between subsystems of the phonology. Recurring across a language family that spans a whole continent and many millennia of time depth, our findings provide one of the most compelling examples yet of Sapir’s ‘drift’ hypothesis of non-accidental parallel development in historically related languages.
Approximate Bayesian computation methods are useful for generative models with intractable likelihoods. These methods are however sensitive to the dimension of the parameter space, requiring exponentially increasing resources as this dimension grows. To tackle this difficulty, we explore a Gibbs version of the ABC approach that runs component-wise approximate Bayesian computation steps aimed at the corresponding conditional posterior distributions, and based on summary statistics of reduced dimensions. While lacking the standard justifications for the Gibbs sampler, the resulting Markov chain is shown to converge in distribution under some partial independence conditions. The associated stationary distribution can further be shown to be close to the true posterior distribution and some hierarchical versions of the proposed mechanism enjoy a closed form limiting distribution. Experiments also demonstrate the gain in efficiency brought by the Gibbs version over the standard solution.