Based on a dataset representing close to ¾ of the world’s languages we investigate differences among languages and between items on the Swadesh list with regard to mean word length from a linguistic typological point of view. Mapping the world-wide distribution of word length shows convergence at a continent-wide level, a Pacific Rim signature, and a tendency for large word length averages to be a recessive trait. The amount of data, which is unparalleled in previous, related studies, allows us to provide more solid estimates and accounts for the interrelationships between word length, phoneme segment inventory size, and population size than was previously possible. Word length differences between items exhibit robust, universal tendencies, which are discussed in relation to other quantities, including stability, synonymy, and attestation.
Cite the source of the dataset as: Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2022. The ASJP Database (version 20).
Cite the source of the dataset as: Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2016. The ASJP Database (version 17).
Cite the source of the dataset as: Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2020. The ASJP Database (version 19).
Cite the source of the dataset as: Wichmann, Søren, André Müller, Annkathrin Wett, Viveka Velupillai, Julia Bischoffberger, Cecil H. Brown, Eric W. Holman, Sebastian Sauppe, Zarina Molochieva, Pamela Brown, Harald Hammarström, Oleg Belyaev, Johann-Mattis List, Dik Bakker, Dmitry Egorov, Matthias Urban, Robert Mailhammer, Agustina Carrizo, Matthew S. Dryer, Evgenia Korovina, David Beck, Helen Geyer, Pattie Epps, Anthony Grant, and Pilar Valenzuela. 2013. The ASJP Database (version 16).
Data and web application of the ASJP Database (version 18). Cite the Database as Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2018. The ASJP Database (version 18).
It is known that phylogenetic trees are more imbalanced than expected from a birth-death model with constant rates of speciation and extinction, and also that imbalance can be better fit by allowing the rate of speciation to decrease as the age of the parent species increases. If imbalance is measured in more detail, at nodes within trees as a function of the number of species descended from the nodes, age-dependent models predict levels of imbalance comparable to real trees for small numbers of descendent species, but predicted imbalance approaches an asymptote not found in real trees as the number of descendent species becomes large. Age-dependence must therefore be complemented by another process such as inheritance of different rates along different lineages, which is known to predict insufficient imbalance at nodes with few descendent species, but can predict increasing imbalance with increasing numbers of descendent species. [Crump-Mode-Jagers process; diversification; macroevolution; taxon sampling; tree of life.].
Since the early 1970s, biologists have debated whether evolution is punctuated by speciation events with bursts of cladogenetic changes, or whether evolution tends to be of a more gradual, anagenetic nature. A similar discussion among linguists has barely begun, but the present results suggest that there is also room for controversy over this issue in linguistics. The only previous study correlated the number of nodes in linguistic phylogenies with branch lengths and found support for punctuated equilibrium. We replicate this result for branch lengths, but find no support for punctuated equilibrium using a different, automated measure of linguistic divergence and a much larger data set. With the automated measure, segments of trees containing more nodes show no greater divergence from an outgroup than segments containing fewer nodes.
Using three worldwide databases, we investigate how average similarity between pairs of languages and cultures are influenced by geographic distance and time of common ancestry. Generally, the similarity between languages or cultures decreases as the geographic distance increases. This occurs even for languages and cultures without a known common ancestor, suggesting the influence of diffusion. At any given distance, related languages are more similar than unrelated languages. However, remotely related cultures are no more similar than entirely unrelated ones, indicating that inherited cultural features tend to be lost more readily over time than inherited linguistic features.
An automated sound correspondence-recognition program developed by the authors is applied to a data set consisting of standardized word lists for over half of the world's languages. Online appendices present the results in a compendium of 692 recurrent sound correspondences that contains information about the frequency of occurrence of each correspondence. Applications of the compendium to historical linguistics are proposed. For example, the catalog of correspondences and frequencies facilitates objective assessment of the commonness or rarity of shared phonological innovations cited as evidence for language-family subgrouping. In another analysis, correspondence frequency is used to measure the degree of similarity between different sounds, yielding models for classifying consonants and vowels that substantially agree with articulatory properties. Correspondence-based similarities are also compared with measurements of sound similarity involving factors such as perceptual confusions, speech errors, and cooccurrence patterns in synchronic phonological rules. Sound similarity discerned from both the perception and production of speech is found to correlate to about the same extent with correspondence-based similarities.
The findings to be presented in this paper were not anticipated, but came about as an unexpected result of looking at how the application of a version of the Levenshtein distance to word lists compares with cognate counting. We were interested in the degree to which the two correlate. The results of this investigation are intrinsically interesting and will be presented in the following section 2, but even more interesting is our finding that differences between counting cognates and measuring the Levenshtein distances vary as a function of average word lengths in the word lists compared. This observation will occupy the remainder of the paper, with section 3 devoted to establishing the sta tis tical significance of the observation across language families, while section 4 establishes the significance within language groups, and section 5 discusses competing explanations. First we briefly explain the specific version of the Levenshtein distance used and the concept of cognate identification. In numerous previous papers, beginning in Holman et al. (2008a), the present authors as well as other members of the network of scholars participating in the project known as ASJP (or Automated Similarity Judgment Pro gram) have applied a computer-assisted comparison of word lists in order to derive a measure of differences among languages. Our method consists in comparing pairs of words to determine the Levenshtein distance, LD, which is defined as the number of substitutions, insertions, and deletions necessary to transform one word into another. The LD is divided by the length of the longer of the two words compared such that any distance will come to lie in the range 0%–100%. This normalized measure, called LDN,2 is averaged over all pairs of words referring to the same concept in lists from two given languages. To enhance discrimination between related and unrelated languages, this average LDN is further divided by the average LDN between words referring to dif ferent concepts in the different lists, to obtain what we call LDND (‘Leven shtein Distance Normalized Divided’). A similarity measure, here called ASJPsim, is defined by subtracting LDND from 100%.
This paper describes a computerized alternative to glottochronology for estimating elapsed time since parent languages diverged into daughter languages. The method, developed by the Automated Similarity Judgment Program (ASJP) consortium, is different from glottochronology in four major respects: (1) it is automated and thus is more objective, (2) it applies a uniform analytical approach to a single database of worldwide languages, (3) it is based on lexical similarity as determined from Levenshtein (edit) distances rather than on cognate percentages, and (4) it provides a formula for date calculation that mathematically recognizes the lexical heterogeneity of individual languages, including parent languages just before their breakup into daughter languages. Automated judgments of lexical similarity for groups of related languages are calibrated with historical, epigraphic, and archaeological divergence dates for 52 language groups. The discrepancies between estimated and calibration dates are found to be on average 29% as large as the estimated dates themselves, a figure that does not differ significantly among language families. As a resource for further research that may require dates of known level of accuracy, we offer a list of ASJP time depths for nearly all the world’s recognized language families and for many subfamilies.
This paper discusses phylogenetic reticulation using linguistic data from the Automated Similarity Judgment Program or ASJP (Holman et al., 2008; Wichmann et al., 2010a). It contributes methodologically to the examination of two measures of reticulation in distance-based phylogenetic data, specifically the δ score of Holland et al. (2002) and the more recent Q -residuals of Gray et al. (2010). It is shown that the δ score is a more adequate measure of reticulation. Our empirical analyses examine possible correlations between δ and (a) the size (number of languages), (b) age, and (c) heterogeneity of language groups, (d) linguistic isolation of individual languages within their respective phylogenies, and (e) the status of speech forms as dialects or recently emerged languages. Among these, only (d) is significantly correlated with δ. Our interpretation is that δ is a realistic measure of reticulation and sensitive to effects of socio-historical events such as language extinction. Finally, we correlate average δ scores for different language families with the goodness of fit between ASJP and expert classifications, showing that the δ scores explain much of the variance.
No abstract given; compares pairs of languages from World Atlas of Language Structures.