Many national censuses collect data on language use, offering valuable insights for tracking language vitality and guiding policy decisions. However, little is known about the reliability of these data. This article proposes a framework for assessing the reliability of census-based speaker counts, focusing on three key aspects: coverage, accuracy, and consistency. We apply this framework to Canadian census data from 2001, 2006, 2011, 2016, and 2021, examining nearly 60 Indigenous languages. Our analysis shows significant improvements in coverage over time, particularly for languages with small speaker populations or those previously grouped under broader 'macrolanguages.' However, we find that the census tends to overestimate speaker numbers for less commonly spoken languages and underestimate them for more widely spoken ones, especially between 2006 and 2016. Overall, the consistency of census data aligns with that of traditional sources. In light of our results, we recommend prioritizing more recent census data and exercising caution when using figures for languages that were previously grouped under macrolanguages.
The Indo-European Cognate Relationships (IE-CoR) dataset is an open-access relational dataset showing how related, inherited words (‘cognates’) pattern across 160 languages of the Indo-European family. IE-CoR is intended as a benchmark dataset for computational research into the evolution of the Indo-European languages. It is structured around 170 reference meanings in core lexicon, and contains 25731 lexeme entries, analysed into 4981 cognate sets. Novel, dedicated structures are used to code all known cases of horizontal transfer. All 13 main documented clades of Indo-European, and their main subclades, are well represented. Time calibration data for each language are also included, as are relevant geographical and social metadata. Data collection was performed by an expert consortium of 89 linguists drawing on 355 cited sources. The dataset is extendable to further languages and meanings and follows the Cross-Linguistic Data Format (CLDF) protocols for linguistic data. It is designed to be interoperable with other cross-linguistic datasets and catalogues, and provides a reference framework for similar initiatives for other language families.
A perennial conflict in historical linguistics centers around the theoretical and practical virtues of tree-like divergence and wave-like diffusion. This paper presents the Dialect Chain Tree, an extension of the tree model that incorporates both tree-like descent and disintegration of dialect chains in a systematic fashion. As such, it provides a formalization and sharpening of Ross’ (1997: 212–228) linkage concept that allows integration into quantitative approaches.
Wurm & Hattori’s Language Atlas of the Pacific Area describes the geographic speaker areas of the languages and language varieties spoken in the Pacific. Thanks to the efforts of the Electronic Cultural Atlas Initiative, this monumental piece of work has been available in digital form for over 15 years. But lacking proper identification of language varieties, this digitized data was largely unusable for today’s research methods. We turned ECAI’s digitized artefacts of the Language Atlas into an open, reusable geo-referenced dataset of speaker area polygons for a quarter of the world’s languages. This allows for much more refined analysis methods to, for example, analyse language contact in the area of the world with the highest linguistic diversity. We also describe a number of tool applications and quality checks which may be useful for methodological development in similar digitization efforts.
In the present paper, we discuss the bibliographical limits for commonplace typological studies and address how to estimate the resources available for an in-depth study using a full-text corpus of grammatical descriptions, considering different metalanguages, temporal stages of description, theoretical perspectives, and quality of grammatical descriptions. In a case study on motion, we illustrate the above perspectives and show how computer-assisted sampling using large-scale keyword searches for information-dense descriptions is a time-saving resource for the linguistic researcher to create genealogically independent samples. The measures discussed in this study allow for a better appraisal of the state of existing information for typological studies, but the problem of wider access to rare publications remains a significant challenge. En el presente art & iacute;culo, discutimos los l & iacute;mites bibliogr & aacute;ficos para estudios tipol & oacute;gicos comunes y abordamos c & oacute;mo estimar los recursos disponibles para un estudio a profundidad utilizando un corpus de texto completo de descripciones gramaticales, considerando diferentes metalenguajes, etapas temporales de descripci & oacute;n, perspectivas te & oacute;ricas y calidad de las descripciones gramaticales. En un estudio de caso sobre el movimiento, ilustramos las perspectivas mencionadas y mostramos c & oacute;mo el muestreo asistido por computadora mediante b & uacute;squedas de palabras clave a gran escala para descripciones densas en informaci & oacute;n es un recurso que ahorra tiempo para el investigador ling & uuml;& iacute;stico al crear muestras geneal & oacute;gicamente independientes. Las medidas discutidas en este estudio permiten una mejor evaluaci & oacute;n del estado de la informaci & oacute;n existente para estudios tipol & oacute;gicos, pero el problema de un acceso m & aacute;s amplio a publicaciones raras sigue siendo un desaf & iacute;o significativo.
Computational methods of language dating make inferences about the divergence times of protolanguages by evaluating the patterns of inheritance in the vocabulary of modemlanguages, given the specification of a model of vocabulary evolution. We consider a model that describes vocabulary evolution as the replacement of traits by new traits from an infinite state space along a tree. This model has been introduced in previous literature but so far it has not been used in many applications. We give a general recursive algorithm for calculating likelihoods and argue that the model gives a more realistic representation of vocabulary evolution over time, compared to existing models like the Stochastic Dollo model. We also provide a case study demonstrating the model's potential applications.
While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here, we outline the Grambank database. With over 400,000 data points and 2400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on the structural diversity of the world's languages, evaluate constraints on linguistic diversity, and identify the world's most unusual languages. An analysis of the consequences of language loss reveals that the reduction in diversity will be strikingly uneven across the major linguistic regions of the world. Without sustained efforts to document and revitalize endangered languages, our linguistic window into human history, cognition, and culture will be seriously fragmented.
Hammarström, Harald & Forkel, Robert & Haspelmath, Martin & Bank, Sebastian. 2022. Glottolog 4.6. Leipzig: Max Planck Institute for Evolutionary Anthropology. (Available online at https://glottolog.org)
The world harbours a diversity of some 6,500 mutually unintelligible languages. As has been increasingly observed by linguists, many minority languages are becom-ing endangered and will be lost forever if not documented. The increased urgency has led to the development of several global endangerment databases and a more fine-grained understanding of the language endangerment progression as well as its possible reversal. In the present paper, we explore the terminological correlates of this development as found in the descriptive linguistic literature, using a corpus of over 10,000 digitized grammatical descriptions. Comparing this with existing en-dangerment databases, we find that simply counting terms related to endangerment does signal endangerment, but the degree of endangerment is more difficult to assess from grammatical descriptions. The label endangered seems to be an umbrella term that covers different situations ranging from moribund languages with less than ten speakers to minority languages with several thousand speakers. For many languages considered endangered in existing databases, explicit terms to this effect cannot be found in their descriptions. The discrepancy is due to incompleteness of the search -term set, gaps in the literature, and projected rather than observed information in the databases. Our explorations illustrate the potential for database curation as-sisted by computational searches both to maintain accuracy of the databases and to investigate assumed language endangerment. Future work includes a larger cloud of search terms, usage of term frequencies, and prescreening of descriptive literature for the existence of a relevant section. From the perspective of descriptive linguistics, this study calls for a more careful correlation between the language endangerment indexes, as developed in the global endangerment databases, and the treatment of the endangerment status of individual languages in descriptive grammars.
Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes an implementation of the Urdu language as a software API, and we deal with orthography, morphology and the extraction of the lexicon. The morphology is implemented in a toolkit called Functional Morphology (Forsberg & Ranta, 2004), which is based on the idea of dealing grammars as software libraries. Therefore this implementation could be reused in applications such as intelligent search of keywords, language training and infrastructure for syntax. We also present an implementation of a small part of Urdu syntax to demonstrate this reusability.
Glottocodes constitute the backbone identification system for the language, dialect and family inventory Glottolog (https://glottolog.org). In this paper, we summarize the motivation and history behind the system of glottocodes and describe the principles and practices of data curation, technical infrastructure and update/version-tracking systematics. Since our understanding of the target domain – the dialects, languages and language families of the entire world – is continually evolving, changes and updates are relatively common. The resulting data is assessed in terms of the FAIR (Findable, Accessible, Interoperable, Reusable) Guiding Principles for scientific data management and stewardship. As such the glottocode-system responds to an important challenge in the realm of Linguistic Linked Data with numerous NLP applications.
Cite the source of the dataset as: Her, One-Soon, Harald Hammarström and Marc Allassonnière-Tang. 2022. Defining numeral classifiers and identifying classifier languages of the world. Linguistics Vanguard. https://doi.org/10.1515/lingvan-2022-0006
While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here we outline the Grambank database. With over 400,000 data points and 2,400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on the structural diversity of the world's languages, evaluate constraints on linguistic diversity, and identify the world's most unusual languages. An analysis of the consequences of language loss reveals that the reduction in diversity will be strikingly uneven across the major linguistic regions of the world. Without sustained efforts to document and revitalize endangered languages, our linguistic window into human history, cognition and culture will be seriously fragmented.
This paper presents a precise definition of numeral classifiers, steps to identify a numeral classifier language, and a database of 3,338 languages, of which 723 languages have been identified as having a numeral classifier system. The database, named World Atlas of Classifier Languages (WACL), has been systematically constructed over the last 10 years via a manual survey of relevant literature and also an automatic scan of digitized grammars followed by manual checking. The open-access release of WACL is thus a significant contribution to linguistic research in providing (i) a precise definition and examples of how to identify numeral classifiers in language data and (ii) the largest dataset of numeral classifier languages in the world. As such it offers researchers a rich and stable data source for conducting typological, quantitative, and phylogenetic analyses on numeral classifiers. The database will also be expanded with additional features relating to numeral classifiers in the future in order to allow more fine-grained analyses.
Human history is written in both our genes and our languages. The extent to which our biological and linguistic histories are congruent has been the subject of considerable debate, with clear examples of both matches and mismatches. To disentangle the patterns of demographic and cultural transmission, we need a global systematic assessment of matches and mismatches. Here, we assemble a genomic database (GeLaTo, or Genes and Languages Together) specifically curated to investigate genetic and linguistic diversity worldwide. We find that most populations in GeLaTo that speak languages of the same language family (i.e., that descend from the same ancestor language) are also genetically highly similar. However, we also identify nearly 20% mismatches in populations genetically close to linguistically unrelated groups. These mismatches, which occur within the time depth of known linguistic relatedness up to about 10,000 y, are scattered around the world, suggesting that they are a regular outcome in human history. Most mismatches result from populations shifting to the language of a neighboring population that is genetically different because of independent demographic histories. In line with the regularity of such shifts, we find that only half of the language families in GeLaTo are genetically more cohesive than expected under spatial autocorrelations. Moreover, the genetic and linguistic divergence times of population pairs match only rarely, with Indo-European standing out as the family with most matches in our sample. Together, our database and findings pave the way for systematically disentangling demographic and cultural history and for quantifying processes of shifts in language and social identities on a global scale.
Languages of diverse structures and different families tend to share common patterns if they are spoken in geographic proximity. This convergence is often explained by horizontal diffusibility, which is typically ascribed to language contact. In such a scenario, speakers of two or more languages interact and influence each other’s languages, and in this interaction, more grammaticalized features tend to be more resistant to diffusion compared to features of more lexical content. An alternative explanation is vertical heritability: languages in proximity often share genealogical descent. Here, we suggest that the geographic distribution of features globally can be explained by two major pathways, which are generally not distinguished within quantitative typological models: feature diffusion and language expansion. The first pathway corresponds to the contact scenario described above, while the second occurs when speakers of genetically related languages migrate. We take the worldwide distribution of nominal classification systems (grammatical gender, noun class, and classifier) as a case study to show that more grammaticalized systems, such as gender, and less grammaticalized systems, such as classifiers, are almost equally widespread, but the former spread more by language expansion historically, whereas the latter spread more by feature diffusion. Our results indicate that quantitative models measuring the areal diffusibility and stability of linguistic features are likely to be affected by language expansion that occurs by historical coincidence. We anticipate that our findings will support studies of language diversity in a more sophisticated way, with relevance to other parts of language, such as phonology.
It has long been recognized that suffixing is more common than prefixing in the languages of the world. More detailed statistics on this tendency are needed to sharpen proposed explanations for this tendency. The classic approach to gathering data on the prefix/suffix preference is for a human to read grammatical descriptions (948 languages), which is time-consuming and involves discretization judgments. In this paper we explore two machine-driven approaches for prefix and suffix statistics which are crude approximations, but have advantages in terms of time and replicability. The first simply searches a large collection of grammatical descriptions for occurrences of the terms ‘prefix’ and ‘suffix’ (4 287 languages). The second counts substrings from raw text data in a way indirectly reflecting prefixation and suffixation (1 030 languages, using New Testament translations). The three approaches largely agree in their measurements but there are important theoretical and practical differences. In all measurements, there is an overall preference for suffixation, albeit only slightly, at ratios ranging between 0.51 and 0.68.
Starting from a large collection of digitized raw-text descriptions of languages of the world, we address the problem of extracting information of interest to linguists from these. We describe a general technique to extract properties of the described languages associated with a specific term. The technique is simple to implement, simple to explain, requires no training data or annotation, and requires no manual tuning of thresholds. The results are evaluated on a large gold standard database on classifiers with accuracy results that match or supersede human inter-coder agreement on similar tasks. Although accuracy is competitive, the method may still be enhanced by a more rigorous probabilistic background theory and usage of extant NLP tools for morphological variants, collocations and vector-space semantics.
Grammatical descriptions of languages of the world form a sub-genre of scholarly documents in the field of linguistics. A document of this genre may be modeled as a concatenation of table of contents, sociolinguistic description, phonological description, morphosyntactic description, comparative remarks, lexicon, text, bibliography and index (where morphosyntactic description is the only mandatory section). Separation of these parts is useful for information extraction, bibliometrics and information content analysis. Using a collection of over 10 000 digitized grammatical descriptions and an associated bibliography with document-level categorizations, we show that standard techniques from text classification can be adapted to classify individual pages. Assuming that the divisions of interest form continuous page ranges, we can achieve the sought after division in a transparent way. In contrast to previous work on similar tasks in other domains, no use is made of formatting cues, no additional annotated data is needed, high-quality OCR is not required, and the document collection is highly multilingual.