This paper demonstrates how to leverage the GeoNames data for seeking patterns in toponymic data using the software package ‘toponym’, which we wrote for the R computational environment. After discussing a distinction between particularistic and pattern-seeking approaches to toponymics, we go on to characterize the data of GeoNames, which are particularly appropriate for the latter type of approach. Then, we present two cases studies. The first case study is on Xincan place names in Guatemala, and the second is on Slavic place names in eastern Germany. These explorations support our hypothesis that toponymics may benefit from new computational tools.
This paper is envisioned as a primarily methodological contribution towards a more sophisticated and systematic approach to conflict research in archaeology and history. Studies of conflicts in these fields have often focused on violence and war. Instead, we offer a more holistic approach to conflict research, taking into account different levels of both escalation and de-escalation that embrace all the possible aspects of a conflict from a mere undeveloped potential over complete annihilation to various countermeasures and stages of resolution. A model taking into account different levels of escalation and de-escalation is presented which embodies our multi-faceted view of conflicts and which also allows for a systematic, comparative analysis of conflict situations anywhere and any time in (pre)history. Through ten relatively detailed European case studies spanning the Bronze Age to the 20th century we demonstrate the comparative potential of our model and suggest ways in which it may help to identify typical patterns in conflict situations.
In this note, we describe how to install and use the ‘toponym’ R package, which is designed for mapping and manipulating toponymic data from the GeoNames database. This introduction will allow even unexperienced users of R to efficiently produce maps and perform simple analyses.
In this note, we describe how to install and use the 'toponym' R package, which is designed for mapping and manipulating toponymic data from the GeoNames database. This introduction will allow even unexperienced users of R to efficiently produce maps and perform simple analyses.
Previous work using lexical data from around the world has suggested that distances between language varieties are distributed such that varieties are typically either rather similar, qualifying as dialects of the same language, or rather dissimilar, qualifying as different languages, with a scarcity of varieties that are around halfway similar. Using a potentially biased sample, Wichmann (2019) observed that there is a bimodal distribution of distances with two roughly normal distributions separated by a valley. Here we test whether a similar distribution is found when using another source of data and an unbiased sample drawn from the cells of a geographical grid (of central Europe). The data consists of 18 lexemes from 274 doculects. Using Bayesian beta regression and leave-one-out cross-validation, we show that the data follows a bimodal distribution which is robust to sampling, and also to at least some aspects of the data (coarse- vs. fine-grained phonetic transcriptions).
The aim of this paper is to show evidence of a statistical dependency of the presence of tones on word length. Other work has made it clear that there is a strong inverse correlation between population size and word length. Here it is additionally shown that word length is coupled with tonal distinctions, languages being more likely to have such distinctions when they exhibit shorter words. It is hypothesized that the chain of causation is such that population size influences word length, which, in turn, influences the presence and number of tonal distinctions.
Are the sound systems of languages ecologically adaptive like other aspects of human behavior? In previous substantive explorations of the climate–language nexus, the hypothesis that desiccation affects the tone systems of languages was not well supported. The lack of analysis of voice quality data from natural speech undermines the credibility of the following two key premises: the compromised voice quality caused by desiccated ambient air and constrained use of phonemic tone due to a desiccated larynx. Here, the full chain of causation, humidity → voice quality → number of tones, is for the first time strongly supported by direct experimental tests based on a large speech database (China’s Language Resources Protection Project). Voice quality data is sampled from a recording set that includes 997 language varieties in China. Each language is represented by about 1200 sound files, amounting to a total of 1,174,686 recordings. Tonally rich languages are distributed throughout China and vary in their number of tones and in the climatic conditions of their speakers. The results show that, first, the effect of humidity is large enough to influence the voice quality of common speakers in a naturalistic environment; secondly, poorer voice quality is more likely to be observed in speakers of non-tonal languages and languages with fewer tones. Objective measures of phonatory capabilities help to disentangle the humidity effect from the contribution of phylogenetic and areal relatedness to the tone system. The prediction of ecological adaptation of speech is first verified through voice quality analysis. Humidity is observed to be related to synchronic variation in tonality. Concurrently, the findings offer a potential trigger for diachronic changes in tone systems.
Multiple factors of the natural environment have been found to impact and mold the phonetic patterns of human speech, among which the potential correlation between sonority and temperature has garnered significant attention. We leverage a large database containing basic vocabularies of 5,293 languages and calculate the average sonority for each language by adopting a universal sonority scale. Our findings confirm a positive correlation between sonority and temperature across macroareas and language families, whereas this relationship cannot be discerned within language families. We suggest that the adaptation of the distribution of speech sounds within languages is a slow process which is moreover insensitive to minor differences in temperature experienced by speakers as they carry their languages to new regions. Nevertheless, at the global level a solid relationship emerges. Furthermore, we delve deeper into the nature of the relationship and contend that it is mainly due to cold temperatures having a weakening effect on sonority. This research provides compelling additional evidence that climatic factors contribute to shaping language and its evolution.
Based on a dataset representing close to ¾ of the world’s languages we investigate differences among languages and between items on the Swadesh list with regard to mean word length from a linguistic typological point of view. Mapping the world-wide distribution of word length shows convergence at a continent-wide level, a Pacific Rim signature, and a tendency for large word length averages to be a recessive trait. The amount of data, which is unparalleled in previous, related studies, allows us to provide more solid estimates and accounts for the interrelationships between word length, phoneme segment inventory size, and population size than was previously possible. Word length differences between items exhibit robust, universal tendencies, which are discussed in relation to other quantities, including stability, synonymy, and attestation.
This study investigates how word categories, namely noun and verb, influence acoustic realizations (duration, F0, intensity) in Standard Mandarin Chinese, a language having phonemically distinctive tones and a simple morpho-logical system. Noun-verb ambiguous words were selected and presented in the final positions of typical syntactic contexts in order to avoid the interference of prosodic boundary, syntactic complexity, contextual predictability, tonal environment, F0 range and syllable properties (consonant, vowel, tone, syllable length). Linear mixed models were fitted to duration, and generative additive mixed models were fitted to F0 and intensity. The results showed that phonetic differences between nouns and verbs were still evident in duration, F0 and intensity after lexical fre-quency, speech rate and some other related factors were taken into consideration in the models. The second syl-lables of nouns were longer than those of verbs, and both syllables of nouns were higher in F0 and greater in intensity than those of verbs. Since the prosodic boundary, frequency and other factors were controlled for, the phonetic differences between nouns and verbs might be attributed to their differences in information load and num-ber of syllables. This study provided evidence that phonetic differences between nouns and verbs might be driven by the grammatical classes themselves and is not an epiphenomenon of other processes.CO 2023 Elsevier Ltd. All rights reserved.
Cite the source of the dataset as: Wichmann, Søren, Eric W. Holman, and Cecil H. Brown (eds.). 2022. The ASJP Database (version 20).
Human history is written in both our genes and our languages. The extent to which our biological and linguistic histories are congruent has been the subject of considerable debate, with clear examples of both matches and mismatches. To disentangle the patterns of demographic and cultural transmission, we need a global systematic assessment of matches and mismatches. Here, we assemble a genomic database (GeLaTo, or Genes and Languages Together) specifically curated to investigate genetic and linguistic diversity worldwide. We find that most populations in GeLaTo that speak languages of the same language family (i.e., that descend from the same ancestor language) are also genetically highly similar. However, we also identify nearly 20% mismatches in populations genetically close to linguistically unrelated groups. These mismatches, which occur within the time depth of known linguistic relatedness up to about 10,000 y, are scattered around the world, suggesting that they are a regular outcome in human history. Most mismatches result from populations shifting to the language of a neighboring population that is genetically different because of independent demographic histories. In line with the regularity of such shifts, we find that only half of the language families in GeLaTo are genetically more cohesive than expected under spatial autocorrelations. Moreover, the genetic and linguistic divergence times of population pairs match only rarely, with Indo-European standing out as the family with most matches in our sample. Together, our database and findings pave the way for systematically disentangling demographic and cultural history and for quantifying processes of shifts in language and social identities on a global scale.
Linguistic Clues to Kiowa-Tanoan Prehistory Michael A. Schillaci (bio), Logan D. Sutton (bio), Søren Wichmann (bio), and Sergi López-Torres (bio) 1. Introduction Linguistic data have great potential for contributing in unique ways to our understanding of prehistory, both regionally and globally. In the American Southwest, there are several notable examples of linguistic research contributing to archaeological reconstructions of cultural history (e.g., Davis 1959; Hill 2002, 2008; Merrill et al. 2009; Ortman, 2012; Shaul 2014; McNeil and Shaul 2018; Ortman and McNeil 2018). In particular, linguistic data have played an important role in the study of Puebloan prehistory. For example, Davis (1959) used the proportion of shared cognates determined through analysis of sound correspondences, phonetic similarity, and semantic affinity to estimate the historical relationships within the Kiowa-Tanoan and Keresan language families, and to generate glottochronological estimates of language divergence (cf. also Hale and Harris 1979). More recently, following Hale (1967) and others (Trager 1942; Davis 1979, 1989), Ortman (2012) examined sound correspondences for vowels and consonants among Kiowa-Tanoan languages. Using the comparative method, Ortman identifies phonological innovations shared among sets of languages. He then uses the pattern of shared innovations to estimate historical relationships among Kiowa-Tanoan [End Page 255] languages, and the sequence of protolanguage divergence. Ortman (2012) also uses shared cognates to reconstruct protolanguage terms for a large set of items including plants, animals, cultigens, and items of material culture that are represented and well dated in the archaeological record. These reconstructed terms were used to date protolanguages within the language family, and to locate their geographic homeland based on current and historical distributions of plant and animal species (also see Ortman and McNeil 2018). In the present research we examine the historical relationships among Kiowa-Tanoan languages using tree-like diagrams generated from phylogenetic analysis of phonological data, as well as measures of linguistic dissimilarity derived from lexical data. We estimate the timing of branching events marking linguistic divergence within the language tree by employing an alternative to glottochronology. Following Ortman (2012), we also estimate the timing of language divergence within the Kiowa-Tanoan language family through a qualitative analysis of shared cognates using a modified "words and things" approach. We also use lexical and cognate data to delineate the Kiowa-Tanoan and Tanoan homelands. Our intention in this paper is to contribute to the study of Kiowa-Tanoan prehistory by presenting chronological and geographic contexts based on linguistic data that will inform archaeological inquiry. The eleven tables appear together after the text of this paper. 2. The Kiowa-Tanoan Language Family The Kiowa-Tanoan language family consists of four primary, undisputed branches with languages that are still spoken today: Kiowa (K), Tewa, Tiwa, and Towa (To). Linguists further divide Tewa into two substantially distinct varieties, Arizona Tewa (AT) and Rio Grande Tewa (RGT), the latter comprising five extant dialects,1 and Tiwa into three main varieties, Taos Northern Tiwa (Ta), Picuris Northern Tiwa (Pi), and Southern Tiwa (ST), with the last consisting of three extant dialects.2 Linguistic specialists therefore tend to treat the family as consisting of seven extant and welldocumented languages, two of which have notable dialect variation (cf. Sutton 2014; Kroskrity 1993; Trager 1967), although much of the literature nominally considers only the four main branches as distinctive, with all variation within Tewa and Tiwa being treated as "dialects." The distinction is one of degree of detail, not of kind, but should be noted, particularly in consideration of time depth of linguistic diversification. [End Page 256] The hyphenated name of the language family reflects both a culturalgeographic distinction and the history of scholarship. The Kiowa language is spoken by the Kiowa people, who historically migrated through the Central and Southern Plains and currently reside in Oklahoma. The Tanoan branches, Tewa, Tiwa, and Towa, are currently spoken across eleven Pueblo communities in the Rio Grande Valley of New Mexico, as well as one community in Arizona and one in El Paso, Texas (figure 1). Another Tanoan language, Piro, which is no longer spoken, was minimally documented by surveyors in the late 1800s in a few place names, two overlapping word lists (approximately...
This work presents an information-theoretic operationalisation of cross-linguistic non-arbitrariness. It is not a new idea that there are small, cross-linguistic associations between the forms and meanings of words. For instance, it has been claimed (Blasi et al., 2016) that the word for tongue is more likely than chance to contain the phone [l]. By controlling for the influence of language family and geographic proximity within a very large concept-aligned, cross-lingual lexicon, we extend methods previously used to detect within language non-arbitrariness (Pimentel et al., 2019) to measure cross-linguistic associations. We find that there is a significant effect of non-arbitrariness, but it is unsurprisingly small (less than 0.5% on average according to our information-theoretic estimate). We also provide a concept-level analysis which shows that a quarter of the concepts considered in our work exhibit a significant level of cross-linguistic non-arbitrariness. In sum, the paper provides new methods to detect cross-linguistic associations at scale, and confirms their effects are minor.
Words in utterance-final positions are often pronounced more slowly than utterance-medial words, as previous studies on individual languages have shown. This paper provides a systematic cross-linguistic comparison of relative durations of final and penultimate words in utterances in terms of the degree to which such words are lengthened. The study uses time-aligned corpora from 10 genealogically, areally, and culturally diverse languages, including eight small, under-resourced, and mostly endangered languages, as well as English and Dutch. Clear effects of lengthening words at the end of utterances are found in all 10 languages, but the degrees of lengthening vary. Languages also differ in the relative durations of words that precede utterance-final words. In languages with on average short words in terms of number of segments, these penultimate words are also lengthened. This suggests that lengthening extends backwards beyond the final word in these languages, but not in languages with on average longer words. Such typological patterns highlight the importance of examining prosodic phenomena in diverse language samples beyond the small set of majority languages most commonly investigated so far.
When exploring diachronic corpora, it is often beneficial for linguists to pinpoint not only the first or the last attestation dates of certain linguistic items, but also the moments in which they become more strongly established in the corpus or, conversely, the moments in which they, despite still being part of the language, become obsolete. In this paper, we propose an algorithm to assist the identification of such periods based on the frequency of items in a corpus. Our simple and generalisable algorithm can be used for the investigation of any linguistic item in any corpus which is divided into time-frames. We also demonstrate the applicability of our method using lexical data from the Corpus of Historical American English (coha), providing case studies on the statistics and characteristics of words that appear in or disappear from this corpus in different periods.
The present work is aimed at (1) developing a search machine adapted to the large DReaM corpus of linguistic descriptive literature and (2) getting insights into how a data-driven ontology of linguistic terminology might be built. Starting from close to 20,000 text documents from the literature of language descriptions, from documents either born digitally or scanned and OCR’d, we extract keywords and pass them through a pruning pipeline where mainly keywords that can be considered as belonging to linguistic terminology survive. Subsequently we quantify relations among those terms using Normalized Pointwise Mutual Information (NPMI) and use the resulting measures, in conjunction with the Google Page Rank (GPR), to build networks of linguistic terms.