
Abstract For many years, the uncontested opinion in the field has been that there is not an empirical linguistic difference, but rather a societal one, between dialects and languages. Recently, however, it has been reported that distances between linguistic varieties are typically bimodal, suggesting that a classification into language vs. dialect might not be exclusively a social construct. Here, I review these papers and attempt to replicate them with a larger and more suitable dataset from Romance languages. Results fail to reveal any empirically salient boundary between languages and dialects. A look at other language families supports this conclusion. Simulations suggest that the distribution of linguistic distances in a family is primarily driven by its phylogenetic structure, with no additional mechanisms needed beyond vertical and horizontal diversification.
Abstract This usage-based investigation analyzes the changing frequency of a periphrastic future construction of Persian over four centuries. The data originates from a diachronic dataset of New Persian, sampled from 55 texts of the 10th to 20th centuries. The statistical results demonstrate a frequency rise in this construction’s usage in the 17th and 18th centuries (phase 1), followed by a frequency drop in the 19th and 20th centuries (phase 2). It is argued that the initial upward trend could have been motivated by systematicity (system pressure), affecting the functionally expanding prefix of imperfective aspect. Subsequently, the two future-marking forms are engaged in a competitive situation from the 18th century, which is observed through the interchangeability feature of the verbs. Ultimately, this competition is resolved, or is in the process of being resolved, due to the economic advantage of the generalized prefix and the higher cost associated with the periphrastic construction.
Colexification—the expression of distinct concepts by a single word form—provides a valuable window onto how lexical meanings evolve. Using phylogenetic comparative models, we analyze lexical data from three language families—Austronesian, Indo-European, and Uralic—to investigate the evolutionary dynamics underlying colexification patterns. We assess the effects of three predictors: semantic relatedness, borrowability, and usage frequency. Our results suggest that more closely related concept pairs tend to be colexified across a larger portion of the family trees and change more slowly. Concept pairs that are more frequent and more prone to borrowing show a tendency to evolve more rapidly and be colexified less often, though these effects are not fully consistent across families. Differences between the language families suggest that areal and cultural factors also play a role.
This usage-based investigation analyzes the changing frequency of a periphrastic future construction of Persian over four centuries. The data originates from a diachronic dataset of New Persian, sampled from 55 texts of the 10th to 20th centuries. The statistical results demonstrate a frequency rise in this construction's usage in the 17th and 18th centuries (phase 1), followed by a frequency drop in the 19th and 20th centuries (phase 2). It is argued that the initial upward trend could have been motivated by systematicity (system pressure), affecting the functionally expanding prefix of imperfective aspect. Subsequently, the two future-marking forms are engaged in a competitive situation from the 18th century, which is observed through the interchangeability feature of the verbs. Ultimately, this competition is resolved, or is in the process of being resolved, due to the economic advantage of the generalized prefix and the higher cost associated with the periphrastic construction.
This article studies variability in the internal structure of verbal bound person-number paradigms within a sample of 174 Western Iranic language varieties. Twenty-one patterns of syncretism are found within the sample, highlighting the neutralization of opposition between cells in the inflectional person/number paradigm. The syncretic patterns in these languages range from complete opposition in the person/number paradigm to partial syncretism in certain feature values, diagonal systems, and the complete neutralization of person, attested in the Koceri variety of Northern Kurdish. The article reflects on the effect of sound change, analogy, and areal and contact factors in developing syncretic person-number patterns within Iranic.
This article traces the evolution of differential object marking (DOM) in Romani, an Indo-Aryan language primarily spoken in Europe in contact with different languages. Drawing on dialectological data from 119 locations in Europe, we demonstrate that D OM in Romani dialects is generally a stable feature constrained by the pronoun vs. noun distinction, animacy, and in some conservative varieties definiteness. By comparing Romani with other Indo-Aryan languages, we propose a diachronic scenario. An indirect object case marker expands into marking direct objects, starting with pronouns, then definite animate nouns, and finally, encompassing all animate nouns. This developmental sequence is reflected in the distribution of DOM-constraining factors in contemporary dialects. By contrast, Romani dialects of Finland are undergoing the process of losing D OM, while Italian varieties have lost it completely by forfeiting nominal case inflection altogether. However, in the varieties of southern Italy this loss is compensated by a new prepositional DOM, a pattern replication from dialectal Italian.
This paper introduces a new corpus-based approach for studying word order change and reconstruction with Bayesian computational phylogenetic methods. We investigate the rates of change in object-verb order in 46 sentences from 36 Indo-European languages extracted from a parallel corpus. A Gaussian mixture model reveals that the rates can be grouped into three components representing syntactic constructions with distinct diachronic dynamics. Contexts with nominal objects are relatively stable, whereas object-verb order in contexts with pronominal objects evolves fast. Complement clauses have a strong diachronic bias towards VO. Stochastic character mapping suggests a VO order of nominal objects and verb in Proto-Indo-European, while the fast rates in contexts with pronominal objects do not allow a reliable reconstruction of ancestral states.
In this paper, we study how phonotactic constraints can play a role in contact-induced morphological simplification. We hypothesize that, in languages with constraints on certain consonant sequences, L2 speakers acquiring the language overgeneralize these constraints and avoid consonant sequences formed by morphological affixing. The resulting process of reduction of such sequences could then lead to loss of morphological affixes. We evaluate this hypothesis using data from Alorese, an Austronesian language spoken in eastern Indonesia, which underwent morphological simplification and has a large proportion of L2 speakers. Results from an agent-based model, which allows for mechanisms to be tested in isolation, indicate that phonotactic mechanisms indeed play a role in contact-induced morphological simplification. Phonotactic mechanisms lead to the strongest simplification when combined with a generalization mechanism, which spreads the simplified forms across different verbs. The model results underline the role of multiple interacting mechanisms in language change.
In this paper, we study how phonotactic constraints can play a role in contact-induced morphological simplification. We hypothesize that, in languages with constraints on certain consonant sequences, L2 speakers acquiring the language overgeneralize these constraints and avoid consonant sequences formed by morphological affixing. The resulting process of reduction of such sequences could then lead to loss of morphological affixes. We evaluate this hypothesis using data from Alorese, an Austronesian language spoken in eastern Indonesia, which underwent morphological simplification and has a large proportion of L2 speakers. Results from an agent-based model, which allows for mechanisms to be tested in isolation, indicate that phonotactic mechanisms indeed play a role in contact-induced morphological simplification. Phonotactic mechanisms lead to the strongest simplification when combined with a generalization mechanism, which spreads the simplified forms across different verbs. The model results underline the role of multiple interacting mechanisms in language change.
This paper presents the design rationale and pilot demonstration of the GramAdapt Social Contact questionnaire; a research tool developed for collecting global comparative sociolinguistic data on language contact scenarios. The questionnaire is qualitative with quantitative potential, inviting language community experts to provide best-assessment answers to questions about social contact in their communities of expertise. The main purpose is to compare contact scenarios, however the questionnaire can also be a broad survey of any given contact situation as it was designed to target factors associated with language contact and change phenomena at large. Two experts of small-scale multilingual communities answer an abridged version of the questionnaire to qualitatively demonstrate this proof of concept. The experts of Mawng and Kunbarlang (northern Australia), and Tundra Enets and Nganasan (northern Siberia) were chosen as these communities defy nation-based models of multilingualism. The responses are broadly successful, thus demonstrating the theoretical contribution and methodological potential of this questionnaire.
Abstract This article focuses on languages of the Kwilu-Ngounie subbranch within a branch of the Bantu language family known as West-Coastal Bantu. Within Kwilu-Ngounie, B70 and B80 languages emerge as paraphyletic in the most comprehensive lexicon-based phylogeny of the branch. We assess whether the impossibility to group them into lexicon-based monophyletic subgroups can be bypassed by using the phonological innovation of word-final loss of Proto-Bantu *ŋg as diagnostic of a new subgroup. It is hard to tell whether this new subgroup is a clade by descent or instead a taxon resulting from a contact-induced innovation affecting related varieties. The unconditioned reflexes of *ŋg across varieties signal that both language-internal lexical diffusion and contact-induced crosslinguistic spread of phonological innovation thwart the Neogrammarian axiom of flawlessly regular sound change. Beyond its relevance for low-level Bantu subgrouping, this article contributes to the methodological issue of conflicting lexical and diachronic phonological evidence for internal classification.
We present a novel approach to identifying individual pairs of phonetic correspondences in a dataset of dialect pronunciations. This continues work identifying shibboleths (i.e., characteristic features of a given dialect), a category that has interested dialectology and that dialectometrical research has examined mostly in the form of categorical data or entire phonetic transcriptions. This article reaches into segmental sequences (phonetic transcriptions) to identify individual phonetic correspondences. We follow earlier work in examining how distinctive and how representative a given phonetic correspondence is for a selected group of varieties. We proceed from string alignments, and innovate in characterizing the important notions via information theory. Despite minor problems, the method improves on the generality of competing approaches and can be shown to be useful in detecting characteristic phonetic correspondences in Tuscan varieties. We argue that this facilitates deeper investigation into the relation between aggregating approaches to dialectology and approaches proceeding from features.
Although code-switching has been quite well studied as a worldwide phenomenon, closer attention to effects on the more localized language involved is needed, especially in the repertoires of younger, well-educated speakers speaking in a multilingual mode. We argue that their language shows creativity going well beyond older instances of borrowing and code-switching into a "third space" grammar, which shows an active and creative synthesis of at least two languages. This study is based on interview data with 37 speakers of Xhosa in Soweto, South Africa. It focuses on (a) the verb suffix -isha as a marker of multilingual Xhosa par excellence, (b) a new lease of life given to the class 14 prefix ubu- in connection with (mainly) Latinate adjectives from English, and (c) the interchangeability of logical connectors across the third space.
This article focuses on languages of the Kwilu-Ngounie subbranch within a branch of the Bantu language family known as West-Coastal Bantu. Within Kwilu-Ngounie, B70 and B80 languages emerge as paraphyletic in the most comprehensive lexicon- based phylogeny of the branch. We assess whether the impossibility to group them into lexicon-based monophyletic subgroups can be bypassed by using the phonological innovation of word-final loss of Proto-Bantu *I]g as diagnostic of a new subgroup. It is hard to tell whether this new subgroup is a clade by descent or instead a taxon resulting from a contact-induced innovation affecting related varieties. The unconditioned reflexes of *I]g across varieties signal that both language-internal lexical diffusion and contact-induced crosslinguistic spread of phonological innovation thwart the Neogrammarian axiom of flawlessly regular sound change. Beyond its rele vance for low-level Bantu subgrouping, this article contributes to the methodological issue of conflicting lexical and diachronic phonological evidence for internal classification.
Computational methods of language dating make inferences about the divergence times of protolanguages by evaluating the patterns of inheritance in the vocabulary of modemlanguages, given the specification of a model of vocabulary evolution. We consider a model that describes vocabulary evolution as the replacement of traits by new traits from an infinite state space along a tree. This model has been introduced in previous literature but so far it has not been used in many applications. We give a general recursive algorithm for calculating likelihoods and argue that the model gives a more realistic representation of vocabulary evolution over time, compared to existing models like the Stochastic Dollo model. We also provide a case study demonstrating the model's potential applications.
Despite the abundance of tonal languages around the world, the diachrony of tone is still poorly understood, especially when compared to segmental sound change. This lacuna has contributed to the untested assumption that tones are inherently unstable and change unpredictably. This paper addresses the questions of whether tones change faster than segments and whether tones show less phylogenetic signal than segments in the Mixtec languages of southern Mexico. To this end, I created a database of tonal and segmental sound changes across a sample of 42 Mixtec languages. I calculated phylogenetic signal with the metric D and estimated rates of change with a hidden Markov model across a posterior sample of phylogenetic trees. The results show that the majority of tone changes show phylogenetic signal and that they generally do not change at a faster rate than segments.
The Totonac branch of the Totonacan (also known as Totonac-Tepehua) family is traditionally broken down into four divisions—Misantla, Northern, Sierra, and Lowland. Misantla is an obvious outlier, but the relationship among the remaining three, which comprise the Central Totonac division, is uncertain due to competing lines of evidence: lexical isoglosses group Sierra and Lowland against Northern while morphological changes appear to set Sierra off against the other two. The spatial distribution of the morphological innovations shows these not to be a coherent set of changes inherited from a common ancestor, but instead a series of successive innovations diffused in a wave-like pattern. This paper also demonstrates that the morphological innovations are more recent than the lexical changes, supporting the prior separation of Sierra-Lowland languages from Northern. The paper also explores the methodological issues associated with the classification of languages in close contact at shallow time depths.
Previous work using lexical data from around the world has suggested that distances between language varieties are distributed such that varieties are typically either rather similar, qualifying as dialects of the same language, or rather dissimilar, qualifying as different languages, with a scarcity of varieties that are around halfway similar. Using a potentially biased sample, Wichmann (2019) observed that there is a bimodal distribution of distances with two roughly normal distributions separated by a valley. Here we test whether a similar distribution is found when using another source of data and an unbiased sample drawn from the cells of a geographical grid (of central Europe). The data consists of 18 lexemes from 274 doculects. Using Bayesian beta regression and leave-one-out cross-validation, we show that the data follows a bimodal distribution which is robust to sampling, and also to at least some aspects of the data (coarse- vs. fine-grained phonetic transcriptions).
Word order is a central issue in the reconstruction of Proto-Indo-European syntax. Categorical approaches have proved to be inadequate because they postulate for the protolanguage a typological consistency which is absent in any of the attested daughter languages. Following recent research, we adopt a gradient approach to word order, which treats word order preferences as a continuous variable. We analyze four word order patterns based on data extracted from treebanks of ancient Indo-European languages. After presenting our results for AdpN/NAdp, GN/NG, AN/NA, and OV/VO, we draw a number of conclusions concerning variation within individual languages, crosslinguistic variation, and variation in diachrony that support the claim that variability should be taken as the normal state across languages, including reconstructed stages. We conclude that a non-discrete approach has the advantage of leading to a reconstruction that better conforms to the situation known from real languages, with variation as a key feature.
Migration events splitting speaker communities and establishing novel contact situations are among the major drivers of language variation and change. While the precise processes that lead to change cannot usually be determined for past events with any certainty, the study of minority and heritage language usage in apparent time may provide insight into the contribution of the linguistic behavior underlying the dynamics. We capitalize on this and compare parts of speech usage in Pear Story renarrations across Gheg Albanian speakers of three generations in Germanspeaking environments, applying methods from information theory. The results suggest that the changing conventions in parts of speech usage across generations and places of residence can be attributed to changing linguistic behavior within the speaker community in the migration setting. These findings highlight the impact of changing sociocultural embedding and the roles of vertical and horizontal transmission in language change