Theories of predictive processing postulate that the brain continuously generates predictions about future linguistic input. Many different linguistic elements have been found to function as predictive cues. One such element is classifiers, i.e., morphemes that classify noun referents according to some semantic property. However, very few studies have so far investigated the predictive function of classifiers, and all existing studies focus on major East Asian languages. The present study investigates whether listeners use classifiers as cues to following nouns in Baniwa (Arawakan, Northwestern Amazonia). Baniwa has a system of 53 classifiers that primarily encode physical shape. In a response time experiment, participants heard numeral-classifier-noun phrases and had to choose which one of two images depicted the noun. We tested two variables: informativity and degree of constraint. With respect to informativity, we found that response times were faster when the classifier could be used to identify the target image. With regard to degree of constraint, classifiers denoting shape yielded faster response times compared to the generic classifier. The results show that classifiers have a predictive function in Baniwa, and, by contributing data from a previously unexplored language family and geographical region, suggest that prediction may be a function of classifiers cross-linguistically.
IntroductionThe directionality of semantic change is problematic in traditional comparative models of language reconstruction. Compared to, e.g., phonological and morphological change, the directions of meaning change over time are potentially endless and difficult to reconstruct. The current paper attempts to reconstruct the mechanisms of lexical meaning change by a quantitative model. We use a data set of 104 core concepts in 160 Eurasian languages from several families, which are coded for colexification as well as cognacy, including semantic change of lexemes in etymologies. In addition, the various meanings are coded for semantic relation to the core concept, including relations such as metaphor, metonymy, generalization, specialization, holonymy, and meronymy. Further, concepts are coded into classes and semantic properties, including factors such as animacy, count/mass, concrete/abstract, or cultural connotations, such as taboo/non-taboo.MethodologyWe use a phylogenetic comparative model to reconstruct the probability of presence at hidden nodes of different colexifying meanings inside etymological trees. We find that these reconstructions come close to meaning reconstructions based on the comparative method. By means of the phylogenetic reconstructions, we measure the evolutionary dynamics of meaning loss of co-lexifying meanings as well as concepts.Results and discussionThese change rates are highly varying, from almost complete stability to complete unstability. Change rates vary between different semantic classes, where for instance wild animals have low change rates and domestic animals and implements have high change rates. We find a negative correlation between taboo animals and change rate, i.e., taboo animals have lower change rates than non-taboo words. Further, we find a negative correlation between animacy and change rate, indicating that animate nouns have lower change rate than inanimate nouns. A further result is a negative correlation between change rate and degree of borrowing (borrowability) of concepts, indicating that lexemes that are more likely to be borrowed are less likely to change semantically. Among semantic relations, we find that metonomy is more frequent than any other change, including metaphor, and that a change from general to more specific is in all cases more frequent than the other way round.
While recent years have seen a substantial increase of studies investigating vocal iconicity in the lexicon of spoken languages, its presence in grammatical structures is poorly understood. This study investigates the presence of vocal iconicity in nominal classification systems by collecting nominal classification devices from the two main system types: 210 non-agreeing languages (126 families) and 151 agreeing languages (123 families). To detect overrepresentations of sound types in class meanings, the nominal classification devices were grouped according to comparable semantic categories, transcribed using comparable phonetic system, and analyzed through Bayesian mixed models. The strongest results were found for associations between nominal classification devices denoting flat and low, front, unrounded vowels, along with several weak associations relating to shape/size/quantity, function, humanness/animacy, and sex. These associations mostly correlate with previous vocal iconicity findings, but crucially, the involved nominal classification devices are mostly semantically typical for non-agreeing, for example, classifier, systems. These findings were attributed to structural differences between nominal classification system types, which result from grammaticalization processes, for example, phonetic erosion and semantic bleaching. Thus, increased formal predictability through grammatical agreement comes at a cost of semantic transparency which, in turn, dismantles the semantic prerequisites needed for vocal iconic associations to be operational.
Languages of diverse structures and different families tend to share common patterns if they are spoken in geographic proximity. This convergence is often explained by horizontal diffusibility, which is typically ascribed to language contact. In such a scenario, speakers of two or more languages interact and influence each other’s languages, and in this interaction, more grammaticalized features tend to be more resistant to diffusion compared to features of more lexical content. An alternative explanation is vertical heritability: languages in proximity often share genealogical descent. Here, we suggest that the geographic distribution of features globally can be explained by two major pathways, which are generally not distinguished within quantitative typological models: feature diffusion and language expansion. The first pathway corresponds to the contact scenario described above, while the second occurs when speakers of genetically related languages migrate. We take the worldwide distribution of nominal classification systems (grammatical gender, noun class, and classifier) as a case study to show that more grammaticalized systems, such as gender, and less grammaticalized systems, such as classifiers, are almost equally widespread, but the former spread more by language expansion historically, whereas the latter spread more by feature diffusion. Our results indicate that quantitative models measuring the areal diffusibility and stability of linguistic features are likely to be affected by language expansion that occurs by historical coincidence. We anticipate that our findings will support studies of language diversity in a more sophisticated way, with relevance to other parts of language, such as phonology.
All languages borrow words from other languages. Some languages are more prone to borrowing, while others borrow less, and different domains of the vocabulary are unequally susceptible to borrowing. Languages typically borrow words when a new concept is introduced, but languages may also borrow a new word for an already existing concept. Linguists describe two causalities for borrowing: need, i.e., the internal pressure of borrowing a new term for a concept in the language, and prestige, i.e., the external pressure of borrowing a term from a more prestigious language. We investigate lexical loans in a dataset of 104 concepts in 115 Eurasian languages from 7 families occupying a coherent contact area of the Eurasian landmass, of which Indo-European languages from various periods constitute a majority. We use a cognacy-coded dataset, which identifies loan events including a source and a target language. To avoid loans for newly introduced concepts in languages, we use a list of lexical concepts that have been in use at least since the Chalcolithic (4000-3000 BCE). We observe that the rates of borrowing are highly variable among concepts, lexical domains, languages, language families, and time periods. We compare our results to those of a global sample and observe that our rates are generally lower, but that the rates between the samples are significantly correlated. To test the causality of borrowing, we use two different ranks. Firstly, to test need, we use a cultural ranking of concepts by their mobility (of nature items) or their labour intensity and "distance-from-hearth" (of culture items). Secondly, to test prestige, we use a power ranking of languages by their socio-cultural status. We conclude that the borrowability of concepts increases with increasing mobility (nature), and with increased labour intensity and "distance-from-hearth" (culture). We also conclude that language prestige is not correlated with borrowability in general (all languages borrow, independently of prestige), but prestige predicts the directionality of borrowing, from a more prestigious language to a less prestigious one. The process is not constant over time, with a larger inequality during the ancient and modern periods, but this result may depend on the status of the data (non-prestigious languages often remain unattested). In conclusion, we observe that need and prestige compete as causes of lexical borrowing.
Authors: Gerd Carling, Sandra Cronhamn, Rob Farren, Rob Verhoeven (Lund University) Date of publication: 2018-02-09 Description includes datasets: Swadesh 100. URL https://diacl.ht.lu.se/WordListCategory/Details/100 Swadesh 200. URL: https://diacl.ht.lu.se/WordListCategory/Details/200 Culture words for South America. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Indo-European. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Austronesia. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Caucasus. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Basque. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Uralic. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Middle-Eastern non-IE. URL: https://diacl.ht.lu.se/WordList/Index/ Culture words for Turkic. URL: https://diacl.ht.lu.se/WordList/Index/
Languages borrow words when there is a need for it, all languages contain loanwords, and no part of the lexicon is entirely “loan-proof”. These are statements about lexical borrowing that are typically found in linguistic textbooks (Hock and Joseph 1996). Further, we know that that there are large discrepancies in the borrowability of different lexical concepts, where core vocabulary domains (sense perception, spatial relations, the body, kinship, and motion) in general are more resistant to borrowing, whereas culture-dependent domains (religion and belief, clothing and grooming, the house, law, social and political relations, agriculture and vegetation) belong to a more loan-intense part of the lexicon. We are also aware that there are large differences in borrowability between languages, something that has multiple connotations, including language history, populations size, language contact, grammatical structure, and so forth (Haspelmath and Tadmor 2009).Our study aims at investigating borrowability more carefully from a historical perspective, using quantitative and statistical methods, focusing on the families of Indo-European, Nakh-Dagestanian, Northwest Caucasian , Kartvelian, Uralic, and Turkic. We use a lexical dataset, which contains 100 lexical meanings each from the domains of basic vocabulary (Swadesh) and culture vocabulary (domains of agriculture, vegetation, food, warfare, hunting, animals, technology), compiled from around 250 languages of the previously mentioned families (around 25,000 lexemes in total). The lexemes of the dataset have been coded (manually, from dictionaries) for cognacy according to a tree model and are distinguished by borrowability versus inheritance at every historical stage in the tree structure. The dataset also systematically includes ancient language forms (including reconstructed ones), and traces continued development of cognates, involving further semantic change. Connected to the dataset, there is language metadata that includes geographic extension, time period, relative population size and family tree topology. Taken together, this makes the dataset a unique and yet unexplored source for investigating large-scale borrowability statistically. In our study, we are specifically aiming the following research questions:• What is the general level of borrowability in our vocabulary (culture words), compared to the average borrowability of the same lexical meaning concluded by the cross-linguistic study of (Haspelmath and Tadmor 2009)?• In our data, are there any internal differences between the borrowability of lexical concepts, depending on semantic domain?• Is there a general connection between loanword directionality and population size of languages?• Is the amount of borrowing generally equal over time (gradually increasing with increasing amounts of documentation), or are some periods and geographic areas more intense in borrowing?• What happens to lexemes upon borrowing? Do they continue to change with the language or are they more likely to be frozen, semantically and/or morphologically? Haspelmath, M. and U. Tadmor (2009). Loanwords in the world's languages: a comparative handbook. Berlin, Mouton de Gruyter.Hock, H. H. and B. D. Joseph (1996). Language history, language change, and language relationship : an introduction to historical and comparative linguistics. Berlin, Mouton de Gruyter. (Less)
The current paper describes the deictic system of Kamaiurá, a language of the Tupí-Guaraní family. The Kamaiurá system of deictic demonstratives and adverbials has a high degree of complexity, including at least 17 different forms, of which several have different functions. The system codes four levels of Participant deixis, with proximal, medial, distal and far distal deixis. Forms can also code anaphora and highly specialized locations of the referent, such as ‘moving away’ and ‘located beside something’. A further peculiar and unusual characteristic of the Kamaiurá system is the coding of Modal and Evidential deixis, which is found among the forms marking far distal deixis. Our study has two foci: the first part describes the system in its independent or exophoric use, and this part is based on deep interviews with native speakers and a deixis elicitation study. The second part of the paper represents the core of our study. Here, we investigate the uses of the deictic system in a recorded frog story, looking at anaphoric and cataphoric usages of the forms as well as how they are used to mark topic and focus in the narrative discourse. The text is very rich in deictic forms, and out of the 17 different forms recorded for Kamaiurá, 9 occur in our frog story. We notice a tendency where the hierarchy of increasing distance from the ego in the independent forms is transferred into increasing focus of the narrative. Epistemic modality of the independent forms is used to mark uncertainty in the narrative, i.e., to indicate lack of terms for a specific item, whereas anaphoric deixis of the independent forms marks general reference in the narrative.
In this paper, we have investigated, by means of quantitative and statistical methods, stability and change in cultural vocabulary of Indo-European in Europe, with a focus on agriculture. For this purpose we have created a culture vocabulary list with lexical head words, organized into subcategories based on their role and function in a cultural system, the purpose of which is to give a representative selection of culture vocabulary terms for a specific system and a certain geographic area. Thereupon, we have collected data from a number of Indo-European languages of Europe, removed languages with too little data, omitted post-colonial borrowings, organized the lexemes into cognate sets and divided lexemes according to whether they are inherited (reconstructed or derived from Proto-Indo-European roots), loaned, or have an uncertain origin. For each term we have kept track of number of cognates, number of lexemes in languages, as well as number of reconstructed Proto-Indo-European roots. The data sets were analyzed by the R statistical tool, basically by means of principal component analysis biplots, but also by calculating standardized residuals for each of the terms and the subgroups. The results demonstrated that there is, from a geographical perspective, relatively little convergence effect on cultural vocabulary. Further, we could see a clear tendency in which manufactured objects (implements, produce) as well as the activities accompanying them (activities) were inherited to a larger extent, whereas objects belonging to the environment (game), as well as the cultural environment (domestic animals, produce) was much more uncertain. The category of predator was most loaned in our set, which could, to a certain extent, be due to the inclusion of partly non-European species. We were also able to identify a stable core vocabulary, consisting mainly of implements, some produce and domestic animal terms, which were rich in cognates and leaning towards being inherited. (Less)
Vocabulary for subsistence and technology may vary a great deal in their degree of borrowability, depending on time, place, inherent subsistence and technology, and the situation of the borrowing. In cross-linguistic typological studies of borrowability, these words tend to group somewhere from middle to high in borrowability, depending on lexical concept (Haspelmath & Tadmor, 2009). We have compiled a set of 100 lexical concepts of importance to hunting, farming, and technology from a perspective of high age and presumed high stability from a cultural perspective. These concepts include, e.g., bovine cattle (BULL, OX, COW), animals of traction (HORSE, DONKEY), important metals (GOLD, IRON, COPPER), important crops (GRAIN, WHEAT), important game (HARE, DEER), essential technological innovations (WHEEL, WAGON). We have compiled a complete data set of lexemes from Indo-European, Caucasian (Kartvelian, Nakh-Dagestanian, Northwest Caucasian), as well as adjacent Uralic and Turkic languages, in all around 300 languages. In particular the Caucasian data is rich and new, based on fieldwork of poorly documented languages. The lexemes have been coded for etymology as well as for borrowing, lexical derivation and semantic change, and are amassed in a lexical cognacy database (Carling, 2017). Preliminary studies on the material indicate, first, that there is a high degree of inherited words for both farming and technology, which are paralleled and independent in both Indo-European and Caucasian families. Interestingly enough, we also find a great deal of vocabulary in the families that apparently have their roots in joint, very ancient migration words. Also, we notice that some words are similar between the families in the way they are derived (e.g., Proto-Kartvelian *borbal ‘wheel’, from *bor- ‘rotation’). An interesting parallel between Indo-European and Caucasian is that semantic change of culture words within etymologies follow almost identical principles, indicating a high cultural component in semantic change. Finally, we notice that borrowability may be high in certain areas and in certain languages, also targeting concepts of very high age, such as farming words. Much of this borrowing is relatively late (e.g., in Caucasian from Persian, Arabic, or Turkic), indicating that cultural impact may have played an important role in changing the vocabulary also for concept for which there must have been an inherent vocabulary. The presentation will look at particular concepts and lexemes, both inherited words, possible ancient loans, migration words, as well as later, obvious borrowings. Further, we will look at statistics on borrowability in general, and type and direction of semantic change of culture concepts of the data. Carling, G. (2017). DiACL - Diachronic Atlas of Comparative Linguistics Online (Publication no. https://diacl.ht.lu.se/). from Lund University https://diacl.ht.lu.se/ Haspelmath, M., & Tadmor, U. (2009). Loanwords in the world's languages: a comparative handbook (M. Haspelmath & U. Tadmor Eds.). Berlin: Mouton de Gruyter. (Less)