
Due to the absence of an objective methodology and empirical research on measuring cohesion and coherence in Turkish, this study aims to establish a framework for the objective assessment of textual coherence and cohesion in this agglutinative language. To achieve this, the coherence, cohesion, and comprehensibility of 25 texts were first evaluated by 60 experts in language education. Next, structural text features, NLP-based similarity models, and large language models with potential to determine coherence and cohesion were identified and applied to the texts. The analyses highlighted the factors that can be employed in evaluating coherence and cohesion, and regression equations were constructed using these variables. Consequently, for the first time, objective and empirical methodologies have been developed to assess the coherence and cohesion of Turkish texts, while accounting for the structural characteristics of agglutinative languages.
This corpus-based multifactorial study aims to explore potential predictors of repetition versus lexical variety in translation of repeated reporting verbs from English into Slovak in literary novels. First, we provide a theoretical overview of research on repetition and reporting verbs in Slovak and Czech translation studies. Next, using a sample of 14 literary novels extracted from InterCorp corpus (v.15), we fit multiple negative binomial regression models with mixed effects to assess the effect that several predictor variables (frequency, semantic category, verb length, number of senses, translators) have on the response variable, i.e. the number of Slovak target-text reporting verbs an English source-text (ST) reporting verb is translated into. The findings revealed that factors such as frequency of use of ST reporting verbs, the semantic category of neutral ST reporting verbs, as well as the translators as a random effect, influence Slovak translators’ decisions of using a wide variety of Slovak reporting verbs instead of preserving the originals’ patterns of repetition. More precisely, the model allowed us to explain 70% of the variation (per conditional r-squared) in the response variable. Against the backdrop of prevailing stylistic norms in Slovak, the findings shed light on the translator’s choices in rendering recurrent reporting verbs introducing direct speech, a stylistically salient feature of literary texts.
Zipf's law and similar frequency laws have been studied in many languages, but their behavior in Uzbek has not been investigated. In this paper, we examine the fit of the generalized Karlin model of Zipf's law to Uzbek texts. For 386 texts consisting of three different genres (prose, poetry, and newspapers), we compute the T_n and H_n statistics and their analogs in the first half of the text and propose a new goodness-of-fit statistic Q_n based on their joint asymptotic behaviour. Our results show that model fit varies systematically with text length and genre: newspapers and poetry, which are typically shorter and more thematically compact, fit the model significantly better than long narrative prose. These results clarify how the Karlin model works in Uzbek texts and provide an empirical baseline for future comparative studies of understudied languages, including agglutinative ones. We also compare the new statistic with two previously proposed tests, based on type and hapax, on ten English reference texts. The results show that Q_n produces virtually identical p-values to the second test and leads to the same accept/reject decisions at standard significance levels, while the first test is systematically more conservative.
This paper investigates syntactic complexity across eight genres in the works of Czech writer Karel & Ccaron;apek using a stylometric approach. The aim of the study is to determine whether syntactic complexity differs among the following genres: novels, short stories, travelogues, poems, newspaper columns, academic essays, children's literature, and personal correspondence. For the analysis, we compute five core complexity metrics: Average Sentence Length (in words and in clauses), Average Clause Length, Mean Dependency Distance (MDD), and Mean Hierarchical Distance (MHD). The results indicate systematic genre-specific differences in syntactic patterning within & Ccaron;apek's literary work. Academic literature and travelogues consistently occupy the upper end of the complexity scale, whereas novels and short stories show comparatively low complexity, in line with their narrative and accessibility-oriented functions. Poetry is characterised by relatively low average complexity but high variability, and children's literature emerges as unexpectedly complex across several metrics. Personal correspondence and newspaper columns take up intermediate positions. Overall, the findings demonstrate that syntactic complexity offers a useful structural signal for distinguishing genres within a single author and underscore the value of syntactic features in stylometric and genre-oriented research.
Text complexity assessment is an important applied problem that lacks a comprehensive solution. Studies vary in the text corpora used, the features analyzed, the analysis algorithms applied, and the assessment methods employed. We present a new methodology for complexity assessment, developed based on a representative collection of Russian-language school textbooks compiled by the authors. Textual complexity is represented numerically by a textbook's target class grade. First, we evaluate the average number of words per sentence and the average number of syllables per word for their ability to predict text complexity and implement machine learning methods, including linear regression and neural networks. We confirm the linear Flesch-Kincaid complexity formula in the sense that it cannot be improved upon. The influence of text segmentation on the assessment results is also examined, and several approaches to utilizing such information are presented in the paper. This approach gave us approximately 10% improvement in the R2 metric compared to the baseline model. We also examine a total of 47 linguistic parameters as predictors of complexity and evaluate their significance.
Zipf's law of abbreviation, namely the tendency of more frequent words to be shorter, has been viewed as a manifestation of compression, i.e. the minimization of the length of forms – a universal principle of natural communication. Although the claim that languages are optimized has become trendy, attempts to measure the degree of optimization of languages have been rather scarce. Here we present two optimality scores that are dualy normalized, namely, they are normalized with respect to both the minimum and the random baseline. We analyze the theoretical and statistical advantages and disadvantages of these and other scores. Harnessing the best score, we quantify the degree of optimality of word lengths per language. This includes parallel texts in 20 languages of 9 families, written in 8 scripts, as well as spoken data for 46 languages of 12 families, two constructed languages, and one isolate. Our analyses indicate that languages are optimized to 62 or 67 percent on average (depending on the source) when word lengths are measured in characters, and to 65 percent on average when word lengths are measured in time. In general, spoken word durations are more optimized than written word lengths in characters. Our work paves the way to measure the degree of optimality of the vocalizations or gestures of other species, and to compare them against written, spoken, or signed human languages.
This study investigates how the word-frequency distributions in spoken language reflect cross-linguistic and cognitive differences in older adults. We analyzed Cookie Theft picture descriptions from 96 older adults: 48 Mandarin speakers (24 cognitively impaired and 24 cognitively normal) and 48 English speakers (24 cognitively impaired and 24 cognitively normal) and modeled their word frequency distributions using three functions: Zipf, Zipf-Mandelbrot, and Exponential model. All three models showed excellent goodness of fit at both group and individual levels, indicating that the basic Zipfian structure of lexical distributions is preserved in late life and is not disrupted by mild cognitive impairment. As for the fitting parameters, however, the decay parameter a in the Exponential and Zipf models consistently distinguished Mandarin from English, suggesting that language-specific lexical patterns are robustly encoded in the slope of the distribution but that adding a shift parameter can dampen how clearly a reflects them. By contrast, differences between cognitive groups were weak and inconsistent, implying that parameter a provides only a coarse and context-dependent reflection of cognitive status in short, constrained picture-description tasks.
The study valuates how lexical diversity differs across language proficiency levels (A1-C1 according to the CEFR). The material used in the research comes from the CzeSL-SGT learner corpus belonging to the Czech National Corpus. This dataset contains more than 8,000 Czech texts written by non-native speakers of different proficiency levels. Moving Average Type-Token Ratio (MATTR) is used to calculate lexical diversity in this study. The results indicate that lexical diversity increases with writers' proficiency. There is also a significant difference between the development of lexical diversity of Slavic and non-Slavic native speakers.
This study explores Zipf's laws of meaning and semanticity in Catalan child language acquisition, focusing on the interaction between syntactic regularities and semantic relationships. Building on previous research on semantic organization and Zipfian distributions in adult speech, using the CHILDES database, we analyse longitudinal corpora of Catalan-speaking children to test whether these statistical laws also emerge early in language development. Statistical and computational analyses show that rank-frequency distributions and related linguistic laws (Zipf's law, the Brevity law, and Heaps-Herdan's law) hold across different age groups and interaction contexts, whereas semantic regularities exhibit a weaker frequency-meaning correlation among younger speakers. However, the measure of semanticity captures the joint evolutionary changes in meaning and structural organization during early language acquisition.
Nabokov, known primarily for his prose, was also a remarkable poet. The article aims to examine the semantic (lexical) features of Nabokov's images in his lyrics. These characteristics include 23 semantic (thematic) classes of words that fill two positions of the figurative model-the position of the Target of the metaphoric transfer and the position of its Source. The material includes 4 lyrical collections by Nabokov, published at different stages of his creative path-an early collection (published when he lived in Russia), 2 later collections from the Berlin period (middle stage) and one collection of mature creative activity when Nabokov lived in the USA. The article examines the relationships between the features of the two positions of the image, the distribution of semantic classes of words in the image system, and changes in the frequencies of these classes over time.
We classify texts using relative word frequencies. The task is to distinguish human-written texts from those generated by a computer using modern algorithms. We study two essay datasets, each containing an equal number of human-written and computer-generated essays. Studying Zipf diagrams shows that the generated texts have a significantly smaller vocabulary compared to human ones. However, the relative frequency of rare words (not included in the 1000 most common) does not allow us to confidently classify the texts. As additional features, we used the relative frequencies of the four most frequent words, as well as the ratio of the number of hapax legomena to the total number of different words. This feature allows to significantly improve the classification. Using these six features allows us to fairly confidently determine whether the text is computer-generated.
Using large language models (LLMs), computers are able to generate a written text in response to a user request. As this pervasive technology can be applied in numerous contexts, this study analyses the written style of one LLM called GPT developed by OpenAI by comparing its generated speeches with those of the recent US presidents. To achieve this objective, the State of the Union (SOTU) addresses written by Reagan to Biden are contrasted to those produced by both GPT-3.5 and GPT-4.o versions. Compared to US presidents, GPT tends to overuse the lemma “we” and produce shorter messages with, on average, longer sentences. Moreover, GPT opts for an optimistic tone, choosing more often for political (e.g., president, Congress), symbolic (e.g., freedom), and abstract terms (e.g., freedom). Even when imposing an author’s style to GPT, the resulting speech remains distinct from addresses written by the target author. Finally, the two GPT versions present distinct characteristics, but both appear overall dissimilar to true presidential messages.
This study addresses a paradox in word order typology. On the one hand, the SOV order has longer dependency distances and therefore higher processing costs compared to verb-medial order. On the other hand, it is the most frequent word order in languages of the world. How come? A study of corpus data annotated with Universal Dependencies provides a simple answer: the costly long distances occur more rarely than one would assume because SOV clauses are infrequent in language use. A quanitative analysis of 150 Universal Dependencies corpora shows that the proportions of verb-final clauses with two overt core arguments are low across languages, including predominantly verb-final languages. Moreover, a series of Bayesian phylogenetic models based on comparable corpora in thirty-two languages show a negative correlation between the proportion of verb-final clauses in a language and the average number of arguments in a clause, while controlling for argument indexing and high-and low-context culture. A closer examination of argument configurations reveals a positive correlation between proportions of verb-final clauses and proportions of subject-less clauses; as for proportions of objectless clauses, the evidence is less clear. The study highlights the importance of the token-based, gradient approach to typology, which gives us insights into what kind of structures language users prefer, and what they avoid.
The syntactic structure of a sentence can be represented as a graph, where vertices are words and edges indicate syntactic dependencies between them. In this setting, the distance between two linked words is defined as the difference between their positions. Here we wish to contribute to the characterization of the actual distribution of syntactic dependency distances, which has previously been argued to follow a power-law distribution. Here we propose a new model with two exponential regimes in which the probability decay is allowed to change after a break-point. This transition could mirror the transition from the processing of word chunks to higher-level structures. We find that a two-regime model-where the first regime follows either an exponential or a power-law decay-is the most likely one in all 20 languages we considered, independently of sentence length and annotation style. Moreover, the break-point exhibits low variation across languages and averages values of 4-5 words, suggesting that the amount of words that can be simultaneously processed abstracts from the specific language to a high degree. The probability decay slows down after the breakpoint, consistently with a universal chunk-and-pass mechanism. Finally, we give an account of the relation between the best estimated model and the closeness of syntactic dependencies as function of sentence length, according to a recently introduced optimality score.
The paper focuses on an ongoing discussion about which manifestations of the “word”, either (1) tokens, (2) types or (3) lemmas, are more appropriate for studying the relationship between the length of words and the lengths of their constituents from the point of view of the Menzerath-Altmann law. Empirically, the word-syllable relationship is studied on the basis of Russian and English texts. In addition, the core vocabularies of both languages are taken into account. The given choice of languages gives the possibility to study the Menzerath-Altmann law in languages with different morphological coding strategies. Although the results are both language and level specific, the general tendency is that the regulation of constituent length is something that is more valid at the language system level, i.e. word types or lemmas are more appropriate units.
Chinese online literature has rapidly gained popularity, displaying characteristics that differ from traditional literature, and has now become an important and unique literary genre. Chinese readers of online literature account for nearly half of the total number of Chinese internet users, approximately 537 million. This study aims to explore the linguistic characteristics of online literature through a combination of qualitative and quantitative methods. The study builds a large-scale corpus from the largest Chinese online literature website Qidian, focusing on six quantitative linguistic indicators: word frequency, lexical richness, lexical density, word vectors, punctuation distribution, and average paragraph length. Additionally, three text mining indicators—thematic concentration, information entropy, and authorial perspective—are introduced. The findings reveal that online literature features a casual and informal language style, characterized by a high frequency of function words, shorter average paragraph lengths, and significant reliance on dialogue. Furthermore, the thematic concentration is generally low, which is related to the verbosity of online works and their emphasis on narrative flow rather than detailed descriptions. Overall, this paper provides empirical data to enhance the understanding of the linguistic features of online literature and suggesting that future research could adopt broader metrics and more comprehensive text mining techniques.
Validity of the Menzerath-Altmann law has been confirmed in a number of works for lan-guages with different morphology. This research is aimed at disclosure of potential correla-tions between the structure of the Tatar word form and the average length of its constituent syllables; taking into consideration that the Turkic language family is rarely represented in quantitative linguistics, we tested the law using data from ten Tatar texts, both poetry and prose, to interpret the results from the point of view of grammar. To assess the goodness of fit of the model we applied the coefficient of determination R2 which for different texts ranged from 0.676 to 0.999. The study revealed that the examined data in Tatar generally abide by G. Altmann’s formula; the average syllable length depended both on its position in the word and on word length, and joining more affixes provided decrease in the average syl-lable length. For individual texts, while the law as a trend was observed, we discovered cer-tain fluctuations when the average syllable length for sufficiently long words was greater than for relatively short ones. Analyzing the whole corpus data we distinguished between tokens (with frequencies taken into account) and types (unique word forms considered once); the research proved the Menzerath-Altmann law valid in both cases, with better result for types.
Sentence-final particles (SFPs) are pervasive in spoken Cantonese for expressing speakers’ attitudes. This study explores the global and local features of SFPs in both spoken Cantonese and Mandarin Chinese using two part-of-speech-based (POS-based) dependency networks. Results show that (1) globally, spoken Cantonese and Mandarin Chinese networks exhibit centralization and scale-free properties, reflecting the communication efficiency of human languages. However, spoken Cantonese manifests weaker centralization properties, as demonstrated by the diversity of its edges. Moreover, SFPs in spoken Cantonese have a greater degree and in-degree than those in Mandarin Chinese, indicating a stronger ability to form syntactic connections with other POSs in the language structure. (2) locally, Cantonese SFPs display more extensive mood expression devices, notably differing in PART-NOUN (discourse:sp), PART-ADV and PART-ADJ dependencies compared to Mandarin Chinese. Additionally, specific examples illustrate how Cantonese SFP usage differs from Mandarin Chinese, showcasing their distinct discourse functions. The findings suggest that communication efficiency is a cross-lingual universal, while spoken Cantonese is distinctive in its use of diverse SFPs to express moods. This study may shed new light on adapting the complex network approach to explore the similarities and differences across human languages.
This study examined the translation style of David Hawkes and the Yangs (Yang Xianyi and Gladys Yang) in their English translations of Hongloumeng, a Chinese Great Classic, by considering the hybrid register nature of fiction. The activity index, a measure from quanti-tative linguistics that calculates the ratio of verb occurrences to the sum of verb and adjec-tive occurrences, was used to analyze the active-descriptive equilibrium patterns across the two Hongloumeng translations and the two sub-registers of fiction. Our analysis is based on a corpus that separates fictional narration and dialogues from the first 80 chapters of the two Hongloumeng translations. The study found that, overall, dialogues tend to be more active than narration and Hawkes’ translation was characterized by a greater level of activity com-pared to the Yangs’ version. Subsequent analysis revealed that Hawkes' translation dis-played a higher level of activity in fictional dialogues while demonstrating a more descrip-tive approach in fictional narration. The results suggest that Hawkes’ translation adheres more closely to the typical stylistic conventions of fiction writing in English. The stylistic differences between the two translations of Hongloumeng are believed to be a result of a combination of factors, including the translators’ language and cultural backgrounds and their choice of translation strategies and approaches, which may have contributed to the var-iations in the final translated products.
The present paper focuses on the statistical modelling of loanword frequencies in different semantic fields of ten selected languages of the world. The main question addressed in this paper is whether the occurrence of loanwords in different semantic fields (data are from the WOLD project) can be modelled by a (simple) power model or not. Within the framework of synergetic linguistics, the basic idea is that the transfer and integration of loanwords is a statistically well-regulated process. In the absence of deductively found models (theoretically, also a birth-and-death process could be a suitable starting point), the proposed modelling aims to shed some new light on the underlying statistical tendencies in linguistic borrowing.