
Cohesive chains offer a sequential view of how dispersed cohesive resources are integrated, and chain distance further links discourse organization to processing constraints. However, the distributional regularities of chain distance remain under-described, obscuring debates about the relative status of referential and lexical chains. Based on long English abstracts, this study combines distribution fitting and sensitivity analysis to model chain-distance distributions and their interaction patterns. Results show that (1) referential-chain distances follow an extended logarithmic distribution, where theta indexes the trade-off between distances 1 and 2, and alpha reflects the mass beyond distance 3; lexical-chain distances follow a right-truncated negative binomial distribution, with k and p primarily governing the relative weights of distances below versus above 4. (2) A stable crossover threshold occurs around distances 4 +/- 1, with referential chains dominating before and lexical chains after. Sensitivity analysis identifies p and theta as primary drivers of threshold shifts, k as the regulator of handover abruptness, and alpha as the main controller of balance-band width. Overall, these patterns indicate that chain distance is self-organized to balance informational demands and cognitive load and that distribution parameters can act as control parameters characterizing the dynamic organization of discourse effectively.
Statistical laws have long been proposed to model word frequency distributions, the most famous being Zipf's law (the inverse proportionality between a word's rank and its frequency) and its extensions. This study evaluates whether a class of Zipfian models can adequately capture both the overall frequency spectra and linguistically interpretable properties (such as word repetition rate, vocabulary growth rate, and entropy) across 193 languages. Using a Bayesian framework, seven Zipfian formulations were compared by assessing their overall fit and their predictive performance on these linguistic measures. The results show that, for nearly all languages, at least one formulation provides a statistically plausible fit to the observed frequency spectrum. Yet, even though different models emerged as best fits for different languages, all can be closely approximated by a power-law function for sufficiently large frequency classes, indicating a strong presence of this law across languages. Furthermore, some models that fit the overall spectrum less well still captured individual linguistic measures effectively. A key methodological contribution of this study lies in expressing linguistic measures as posterior predictive distributions rather than point estimates, allowing uncertainty to be represented directly and offering a richer, probabilistic understanding of lexical structure and variation across languages.
This study aims to revisit the Menzerath-Altmann Law (MAL) at the sentence-clause-word level and the paragraph-sentence-clause level across nine languages from the Indo-European language family: Albanian (Albanian), Armenian (Armenic), Lithuanian (Balto-Slavic), Gaelic (Celtic), Danish (North Germanic), German (West Germanic), Urdu (Indo-Iranian), Italian (Italian Romance), and French (Western Romance), and across four registers: press, general prose, academic prose, and fiction. The study is based on nine one-million-token Stanza-annotated multilingual balanced corpora that followed the Brown Corpus sampling frame, which made the corpora highly comparable. The results show that the frequency distributions of the sentence and paragraph lengths are strongly right-skewed, regardless of language and register. At the sentence-clause-word level, MAL holds robustly across the nine Indo-European languages (R2 = 0.907 to 0.990). Nevertheless, at the paragraph-sentence-clause level, MAL is not consistently valid: only Albanian and Gaelic have R2 >= 0.7. MAL fits are register-sensitive. At the sentence-clause-word level, the MAL fit of fiction is significantly weaker than those of press, general prose, and academic prose. At the paragraph-sentence-clause level, only press shows moderately good fits (e.g., Danish, Gaelic, Lithuanian). The study also shows that MAL's cross-linguistic differences increase and goodness-of-fit decreases from sentence-clause-word level to paragraph-sentence-clause level.
The Menzerath-Altmann law describes an inverse relationship between the size of a linguistic construct and the average size of its constituents. While its validity has been widely confirmed for lower-level language units, its application to higher levels remains less explored and sometimes inconclusive. For the first time, this study not only investigates but also corroborates the validity of the Menzerath-Altmann law across a fine-grained hierarchical structure of linguistic units. In particular, the following units are considered: sentence - independent clause - clause - phrase - subphrase - chunk - word - syllable - phoneme. Using a corpus of written Czech, we confirm the validity of the law across this hierarchy.