
Cohesive chains offer a sequential view of how dispersed cohesive resources are integrated, and chain distance further links discourse organization to processing constraints. However, the distributional regularities of chain distance remain under-described, obscuring debates about the relative status of referential and lexical chains. Based on long English abstracts, this study combines distribution fitting and sensitivity analysis to model chain-distance distributions and their interaction patterns. Results show that (1) referential-chain distances follow an extended logarithmic distribution, where theta indexes the trade-off between distances 1 and 2, and alpha reflects the mass beyond distance 3; lexical-chain distances follow a right-truncated negative binomial distribution, with k and p primarily governing the relative weights of distances below versus above 4. (2) A stable crossover threshold occurs around distances 4 +/- 1, with referential chains dominating before and lexical chains after. Sensitivity analysis identifies p and theta as primary drivers of threshold shifts, k as the regulator of handover abruptness, and alpha as the main controller of balance-band width. Overall, these patterns indicate that chain distance is self-organized to balance informational demands and cognitive load and that distribution parameters can act as control parameters characterizing the dynamic organization of discourse effectively.
Statistical laws have long been proposed to model word frequency distributions, the most famous being Zipf's law (the inverse proportionality between a word's rank and its frequency) and its extensions. This study evaluates whether a class of Zipfian models can adequately capture both the overall frequency spectra and linguistically interpretable properties (such as word repetition rate, vocabulary growth rate, and entropy) across 193 languages. Using a Bayesian framework, seven Zipfian formulations were compared by assessing their overall fit and their predictive performance on these linguistic measures. The results show that, for nearly all languages, at least one formulation provides a statistically plausible fit to the observed frequency spectrum. Yet, even though different models emerged as best fits for different languages, all can be closely approximated by a power-law function for sufficiently large frequency classes, indicating a strong presence of this law across languages. Furthermore, some models that fit the overall spectrum less well still captured individual linguistic measures effectively. A key methodological contribution of this study lies in expressing linguistic measures as posterior predictive distributions rather than point estimates, allowing uncertainty to be represented directly and offering a richer, probabilistic understanding of lexical structure and variation across languages.
This study aims to revisit the Menzerath-Altmann Law (MAL) at the sentence-clause-word level and the paragraph-sentence-clause level across nine languages from the Indo-European language family: Albanian (Albanian), Armenian (Armenic), Lithuanian (Balto-Slavic), Gaelic (Celtic), Danish (North Germanic), German (West Germanic), Urdu (Indo-Iranian), Italian (Italian Romance), and French (Western Romance), and across four registers: press, general prose, academic prose, and fiction. The study is based on nine one-million-token Stanza-annotated multilingual balanced corpora that followed the Brown Corpus sampling frame, which made the corpora highly comparable. The results show that the frequency distributions of the sentence and paragraph lengths are strongly right-skewed, regardless of language and register. At the sentence-clause-word level, MAL holds robustly across the nine Indo-European languages (R2 = 0.907 to 0.990). Nevertheless, at the paragraph-sentence-clause level, MAL is not consistently valid: only Albanian and Gaelic have R2 >= 0.7. MAL fits are register-sensitive. At the sentence-clause-word level, the MAL fit of fiction is significantly weaker than those of press, general prose, and academic prose. At the paragraph-sentence-clause level, only press shows moderately good fits (e.g., Danish, Gaelic, Lithuanian). The study also shows that MAL's cross-linguistic differences increase and goodness-of-fit decreases from sentence-clause-word level to paragraph-sentence-clause level.
The Menzerath-Altmann law describes an inverse relationship between the size of a linguistic construct and the average size of its constituents. While its validity has been widely confirmed for lower-level language units, its application to higher levels remains less explored and sometimes inconclusive. For the first time, this study not only investigates but also corroborates the validity of the Menzerath-Altmann law across a fine-grained hierarchical structure of linguistic units. In particular, the following units are considered: sentence - independent clause - clause - phrase - subphrase - chunk - word - syllable - phoneme. Using a corpus of written Czech, we confirm the validity of the law across this hierarchy.
Research on translation universals has traditionally focused on isolated linguistic features along paradigmatic dimensions due to ease of interpretation. However, syntagmatic approaches, which examine how linguistic elements combine sequentially, remain underexplored. This corpus-based study addresses this gap by analysing R-motifs, defined as recurring sequences of part-of-speech tags, across four genres in translated and native Chinese texts. We investigate both the rank-frequency distributions of R-motif types and motif lengths as potential indicators of translation universals. Our analysis shows that R-motif frequencies in both text types follow the right-truncated Zeta distribution, whereas motif length distributions conform to the P & oacute;lya model. Random Forests are used to establish the text classification model where texts are represented by the POS R-motif distribution parameters and attributes. The experiments show that the combination of features from distribution parameters and attributes can detect the translationese efficiently. Future research may extend this approach by exploring more granular features beyond part-of-speech sequences.
Machine translation (MT) systems are typically evaluated by comparing outputs to human references using metrics that approximate adequacy and fluency, but these metrics are not designed to measure stylistic fidelity, i.e. how closely an output matches the target-language stylistic profile of a high-quality human literary translation. We test whether stylometric distance, operationalized with Burrows' Delta over the 500 most frequent words, can serve as a convergent validator of adequacy signals while providing interpretable, reference-free diagnostics. Using nine contemporary Greek short stories with author-produced English self-translations and MT outputs, segmented into non-overlapping five-sentence windows, we compare an inverted, min-max normalized Burrows' Delta score (inv Delta B) against standard reference-based MT metrics (BLEU, chrF2, TER, BLEURT, COMET, BERTScore) and against an adequacy composite (TQI_win). We find strong convergence between stylometric proximity and adequacy signals, particularly at decision-relevant extremes, but stylometry underperforms adequacy metrics when used alone and provides no incremental predictive benefit beyond semantic-embedding baselines. We conclude that stylometry is best used as a complementary, explainable diagnostic and as a constrained reference-free monitor and not as a substitute for adequacy-oriented MT evaluation.
This pilot study introduces methods from dynamical systems theory to the analysis of sign language, highlighting their potential to reveal patterns of stability and variability in linguistic signals. We present two complementary measures: local Lyapunov exponents, which indicate how sensitive sequences of linguistic units are to small changes, and topological entropy, which quantifies the overall temporal complexity (chaoticity) of the sign language. The methods are illustrated using a single sign language text, analysed at two levels, sentences and individual signs, measured both in numbers of signs or pseudosyllables and in seconds. Results show higher complexity and lower stability at the level of individual signs. Lyapunov exponents capture local fluctuations and sensitivity in linguistic structure, suggesting moments where planning or motor execution may influence production, while topological entropy reflects the broader organization and predictability of discourse. Together, these measures provide a dynamic, multi-level perspective on language organization, indicating how micro-level variability interacts with macro-level structure, and offering new insights into the temporal and structural dynamics of sign language communication.
Zipf's law, Zipf-Mandelbrot law and Heaps' law have been validated across languages and are viewed as universal linguistic principles. Recent studies increasingly investigated their parameter implications. However, their applicability to ancient languages remains underexplored, and the exponent of Heaps' law has received limited attention. Our study explores how well these laws hold in Classical Chinese and whether their exponents can serve as diachronic indicators of lexical diversity. The results indicate that Classical Chinese exhibits distributional patterns consistent with the laws. The exponent of Zipf's law decreases diachronically, whereas that of Heaps' law and lexical diversity increase. The exponent of Zipf's law correlates negatively with lexical diversity, that of Heaps' law positively, and all the three show strong pairwise correlations. The parameters of Zipf-Mandelbrot law exhibit no clear monotonic trend and correlate only internally. Our findings provide support for the three laws in Classical Chinese and demonstrate both the exponents of Zipf's law and Heaps' law function as diachronic indicators of lexical diversity in Classical Chinese. However, the study is limited by its language scope, metric choice, missing polysemy analysis, untested mechanisms and unit-related issues. Future research could further extend these points.
Collocation analysis is a widespread method in corpus linguistics. A key metric used for collocation discovery is pointwise mutual information (PMI), determined by how frequently a collocation occurs relative to its expected frequency under the assumption of independence. However, PMI suffers from several limitations, especially its well-known bias for collocations involving low-frequency words. In this paper, we propose a method to determine the significance of the PMI statistic by calculating its p-value, following two probability models for collocations involving the binomial distribution and the Poisson distribution. We demonstrate the effectiveness of this method by investigating collocations involving the Greek word theta epsilon & oacute;& varsigma; , 'god', in ancient historiography. This example illustrates that the PMI statistic alone does not reveal the significance of a collocation, but rather that p-values provide a consistent threshold of statistical significance and thereby overcome many of the well-known limitations of PMI.
Synonymy is a widespread yet puzzling linguistic phenomenon. Absolute synonyms should theoretically not exist, as they do not expand language's expressive potential. However, it was suggested that even if synonyms denote the same concept, they may reflect different perspectives or carry distinct cultural associations, claims that have rarely been tested quantitatively. In Hindi, prolonged contact with Persian produced many Perso-Arabic loanwords coexisting with their Sanskrit counterpart, forming numerous synonym pairs. This study investigates whether centuries after these borrowings appeared in the Subcontinent their origin can still be distinguished using distributional data alone and regardless of their semantic content. A Random Forest trained on word embeddings of Hindi synonyms successfully classified words by Sanskrit or Perso-Arabic origin, even when they were semantically unrelated, suggesting that usage patterns preserve traces of etymology. These findings provide quantitative evidence that context encodes etymological signals and that synonymy may reflect subtle but systematic distinctions linked to origin. They support the idea that synonymous words can offer different perspectives and that etymologically related words may form distinct conceptual subspaces, creating a new type of semantic frame shaped by historical origin. Overall, the results highlight the power of context in capturing nuanced distinctions beyond traditional semantic similarity.
The frequency of the preferred order for a noun phrase formed by demonstrative, numeral, adjective and noun has received significant attention over the last two decades. We investigate the actual distribution of the 24 possible orders. There is no consensus on whether it is well-fitted by an exponential or a power law distribution. We find that an exponential distribution is a much better model. This finding and other circumstances where an exponential-like distribution is found challenge the view that power-law distributions, e.g., Zipf's law for word frequencies, are inevitable. We also investigate which of two exponential distributions gives a better fit: an exponential model where the 24 orders have non-zero probability (a geometric distribution truncated at rank 24) or an exponential model where the number of orders that can have non-zero probability is variable (a right-truncated geometric distribution). When consistency and generalizability are prioritized, we find higher support for the exponential model where all 24 orders have non-zero probability. These findings strongly suggest that there is no hard constraint on word order variation and then unattested orders merely result from undersampling, consistently with Cysouw's view.
Consider a linguistic structure formed by n elements, for instance, subject, direct object and verb (n=3) or subject, direct object, indirect object and verb (n=4). We investigate whether the frequency of the n! possible orders is constrained by two principles. First, entropy minimization, a principle that has been suggested to shape natural communication systems at distinct levels of organization. Second, swap distance minimization, namely a preference for word orders that require fewer swaps of adjacent elements to be produced from a source order. We present average swap distance, a novel score for research on swap distance minimization. We find strong evidence of pressure for entropy minimization and swap distance minimization with respect to a die rolling experiment in distinct linguistic structures with n=3 or n=4. Evidence with respect to a Polya urn process is strong for n=4 but weaker for n=3. We still find evidence consistent with the action of swap distance minimization when word order frequencies are shuffled, indicating that swap distance minimization effects are beyond pressure to reduce word order entropy.
Dependency direction is widely recognized as an important measure in quantitative linguistics, with wide applications in fields such as language typology and stylometrics. However, little research has addressed the diachronic evolution of dependency direction. Using the Corpus of Historical American English (COHA), this study investigates the evolution of dependency direction in modern English across four genres (fiction, magazines, news, and academic texts). The findings show that the proportion of head-final dependencies has increased in informational genres (academic and media texts), while no significant trend was found in fiction. Cross-genre comparisons reveal that fiction exhibits the lowest proportion of head-final dependencies, while media texts show comparatively higher levels. These findings highlight the dynamic and genre-sensitive nature of dependency directionality in modern English. We further explore the possible reasons for these diachronic and cross-genre variations and discuss the implications for natural language processing and academic writing.
Evaluating statistical fluctuations in natural language data is essential for assessing the consistency between observed data and the predictions of language models. In this study, fluctuations in word frequencies in Japanese texts are quantified, revealing that they are underestimated when those expected from a Poisson distribution are used. The evaluated fluctuations are incorporated into the data points. The consistency with Zipf's law is then examined using ${\chi <^>2}$chi 2 and Kolmogorov - Smirnov (KS) tests, in order to investigate how the outcomes of these statistical tests differ from those obtained under the assumption of Poisson errors. The results indicated that the fluctuations evaluated in this study should be used for precise comparisons between natural language data and language models.
A tradition in western thought ranging from Aristotle to theories of mental content in present-day linguistics and cognitive science says that the meaning of a word is its signification of a mental concept, and a mental concept is a representation of the mind-external environment causally generated by the cognitive agent's interaction with that environment. In this, the role of perception via the human sensory modalities in generation of the mental representations that underlie linguistic meaning is fundamental. The present discussion assesses the viability of two artificial neural network architectures, the Self Organizing Map and the Topology Preserving Network, as mechanisms for generation of suitable representations.
Contextual word embeddings have proven valuable for analysing word meanings, yet their application to Chinese polysemy remains underexplored. This study examines how polysemous and monosemous Chinese words behave in embedding space using the Chinese-ROBERTA-wwm-ext model. We analyse 59 words from the Chinese Wikipedia corpus through theoretically-grounded metrics based on information theory and distributional semantics, including embedding magnitude (information content) and standard deviation (contextual variability). Our findings reveal no statistically significant differences between polysemous and monosemous words in their distributional patterns, providing crucial cross-linguistic validation of previous findings in English. This cross-linguistic consistency suggests a potentially language-independent principle in how contextual embeddings encode lexical meaning, where meaning variation exists on a continuous spectrum rather than in discrete categories. This aligns with similar findings in English, indicating that the relationship between dictionary-defined meanings and distributional semantics might be fundamentally more complex than previously theorized.
To explore, mathematically, how mean dependency distance (MDD) and mean hierarchical distance (MHD) interact to create a trade-off relationship based on the inherent two-dimensional structure of language, this study introduces 'Dependency Structure Matrix' into quantitative syntax analysis, proposes 'Minimal Change Matrix Pair', and determines the theoretical and empirical method of adding a node that maximizes the growth rate (GR) of MDD and MHD in contemporary written Japanese, respectively. We found that the direct cause of the discrepancies between the theoretical and empirical maximum of MDD GR and MHD GR lies in the interaction between DD and HD of turning-point nodes that should lie on the diagonal of the smaller matrix and nodes that exceed the threshold of predicate valency and move to deeper layers, through which the trade-off between MDD and MHD is maintained. This interaction is numerically reflected in the fact that increases in HD tend to slow down MDD's growth, and vice versa. The underlying cause of such discrepancies is assumed to be that Japanese native speakers cannot tolerate increased cognitive load in either the linear or hierarchical dimension alone, and therefore tend to alleviate the cognitive burden in one dimension by increasing syntactic complexity in the other.
This study addresses the calculation and evaluation of syntactic distance, which is a quantitative measure of structural similarity or divergence between languages. Building on existing alignment-based, feature-based and data-driven approaches, we introduce a novel hypergraph-based metric that assesses syntactic distance through structural alignment while explicitly incorporating word order features. The approach is then applied to a multilingual parallel corpus annotated within the Universal Dependencies (UD) framework, yielding syntactic distances between English and 19 non-English languages. Empirical evaluation further demonstrates the robustness and effectiveness of the proposed measure. Compared with approaches that ablate the hypergraph formalism, ignore word order or rely solely on data-driven metrics, the new metric proves robust under random sampling variation and effectively captures syntactic distance: statistical analyses show that intra-group language pairs exhibit significantly shorter syntactic distances than inter-group pairs. This approach thus provides a novel, formally grounded perspective on language distance based purely on structural properties.
Given a vocabulary $V$V and a text, viewed as a set of sentences $S$S, a geometrical object or, more precisely, a simplicial complex, is defined. This simplicial complex $K(V,S)$K(V,S) describes how $V$V and $S$S are linked together and can be interpreted in terms of information retrieval. After removing the indiscernible elements of $V$V (with respect to $S$S), and using the tools of Formal Concept Analysis, we introduce the concept sub-complex of $K(V,S)$K(V,S). The concept sub-complex preserves the homological knowledge about $K(V,S)$K(V,S), and it is optimal from an information retrieval point of view.