
Abstract This study investigates the constructional organization of modal interrogatives by examining how modal meaning shapes response behavior. Previous research on question-response systems has largely focused on canonical polar questions, treating response particles such as yes and no as markers of propositional polarity. However, modal interrogatives introduce additional interpretive dimensions in which responses may perform discourse actions such as authorization, refusal, or epistemic endorsement. To examine whether modal interrogatives exhibit construction-specific response patterns, this study analyzes may and might interrogatives using corpus-based collostructional analysis and large language model surprisal measures. Collostructional analysis identifies distributional associations between interrogative constructions and lexical items, while LLM-based surprisal provides a probabilistic diagnostic of semantic compatibility in constructional contexts. The results reveal systematic constructional differentiation between the two constructions. The may interrogative strongly attracts first-person subjects, interaction-oriented verbs, and affirmative responses, signifying its prototypical permission-seeking use. In contrast, the might interrogative preferentially combines with non-first-person subjects, stative predicates, and evaluative responses associated with epistemic assessment. Surprisal values align with these distributional preferences, indicating that lexically compatible realizations are more strongly predicted in context. These findings suggest that modal interrogative-response pairings constitute discourse-level constructions characterized by both distributional entrenchment and probabilistic expectancy.
Abstract English mental adjectives can have their complements marked by a variety of prepositions. Angry can pattern with about , at , and with . Drawing on data from COCA, this paper investigates the alternation of these prepositional phrase complements (PPCs). Exploratory multivariate analysis reveals animacy as primary predictor. In human contexts, competition is mainly between at and with ; in nonhuman contexts between at and about . Next, modelling these localized alternations, predicate type and genre emerge as the main determinants, while semantic distinctions traditionally emphasized in the literature, such as intensity and instantaneity, fail to reach significance. The effects associated with these predictors differ across the localized contexts, and predictive accuracy remains moderate. Localized probabilistic modelling is therefore best understood as an analytical approximation rather than a full account of PPC competition, underscoring the need for more integrated approaches to alternation involving multiple competing forms.
Frequency-based learner corpus approaches struggle to separate systematic grammatical patterning from artifacts of prompts or register. This study introduces the Linguistic Clusters (LC) Framework, a multi-metric procedure for identifying stable distributional patterns in learner passives. The framework implements a three-layer architecture integrating instance-level filters, schema-level acceptance criteria, and validation procedures. Applied to over 57,000 valid passive constructions from the International Corpus of Learner English, the LC Framework identified eight validated schemas from argumentative essays across diverse first-language backgrounds. Permutation testing confirmed non-random TAM & times; complement structure and demonstrated that schema acceptance is structurally rather than lexically driven. Sensitivity analysis showed that the majority of schemas remain stable across threshold variations, while ablation analysis attributed most false-positive reduction to instance-level filters. The validated schemas cluster around modal passives with spatial prepositional complements, revealing a frequency-formulaicity dissociation: modal-be passives account for a small fraction of corpus constructions yet comprise the majority of validated schemas. These findings support usage-based accounts that treat L2 constructional development as driven by the convergence of multiple distributional cues, not by raw token frequency alone.
We present our work on the register-dependence of metaphor, where register is defined as variation of linguistic behaviour appropriate for a given communication situation, including the purpose of the interaction. We investigated this dependence on the basis of a new German corpus of six text types with a wide range of register properties, which we annotated for metaphor. Specifically, we analysed the relation between register and metaphor in terms of investigating correlations between metaphor and specific register features, as they are exhibited by the text types. Focussing on metaphors that can serve as explicit register markers and annotating a subgroup of them in our corpus, we found significant differences between the text types in metaphor usage, which we interpreted in terms of specific register properties of the text types.
Previous research on the phonetic outcomes of code-switching has most commonly found a degree of cross-linguistic interference, although such studies have relied heavily on read-aloud paradigms. Moreover, while prior studies have mostly examined the effect at the point-of-switch, less work has examined potential impacts beyond the point-of-switch. The current study investigates whether the distance from the point-of-switch impacts phonetic outcomes in bilingual vowel production in a Mandarin-English corpus of bilingual speech. English and Mandarin vowels, /(sic)/ and /e/ respectively, produced by 30 Singaporean speakers were analyzed, along with the distance from the point-of-switch. Results showed evidence of cross-linguistic interference, with English tokens becoming more Mandarin-like, although these effects were modulated by the distance from the point-of-switch. This effect followed a nonlinear trajectory, with the strongest effect at the point-of-switch, indicating a short-term time course of phonetic interference. These results support a transient, decay-oriented cross-linguistic interference model driven by co-activation of both languages and adding to current understandings of language selection and separation mechanisms in bilinguals.
This study investigates the quantitative properties of Mandarin measure words using corpus data and word embeddings to address three research questions: (1) how does constructional productivity relate to construction frequency and construction frequency distributions; (2) how semantically cohesive are different measure words; and (3) do classifiers and massifiers differ systematically in productivity and semantic cohesion. We estimated the productivity of measure word constructions with the Good-Turing estimate of unseen types and observed that productivity decreases linearly with the number of types, consistent with previously reported results for Mandarin compounds, and indicating that individual measure word constructions are sampled from the same population of measure word constructions (a measure word schema). The rank-frequency distributions of measure words have power-law like distributions, most of which are best characterized with breakpoint linear regression models. These yield quantitative measures that are predictive for construction productivity. Construction productivity also covaried with semantic measures gauging the semantic cohesion measure word constructions, such that semantically more cohesive measure words were associated with lower Good-Turing estimates of unseen types. Furthermore, massifiers were observed to be more productive and more semantically cohesive than classifiers. Our study demonstrates that the quantitative measures that have proved useful for gauging the productivity of derivation and compounding (morphological constructions) also provide insight into the productivity of measure words (syntactic constructions).
This study investigates the internal structure and ordering of functional categories in the nominal domain of Turkish, contributing to the broader discussion on the universality of syntactic hierarchies. Building on the cartographic approach, which posits a rigid, universal ordering of functional projections across languages, this study aims to empirically test these claims within the nominal domain. To this end, we used a corpus-based dataset and theoretical diagnostics to analyze the distribution and relative positioning of numerals, classifiers, demonstratives, possessors, quantifiers, and adjectives within the nominal domain in Turkish. The observed ordering of functional categories in Turkish broadly aligns with the cross-linguistic patterns proposed in previous literature, supporting the view that the nominal domain is governed by universal principles. At the same time, this study provides new insights into both universal and language-specific aspects of the cartographic architecture in the nominal domain, contributing to the theoretical discussion on the hierarchical organization of functional categories.
Entities in discourse vary in salience: main participants, objects and locations stay prominent, while others are quickly forgotten, raising questions about how humans signal and infer discourse-level salience. Using a graded operationalization of discourse-level salience based on summary-worthiness in multiple summaries, this paper investigates whether predictors of utterance-level prominence extend to the discourse level, and how they interact across 24 spoken and written genres of English. We examine features including grammatical function, definiteness, entity type, linear order, discourse relations and hierarchy, and referential structure, as well as the impact of genre. Our results show that utterance-level predictors significantly correlate with discourse-level salience, but interact with and are modulated by entity-level factors such as frequency and dispersion across the document. Multifactorial models reveal that no single factor determines salience; rather, discourse-structural and semantic features prove more robust than morphosyntactic ones, with substantial variation by genre and communicative intent.
The collocational patterns of a word and its phonetic properties are often considered independent of each other. This corpus-based study tracks change in the word just over real and apparent time in New Zealand and Australian English and shows that these two phenomena take place in tandem. Four major changes are reported. First, the frequency of just has increased in both varieties, mirroring reports in Canadian and UK varieties of English. Second, the collocations of just demonstrate that it has undergone pragmatic change, evidenced by, for example, its increased co-occurrence with sort of, like, and an existential subject (it's just & mldr;). Third, acoustic phonetic analyses of changes in the vowel in just indicate that it does not pattern with other content or grammatical words belonging to its canonical vowel class and has instead moved towards a closer and fronter realization. Finally, we show that this phonetic change is led by variants of just which occur in newer (more recent) collocational contexts. Together these findings provide an example of word-specific change and evidence that a word's phonetic realization depends in part on its grammatical status, and that phonetic change over time can be linked with shifts in collocational patterns and pragmatic function.
This study employs a BERT-assisted behavioural profiles approach to examine the diachronic interaction of two Chinese extreme degree resultative constructions: [V-HUAI 'go bad'] and [V-SI 'die']. Drawing data from the CCL Ancient Chinese Corpus, we trace their semantic developments and interactions across three historical periods - spanning approximately 2,000 years. The findings reveal a converging tendency with both constructions undergoing multi-stage evolution. Clustering analyses show semantic convergence stems from their shared ability to express outcomes of the same causal events while differentiation arises from HUAI's encoding of gradable scales. Notably, the study reveals a collaborative interplay between attraction and differentiation at distinct levels: attraction contributes to the schematic generalization and formation of constructions, while differentiation operates at the instance level, serving to stabilize the constructional network. This research enriches evolutionary construction grammar by clarifying the interplay between semantic-driven attraction and differentiation, which governs the diachronic evolution of extreme-degree resultative constructions.
Rational numbers can be expressed as decimals, fractions, and percentages, each serving distinct functions. The differences among the three forms have been examined in terms of structural distinction and the cognitive effort required for their processing. Using a multivariate approach, this study integrates cognitive and linguistic factors to explore how these representations alternate in diverse linguistic contexts. A total of 2,217 rational number tokens were coded for eight features. Their relative importance in predicting notation choice was assessed using random forests, with conceptual interpretation and syntactic structure as the most influential predictors, followed by entity type, genre, grammatical function, semantics of verbs, and definiteness, while approximation contributed little predictive information. The findings suggest that the three notations function as members of a prototype-based category, each showing a typical usage core with peripheral cases. Individuals select numerical forms in an effort-saving manner, favoring fractions over decimals and percentages for precise enumeration. Furthermore, these forms are also found to evolve and develop varied grammatical functions and syntactic structures in actual use.
Understanding how different dimensions of syntactic complexity interact with processing demands is fundamental to language acquisition, pedagogy, and assessment. While structural measures like Clausal Density (CD) and Embedding Depth (ED) have been examined separately, their relationship with processing load, as proxied by Mean Dependency Distance (MDD), remains underexplored. This study investigates this interplay in 42,654 complex sentences from the Brown and LOB corpora, using statistical modeling to examine the nature of the CD-MDD relationship, ED's independent contribution to MDD, and their interaction. Results reveal a non-linear CD-MDD relationship peaking at CD = 3, a distinct processing cost at the transition to double embedding (ED = 2), and a significant interaction where deeper embedding reverses the CD-MDD relationship. These findings demonstrate a dynamic trade-off in syntactic complexity, challenging additive models and supporting cognitive constraint theories. The results have direct implications for refining complexity measures in second language acquisition and for informing psycholinguistically-grounded pedagogy and assessment.
Cross-linguistic work on the causal-noncausal alternation has shown that there exists a strong correlation between the frequency of use of a given verb in causal versus noncausal contexts and its preferred coding pattern for the alternation, with verbs typically occurring in causal contexts preferably selecting the anticausative pattern. These studies assume that the correlation between frequency and coding is diachronic, but this assumption has so far not been tested on historical data. This paper aims to fill this gap, by contrastively exploring the interplay between frequency effects and the coding of the causal-noncausal alternation in diachronic corpora of Italian and Spanish. By resorting to inferential statistic techniques, our findings largely confirm the hypothesis advanced by frequentist approaches, while at the same time also offering a more nuanced understanding of the role of frequency, which is only one among several factors that shape speakers' choice of anticausativization patterns in Romance languages.
When corpus researchers have to down-size their data, different techniques may be used to optimize the sub-sample. In alternation studies, one strategy is the selection of instances based on the observed realization of the outcome variable. This is referred to as a case-control design in the health sciences, which have developed a rich methodology surrounding this approach. The present paper argues that corpus linguists should be sounding out the potential for methodological transfer. It pursues three goals. The first is to overcome terminological barriers by making transparent the peculiar jargon associated with this research design. Further, I will provide an overview and illustration of some principles of study design and data analysis that form the core of this method. Finally, distinctive features of case-control down-sampling will be identified, which allows for a focused exploration of the extensive literature related to this approach and also provides guideposts for future methodological work.
Pijpops and Van de Velde (2016. Constructional contamination: How does it work and how do we measure it? Folia Linguistica 50(2). 543-581) demonstrated that an alternation within a construction may be influenced by a different construction that is coincidentally formally similar through the process of constructional contamination. This mechanism has been explained as a consequence of the storage of unanalyzed chunks produced by a contaminating construction, which are later recycled in a target construction. Consequently, constructional contamination has hitherto been measured by calculating the frequency of individual lexemes (which are assumed to be part of exemplar chunks eligible for recycling) in the contaminating construction. In this paper, we propose a different measurement for constructional contamination. Using the large language model RobBERT, we define "prototype embeddings" for both the target construction and the contaminating construction involved in constructional contamination. For each instance of the target construction, we calculate its similarity to these prototypes. We then show that these similarity variables are significantly correlated with the alternation within the target construction: the more an instance of the target construction is semantically similar to the contaminating construction, the more likely it is to appear with the alternant that is exclusively used in that contaminating construction. Our results are in line with a scenario in which horizontal links, which hold between formally similar constructions, can directly trigger constructional contamination. In addition, we find that the more semantically similar an instance is to the prototype of the target construction itself, the more likely it is to appear with the alternant that is not found in the contaminating construction, supporting a view of the constructional network in which constructions are themselves conceptualized as links rather than nodes.
Mandarin differs essentially in colexification from many other languages due to its preference for controlled visual activity verb kan "look" over uncontrolled visual experience verb jian "see." This triggers inquiries as to whether Mandarin is historically inclined to controlled visual activity. Based on a dynamic behavioral profile analysis, this study traces the semantic change in kan and jian from Archaic Chinese to Contemporary Chinese, with a particular focus on the evidential discourse markers derived from them. Through descriptive statistics and multiple correspondence analysis (MCA), the data reveal that: (1) kan and jian have distinct behavioral profiles, and they primarily convey controlled visual activity and uncontrolled visual experience respectively in history; (2) kan has undergone semantic expansions with more varied usages, while jian has been decreasingly used; and (3) except in the Archaic period, the evidential discourse markers derived from kan exhibit greater semantic and structural diversities compared with those derived from jian. The comparison and contrast between the two basic vision verbs help to reveal Mandarin's commonalities in semantic change while highlighting its peculiarity in emphasizing controlled visual activity over time.
This study describes and evaluates a multi-method approach for identifying and extracting collocations to develop a learner Italian collocation dictionary. The approach integrates part-of-speech tagging and dependency parsing to extract six syntactic relations from a reference corpus of Italian. The initial set of candidates was gradually reduced using frequency, dispersion, and association measures. This set was then evaluated by comparing it with existing collocation dictionaries and gathering expert judgments on which collocations should be included. Combining these two evaluations, further refined the list. Moreover, the effect of statistical measures on expert judgments was investigated. Results revealed that dispersion and association measures positively influenced human evaluations, while higher frequency often correlated with negative ratings. This triangulation of corpus-based and statistical methods, human judgements and comparison with existing dictionaries captures collocations widely used across genres, suitable for inclusion in a learner dictionary, offering a useful tool for learners while contributing to corpus-based collocation research.
This paper presents a novel treebank-driven approach to comparing syntactic structures in speech and writing using dependency-parsed corpora. Adopting a fully inductive, bottom-up method, we define syntactic structures as delexicalized dependency (sub)trees and extract them from spoken and written Universal Dependencies (UD) treebanks in two syntactically distinct languages, English and Slovenian. For each corpus, we analyze the size, diversity, and distribution of syntactic inventories, their overlap across modalities, and the structures most characteristic of speech. Results show that, across both languages, spoken corpora contain fewer and less diverse syntactic structures than their written counterparts, with consistent cross-linguistic preferences for certain structural types across modalities. Strikingly, the overlap between spoken and written syntactic inventories is very limited: most structures attested in speech do not occur in writing, pointing to modality-specific preferences in syntactic organization that reflect the distinct demands of real-time interaction and elaborated writing. This contrast is further supported by a keyness analysis of the most frequent speech-specific structures, which highlights patterns associated with interactivity, context-grounding, and economy of expression. We argue that this scalable, language-independent framework offers a useful general method for systematically studying syntactic variation across corpora, laying the groundwork for more comprehensive data-driven theories of grammar in use.
A growing body of literature has demonstrated that semantics can codetermine fine phonetic detail. However, the complex interplay between phonetic realization and semantics remains understudied, particularly in pitch realization. The current study investigates the tonal realization of Mandarin disyllabic words with all 20 possible combinations of two tones, as found in a corpus of Taiwan Mandarin spontaneous speech. We made use of Generalized Additive Mixed Models (GAMMs) to model f0 contours as a function of a series of predictors, including gender, tonal context, tone pattern, duration, word position, bigram probability, speaker, and word. In the GAMM analysis, word and sense emerged as crucial predictors of f0 contours, with effect sizes that exceed those of tone pattern. For each word token in our dataset, we then obtained a contextualized embedding by applying the GPT-2 large language model to the context of that token in the corpus. We show that the pitch contours of word tokens can be predicted to a considerable extent from these contextualized embeddings, which approximate token-specific meanings in contexts of use. The results of our corpus study show that meaning in context and phonetic realization are far more entangled than standard linguistic theory predicts.
Hewram & icirc; (ISO 639-3) is an Iranic language featuring tense-sensitive alignment, characterised in terms of indexing by employing two sets of bound person markers for expressing direct objects. In TAM categories based on past stem verbs, the direct object is expressed by suffixal person/number morphology. However, the index may be absent under certain conditions. This article explores object indexing in Hewram & icirc;, considering its token frequency in discourse and factors conditioning differential object indexing (DOI). The data come from two spoken corpora of over 35,000 words from the Tekht variety of Hewram & icirc;. The corpus data were annotated for factors such as animacy, identifiability, person of the referent and textual givenness. The data provide some support for the complementarity hypothesis, meaning that agreement markers slightly favour zero object arguments. The most decisive factor triggering the lack of indexing with overt Os is the co-optation of the agreement slot by a higher-ranked argument. But when taking semantic-referential features into account, the person of the object referent strongly predicts differential object indexing. With the direct object being 3pl and the NP modified, animacy predicts DOI.