
Abstract Registers are text varieties associated with situations of use. It has now been documented, however, that texts within registers are not situationally homogeneous — just as they are not linguistically uniform. This led Biber and Egbert (2023) to propose that linguistic variation within registers functionally corresponds to situational variation among their texts. The situational factor consistently shown to vary within registers is communicative purpose. However, previous studies have examined variability in purpose within a single register and always relied on coding schemes developed for that register. As a result, there is no single taxonomy of communicative purposes to be applied across registers. This paper proposes such a framework, offering guidance for its future adaptations to other corpora, and demonstrates how its application enables (a) analyses of functional correspondence between communicative and linguistic variation among texts of several registers; (b) register comparisons for the extent of internal variation.
Abstract The variable realisation of there + be +plural arguments (e.g. there’s a lot more Chinese people ) is well-documented across English varieties, but little research has examined its use among ethnic minorities or how intersecting social factors shape its variation. This study addresses these gaps, analysing 1,549 tokens from 206 Australians of Anglo, Italian, Greek, and Chinese backgrounds, recorded in the 1970s and 2010s. The study finds an apparent-time increase in there’s during the 1970s, led by teenagers across ethnic groups, followed by broad adoption by the 2010s. Usage patterns reflect intersections of age, ethnicity, and class, with working-class Adult Anglos consistently favouring there’s . Additionally, the grammatical conditioning of ( there’s ) has shifted: the definiteness effect, where definite and numerical determiners (e.g. these, two ) favour there’s , emerged among 1970s teenagers and has since been maintained. Changes in frequency are accompanied by shifts in linguistic and social conditioning, illustrating how variation evolves within complex sociolinguistic ecologies.
Abstract This paper provides a detailed account of the Turkish Learner Corpus (TURLEC). Building on the first author’s doctoral dissertation project, which aimed to identify proficiency descriptors for four skills (listening, reading, writing, and speaking) for learners of Turkish as a second language (L2) at various CEFR levels, the main motivation to build a learner corpus is to outline the language learners actually use at different proficiency levels. With the written and spoken texts of learners of Turkish L2 at university level coming from various countries with numerous L1 backgrounds, TURLEC comprises 735 texts and approximately 97,000 tokens. After rigorous anonymisation, annotation, and error-tagging efforts, TURLEC contains ~21,000 word forms with 3,561 lemmas, which will be profiled based on the CEFR levels. TURLEC is the first learner corpus built to offer a vocabulary profile for L2 Turkish, which is an ever-growing field of study with an increasing number of students.
Tokenisation is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring the reliability of qualitative approaches. This paper examines how discrepancies in tokenisation affect the representation of language data and the validity of analytical findings. Investigating the challenges posed by emojis and homoglyphs, the study highlights the necessity of pre-processing these elements to maintain corpus fidelity to the source data. The research presents methods for ensuring that digital texts are accurately represented in corpora, thereby supporting reliable linguistic analysis and facilitating the repeatability of linguistic interpretations. The findings emphasise the necessity of a detailed understanding of both linguistic and technical aspects involved in digital textual data to enhance the accuracy of corpus analysis, and have significant implications for both quantitative and qualitative approaches in corpus-based research.
Abstract This editorial introduces a special issue of the International Journal of Corpus Linguistics on corpus perspectives on legal discourse. It first situates corpus-based research within the broader empirical tradition of language and law studies, tracing the emergence of a ‘corpus turn’ from the mid-1990s onwards and its contribution to the analysis of legal texts. It then identifies established and emerging trends exemplified by the five studies in the issue, which span jurisdictions in Europe, North America and Asia and draw on both common law and civil law traditions. Four trends are discussed: the continued dominance of written legal discourse as the core object of corpus analysis; comparison as a foundational methodological design principle; the role of intertextuality, interdiscursivity and genre networks in situating legal texts within broader institutional and societal contexts; and the spectrum from impersonal, informational legal language to more involved, evaluative discourse in judicial settings.
Abstract The starting point of this paper is a central problem in constructicography: on the one hand, constructicon projects aim at describing a broad spectrum of constructions based on usage data; on the other hand, the established method for determining salient fillers of construction elements (slots) requires comprehensive annotations of entire corpora. In this paper, we propose a new method to tackle this problem. Using the pre-trained language model BERT and a limited number of example sentences, the method predicts typical fillers of slots, so-called collo-profiles. A critical evaluation of the method reveals not only its potentials and limitations but also differences between such a prediction-based method and classical count-based methods.
This paper provides a detailed account of the Turkish Learner Corpus (TURLEC). Building on the first author's doctoral dissertation project, which aimed to identify proficiency descriptors for four skills (listening, reading, writing, and speaking) for learners of Turkish as a second language (L2) at various CEFR levels, the main motivation to build a learner corpus is to outline the language learners actually use at different proficiency levels. With the written and spoken texts of learners of Turkish L2 at university level coming from various countries with numerous L1 backgrounds, TURLEC comprises 735 texts and approximately 97,000 tokens. After rigorous anonymisation, annotation, and error-tagging efforts, TURLEC contains similar to 21,000 word forms with 3,561 lemmas, which will be profiled based on the CEFR levels. TURLEC is the first learner corpus built to offer a vocabulary profile for L2 Turkish, which is an ever-growing field of study with an increasing number of students.
The starting point of this paper is a central problem in constructicography: on the one hand, constructicon projects aim at describing a broad spectrum of constructions based on usage data; on the other hand, the established method for determining salient fillers of construction elements (slots) requires comprehensive annotations of entire corpora. In this paper, we propose a new method to tackle this problem. Using the pre-trained language model BERT and a limited number of example sentences, the method predicts typical fillers of slots, so-called collo-profiles. A critical evaluation of the method reveals not only its potentials and limitations but also differences between such a prediction-based method and classical count-based methods.
The utility of historical language databases is hindered by their complexity, text variety, and lack of register information. We investigate the feasibility of deriving linguistically motivated and reliable register predictions for unannotated, long historical texts. We fine-tune BERT-based deep learning models using register-annotated data from the Corpus of Founding Era American English and predict registers for different text parts of unannotated Eighteenth Century English Online documents. We determine the model's effectiveness in capturing pervasive linguistic features across different text sections (e.g. beginnings vs. endings), analyzing how document internal variation affects model performance and identifying which sections are most effectively predicted. Additionally, we employ the Stable Attribution Class Explanation method to extract and compare keywords from various text parts to determine the quality of the predictions. Our findings indicate that text beginnings consistently provide more reliable classifications.
Corpora containing newspaper articles are widely used in corpus-based discourse research. However, these corpora often contain duplicate texts due to issues such as syndication, the existence of agency copy, and largely similar print and online versions of the same article. While there are several tools that have automated deduplication functions, only a limited number of these tools allow corpus compilers to inspect duplicates before removal. In this paper, we will introduce the Document Similarity tool, which was developed to automatically remove identical texts and then provides near-identical candidates for users to examine side-by-side, giving corpus compilers more control of the deduplication process. We will describe several experiments using corpora that have been deduplicated using the Document Similarity tool in order to demonstrate its utility for corpus creation in corpus-based discourse analysis of newspapers and potentially other text types.
We compare to what extent the choice of a variant can be predicted from language-internal factors for two linguistic variables: the English dative alternation and the omission of the infinitival marker att in a Swedish future-tense construction. Previous research has shown that for the dative alternation, near-ceiling performance can be achieved by fitting a regression model with manually selected predictors. Similar attempts for att-omission have been unsuccessful. To test whether the two variables differ in predictability or whether optimal methods have not been found for att-omission, we apply a large language model (LLM) to the same task. For the dative alternation, LLM and regression perform equally well. For att-omission, LLM outperforms regression, but still performs worse than for the dative alternation, thus suggesting that att-omission is inherently less predictable. We argue that LLMs can be useful for estimating how much the choice of a variant depends on language-internal factors.
This study presents a corpus-assisted discourse analysis examining the semantic frames of (Dis)Interest and (Un)Importance in leave to appeal decisions of the HKSAR appellate courts. With the use of two corpora - Approve (67,694 tokens) and Dismiss (143,462 tokens) - the research investigated the construction and use of (Dis)Interest and (Un)Importance frames. A further distinction of performative versus descriptive use informed the analysis. In terms of frame construction, decision writers used four frame elements: Trigger, Explanation, Degree, and Experiencer. The findings show that (Un)Importance frames appeared three times more frequently than (Dis)Interest frames. Reflecting the inherent nature of the outcome, Approve writers used significantly more performative Interest and Importance; Dismiss writers used significantly more performative Disinterest. Dismiss writers also used significantly more descriptive Importance in refutation structures. Qualitative analysis revealed that Approve decisions featured more straightforward performative framing, whereas Dismiss decisions displayed a complex interplay of performative and descriptive elements.
Stance is deep-rooted in law, where legal values can never stand in a vacuum. Despite a growing body of literature on stance in legal genres, cross-genre examinations conducted from a corpus-based perspective still leave room for improvement. This study conducts a cross-genre examination of how legal professionals express stance across three legal genres, i.e. legislation, judgments, and legal academic articles. By adopting a corpus-based approach, evidence-based insights are provided into the general profile of stance expressions in legal settings, as well as the distribution of stance features across the three legal genres. Additionally, this study delineates a continuum of stance in law, which illustrates the variation in stance expressions by categorizing them as objective or subjective, certain or uncertain, direct or indirect, and explicit or implicit. The findings suggest that stance may serve as a discourse anchor to help frame legal rules, construct legal facts, and convey legal values.
The aim of this paper is to combine a quantitative analysis of indicators of pragmatic argumentation with a qualitative investigation of the argument scheme in a corpus of Supreme Court of Ireland's judgments. The quantitative analysis indicates that the Supreme Court's argumentation tends to support judicial standpoints by focusing on the negative impact of alternative lines of argument, whether from other judgments or the parties' submissions. Alternatively, the argumentation draws the relevant audience's attention to legal values and principles underlying legislation or the Constitution. The qualitative study of Heneghan v. Minister for Housing, furthermore, shows how the Court's pragmatic argumentation combined its positive and negative variant, and responded to the relevant critical questions. Overall, the use of corpus-informed tools played a central role in the study of indicators of argumentation as "'entry points" into the construction of judicial argumentation (Go & zacute;d & zacute;-Roszkowski, 2021), which is fruitfully integrated with insights from argumentation theory.
This paper sets out a bilingual (English and French) corpus-based approach to quantify divergence in terminology between legal discourse and news articles, triangulating a series of complementary indicators of frequency difference, predominant terms and absent terms. This methodology is then applied to purpose-built corpora consisting of EU legal discourse and newspaper articles on the subject of migration in English and French, illustrating the relevance of an approach to measuring shifts in terminological distance. The results of such a study can provide insights into the level of comprehensibility of legal discourse, which is fundamental to ensuring access to justice. This context makes it vitally important to develop such a methodology, which empirically measures whether the terminology used in EU legal discourse is continuing to diverge from language used in non-specialist settings.
This study applies full Multidimensional Analysis (MDA) to examine linguistic variation in the Polish Eurolect — a hybrid variety shaped by translation and institutional constraints within the European Union — by comparing it to the national variety. Using a corpus of key institutional registers (legal acts, judgments, administrative reports, and citizen-oriented websites), we identify four dimensions of variation: Argumentative vs Informational, Engaged Instruction vs Distanced Authority, Prescriptive vs Narrative, and Lexical Richness. The findings reveal notable differences between how supranational and national institutions communicate. EU legal acts and judgments show greater prescriptiveness, legal referencing, and argumentative structuring compared to their Polish counterparts. EU websites have less engagement and explanatory strategies while EU reports favour a less distanced style. The findings map variation and group institutional registers, thereby visualizing similarities and differences between supranational and national institutional communication.