Existing research often treats parliamentary discourse as a homogeneous whole, overlooking topic-specific patterns. Parliamentary speeches address a wide range of topics, some of which evoke stronger emotions than others. While everyone has intuitive assumptions about what the most emotive topics in a parliament may be, there has been little research into the emotions typically linked to different topics. This paper strives to fill this gap by examining emotion expression among the topics of parliamentary speeches delivered in Eduskunta, the Finnish Parliament, between 2000 and 2020. An emotion analysis model is used to investigate emotion expression in topics, from both synchronic and diachronic perspectives. The results strengthen evidence of increasing positivity in parliamentary speech and provide further insights into topic-specific emotion expression within parliamentary debate.
This study applies corpus linguistic and machine learning methods to analyze 300 Finnish children’s books intended to be read between the ages of 7 and 15, and examines the linguistic characteristics and distinctiveness of texts written for particular ages. The corpus we use, the Turku Children’s Book Lexicon (TCBC), consists of three sub-corpora based on the intended reading ages of the books (7-8, 9-12, 13-15) and we use two different methodologies to analyze the linguistic characteristics of texts in these age groups. First, we use Key Feature Analysis (Egbert and Biber 2023) to examine the typical functional features of texts on the corpus level. Second, for examining the distinctiveness of the text classes on the level of individual texts, we use interpretable machine learning methods and Support Vector Machines (SVM) in particular. In our analysis, we use a total of 148 features. We show that the texts in each age group have distinct linguistic characteristics, with differences in phrasal and clausal complexity, morpho-lexical richness, and lexical variation. However, classifying individual texts proves quite difficult and machine learning models struggle especially with distinguishing texts of the middle age group from the others. We also show empirical evidence that traditionally used readability metrics are insufficient for characterizing text classes. This study provides crucial information for a wide range of fields applying age-specific text data, such as psycholinguistic reading studies and the development of education technologies.
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for 24 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder-decoder models, as well as a handful of monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks spanning four diverse task categories (retrieval, bitext mining, pair classification, and summarization) in both English and multilingual settings, and reveal that nearest-neighbor overlap and magnitude differences in independent component analysis (ICA) between paired text instances strongly correlate (even up to 0.97) with performance on the given task. Ultimately, we show that embedding tasks display varying degrees of linearity and reliance on retention of local information. Our results further the understanding of embeddings, their relation to model performance, and shed light on possible future training objectives and optimizing conditional embeddings.
The utility of historical language databases is hindered by their complexity, text variety, and lack of register information. We investigate the feasibility of deriving linguistically motivated and reliable register predictions for unannotated, long historical texts. We fine-tune BERT-based deep learning models using register-annotated data from the Corpus of Founding Era American English and predict registers for different text parts of unannotated Eighteenth Century English Online documents. We determine the model's effectiveness in capturing pervasive linguistic features across different text sections (e.g. beginnings vs. endings), analyzing how document internal variation affects model performance and identifying which sections are most effectively predicted. Additionally, we employ the Stable Attribution Class Explanation method to extract and compare keywords from various text parts to determine the quality of the predictions. Our findings indicate that text beginnings consistently provide more reliable classifications.
This article presents a study on the use of artificial intelligence to analyze Finnish-language newspapers published in North America between 1876 and 1923. Using GPT-4 and LLaMA 3.1, we develop and evaluate a text segmentation method on large-scale digitized historical data, and we propose a hierarchical taxonomy for register (genre) classification based on manual annotation. We present a methodology for identifying text boundaries and, separately, a manually developed register taxonomy used to classify segments. Our analysis is grounded in a manually annotated stratified corpus of 500 documents selected from 312,300 digitized pages. Results demonstrate strong segmentation accuracy, contributing both to digital humanities methodology and the historical understanding of the Finnish-American immigrant press.
It is sometimes said that all politicians sound the same with their speeches mired in political jargon full of clich & eacute;s and false promises. To investigate how distinct the plenary speeches of political parties truly are and what linguistic features make them distinct, we trained a BERT classifier to predict the party affiliation of Finnish members of parliament from their plenary speeches. We contrasted and compared model performance to human responses to see how humans and the model differ in their ability to distinguish between the parties. We used the model explainability method SHAP to identify the linguistic cues that the model most relies on. We show that a deep learning model can distinguish between parties much more accurately than the respondents to the questionnaire. The SHAP explanations and questionnaire responses reveal that whereas humans tend to rely on mostly topical cues, the model has learned to recognize other cues as well, such as personal style and rhetoric.
This work introduces TCBLex, a lexical database of Finnish literary works read by children between the ages of 7 and 15. We explain in detail the work done to build the corpus TCBLex is based on, including how books were sampled and collected, turned into text files, and finally processed. We also touch on legal considerations and how it is possible to build such a corpus in the EU. TCBLex contains over 11 million tokens that are annotated with parts-of-speech tags and lemmatized. We provide 14 different sub-lexicons in total, covering individual intended reading ages, age groups, as well as different genres. We also provide versions with additional morphological information, such as the cases and tenses of words. TCBLex provides various psycholinguistically interesting lexical statistics for both word types and lemmas, such as different frequency metrics, distributions, word lengths, numbers of syllables, morphological paradigm sizes, and for the first time in a Finnish lexicon, ages when words and lemmas are first encountered in books. TCBLex is freely available at https://doi.org/10.5281/zenodo.15655580 .
The pervasiveness of the internet has given web language use a central role in society. However, the lack of multilingual corpora and scalable methods has led to the focus on English in web language research. To address this gap, the present paper sets itself in the register research tradition and explores French and Swedish web registers from a cross-linguistic angle. Methodologically we combine keyword analysis with multilingual deep learning, suggesting an approach that enables computational comparisons across languages. Specifically, we extract keywords for French and Swedish web registers, then associate the keywords with fastText word embeddings, and finally, cluster these key embeddings. The findings indicate that there are topical and functional clusters, and they are linguistically motivated and multilingual. The same clusters occur within the same registers in both languages pointing to shared topical and functional similarities - the registers are strikingly similar. The dissimilarities, in contrast, indicate that certain registers like Narrative blogs are to some extent different in French and Swedish. Moreover, grammatical specificities such as the location of adjectives explain some dissimilarities.
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora, extending prior work of the HPLT project. The monolingual portion of the data contains 8T tokens covering 193 languages, while the parallel data contains 380M sentence pairs covering 51 languages. We document the entire data pipeline and release the code to reproduce it. We provide extensive analysis of the quality and characteristics of our data. Finally, we evaluate the performance of language models and machine translation systems trained on HPLT v2, demonstrating its value.
Pretraining data curation is a cornerstone in Large Language Model (LLM) development, leading to growing research on quality filtering of large web corpora. From statistical quality flags to LLM-based labelling systems, datasets are divided into categories, frequently reducing to a binary: those passing the filters are deemed as valuable examples, others are discarded as useless or detrimental. However, a more detailed understanding of the contribution of different kinds of texts to model performance is still largely lacking. In this article, we present the first study utilising registers or genres - a widely used standard in corpus linguistics to model linguistic variation - to curate pretraining datasets and investigate the effect of register on the performance of LLMs. We train small generative models with register classified data and evaluate them using standard benchmarks, and show that the register of pretraining data substantially affects model performance. We uncover surprising relationships between the pretraining material and the resulting models: using the News register results in subpar performance, and on the contrary, including the Opinion class, covering texts such as reviews and opinion blogs, is highly beneficial. While a model trained on the entire unfiltered dataset outperforms those trained on datasets limited to a single register, combining well-performing registers like How-to-Instructions, Informational Description, and Opinion leads to major improvements. Furthermore, analysis of individual benchmark results reveals key differences in the strengths and drawbacks of specific register classes as pretraining data. These findings show that register is an important explainer of model variation and can facilitate more deliberate future data selection practices.
This paper describes the process of creating a TEI XML corpus of late medieval Latin documents for NLP tasks from books in PDF format. The documents of the Apostolic Penitentiary (a tribunal of the Catholic Church responsible for granting absolutions, dispensations, and indulgences) have been originally published as printed books. For the purposes of this corpus, they were derived from PDF files used for proofreading before printing. These editions, containing 1,511 documents and 211,398 words, are designed by and for human scholars engaged in close reading. As a result, they encode implicit semantic information through typographical features such as page layout and italics, which are sometimes inconsistent. Although human readers, equipped with holistic understanding, can interpret such variations, NLP tools require unambiguous, text-only input.We report in detail the process of transforming the PDF editions into a structured, machine-readable, and openly accessible corpus. Our approach combines a rule-based workflow using regular expressions with close reading and manual corrections. Such conversion procedures, which are highly time-consuming and require in-depth knowledge of medieval Latin philology and manuscript studies, are regrettably seldom made explicit, despite their vital role in ensuring the reproducibility and scalability of research.
A register, defined as a text variety with specific situational characteristics and a communicative purpose (Biber & Conrad 2019), is also recognized as a cultural construct (Biber & Egbert 2023). Registers merit thorough investigation due to their pivotal role in reflecting linguistic and cultural landscapes. However, existing studies predominantly focus on Indo-European languages. This study investigates Turkish web registers through the introduction of the Turkish Corpus of Online Registers (TurCORE). Comprising 2,780 web texts, TurCORE was manually annotated using a register taxonomy targeting the entire unrestricted web and identifying 24 web register categories. By employing Text Dispersion Keyword Analysis (Egbert & Biber 2019), the research examines the register characteristics with a specific focus on news reports, interactive discussions, and recipes, drawing comparisons with their English equivalents. Results reveal parallels between Turkish and English news reports while Turkish interactive discussions and recipes exhibit distinctive language- and culture specific features.
This volume highlights the ways in which recent developments in corpus linguistics and natural language processing can engage with topics across language studies, humanities and social science disciplines.New approaches have emerged in recent years that blur disciplinary boundaries, facilitated by factors such as the application of computational methods, access to large data sets, and the sharing of code, as well as continual advances in technologies related to data storage, retrieval, and processing. The “march of data” denotes an area at the border region of linguistics, humanities, and social science disciplines, but also the inevitable development of the underlying technologies that drive analysis in these subject areas.Organized into 3 sections, the chapters are connected by the underlying thread of linguistic corpora: how they can be created, how they can shed light on varieties or registers, and how their metadata can be utilized to better understand the internal structure of similar resources. While some chapters in the volume make use of well-established existing corpora, others analyze data from platforms such as YouTube, Twitter or Reddit. The volume provides insight into the diversity of methods, approaches, and corpora that inform our understanding of the “border regions” between the realms of data science, language/linguistics, and social or cultural studies