THE STYLE AND STRATEGY OF ELECTORAL COMMUNICATION IN THE CAMPAIGN SPEECHES OF JAROSŁAW KACZYŃSKI AND DONALD TUSK FROM A QUANTITATIVE PERSPECTIVE The article presents a quantitative analysis of the differences in the communication styles of two influential Polish politicians, Jarosław Kaczyński and Donald Tusk, during the 2023 election campaign. The study is based on the analysis of a corpus of speeches by both politicians, which were automatically transcribed and grammatically annotated, as well as enriched with metadata. The analysis employed the key feature analysis method suggested by Biber and Egbert, which is a simplified version of the so-called multidimensional analysis. The study utilised grammatical, lexical, and syntactic features, as well as full-text measures of syntactic complexity and lexical richness. The results of the study show how both politicians construct their speeches, how they address their audiences, how they refer to each other, and how they attempt to attract the attention of voters.
Despite extensive research, populism remains one of the most strongly contested concepts relating to political language. This study aims to contribute to its understanding by putting forward an approach to identifying markers of populist communication, which draws on the Corpus-Assisted Discourse Studies toolkit. In order to test the “populist Zeitgeist” hypothesis (Mudde, 2007), and to fill the gap in research on populism in Eastern Europe, this methodology is then applied to a corpus of Polish mainstream politicians’ election campaign speeches. The results point to the populist contagion, provide counter-evidence to the claims of a transitory nature of populism and show significant variation, despite the relatively subtle ideological differences between the analysed parties. This suggests that populism should not be considered merely as an attachment to a “host” (usually extreme) ideology, but rather as a complex discursive phenomenon in its own right.
In this paper, we present PAWUK, the Polish Automatic Web corpus of UKrainian language. It is a linguistic corpus containing Ukrainian texts acquired from the Internet (selected web pages and social media accounts) and has been updated daily since 2022. It is automatically annotated with morphosyntactic tags, syntactic dependencies and named entities using Stanza framework with a model custom-built for Ukrainian to produce both Universal Dependencies tags and VESUM morphological tags. Users can interact with the corpus through a publicly available web interface.
Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish Large Language Model), the largest open-source family of foundation models tailored specifically for the Polish language. Developed by a consortium of major Polish research institutions, PLLuM addresses the need for high-quality, transparent, and culturally relevant language models beyond the English-centric commercial landscape. We describe the development process, including the construction of a new 140-billion-token Polish text corpus for pre-training, a 77k custom instructions dataset, and a 100k preference optimization dataset. A key component is a Responsible AI framework that incorporates strict data governance and a hybrid module for output correction and safety filtering. We detail the models' architecture, training procedures, and alignment techniques for both base and instruction-tuned variants, and demonstrate their utility in a downstream task within public administration. By releasing these models publicly, PLLuM aims to foster open research and strengthen sovereign AI technologies in Poland.
DERIVATIONAL NETWORK ON THE BASIS OF SŁOWNIK GRAMATYCZNY JĘZYKA POLSKIEGO: SUGGESTION FOR CLASSIFICATION The article presents a suggestion of adding a derivational network to the Słownik gramatyczny języka polskiego (SGJP, Grammatical Dictionary of Polish). To date, SGJP regularly noted selected wordbuilding relations (e.g., names of features created by adding the ending -ość to adjectives), several (such as diminutives) were marked inconsistently and many were not introduced there at all. To fill this gap, the authors adjusted the classification of derivatives developed by Renata Grzegorczykowa to the needs of SGJP (with slight modifications discussed in the first part of the text). Next, an experiment was conducted where lexemes from a list of 2500 most common SGJP words were connected with their derivatives and the latter were then classified according to the assumed rules.
The paper presents the results of a quantitative analysis of a corpus comprising Polish prime ministers' speeches given between 1919 and 2019. Its aim is to trace potential changes Polish political discourse has undergone over the last 100 years, which might also reflect wider extra-linguistic changes. To this end, it employs tools associated with corpus linguistics, specifically keywords analysis. The study identifies three main linguistically distinct periods in the recent history of Polish political discourse, but it also points to trends characterized by stability rather than change, reflecting the requirements of the genre. The use of quantitative methods puts into perspective some of the findings of previous qualitative studies by providing a broader and more nuanced diachronic view of Polish political discourse.
The paper presents Korpusomat, a new, free web-based platform for effortless building and analysing linguistic data sets (corpora). The aim of Korpusomat is to bridge the gap between corpus linguistics, which requires tools for corpus analysis based on various linguistic annotations, and modern multilingual machine learning-based approaches to text processing. A special focus is placed on multilinguality: the platform currently serves 29 languages, but more can be easily added per user request. We discuss the use of Korpusomat in multidisciplinary research, and present a case study located at the intersection between discourse analysis and migration studies, based on a corpus generated and queried in the application. This provides a general framework for using the platform for research based on automatically annotated corpora, and demonstrates the usefulness of Korpusomat for supporting domain researchers in using computational science in their fields.
This article presents TwitterEmo, a new dataset for emotion and sentiment analysis in Polish. TwitterEmo provides a non-domain-specific and colloquial language dataset, which includes Plutchik’s eight basic emotions and sentiment annotations for 36,280 tweets collected over a one-year period. Additionally, a sarcasm category is included, making this dataset unique in Polish computational linguistics. Each entry was annotated by at least four annotators. We present the results of the evaluation using several language models, including HerBERT and TrelBERT. The TwitterEmo dataset is a valuable resource for developing and training machine learning models, broadening possible applications of emotion recognition methods in Polish, and contributing to social studies and research on media bias.
The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements made on several fronts over the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 67 new languages, including 30 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g. missing gender and macron information. We have also amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.
Jordan Kodner, Salam Khalifa, Khuyagbaatar Batsuren, Hossep Dolatian, Ryan Cotterell, Faruk Akkus, Antonios Anastasopoulos, Taras Andrushko, Aryaman Arora, Nona Atanalov, Gábor Bella, Elena Budianskaya, Yustinus Ghanggo Ate, Omer Goldman, David Guriel, Simon Guriel, Silvia Guriel-Agiashvili, Witold Kieraś, Andrew Krizhanovsky, Natalia Krizhanovsky, Igor Marchenko, Magdalena Markowska, Polina Mashkovtseva, Maria Nepomniashchaya, Daria Rodionova, Karina Scheifer, Alexandra Sorova, Anastasia Yemelina, Jeremiah Young, Ekaterina Vylomova. Proceedings of the 19th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology. 2022.
The paper presents a tool for automatic marking up of quantifying expressions, their semantic features, and scopes. We explore the idea of using a BERT based neural model for the task (in this case HerBERT, a model trained specifically for Polish, is used). The tool is trained on a recent manually annotated Corpus of Polish Quantificational Expressions (Szymanik and Kieraś, 2022). We discuss how it performs against human annotation and present results of automatic annotation of 300 million sub-corpus of National Corpus of Polish. Our results show that language models can effectively recognise semantic category of quantification as well as identify key semantic properties of quantifiers, like monotonicity. Furthermore, the algorithm we have developed can be used for building semantically annotated quantifier corpora for other languages.
The paper presents a manually annotated corpus of Polish quantificational expressions. The quantifier annotation was conducted on top of existing gold-standard data for Polish as its separate layer. This paper releases the data and gives an overview of the corpus and related tools. As far as we know, this is the first large-scale annotation of generalized quantifiers together with their crucial semantic properties, including monotonicity profile. We also discuss the potential further use of the corpus in linguistics and cognitive science.
Tiago Pimentel, Maria Ryskina, Sabrina J. Mielke, Shijie Wu, Eleanor Chodroff, Brian Leonard, Garrett Nicolai, Yustinus Ghanggo Ate, Salam Khalifa, Nizar Habash, Charbel El-Khaissi, Omer Goldman, Michael Gasser, William Lane, Matt Coler, Arturo Oncevay, Jaime Rafael Montoya Samame, Gema Celeste Silva Villegas, Adam Ek, Jean-Philippe Bernardy, Andrey Shcherbakov, Aziyana Bayyr-ool, Karina Sheifer, Sofya Ganieva, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Andrew Krizhanovsky, Natalia Krizhanovsky, Clara Vania, Sardana Ivanova, Aelita Salchak, Christopher Straughn, Zoey Liu, Jonathan North Washington, Duygu Ataman, Witold Kieraś, Marcin Woliński, Totok Suhardijanto, Niklas Stoehr, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Richard J. Hatcher, Emily Prud'hommeaux, Ritesh Kumar, Mans Hulden, Botond Barta, Dorina Lakatos, Gábor Szolnok, Judit Ács, Mohit Raj, David Yarowsky, Ryan Cotterell, Ben Ambridge, Ekaterina Vylomova. Proceedings of the 18th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology. 2021.
The article describes the well-known and widely used National Corpus of Polish in a new setup. The update consists of the annotation scheme modification in the morphosyntactic layer (especially in its parts related to the grammatical gender), as well as adding new layers of annotation: the syntactic layer and the named entities layer. All three layers are indexed in the MTAS corpus search engine and can be referenced in CQLcorpus queries.
The paper describes the process of building the electronic corpus of 17th- and 18th-century Polish texts, a relatively large, balanced, structurally and morphologically annotated resource of the Middle Polish language, available for searching at https://www.korba.edu.pl . The corpus consists of samples extracted from over seven hundred texts written and published between 1601 and 1772, summing up to a total size of 13.5 million tokens which makes it one of the largest historical corpora for a Slavic language.
The article presents the Korpusomat web application for creating user’s own annotated linguistic corpora. The application offers an automatic annotation of texts and the ability to search it based on the annotation of inflectional and syntactic features of words and named entities. All annotation layers are presented along with examples of their application in linguistic analysis. The Korpusomat also offers statistical summaries of the collected data, as well as the possibility of sharing the created corpora with other users.