In argument mining research, educational texts such as student essays have been a popular target genre from the beginning. The annotated corpora available so far, however, focus on essays written by older (or adult) students with relatively high proficiency levels. In our work, we expand the range of texts to German essays by school students in grade 9, which display very different qualities on all levels of analysis. We show that a common approach to representing argument structure as trees is not sufficient to capture the constellations of argument components in those essays, and we propose a suitable extension of that scheme, which we applied to an initial corpus of 50 essays. Furthermore, we conducted experiments with large language models on a fine-grained version of the argument component type classification task and show that medium-sized open-source models can achieve promising classification results using only two labeled essays in a few-shot prompting approach.
Argumentation mining comprises several subtasks, among which stance classification focuses on identifying the standpoint expressed in an argumentative text toward a specific target topic. While arguments-especially about controversial topics-often appeal to emotions, most prior work has not systematically incorporated explicit, fine-grained emotion analysis to improve performance on this task. In particular, prior research on stance classification has predominantly utilized non-argumentative texts and has been restricted to specific domains or topics, limiting generalizability. We work on five datasets from diverse domains encompassing a range of controversial topics and present an approach for expanding the Bias-Corrected NRC Emotion Lexicon using DistilBERT embeddings, which we feed into a Neural Argumentative Stance Classification model. Our method systematically expands the emotion lexicon through contextualized embeddings to identify emotionally charged terms not previously captured in the lexicon. Our expanded NRC lexicon (eNRC) improves over the baseline across all five datasets (up to +6.2 percentage points in F1 score), outperforms the original NRC on four datasets (up to +3.0), and surpasses the LLM-based approach on nearly all corpora. We provide all resources-including eNRC, the adapted corpora, and model architecture-to enable other researchers to build upon our work.
Segmenting text into so-called "elementary discourse units" (EDUs) is a task that is relevant for several NLP applications, including discourse parsing or argument mining. In recent years, EDU segmentation has been addressed as part of a shared task on multilingual discourse parsing ("DISRPT"), where BERT-based encoder models proved particularly successful. The German language has been represented in DISRPT with the Potsdam Commentary Corpus, but recently, more German data with EDU segmentation has been published. In this paper, we conduct detailed tests on the German-language datasets that are currently available. We test a multilingual off-the-shelf model, several BERT-based encoders, and the current generation of LLMs. The results are analyzed both qualitatively and quantitatively and are compared to the multilingual state-of-the-art. We are making the best-performing model available as a tool that can be used by the community.
In this paper, we are testing sentence alignment on complex, semi-parallel corpora, i.e., different versions of the same text that have been altered to some extent. We evaluate two hypotheses: To make alignment algorithms more efficient, we test the hypothesis that matching pairs can be found in the immediate vicinity of the source sentence and that it is sufficient to search for paraphrases in a 'context window'. To improve the alignment quality on complex, semi-parallel texts, we test the implementation of a segmentation into Elementary Discourse Units (EDUs) in order to make more precise alignments at this level. Since EDUs are the smallest possible unit for communicating a full proposition, we assume that aligning at this level can improve the overall quality. Both hypotheses are tested and validated with several embedding models on varying degrees of parallel German datasets. The advantages and disadvantages of the different approaches are presented, and our next steps are outlined.
Par le biais d’un corpus de tous les tweets émis par des membres du parlement allemand sur une période de six ans, nous étudions la complexité formelle de l’argumentation trouvée dans des paires de tweets, où le second répond au premier. Nos mesures de complexité sont fondées sur le nombre d'unités argumentatives et les relations entre elles. Nous proposons une méthode pour identifier une argumentation potentiellement complexe, puis pour annoter manuellement un ensemble de paires de tweets en vue de leur structure argumentative. Malgré la faible longueur des tweets, nous constatons que les politiciens construisent des argumentations plutôt élaborées mais aussi que les structures trouvées impliquent des modifications du plan d’annotation utilisé précédemment dans des corpus similaires, mais constitués de genres textuels plus traditionnels.
Using a corpus of all tweets issued by German members of parliament in a recent 6-year period, we study the formal complexity of the argumentation found in tweet pairs, where the second is a reply to the first. Our complexity measures are based on the number of argumentative units and the relations among them. We suggest a method for identifying potentially complex argumentation and then manually annotating a set of tweet pairs for their argument structure. Despite the length limit of tweets, we find that politicians indeed perform rather elaborate arguments and that the structures require amendments to an annotation scheme that has previously been used in similar corpus projects, though on more traditional types of texts.
This paper addresses the problem of crossregister generalization in argument mining within political discourse. We examine whether models trained on adversarial, spontaneous U.S. presidential debates can generalize to the more diplomatic and prepared register of UN Security Council (UNSC) speeches. To this end, we conduct a comprehensive evaluation across four core argument mining tasks. Our experiments show that the tasks of detecting and classifying argumentative units transfer well across registers, while identifying and labeling argumentative relations remains notably challenging, likely due to register-specific differences in how argumentative relations are structured and expressed. As part of this work, we introduce ArgUNSC, a new corpus of 144 UNSC speeches manually annotated with claims, premises, and their argumentative links. It provides a resource for future in- and cross-domain studies and novel research directions at the intersection of argument mining and political science.
Communication aiming to persuade an audience uses strategies to frame certain entities in 'character roles' such as hero, villain, victim, or beneficiary, and to build narratives around these ascriptions. The Character-Role Framework is an approach to model these narrative strategies, which has been used extensively in the Social Sciences and is just beginning to get attention in Natural Language Processing (NLP). This work extends the framework to scientific editorials and social media texts within the domains of ecology and climate change. We identify characters' roles across expanded categories (human, natural, instrumental) at the entity level, and present two annotated datasets: 1,559 tweets from the Ecoverse dataset and 2,150 editorial paragraphs from Nature & Science. Using manually annotated test sets, we evaluate four state-of-the-art Large Language Models (LLMs) (GPT-4o, GPT-4, GPT-4-turbo, LLaMA-3.1-8B) for character-role detection and categorization, with GPT-4 achieving the highest agreement with human annotators. We then apply the best-performing model to automatically annotate the full datasets, introducing a novel entity-level resource for character-role analysis in the environmental domain.
The facilitating effect of connectives on discourse processing has been found to be smaller in result relations, compared to other relations (e.g., concession). In addition, connectives are hypothesized to facilitate more in some languages than in others due to typological differences between languages. Speakers of analytic languages (such as English) are assumed to rely more on contextual cues and therefore be less affected by the presence of a connective than speakers of synthetic languages (such as German), who are presumed to rely more on lexical information. We present two self-paced reading studies examining how the effect of a connective depends on the relation type and the language. We find that the presence of a connective facilitates reading more in concession relations than in result relations. This interaction between relation type and relation marking was only found in German.
This mixed-methods pilot study examines the effects of structured classroom debate training on the written persuasiveness of pro–con argumentative essays produced by ninth-grade students in non-academic-track schools. The research forms part of the “Fair Debating and Written Argumentation” project, the first in a German-speaking context to systematically integrate oral and written argumentation instruction. In the QASA (Qualitative Argumentation Structure Analysis) substudy, 18 essays from nine students were analyzed before and after a six-session debate intervention. Quantitative fivepoint ratings, based on a validated writing assessment framework and qualitative argumentation structure mapping, indicated that most students increased the number and variety of arguments, incorporated more examples, and improved their overall persuasive coherence. These findings align with international evidence demonstrating that structured debate fosters critical thinking and supports the transfer of skills from oral to written argumentation. Implications for inclusive writing instruction, formative feedback through argument structure diagrams, and the design of integrated oral–written argumentation curricula are discussed.
We explore the capability of four open-weight large language models (LLMs) in argumentation mining (AM). We conduct experiments on three different corpora; persuasive essays (PE), argumentative microtexts (AMT) Part 1 and Part 2, based on two argumentation mining subtasks: (i) argument component type classification (ACTC), and (ii) argumentative relation classification (ARC). This work aims to assess the argumentation capability of openweight LLMs, including Mistral 7B, Mixtral 8x7B, LLaMA2 7B and LLaMA3 8B in both, zero-shot and few-shot scenarios. Our results demonstrate that open-weight LLMs can effectively tackle argumentation mining subtasks, with context-aware prompting improving relation classification performance, though the models' effectiveness varies across different argumentation patterns and corpus types, suggesting potential for specialized adaptation in future argumentation systems. Our analysis advances the assessment of computational argumentation capabilities in open-weight LLMs and provides a foundation for future research.
Framing is an extensively used research method in the context of climate change communication research. Frames are an essential part of communication by which a speaker can put boundaries on complex issues and create a tailored lens through which to communicate while remaining true to reality. Though there is a rich literature looking at frames present in a variety of media and social media content, rather little attention has been given so far to scientific editorials. We extend upon previously published research on climate change related editorials from the journals Nature and Science by extending it with new data, new coding structures, new analyses, and results on automatic classification. Specifically, we show that applying frame analysis not as regularly done on text level but on the level of paragraphs (and moving to the level of individual sentences when necessary) leads to a more nuanced account of the underlying communication strategies. Next, in addition to analysing customary issue frames, we also code the rhetorical dimensions present in the paragraphs, informed by Entman’s four frame functions. In analysing the resulting data, we find a pervasive use of the Governance and Scientific issue frames in both journals, and for the rhetorical dimension a dominance of frames that represent the Describe Problem and the Moral Judgement perspectives. Finally, as an add-on to the qualitative work, we assess the capabilities of automatic text classification methods for our three tasks of determining (i) the degree of paragraph topicality, (ii) the issue frames, and (iii) the rhetorical frames. We compare various supervised und unsupervised methods and find that a RoBERTa model achieves decent performance on the frequent classes, while the rare classes pose problems. Similarly, zero-shot learning with LLMs is not yet able to provide a reliable classification, showcasing the results close to those of an SVM baseline. At the end of the paper, we situate our approach in the quest for even more fine-grained analyses and computational models of framing.
We probe a new approach to linguistic areas. Instead of similarity of a feature across languages of the area, we focus on its adaptation to the area. Adaptation is a set of changes and/or retentions in a language towards, but not necessarily into, similarity with the other languages of the area. Technically, we estimate adaptation by comparing the distance between the focus language from the area and a geographically and genealogically closely related language outside of the area (its benchmark language) as tertium comparationis. If the focus language is closer to the area than its benchmark, we interpret it as evidence for adaptation towards the other languages of the area. Adaptation includes all possible scenarios of change and non-change. We test word order and find that all languages of the CB area show effects of adaptation, with Baltic Romani and both Baltic languages being in the center of the area.