We present a collection of human association norms to German personal name compounds (PNCs) such as Tore-Klose ('goal-Klose') and corresponding full names (Miroslav Klose), thus providing a novel testbed for PNC evaluation, i.e., analogical vs. contrastive positive vs. negative perception effects. The associations are obtained in an online experiment with German native speakers, analyzed regarding our novel intertwined PNC-person association setup, and accompanied by an LLM synthetic generation approach for augmentation.
Die Annotation politischer Diskurse ist traditionell zeit- und personalintensiv. Die jüngsten Entwicklungen im Bereich NLP versprechen allerdings eine Automatisierung komplexer Annotationsaufgaben. Feingetunte transformatorbasierte Sprachmodelle übertreffen inzwischen menschliche Annotator:innen bei einigen Annotationsaufgaben, aber sie setzten große manuell annotierte Trainingsdatensätze voraus. In unserem Beitrag untersuchen wir, inwieweit ein ursprünglich manuell annotierter Datensatz mit den heutigen NLP-Methoden automatisch repliziert werden kann, indem wir unüberwachtes maschinelles Lernen sowie Zero- und Few-Shot-Learning einsetzen.
The opinions of political actors (e.g., politicians, parties, organizations) expressed through claims are the core elements of political debates and decision-making. Political actors communicate through different channels: parties publish manifestos for major elections, while individual actors make statements on a day-to-day basis as reflected in the media. These two channels offer different approaches for analysis: Manifestos, on the one hand, are useful to characterize the parties’ positions at a global ideological level over time. In contrast, individual statements can be collected to analyze debates in particular policy domains on a fine-grained level, in terms of individual actors and claims. In this article, we summarize a series of studies we have carried out. We apply NLP-driven (semi-)automatic analyses on these two channels and compare their potentials and challenges. The fine-grained analysis yields rich insights into the communication but comes at the cost of three challenges: (a) a substantial hunger for manual annotation, introducing practical hurdles for analysis both within and across languages; (b) difficulties in claim classification arising from the uneven frequency distribution over the theory-based annotation schemas; (c) the need to map actor mentions onto canonical versions. Manifesto-based analysis avoids these challenges to a substantial extent when a more coarse-grained analysis of party positions is sufficient. We highlight the benefits and challenges of both approaches, and conclude by outlining perspectives for addressing the challenges in future research.
We present a comprehensive computational study of the under-investigated phenomenon of personal name compounds (PNCs) in German such as Willkommens-Merkel ('Welcome-Merkel'). Prevalent in news, social media, and political discourse, PNCs are hypothesized to exhibit an evaluative function that is reflected in a more positive or negative perception as compared to the respective personal full name (such as Angela Merkel). We model 321 PNCs and their corresponding full names at discourse level, and show that PNCs bear an evaluative nature that can be captured through a variety of computational methods. Specifically, we assess through valence information whether a PNC is more positively or negatively evaluative than the person's name, by applying and comparing two approaches using (i) valence norms and (ii) pretrained language models (PLMs). We further enrich our data with personal, domain-specific, and extra-linguistic information and perform a range of regression analyses revealing that factors including compound and modifier valence, domain, and political party membership influence how a PNC is evaluated.
Annotation of political discourse is resource-intensive, but recent developments in NLP promise to automate complex annotation tasks. Fine-tuned transformer-based models outperform human annotators in some annotation tasks, but they require large manually annotated training datasets. In our contribution, we explore to which degree a manually annotated dataset can be automatically replicated with today's NLP methods, using unsupervised machine learning and zero- and few-shot learning.
We document the results of a retro-archiving and literary analysis of Dana Buchzik’s blog Ze zurrealism itzelf . The starting point of our study is to consider how the blog has developed since 2010, what poetic processes characterise it, and how its genesis and poetics can be reconstructed using what is available of the blog in the Internet Archive. We describe the structure of the archival object and use the data available to reconstruct older versions of the blog. We cover the specific types of changes in the course of the blog’s textual genesis, describe them as elements of its poetics, and compare thematic and poetological aspects of three selected reconstructed versions. The methods of retro-archiving and reconstruction as well as the epistemology of the archival object serve as a basis for assessing types of changes and literary versions and constitute a workflow that can be used to analyse other blogs. This workflow is also relevant for the preparation of a text-genetic digital edition of a blog.
Newspaper reports provide a rich source of information on the unfolding of public debates, which can serve as basis for inquiry in political science. Such debates are often triggered by critical events, which attract public attention and incite the reactions of political actors: crisis sparks the debate. However, due to the challenges of reliable annotation and modeling, few large-scale datasets with high-quality annotation are available. This paper introduces DebateNet2.0, which traces the political discourse on the 2015 European refugee crisis in the German quality newspaper taz. The core units of our annotation are political claims (requests for specific actions to be taken) and the actors who advance them (politicians, parties, etc.). Our contribution is twofold. First, we document and release DebateNet2.0 along with its companion R package, mardyR. Second, we outline and apply a Discourse Network Analysis (DNA) to DebateNet2.0, comparing two crucial moments of the policy debate on the "refugee crisis": the migration flux through the Mediterranean in April/May and the one along the Balkan route in September/October. We guide the reader through the methods involved in constructing a discourse network from a newspaper, demonstrating that there is not one single discourse network for the German migration debate, but multiple ones, depending on the research question through the associated choices regarding political actors, policy fields and time spans.
The research project »text sound«: mixed-methods-analysis of lyric poetry in text and tonal sound (funded by the Federal Ministry for Education and Research, BMBF) aims to undertake a systematic and diachronic investigation of the relationship between literary texts, especially lyric poetry from the Romantic period, and their phonetic realisation in recitations or musical performances. Ideas of orality, sound and voice, which are particularly associated with poetry, are investigated empirically and also theorised in the line with modern approaches to the analysis of lyric poetry. Of particular importance is the experimental approach of speech synthesis, i.e. using computers to artificially produce a human sounding voice; this approach makes it possible to explore an ideal-typical realisation of the text and to test the aesthetic peculiarity of human realisations.
We present the steps taken towards an exploration platform for a multi-modal corpus of German lyric poetry from the Romantic era developed in the project "textklang". This interdisciplinary project develops a mixed-methods approach for systematic investigations of the relationship between written text (here lyric poetry) and its potential and actual sonic realisation (in recitations and musical performances). The multi-modal "textklang" platform will be designed to technically and analytically combine three modalities: the poetic text, the audio signal of a recorded recitation and, at a later stage, music scores of a musical setting of a poem. The methodological workflow will enable scholars to develop hypotheses about the relationship between textual form and sonic/prosodic realisation based on theoretical considerations, text interpretation and evidence from recorded recitations. The full workflow will support hypothesis testing either through systematic corpus analysis alone or with additional contrastive perception experiments. For the experimental track, researchers will be enabled to manipulate prosodic parameters in (re-)synthesised variants of the original recordings. The focus of this paper is on the design of the base corpus and on tools for systematic exploration - placing special emphasis on our response to challenges stemming from multi-modality and the methodologically diverse interdisciplinary setup.
Structured argumentation is a well-studied topic in the literature on computational models of argument, and several logic-based formalisms for defining structured argumentation have been proposed, including Assumption-Based Argumentation (ABA). The "semantics" of structured argumentation is usually defined by means of notions of "extensions", sanctioning some sets of arguments as dialectically acceptable together (and others as dialectically unacceptable). Arguments in structured argumentation are usually defined as trees and extensions as sets of such tree-based arguments with various properties depending on the particular argumentation semantics. However, these arguments and extensions may have redundancies as well as circularities, which are conceptually and computationally undesirable. Focusing on the specific case of ABA, I will discuss novel notions of arguments and extensions, both defined in terms of graphs. I will show that this avoids the redundancies and circularities of standard accounts, and set out the relationship to standard tree-based arguments and admissible/grounded extensions (as sets of arguments). Finally, I will discuss how sets of tree-based arguments and graph-based arguments may be applicable to structure natural language explanations for NLP tasks. We present progress from the CUEPAQ project. The goal of the project is to explore the role of linguistic cues (stylistic and interpretational) in determining personalized argument quality. The project works under the assumption that argument quality is subjective to users, or, at least, certain user groups. Based on this, we aim to produce preference profiles that point out patterns in the judgment of argument quality. For a controlled investigation of linguistic cues, we develop the minimal pair corpus that contains arguments from varying domains and variations of these arguments that are constructed based on the linguistic idea of minimal pairs. More concretely, we focus on variations in the premise of traditional premise/conclusion pairs while paying particular attention to the effect of the change on the relation between premise and conclusion. In example (1), (1a), and (1b) illustrate such a contrast. While (1a) invites the underlying inference, (1b) changes the view on the relationship between premise and conclusion to an ”attack” relation, i.e., the premise does not support the conclusion, thus, changing the quality of the argument based on the choice of words in the premise. For this workshop, we present progress on creating the minimal pair corpus. This includes investigations of various linguistic cues and their potential impact on argument quality. We focus on linguistic features that are associated with structuring beliefs and features that invite implicatures, such as the one illustrated in (1). More generally, we take a look at features that change the surface form of the argument minimally but that have semantic and pragmatic repercussions affecting the attitude towards particular propositions and inferences. Based on this, we present ongoing and planned work that aims at exploring the generalizability of the effect of these features and their meaning on the perceived argument quality. We also explore how to take into consideration pre-existing beliefs and stances and their effects on argument quality. In the ACQuA project, we develop algorithms and technology to understand and answer comparative information needs like ’Which is better, Bali or Phuket?’. However, subjectivity in perspectives can be problematic for a "fair" comparison—some answers on the Web might just irrationally prefer a particular option over the other. We for subjective a less some like year that I visit Phuket, there is much more on the beach’. By creating a dataset of questions and answers manually labeled as biased or not, we want to be able to train bias classifiers as a component of a comparative question answering system that is able to inform its users about possible “subjectivities” in the retrieved answers. Generating an argumentative conclusion from a set of textual premises is a challenging task, due to an extensive range of possible conclusions. In order to provide a conclusion generation model with guidance towards generating conclusions from a particular perspective, we explore the impact of conditioning the model on information about the desired framing. We experiment with two kinds of frame information: The first kind is the fine-grained issue-specific framing, describing the discussed perspective with a free-text label and, therefore, using issue tailored frame labels. The other kind is the generic framing. These frame sets are issue-agnostic and, therefore, more course-grained. We decide to experiment with the Media-Frames-set, containing 15 different frame classes. Using a dataset with given issue-specific frame labels for each argument, we explore the use of these frame information as well as the inferred generic frame information for guiding the conclusion generation process. Beyond enriching the model’s input with frame information, we investigate the impact of strategies to further improve the generated conclusion by an informative label smoothing method that dynamically smooths one-hot-encoded reference conclusion vectors as a regularization mechanism, also by using prior frame knowledge about word frequencies for each generic frame, and a conclusion re-ranking strategy based on reference-less scores at inference time. We evaluate the benefits of our methods using metrics for automatic evaluation complemented with an extensive manual study. Our results show that frame-guided conclusion generation is beneficial: it increases the ratio of valid and novel conclusions by 23%-points compared to a baseline without frame information. Our work indicates that by injecting different types of frame information, conclusion generation can be directed towards desired aspects, and, at the same time, it can be manually confirmed to yield more valid and novel conclusions. A natural different points of and form opinions the exchange The analysis of policy debates on key issues (e.g., migration, pension, Covid-19) allows researchers to trace how political decisions are construed and reached through deliberation and argumentation: What measures do political and institutional actors propose, discuss, and implement in different countries? What is the underlying reasoning given to support the presented policies? What coalitions are formed or changed through dynamic discursive constellations? Unraveling these questions requires to extract and combine political entities (deliberate actors), their corresponding policy propositions (political claims), and the supporting justifications (frames) from multilingual text sources. We further aim to approach these questions a) on a cross-sectional, domestic level by annotating the manifestos of various German parties and analyzing similarities between their programs and their proximity according to their policy positioning; and b) from a comparative and longitudinal perspective by contrasting the newspapers debates around central themes (e.g., lockdown, mask and vaccine mandates within the COVID-19 discourse) in Germany and the UK. The goal of both approaches is to carve out differences and similarities both within and between national debates under a relational, network based viewpoint. This work touches on several computational tasks such as claim and justification detection and classification (categorize both policies and justifications), and the scaling of text with large language models to capture party stances along different political dimensions. In our presentation, we will first present results regarding text-based party similarities and then, sketch our roadmap to open up the comparative perspective described above. The ReCAP II project contributes to inference and summarization in argumentation in various ways: We cluster similar arguments that may come from heterogeneous sources, both to avoid redundancy, as well as to consider the frequency of arguments as an indication of their strength in convincing. In an effort to build complex argumentation machines, we develop automated, manual, and even hybrid argument mining approaches that allow for using structural information for subsequent tasks. Given an argument graph as a query to a retrieval system, we perform a case-based retrieval that incorporates the structure as well as the semantics. We intend to further increase the relevance of the retrieved arguments by generalizing/specializing them towards the query. In the ACQuA project, we develop algorithms to understand and answer comparative information needs like “Is a cat or a dog a better friend?” by retrieving and combining facts, opinions, and arguments from web-scale resources. Ideally, an answer explains why under what circumstances which comparison alternative should be chosen. Retrieval-based comparative question answering starts with identifying the important constituents: (1) the objects that should be compared (e.g., ‘cat’ and ‘dog’ in the above example), (2) the aspects that indicate which properties should be emphasized in a comparative answer (e.g., ‘friend’ ), and (3) predicates that guide the direction of the comparison (e.g., ‘better’ instead of ‘worse’ ). When deriving a comparative answer by combining different sources (e.g., different web pages), the following steps can be important: (1) relevance assessment of the individual sources (e.g., a web forum on pets might be more relevant than a page on cat or dog movies), (2) quality assessment and stance detection (e.g., pro ‘cat’ or pro ‘dog’ ) of the retrieved arguments, (3) argument clustering based on the semantic similarity, stance, and quality, (4) re-ranking based on the predicted stance and quality, and (5) answer generation from the final ranking. So far, our fine-tuned RoBERTa-based token classifier (trained and evaluated on 3,500 manually labeled comparative
Many tasks in text-based computational social science (CSS) involve the classification of political statements into categories based on a domain-specific codebook. In order to be useful for CSS analysis, these categories must be fine-grained. The typically skewed distribution of fine-grained categories, however, results in a challenging classification problem on the NLP side. This paper proposes to make use of the hierarchical relations among categories typically present in such codebooks: e.g., markets and taxation are both subcategories of economy, while borders is a subcategory of security. We use these ontological relations as prior knowledge to establish additional constraints on the learned model, thus improving performance overall and in particular for infrequent categories. We evaluate several lightweight variants of this intuition by extending state-of-the-art transformer-based text classifiers on two datasets and multiple languages. We find the most consistent improvement for an approach based on regularization.
The analysis of public debates crucially requires the classification of political demands according to hierarchical claim ontologies (e.g. for immigration, a supercategory “Controlling Migration” might have subcategories “Asylum limit” or “Border installations”). A major challenge for automatic claim classification is the large number and low frequency of such subclasses. We address it by jointly predicting pairs of matching super- and subcategories. We operationalize this idea by (a) encoding soft constraints in the claim classifier and (b) imposing hard constraints via Integer Linear Programming. Our experiments with different claim classifiers on a German immigration newspaper corpus show consistent performance increases for joint prediction, in particular for infrequent categories and discuss the complementarity of the two approaches.
Jonas Kuhn合作论文数Institute for Natural Language Processing, University of Stuttgart25