We present a study in the context of computational social science that explores the topics debated in the context of the 2021 German Federal Election by using the topic modeling technique BERTopic. The corpus consists of German language tweets posted by political party accounts of the major German parties, as well as tweets by the general public mentioning the party accounts. We examined the textual content of the tweets but also included the text in images that were posted into the analysis by extracting the text using optical character recognition (OCR). Our results show that the most frequently discussed topics are party-oriented policies (including call-to-action content), climate policy and financial policy, with these topics being discussed in tweets by both, the political party accounts and tweets by accounts mentioning them. In addition, we observed that some topics were discussed consistently throughout the year, such as the COVID-19 pandemic, climate policy or digitization, while other topics, such as the return to power of the Taliban in Afghanistan or Israel were debated to a greater extent at limited time frames during the election year.
This article presents a method of emotion analysis for German drama from the 17th to the 19th century that significantly goes beyond previous research approaches in computational literary studies. It is based on annotations of 17 dramatic texts resulting in 11,939 annotations which were used as training material to fine-tune a German language BERT model that achieves an average accuracy of 73% for the single-label emotion classification of fourteen emotion types in cross-validation. We apply the emotion classification on a corpus of 141 comedies and 92 tragedies to compare these genres. For tragedies, the mean proportion percentages of ‘suffering’ and ‘abhorrence’ are higher than for comedies. Inversely, mean percentages of ‘anger’ and ‘joy’ are higher for comedies than for tragedies. A new finding is the surprisingly high proportion of ‘anger’ in comedies. Emotion distribution of the last scenes in dramatic texts also proves the quality of the classified data in terms of literary studies. In addition, the emotion distribution for several subgenres of comedy is investigated including non-canonical works of wide circulation which reached the recipients directly through the depicted emotions in the Kasperl Plays. Comedies from 1740 to 1770 are characterized by a pairing of higher amounts of ‘friendship’ and ‘love’. Satirical comedies from the same period stand out due to high rates of ‘anger’ as well as ‘suffering’. The very successful Kasperl plays turn out to be characterized by a comparatively large percentage of ‘schadenfreude’ and ‘joy’.
This article presents the annotated corpus of the DFG research project Emotions in Drama . The main aim of the project was the training of transformer-based language models for the classification of emotions in over 300 dramas from the period 1650 to 1815. For this purpose, a corpus of 18 German dramas was annotated by two annotators with 13 individual emotions each, resulting in 21.000 emotion annotations and over 46.000 source / target annotations. We report agreement metrics, the distribution of emotions and the distribution of source and target annotations in the overall corpus as well as distributions by character gender. The corpus is a unique resource for the digital humanities, since there is no other corpus with context-sensitive emotion annotations based on historical emotion definitions for dramas from this period.
Zwischen dem Ende des Dreißigjährigen Krieges und dem Anfang der Restaurationsepoche entwickelt sich das Drama rasant und wird im deutschsprachigen Gebiet zur publikumswirksamsten Gattung dieses Zeitraums (Baumann 1985; Jahn 1996; Krämer 1998; Urchueguía 2015; Meid 2009: 327-501). Es wird zu einer ‚Schule der Affekte’ ( palaestra affectum, Rotermund 1972: 25), in der man lernen soll, gewünschte Emotionen zu empfinden und mit unerwünschten wie Angst, Neid oder Leid angemessen umzugehen (Schings 1971; Schings 1980; Wiegmann 1987; Meier 1993; Schulz 1988; Zeller 2005; Lukas 2005; Schonlau 2017). Heutige Konzepte von Emotionen bilden sich dabei erst nach und nach heraus, so dass sich einzelne Figurenaussagen und die Motivation der Handlung für Rezipient*innen der Gegenwart nicht immer unmittelbar erschließen. Die derzeitigen Bemühungen um eine digitale Konservierung von Dramentextbeständen und ihre Analyse müssen folglich um eine Erschließung der Emotionsstrukturen ergänzt werden, wenn dieser Teil des kulturellen Erbes so im kulturellen Gedächtnis konserviert werden soll, dass er auch verständlich bleibt (vgl. zum Begriff des kulturellen Erbes Tauschek 2013). Gemeint ist dabei nicht ein Alltagsverständnis von Gedächtnis im Sinne dessen, „was das Bewußtsein bewußt erinnert" (Luhmann 1999: 44). Es geht vielmehr um das Gedächtnis sozialer Systeme, das durch Kommunikation in der aktuellen Gegenwart benutzt und reproduziert wird. Plädiert wird vorwiegend dafür, die Emotionen in Dramen des Untersuchungszeitraumes im Speichergedächtnis zu bewahren (vgl. Assmann 2003: 133 ff.). Für das Funktionsgedächtnis, das nur das aktuell Anschlussfähige und Zukunftsorienterte verfügbar hält, sind jedoch zumindest diejenigen Emotionsverteilungen und -verläufe von Interesse, die auch heute noch die Dramaturgie von Dialogformaten bestimmen. Literaturwissenschaftler haben bisher schwerpunktmäßig wenige Texte einzelner Gattungen hinsichtlich der Frage untersucht, welche emotionale und rationale Wirkung durch sie erreicht werden soll (Pikulik 1965; Sauder 1974-1980; Schings 1971; Nolting 1986). Nur wenige Wissenschaftler*innen erforschen Emotionen der Figuren selbst (Anz 2011; Schulz 1988; Mönch 1993, 344– 350; Schonlau 2017). Das führt dazu, dass sehr wenig darüber bekannt ist, welche Figurenemotionen in Dramen dargestellt werden und wann sie im Handlungsverlauf zum Einsatz kommen. Obwohl die Untersuchung von Gefühlen und Emotionen in den letzten Jahren in der computergestützten Literaturwissenschaft für zahlreiche Textgenres an Bedeutung gewonnen hat (siehe Mohammad 2011; Nalisnick/Baird 2013; Reagan et al. 2016; Zehe et al. 2016; Schmidt/Burghardt 2018; Mc Hardy/Adel/Klinger 2019, Kim/Klinger 2019; Jacobs 2019; Pianzola/Rebora/Lauer 2020), wurden speziell emotionale Aspekte in Dramentexten bisher nur vereinzelt untersucht. Der Fokus lag dabei auf der Analyse von Valenz oder Polarität, d.h. der positiven respektive negativen Konnotation von Sätzen und Textteilen, und zumeist auch auf einzelnen Autoren bzw. Werken (Mohammad 2011; Nalisnick/Baird 2013; Schmidt/Burghardt 2018; Schmidt/Burghardt/Dennerlein 2018a; 2018b; Schmidt/Burghardt/Wolff 2019; Schmidt/Burghardt/Dennerlein/Wolff 2019a; 2019b). Die Untersuchung von Emotionsverteilungen und -verläufen in einzelnen Texten, in verschiedenen Genres und Epochen ist das Ziel des Projekts „Emotions in Drama“, das seit April 2020 Emotionen in Dramen dieses Zeitraums in einem kombinierten Verfahren aus händischer Annotation und ihrer Vorhersage mittels Deep Learningbasierter Sprachmodelle erforscht. Im Folgenden sollen konzeptionelle Überlegungen und erste Ergebnisse dieses Projekts vorgestellt werden.
We present an exploratory study performing distant viewing via computer vision methods in the genre of fantasy movies. As a case study we use 10 modern fantasy movies of the Harry Potter franchise (also referred to as Wizarding World franchise). We apply methods and state-of-the-art models for color and brightness analysis, object detection, location classification as well as facial emotion recognition. We present descriptive results as well as inference statistics. Furthermore, we discuss the results and the quality of the methods for this unique use case and give examples. We were able to find significant differences in our statistical analysis in the results of the methods across the movies with the movies of the Harry Potter series getting darker and negative emotional expressions on faces becoming more frequent.
We present the results of a project performing sentiment analysis on tweets from German politicians and party accounts for the 2021 German federal election. We collected over 58,000 tweets from the Twitter accounts of the seven parties represented in the German Bun-destag, of which a selection of 2,000 tweets were annotated by three annotators. Based on the annotated data, we implemented multiple sentiment analysis approaches and evaluated the sentiment classification performance. We found that transformer-based models like bidirectional encoder from transformers (BERT) performed better than traditional machine learning models such as Naive Bayes and lexicon-based models like GerVADER. The best performing BERT model achieved an accuracy of 93.3% and macro f1 score of 93.4%. Applying sentiment analysis on the overall corpus via this method showed that overall, negative sentiment was most frequent and that there were multiple major shifts in sentiment a few months before and after the election. Furthermore, we found that tweets from opposition parties had on average more negative sentiment than those from governing parties.
In this paper, we present a survey study with 119 participants conducted in German, which investigates respondents’ Facebook behavior. In particular, the survey provides insight into how the individual factors gender, user type and trust in social media influence information behavior with respect to false information on Facebook. Our participants’ Facebook use is predominantly passive, the trust in social media is mediocre and most users claim to encounter false information on a weekly basis. If the truthfulness of information is verified it is mostly done by checking alternative sources and for the most part, users do not react actively to false information on Facebook. Of the different categories of Facebook users studied, more active and intensive users of Facebook (posters and heavy users) encounter false information the most. These users are the only user group to report posts with false information to Facebook or interact with the post. Participants with higher trust in social media tend to check the comments of a post to verify information.
We present the results of an evaluation study in the context of lexicon-based sentiment analysis resources for German texts. We have set up a comprehensive compilation of 19 sentiment lexicon resources and 20 sentiment-annotated corpora available for German across multiple domains. In addition to the evaluation of the sentiment lexicons we also investigate the influence of the following preprocessing steps and modifiers: stemming and lemmatization, part-of-speech-tagging, usage of emoticons, stop words removal, usage of valence shifters, intensifiers, and diminishers. We report the best performing lexicons as well as the influence of preprocessing steps and other modifications on average performance across all corpora. We show that larger lexicons with continuous values like SentiWS and SentiMerge perform best across the domains. The best performing configuration of lexicon and modifications considering the f1-value and accuracy averages across all corpora achieves around 67%. Preprocessing, especially stemming or lemmatization increases the performance consistently on average around 6% and for certain lexicons and configurations up to 16.5% while methods like the usage of valence shifters, intensifiers or diminishers rarely influence overall performance. We discuss domain-specific differences and give recommendations for the selection of lexicons, preprocessing and modifications.
We present SentText, a web-based tool to perform and explore lexicon-based sentiment analysis on texts, specifically developed for the Digital Humanities (DH) community. The tool was developed integrating ideas of the user-entered design process and we gathered requirements via semi-structured interviews. The tool offers the functionality to perform sentiment analysis with predefined sentiment lexicons or self-adjusted lexicons. Users can explore results of sentiment analysis via various visualizations like bar or pie charts and word clouds. It is also possible to analyze and compare collections of documents. Furthermore, we have added a close reading function enabling researchers to examine the applicability of sentiment lexicons for specific text sorts. We report upon the first usability tests with positive results. We argue that the tool is beneficial to explore lexicon-based sentiment analysis in the DH but can also be integrated in DH-teaching.
We present first results of an exploratory study about sentiment analysis via different media channels on a German historical play. We propose the exploration of other media channels than text for sentiment analysis on plays since the auditory and visual channel might offer important cues for sentiment analysis. We perform a case study and investigate how textual, auditory (voice-based), and visual (face-based) sentiment analysis perform compared to human annotations and how these approaches differ from each other. As use case we chose Emilia Galotti by the famous German playwright Gotthold Ephraim Lessing. We acquired a video recording of a 2002 theater performance of the play at the “Wiener Burgtheater”. We evaluate textual lexicon-based sentiment analysis and two state-of-the-art audio and video sentiment analysis tools. As gold standard we use speech-based annotations of three expert annotators. We found that the audio and video sentiment analysis do not perform better than the textual sentiment analysis and that the presentation of the video channel did not improve annotation statistics. We discuss the reasons for this negative result and limitations of the approaches. We also outline how we plan to further investigate the possibilities of multimodal sentiment analysis.
In this paper, we present a survey study with 119 participants conducted in German, which investigates respondents’ Facebook behavior. In particular, the survey provides insight into how the individual factors gender, user type and trust in social media influence information behavior with respect to false information on Facebook. Our participants’ Facebook use is predominantly passive, the trust in social media is mediocre and most users claim to encounter false information on a weekly basis. If the truthfulness of information is verified it is mostly done by checking alternative sources and for the most part, users do not react actively to false information on Facebook. Of the different categories of Facebook users studied, more active and intensive users of Facebook (posters and heavy users) encounter false information the most. These users are the only user group to report posts with false information to Facebook or interact with the post. Participants with higher trust in social media tend to check the comments of a post to verify information.
We present first results of the project “Emotions in Drama” in which we explore the annotation of emotions and the application of computational emotion analysis, predominantly deep learning-based methods, in the context of historical German plays of the time around 1800. We performed a pilot annotation study with five plays generating over 6,500 annotations for up to 13 sub-emotions structured in a hierarchical scheme. This emotion scheme includes common types like joy , anger or hate but also concepts that are specifically important for German literary criticism of this period like friendship , compassion or Schadenfreude . We evaluate the performance of various methods of emotion-based text sequence classification including lexicon-based methods, traditional machine learning, fastText as static word embedding, various transformer models based on BERT - or ELECTRA -architectures and pretrained with contemporary language, transformer-based methods pretrained or finetuned for historical and/or poetic language as well as the finetuning of BERT models via our own corpora and plays. We do achieve state-of-the-art results with hierarchical levels with two or three classes, i. e. the classification of valence (positive/negative). The best models are the transformer-based models gbert-large and gelectra-large by deepset pretrained on large corpora of contemporary German, which achieve accuracy values of up to 83%. Lexicon-based methods, traditional machine learning as well as static word embeddings are consistently outper-formed by transformer-based models. Models trained on historical texts show small and inconsistent improvements. The performance becomes significantly smaller for settings with multiple sub-emotions like 6 or 13 due to the general challenge and class imbalances in which the models achieve 57% and 47% respectively. We discuss how we intend to continue our annotations and how to improve the prediction results via various optimization techniques in future work.
We present results of a project on emotion classification on historical German plays of Enlightenment, Storm and Stress, and German Classicism. We have developed a hierarchical annotation scheme consisting of 13 sub-emotions like suffering, love and joy that sum up to 6 main and 2 polarity classes (positive/negative). We have conducted textual annotations on 11 German plays and have acquired over 13,000 emotion annotations by two annotators per play. We have evaluated multiple traditional machine learning approaches as well as transformer-based models pretrained on historical and contemporary language for a single-label text sequence emotion classification for the different emotion categories. The evaluation is carried out on three different instances of the corpus: (1) taking all annotations, (2) filtering overlapping annotations by annotators, (3) applying a heuristic for speech-based analysis. Best results are achieved on the filtered corpus with the best models being large transformer-based models pretrained on contemporary German language. For the polarity classification accuracies of up to 90% are achieved. The accuracies become lower for settings with a higher number of classes, achieving 66% for 13 sub-emotions. Further pretraining of a historical model with a corpus of dramatic texts led to no improvements.