Mobile devices are increasingly central as sources of up-to-date information, making the precise recording of information behavior on these devices more relevant for research. Established methods of automated data collection are reaching their limits when capturing in-app communication, such as political content within social media applications. Based on two case studies, we present two approaches that enable the identification of exposure to relevant content and, to some extent, the collection of in-app content. The first case study focuses on identifying relevant exposure across different apps on a mobile device using app tracking and screen recordings. The second case study focuses on linking exposure to seen content and deriving respective content features to obtain an enriched dataset. We discuss the advantages and limitations of both approaches and present conceptual frameworks for processing and analyzing such data.
Advancements in Large Language Models have been showing important research opportunities within the field of communication studies. It offers the capacity to conduct large-scale content classification and annotation with low computational expertise and reduced manual coding efforts, potentially allowing more possibilities for researchers in social sciences to explore understudied topics (Bail, 2023; Chang et al., 2024). Because of its functioning and vast domains and language training, LLMS also potentially unlocks more generalizable, complex, and diverse analyses across various communication materials than previous computational tools and approaches (Chang et al., 2024). These materials encompass a wide spectrum, ranging from journalistic content to the digital discourse of political actors and social media conversation threads. At the same time, LLMs also raise important concerns with potential biases, data privacy, models’ transparency, environmental impact, and power imbalances (Jameel et al., 2020; Fecher et al., 2023). Although highly discussed recently, as a recent topic LLMs still need deeper theoretical elaboration and dialogue between empirical investigations specifically for communication scholars (Gil de Zúñiga et al., 2024; Guzman and Lewis, 2020). Our panel assembles a collection of case studies that harness LLMs to tackle text classification and annotation tasks related to media and communication problems, issues, and topics. These research papers engage in an exploration of: (a) pipeline structuring: diverse methodologies for structuring effective pipelines tailored to this form of analysis; (b) tools and models comparison: comparisons of the various LLMs tools and models available for text classification and annotation, highlighting their strengths and weaknesses; (c) optimal variables and tasks: identifying the variables and tasks where LLMs demonstrates exceptional performance and reliability; (d) limitations: discussions on the existing limitations of these tools, including limitations related to specific tasks, variables, languages and data formats; (e) prompt development: strategies for developing, adapting and adjusting prompts that allows better results for specific tasks; and (e) ethical and political dimensions: an examination of the ethical and political considerations inherent in the deployment of LLMs in communication research. This panel puts together valuable efforts of different research groups across the world to not only use, but also reflect on the use of LLMs in Communication studies. They show important avenues for the field to think about different approaches to validity, ethics and truthful cooperation between humans and computational models without erasing the challenges of doing so, and the disagreements - not only between humans, but between humans and their computational language models too.
Search engines are both frequently used and widely trusted sources of current political information. While research has examined the different stages of information seeking via search, the question of how political preferences and sociodemographic factors impact political online search has received less attention. In particular, the way in which attitudinal factors, such as sentiment toward a politician or interest in their personal life motivate information seeking have not been widely studied. We present findings from a panel study that combines browser tracking data with survey results for 1863 German participants. Our results suggest that both sympathy and antipathy toward a selection of popular politicians is predictive of searching for them. Queries also vary in their composition, with searches for female politicians highlighting their personal lives and physical attributes somewhat more often than for males. We conclude that the role of “soft” factors in motivating information seeking on political actors remains underexplored.
This study examines how German social media commentators express their views on democratic values and nostalgically construct the past in public Facebook (FB) comments. Using a mixed-methods approach that combines qualitative content analysis with Large Language Model (LLM)-driven data interviews, we analyze a diverse set of comments posted between April 2010 and November 2019 in German-language FB groups. The study aims to provide insights into how right-wing commentators articulate their beliefs about democratic values and history, highlighting their political aspirations. Our analyses highlight the significant role of nostalgia in these discourses, with the past portrayed as a more unified and less complicated time, often linked to nativist ideologies. In these narratives, the past is constructed as morally superior to a corrupt and irresponsible present, with significant historical events such as World War II and the “refugee crisis” serving as markers of Germany’s alleged decline. By drawing on Hannah Arendt’s theories of the public sphere, we explore how these discourses challenge pluralism and democratic principles. Additionally, we introduce an innovative methodological approach by combining LLM with qualitative analysis to better understand political discourse on social media.
Mining detailed content from mobile human-computer interactions often relies on broad log files, which limits the specificity of analyses. Our study presents a unique approach capitalizing on Optical Character Recognition to continuously detect keywords across all applications and media formats, correlating this with system logs for context and duration tracking. Such a strategy even allows tracing of topics within pictorial content like memes on social media. A privacy concept based on a whitelist for keywords and topics and anonymized log files addresses typical concerns of potential study participants regarding their personal data. In our four-month study involving 25 participants, we generated an expansive nine-million-point dataset. This detailed dataset not only validates the efficacy of our approach but also exemplifies its capacity for rich, cross-app, long-term interaction analysis. In order to clarify data protection issues, we conducted qualitative interviews with eight of the participants on a voluntary basis.
The study of right-wing alternative news sources has moved to the center of scholarly attention recently. Such sources cater to news consumers characterized by extreme political views and mistrust toward mainstream news. However, research into the predictors of alternative news consumption is still scarce. We approach this gap by combining online tracking and survey data from 2,009 German panel participants. We find conspiratorial thinking and pro-Russian stance to be predictors of alternative news consumption, both in absolute and in relative terms. Our analysis thus contributes to a more nuanced understanding of alternative news consumers.
The aim of this article is to more precisely define the field of research on the automation of communication, which is still only vaguely discernible. The central thesis argues that to be able to fully grasp the transformation of the media environment associated with the automation of communication, our view must be broadened from a preoccupation with direct interactions between humans and machines to societal communication. This more widely targeted question asks how the dynamics of societal communication change when communicative artificial intelligence—in short: communicative AI—is integrated into aspects of societal communication. To this end, we recommend an approach that follows the tradition of figurational sociology.
Scholars of media and communication never stop emphasizing the importance of journalism in informing the public and in serving as a conduit for critical societal debates. The metaphor of a theater performance is sometimes used in this context. In this perspective, news is a stage on which political actors play their parts, and citizens represent a frequently critical (and in the age of social media sometimes hostile) audience. Meanwhile, journalists set the stage and arrange the props, but mostly stay in the background, allowing public figures to grab the limelight. Arguably, journalism research can also be regarded as a performance of sorts, one that also has a front stage – polished papers, well-timed conference presentations – and a backstage – projects that are behind schedule, messy data analyses, the many perils of peer review. And yet, apart from giving each other occasional, and usually confidential tours of their own behind-the-scenes area, the degree to which journalism scholars make the research process and its products available beyond publications is still limited, possibly for the same reasons as in journalism – it is quite messy back there and we never seem to find the time to tidy up properly. Much data sits undocumented and forgotten on hard drives and many strategies for analysis remain improvised rather than thoroughly planned, particularly when relying on computational methods where the methodological playbook is still in the process of being written. Of course, the appendix was not invented yesterday, and data sharing in media and communication research more broadly is increasingly common, but it seems appropriate to say that digital journalism research – in spite of the copious volumes of data we increasingly regard as normal – is not yet as open and transparent regarding its methods, data, and research practices as seems desirable.
One in four German internet users claims that search engines are their main gateway to news and a majority of Germans reports to primarily use their smartphone over their laptop/desktop computer to access news online. Yet, search-engine providers such as Google have repeatedly pointed out to actively favor specific forms of technical content optimization for mobile devices (e.g., Accelerated Mobile Pages), raising the question of whether this preference results in biases toward mobile-optimized content. In light of regulatory changes, this study investigates source diversity and source prominence in news items related to a range of issues presented to users when searching for news-related subjects via a smartphone or laptop/desktop computer in Germany. Using agent-based testing, 75,767 searches were performed on Google in mid-2020, applying a dynamic set of search terms on a range of different topics. Results indicate significant discrepancies in source diversity and source prominence between smartphone and laptop/desktop computers that can largely be attributed to mobile-optimized content likely to reach primarily younger news consumers who favor mobile devices. Overall, however, source diversity is equally high on both types of devices, highlighting the necessity for future research to focus on algorithmic influences on societies shared understandings of relevance beyond source diversity.
How is political news shared online? This fundamental question for political communication research in today’s news ecology is still poorly understood. In particular, very little is known about whether and how news sharing differs from news viewing. Based on a unique dataset of ≈ 870,000 URLs shared ≈ 100 million times on Facebook, grouped by countries, age brackets, and months, we study the correlates of viewing versus sharing of political versus non-political news. We first identify websites that at least occasionally contain news items, and then analyze metrics of the news items published on these websites. We enrich the dataset with natural language processing and super- vised machine learning. We find that political news items are viewed less than non-political news items, but are shared more than one would expect based on their views. Furthermore, the source of a news item and textual features, which are often studied in clickbait research and in commercial A/B testing, matter. Our findings are conditional on age, but are very similar across four different countries (Italy, Germany, Netherlands, Poland). While our research design does not allow for causal claims, our findings suggest that future work is well-advised to both theoretically and methodologically differentiate between factors that may explain (a) viewing versus sharing of news, and (b) political versus non-political news.
The objective of this article is to define more precisely the field of research into the automation of communication, which is currently only vaguely discernible. The central thesis is that, in order to be able to grasp the transformation of the media environment associated with the automation of communication, the view must be broadened from the “direct interaction of humans and machines” to “societal communication”. Broadening our perspective as such allows us to ask how the dynamics of societal communication as a whole change when “communicative AI” becomes part of societal communication. To support this thesis, the article first takes a closer look at the automation of communication as a phenomenon. Against this background, the concept of communicative AI is then developed in more detail as a “sensitizing concept” that sensitizes to both the breadth and depth of the phenomenon. On this basis, the bridging concept of the “hybrid figuration” is developed in order to grasp the agency of communicative AI and to be able to connect to “definitive concepts” of social science and media and communication research. It becomes apparent that with communicative AI as a field of research, the basic concepts of communication and media research—“communication” and “media”—are themselves challenged. The article is concluded by a conclusion that shows the research perspectives resulting from such an approach.
We describe a novel computational dictionary for the study of right-wing populist conspiracy discourse (RPC) on the internet, specifically in the context of contemporary German politics. After first presenting our definition of conspiracy discourse and grounding it in antecedent research on mediated rhetoric at the intersection of right-wing populism and conspiracy theory, we proceed by outlining our approach to dictionary construction, relying on a combination of manual and automated methods. We validate our dictionary via parallel manual coding of 2,500 sentences using the categories contained in the dictionary as labels and compare the consensus result with the label assigned to each sentence by the dictionary, achieving satisfactory results. We then test our approach on two different datasets composed of alternative news articles and Facebook comments that spread conspiracy theories. Finally, we summarize our observations both on the methodological premises of the approach and on the object of populist right-wing conspiracy discourse and its dynamics more broadly. We close with an outlook on the potentials and limitations of the dictionary-based approach and future directions in applications of content analysis to the study of conspiracy discourse.
Twitter continuously tightens the access to its data via the publicly accessible, cost-free standard APIs. This especially applies to the follow network. In light of this, we successfully modified a network sampling method to work efficiently with the Twitter standard API in order to retrieve the most central and influential accounts of a language-based Twitter follow network: the German Twittersphere. We provide evidence that the method is able to approximate a set of the top 1% to 10% of influential accounts in the German Twittersphere in terms of activity, follower numbers, coverage, and reach. Furthermore, we demonstrate the usefulness of these data by presenting the first overview of topical communities within the German Twittersphere and their network structure. The presented data mining method opens up further avenues of enquiry, such as the collection and comparison of language-based Twitterspheres other than the German one, its further development for the collection of follow networks around certain topics or accounts of interest, and its application to other online social networks and platforms in conjunction with concepts such as agenda setting and opinion leadership.
In this paper we analyze Twitter as a news channel in which the network of followers and followees significantly corresponds with the message content. We classified our data into twelve topics analogous to traditional newspaper sections and investigated whether the spread of information depended upon the Twitter network of followers and followees. To test this, we mapped the social network related to each topic and calculated the occurrence of retweet and mention mes-sages whose senders and receivers were interconnected as followers and followees. We found that on average 10% of retweets (RT-messages) and 5% of direct mentions between users (AT-messages) in Twitter hashtags are sent and received by users interconnected as followers and followees. These figures vary considerably from topic to topic, ranging from 15%-19% within Technology, Special Events and Politics to 3%-5% within the categories Personalities and Twitter-Idioms. The results show that hard-news messages are retweeted by a considerably larger community of users interconnected as followers and followees. We then performed a statistical correlation analysis of the dataset to validate the classification of hashtag in news sections based on retweet connectivity.
Computational methods offer a new perspective on the evolving agendas of right-wing movements and parties online. This article showcases computational approaches to text analysis (specifically so-called topic models) to diachronically investigate nativist right-wing issues in social media by comparing comments posted on the Facebook page of the Pegida movement to those of the Alternative for Germany. After describing topic modelling as an increasingly popular method and drawing on the literature on right-wing nativism online, we investigate a set of shared issues relevant to the mobilization of commentators, including opposition to Islam, migration, the government and the media. We furthermore show contrastively how issue prevalence differs between the two groups, and how issue shares change over time, in some instances converging on a shared nativist core. We close with a series of suggestions on the utility of computation content analysis for the study of rapidly evolving political agendas.
Understanding how citizens keep themselves informed about current affairs is crucial for a functioning democracy. Extant research suggests that in an increasingly fragmented digital news environment, search engines and social media platforms promote more incidental, but potentially more shallow modes of engagement with news compared to the act of routinely accessing a news organization’s website. In this study, we examine classic predictors of news consumption to explain the preference for three modes of news engagement in online tracking data: routine news use, news use triggered by social media, and news use as part of a general search for information. In pursuit of this aim, we make use of a unique data set that combines tracking data with survey data. Our findings show differences in predictors between preference for regular (direct) engagement, general search-driven, and social media–driven modes of news engagement. In describing behavioral differences in news consumption patterns, we demonstrate a clear need for further analysis of behavioral tracking data in relation to self-reported measures in order to further qualify differences in modes of news engagement.
Since its inception, the internet has been as much technological as social, practical as ideological in character. This article examines academic discourse and asks how research on the multifaceted internet has evolved over the past 25 years. In order to investigate the formation of this academic field, we collected articles published in major academic journals dedicated to new media and digital communication as well as mainstream periodicals in communication studies over the past quarter of a century. Relying on a combination of (semi)automated content analysis and citation analysis, we find that articles related to the internet and its manifold aspects are cited more often than research on other topics. The literature review suggests that as the socio-material infrastructure of the internet has become deeply enmeshed in society its study has evolved from a niche pursuit to the discipline's core area of inquiry.