
Understanding cross-national news coverage is central to international communication research. However, quantifying how countries appear in the news remains methodologically challenging: simple keyword counts lack context, while manual coding does not scale. This paper introduces NEClass, a reproducible, open-source pipeline that combines standard Named Entity Recognition (NER) with a fine-tuned Large Language Model (LLM) to classify entities' country affiliations. Unlike static dictionary approaches, NEClass leverages the latent world knowledge of modern LLMs and local context to resolve ambiguities dynamically. I demonstrate the pipeline's validity regarding country attribution (weighted F1 of 0.95), show that it performs more reliably than established baseline approaches (including keyword matching and zero-shot LLM prompting), and illustrate its utility through a case study of the Eurozone crisis coverage in German and Italian newspapers. Ultimately, NEClass provides researchers with a robust and accessible toolkit to operationalize international attention patterns beyond the limitations of purely keyword-based metrics.
Concerns about publication bias and the replicability of psychological research have motivated the development of methods that evaluate the credibility of published significant results. Two prominent approaches are p-curve and z-curve, which use the distribution of significant p-values to estimate the evidential value of a set of studies. Although both methods rely on selection models, they differ critically in how they handle heterogeneity in true effect sizes and statistical power. We systematically vary heterogeneity in effect sizes (tau) across a realistic range and show that both methods perform well when heterogeneity is small, but only z-curve performs well when heterogeneity is moderate. We illustrate the practical consequences using data from the Reproducibility Project, showing that only z-curve correctly predicts that many replication attempts will fail. Because z-curve performs as well as or better than p-curve across conditions, we recommend using z-curve to evaluate the credibility and evidential value of meta-analytic findings, particularly when meta-analyses synthesize conceptual rather than direct replications.
Applied data analysts who recognize the power of numbers express concern for erasing smaller demographic groups with their statistical analyses. Data analysts ought to strive for representative data categories when numbers impact participants' lives, yet underrepresented groupmembers are often lumped into nonrepresentative categories or excluded from analysis. We conducted interviews with 31 applied data analysts to thematically investigate how they prioritized group inclusion. We learned analysts employed mostly a QuantCrit approach toward social group inclusion by (a) disaggregating demographic social groups when appropriate, (b) educating leadership, (c) acknowledging demographic missingness and error, (d) providing response-fluid survey options, (e) engaging in reflexivity, and (f) codesigning demographic measures. They also employed qualitative methods to present marginalized groups' stories and counterstories. Contributions are five-fold: First, the results provide evidence that qualitative data representing smaller demographic groups were perceived as more effective in attracting funding and influencing change than numerical data. Second, inclusion conceptual boundaries are explored by studying professionals who confront normative quantitative data practices. Third, applied professionals treated demographic measurement structure with flexibility to create culturally and locally relevant demographic measures. Fourth, we share the research utility of QuantCrit tenets. Last, and most importantly, reflexivity is required at every data stage.
Our field has been questioning whether self-reported media use actually captures the time spent using media, resulting in a shift from subjective to objective measurements. We advocate for a new paradigm that, instead of abandoning subjectivity, places it at the center of our research. To this end, we introduce a novel theoretical and methodological framework called 'felt time', which captures nuances in how individuals perceive the passage of time while interacting with media. In two studies (N1 = 68, N2 = 1,030), we develop and validate a measure called the Felt Time Scale (FTS) and its short version (FTS-S), revealing that studies relying on subjective self-reports or the more objective, tracked screen time may have overlooked a critical dimension of users' media experience: felt time. We validate the scale in the context of screen media use and show empirically that screen engagement situationally alters how people feel time, producing unique effects. We conclude by proposing future directions that integrate this perspective, aiming to foster a more comprehensive understanding of screen time.
Data donations are a method to access user-level digital trace data, as they provide fine-grained measures of content exposure on social media. The interest in data donation as a data collection method is accompanied by a broad uncertainty about the reasons that drive the donation of data by users. The current literature lacks comparative analysis across various platforms. This study investigates platform-specific predictors for data donation behavior of a non-probability quota sample of German social media users (N = 2,296) for YouTube, Facebook, Instagram, and TikTok and the resulting non-response biases. Based on the analysis of 340 data donation packages, we find that participants are less likely to donate TikTok data compared to the other platforms. Gender is the main driver during the willingness step for drop offs, while political leaning is a key predictor for all platforms except Facebook during the donation stage. Data donors tend to self-report less active social media usage with news and political content than those who donate data. Our findings highlight the importance of considering platform-specific differences in expected donation rates, biases, and the potential for discrepancies between indicated willingness and actual donation behavior when designing and interpreting data donation studies.
The proliferation of social bots poses significant challenges for authentic online communication, motivating a growing body of research in communication studies. Yet, conceptual ambiguities and methodological inconsistencies continue to undermine the reliability and replicability of social bot research. This study reviews recent work on social bots, revealing a frequent mismatch between detection methods and the types of bots being investigated. To address these issues, we propose a typology based on three dimensions: intention (benign vs. malicious), coordination (independent vs. coordinated), and operation (rule-based vs. generative). We further build a multiclass dataset of 4,071 bots and 6,386 humans from Bluesky and evaluate four major detection methods: rule-based approaches, supervised machine learning, unsupervised approaches, and large language model-based techniques. By highlighting the strengths and limitations of each, we advocate for multimethod strategies to better respond to evolving bot behaviors. We conclude with recommendations for standardizing detection practices and enhancing methodological rigor in social bot research.
Communication is commonly considered a process that is dynamically situated in a temporal context. However, there remains a disconnection between such theoretical dynamicality and the non-dynamical character of communication scholars' preferred methodologies. In this paper, we argue for a new research framework that uses computational approaches to leverage the fine-grained timestamps recorded in digital trace data. In particular, we propose to maintain the hyper-longitudinal information in the trace data and analyze time-evolving 'user-sequences,' which provide rich information about user activity with high temporal resolution. To illustrate our proposed framework, we present a case study that applied six approaches (e.g., sequence analysis, process mining, and language-based models) to real-world user-sequences containing 1,262,775 timestamped traces from 309 unique users, gathered via data donations. Overall, our study suggests a conceptual reorientation towards a better understanding of the temporal dimension in communication processes, resting on the exploding supply of digital trace data and the technical advances in analytical approaches.
Digital trace data are increasingly used across the social and behavioral sciences. They allow researchers to access large volumes of highly detailed and continuous information. Such scale and speed cannot be achieved when using traditional sources, such as surveys. Digital traces are also believed to overcome some of the limitations that surveys are criticized for. However, while their use undoubtedly presents researchers with new possibilities, it also introduces new quality challenges that have been increasingly acknowledged. Accounting for these limitations is crucial, as they can lead to biased results and incorrect research findings. Therefore, in this paper, we apply hidden Markov models (HMMs) to digital trace data on Facebook use to assess the nature and incidence of error in measures of Facebook use frequency. HMMs are an attractive method that allows for the estimation and correction of error without the availability of (error-free) gold-standard data, if the assumptions regarding the underlying construct of interest and the nature of the error are met. Our results suggest that the measures derived from digital trace data severely underestimate the frequency of Facebook use for a third of our sample, in particular when not all relevant devices are tracked.
Computational Social Science (CSS) increasingly engages in critical discussions about bias in and through computational methods. Two developments drive this shift: first, the recognition of bias as a societal problem, as flawed CSS methods in socio-technical systems can perpetuate structural inequalities; and second, the field's growing methodological resources, which create not only the opportunity but also the responsibility to confront bias. In this editorial to our Special Issue on CSS and bias, we introduce the contributions and outline a research agenda. In defining bias, we emphasize the importance of embracing epistemological pluralism while balancing the need for standardization with methodological diversity. Detecting bias requires stronger integration of bias detection into validation procedures and the establishment of shared metrics and thresholds across studies. Finally, addressing bias involves adapting established and emerging error-correction strategies from social science traditions to CSS, as well as leveraging bias as an analytical resource for revealing structural inequalities in society. Moving forward, progress in defining, detecting, and addressing bias will require both bottom-up engagement by researchers and top-down institutional support. This Special Issue positions bias as a central theme in CSS - one that the field now has both the tools and the obligation to address.
This study explores potential biases in current language models for detecting different forms of incivility in online discussions. Our work builds on prior research demonstrating that numerous classification approaches disproportionately focus on overt forms of uncivil language, while more subtle forms of incivility remain underrecognized. Such classification bias may lead to the unfair treatment of social groups that are more frequently targeted by implicit forms of incivility in online debates, such as stereotyping and discrimination, including, for instance, female politicians. On a dataset of 24,681 user comments from YouTube and X, we evaluate state-of-the-art language models including BERT, GPT-4, and Meta Llama 3, for their ability to reliably identify different subtypes of incivility directed at female and male politicians, combining comparative performance evaluation and in-depth error analysis. Our results suggest that stereotyping and discriminatory comments are less reliably classified across all models and learning strategies than, for example, vulgar language and insults, indicating a continued risk of bias toward overt forms of incivility in modern language models. As a result, incivility directed at social groups that are more frequently targeted with stereotyping and discrimination may remain underdetected. With this study, we aim to offer valuable implications for the use of current language models in both incivility research and online moderation, supporting the development of more transparent and fairer artificial intelligence grounded in democratic principles.
This paper discusses automated gender classification for social media profiles. It focuses on epistemological and methodological issues that researchers need to consider when using gender classification techniques. It begins with an examination of the category of gender from a queer feminist perspective in the first part. The second part discusses decision points and their potential impact on results at different stages of the research process: From identifying the research angle, to choosing a method, to analysis and validation, to reporting results. Several existing approaches are critically discussed. In the third part, an empirical case study is presented that follows a multi-step procedure of gender classification that aims to overcome a binary gender logic using different methods, such as dictionaries of gendered attributes and personal pronouns, as well as name-gender-inference. By reflecting on the challenges of each step of the analysis, the procedure follows an approach of discrimination-aware data analysis, which can be conceptualized as a form of intervention within the context of queer data practices.
Although it is widely acknowledged that political communication is multimodal and that textual and visual modalities have an interactive effect on how content is interpreted, research thereto is still limited. In particular, methods for creating multimodal datasets are rarely discussed; their validation is typically omitted. This study addresses this gap by (1) designing and validating a reliable method to create a multimodal dataset, and (2) use this dataset to study modality patterns in political news coverage. We demonstrate this by means of a case study examining news coverage by two online news outlets during the 2023 Dutch elections and check the robustness of our method by applying it to news coverage from two outlets in the United Kingdom during the 2024 elections. Besides creating a solid methodology, we were able to unravel interesting media patterns for the political actors in both countries. We found that the monomodal news coverage of these actors in either the textual or visual modality differed considerably from their multimodal news coverage, where they were visible in both. These results stress the need for a more inclusive approach to investigate multimodal political news coverage, as textual and visual information affect a person's interpretation of political news differently.
This article explores the (in)ability of automated tools to measure the deliberative quality of online user comments along the standards set out by Habermas: interactivity, diversity, rationality, and (in)civility. Utilizing a stratified sample of manually coded comments (n = 3,862) responding to news videos on YouTube and Twitter, we examined the performance of rule-based measures (i.e. dictionaries), machine-learning classifiers (conventional and transformer-based) and measurements by generative AI (Llama 3.1, GPT-4o, GPT-4T). We present results for over 50 metrics side-by-side to judge the opportunity costs of choosing one method over another. The results revealed strong variation across different groups of models. Overall, our expectation that more modern methods (transformers and generative AI) outperform the older, simpler ones was confirmed. However, the absolute differences between these model groups strongly depended on the measured concept, and we observed strong variance in performance among models of the same group. We provide recommendations for future research that balance ease of use with the performance of automated measurements, along with important cautions to consider.
What is the prevalence of moral themes in textual corpora? Most previous moral mining research has been informed by Moral Foundations Theory (MFT). Here, we develop and evaluate an alternative moral mining tool based on the theory of Morality as Cooperation (MAC), using crowd-sourced annotations and the web-based hybrid content annotation platform, the Moral Narrative Analyzer (MoNA). We compare the empirical performance of the extended Morality as Cooperation Dictionary (eMACD) with previous dictionaries, including the extended Moral Foundations Dictionary (eMFD), across ten validation analyses, encompassing diverse corpora that include news media, journal entries, presidential speeches, social media, and movies. We find that eMACD outperforms previous dictionaries in most cases. We conclude that eMACD is an important addition to the methodological toolkit of researchers examining moral content in textual data, and we provide eMACDscore - a Python-based moral mining tool - to facilitate future research.
Understanding visual narratives is crucial for examining the evolving dynamics of media representation. This study introduces VisTopics, a computational framework designed to analyze large-scale visual datasets through an end-to-end pipeline encompassing frame extraction, deduplication, and semantic clustering. Applying VisTopics to a dataset of 452 NBC News videos resulted in reducing 11,070 frames to 6,928 deduplicated frames, which were then semantically analyzed to uncover 35 topics ranging from political events to environmental crises. By integrating Latent Dirichlet Allocation with caption-based semantic analysis, VisTopics demonstrates its potential to unravel patterns in visual framing across diverse contexts. This approach enables longitudinal studies and cross-platform comparisons, shedding light on the intersection of media, technology, and public discourse. The study validates the method's reliability through human coding accuracy metrics and emphasizes its scalability for communication research. By bridging the gap between visual representation and semantic meaning, VisTopics provides a transformative tool for advancing the methodological toolkit in computational media studies. Future research may leverage VisTopics for comparative analyses across media outlets or geographic regions, offering insights into the shifting landscapes of media narratives and their societal implications.
Information shapes citizens' political decision-making. This process is amply studied by social scientists, who have human annotation as a crucial instrument in their toolkit. To address concerns regarding the validity and reliability of annotation tasks, we establish strict standards. Due to the democratization of data and the advances in NLP more data can be analyzed or classified, making these standards are more important than ever: An algorithm trained on biased data will reproduce and often exacerbate bias. Our tools to create valid and reliable annotation data currently do not allow for dealing with different perspectives of annotators. In two pre-registered experiments in the United States and the Netherlands, we show that personal characteristics of annotators, like political ideology or knowledge, interfere with annotators' judgement of political stances. Our results show that to improve annotated data for automated text analyses, and for stance detection models in particular, we need to critically evaluate how we create our gold standards.
As the exploration of digital behavioral data revolutionizes communication research, understanding the nuances of data collection methodologies becomes increasingly pertinent. This study focuses on one prominent data collection approach, web scraping;specifically, its application in the growing field of research relying on web browsing data. We investigate discrepancies between content obtained directly during user interaction with a website (in-situ) and content scraped using the URLs of participants' logged visits (ex-situ) with various time delays (0, 30, 60, and 90 days). We find substantial disparities between the methodologies, uncovering that errors are not uniformly distributed across news categories regardless of the classification method (domain, URL, or content analysis). These biases compromise the precision of measurements used in the existing literature. The ex-situ collection environment is the primary source of the discrepancies (33.8%), while the time delays in the scraping process play a smaller role (adding similar to 6.5% points in 90 days). Our research emphasizes the need for data collection methods that capture web content directly in the user's environment. However, acknowledging its complexities, we further explore strategies to mitigate biases in web-scraped browsing histories, offering recommendations for researchers who rely on this method and laying the groundwork for developing error-correction frameworks.
Measurements of political polarization online have so far been largely focused on visible traces accessible through platformAPIs,while neglecting invisible traces not recorded or otherwise unavailable via online channels, which can reveal key aspects of political engagement online. Our study aims to address this gap by investigating the polarization measurement bias that arises when only visible engagement is considered. Using a combined dataset that links survey responses with YouTube digital traces froma representative sample of Hungarian Internet users (N=758), we uncover disparities at both user and channel levels.We find that users who visibly engage through commenting are more politically polarized. People tend to comment and subscribe to ideologically concentrated content, while viewingideologically diverse content. We also notice that ideologically heterogeneous channels are more likely to share viewers than subscribers or commenters. Thus, the segregation ofpolitical channels may be overstated or simplified when relying solely on public comment data. Our results suggest that research using only visible engagement may overestimate the extent of polarization and the prevalence of echo chambers online. We highlight the benefits of data donation to address measurement bias in online political communication, and contribute to the polarization literature by providing a fresh evaluation of potential biases in platform-focused research.
The goal of the current study is to compare different methods for automated object detection (i.e. tag detection, shape detection, matching, and machine learning) with manual coding on different types of objects (i.e. static, dynamic, and dynamic with human interaction) and describe the advantages and limitations of each method. We tested the methods in an experiment that utilizes mobile eye tracking because of the importance of attention in communication science and the challenges posed by this type of data when analyzing different objects because visual parameters are consistently changing within and between participants. Machine learning was found to be the most reliable method to detect all types of objects and was slightly more conservative compared to manual coding. Feature-based matching worked well for static objects. We discuss the advantages and challenges of each method along with key considerations for researchers depending on their research objective, the type of object, and the object detection method they will use.