
Organizations are adopting generative AI faster than the evidence base needed to govern it. Existing evaluation tools such as benchmarks, alignment scores, and safety tests were built for model development, not for judging whether systems will create value, introduce friction, or shift risk in specific real-world settings. As a result, there is little systematic evidence about how AI behaves once it is embedded in everyday work. This paper proposes a real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time. Instead of treating variability across users, tasks, and settings as noise to be controlled away, the framework treats that variation as the central source of deployment-relevant evidence. It sets out four design principles for producing decision-ready evidence at scale and proposes a shared evaluation architecture combining a structured observation environment, a metrics hub, and reusable consortium models that summarize system behavior across contexts. Rather than replacing traditional benchmarks, this framework adds a sociotechnical evidence layer that connects model capabilities to the organizational and practitioner level outcomes where deployment decisions are actually made.
The increasing adoption of Large Language Models (LLMs) as a text analysis method in social science presents a critical yet under-examined trade-off between model performance and environmental sustainability. This research provides a systematic evaluation comparing the performance, energy consumption, processing time, and CO 2 emissions of various computational text analysis methods (CTAM), including dictionaries, trained classifiers, and self-hosted open LLMs when performing sentiment analysis of parliamentary speeches, classification of open-ended survey responses, and named entity recognition of newspapers. The analysis is limited to self-hosted deployment in local and server environments where per-task energy consumption is directly measurable. Although self-hosted LLMs demonstrate strong performance in sentiment analysis, closely aligning with human judgment, they require significantly more energy and time than non-LLM approaches. For classification and named entity recognition, pretrained task-specific models achieve better F1 scores with a lower carbon footprint, challenging the primacy of larger models. To navigate this trade-off, we propose a CO 2 -Adjusted F1 Score that penalizes emissions while rewarding performance. Applying this metric, we show that smaller, task-specific models may be preferred over larger general-purpose LLMs for efficient text analysis. We highlight the necessity for thoughtful and responsible model selection, promoting a “right-fit” approach for CTAM.
To collect digital trace data, researchers continue to rely on for-profit companies’ Application Programming Interfaces (APIs). These APIs often return samples of the data based on intransparent sampling procedures and algorithms. In this paper, we extend research on the reliability of digital trace data from APIs by examining the effect of the download location on inconsistencies across returned samples: Do we get different values from digital trace data APIs depending on where we download the data from? We compare samples from Google Trends, YouTube Data, and the New York Times (NYT) API from four countries across three continents (Austria, Germany, the U.S., and Australia) for the same query parameters (i.e., search term, region, and time range). Our results show that the download location impacts the returned samples for all three APIs, depending on the query. We find large inconsistencies for samples from Google Trends and the YouTube Data API, while the NYT API returns identical article sets from each download location for most queries. We conclude with practical recommendations for researchers using these APIs. Our findings serve as a cautionary reminder for social scientists relying on sampling-based APIs as they point to yet another limitation regarding their reliability and reproducibility.
This study examines how the #MeToo movement reshaped long-term public discourse on sexual violence in South Korea. While research has documented the surge of attention following prosecutor Seo Ji-hyun’s 2018 disclosure, little is known about how the movement reorganized the broader discursive field beyond the initial moment of crisis. Drawing on discursive institutionalism, we argue that #MeToo acted as a discursive shock that consolidated previously fragmented conversations into more coherent and durable interpretive frameworks. Using an 8-year corpus of 351,582 Korean-language tweets (2015–2023), we apply Dirichlet Multinomial Regression topic modeling and time-series analysis to identify shifts in both topical content and discursive structure. Our findings show a transition from episodic, emotion-driven narratives to stable, thematic frames emphasizing human rights, systemic discrimination, and institutional accountability. We further demonstrate increasing convergence between social media and mainstream news, indicating diffusion and stabilization of these new frames. The results suggest that digital activism can produce enduring transformations in public meaning, illuminating how social movements institutionalize discourse over time.
Scholars have studied the relationship between incidental exposure (i.e., encountering political information unintentionally) and political participation. Nevertheless, the dynamic processes relating incidental exposure to political participation remain unclear. Building upon the political incidental news exposure (PINE) and social media political participation (SMPP) models, this study ( N = 702) examines first- and second-level incidental exposure and its relationship to low-effort online and high-effort political participation. Furthermore, it examines whether perceived satisfaction from low-effort online participation moderates the relationship between incidental exposure and high-effort participation. Results suggest that second-level incidental exposure relates to greater low-effort online participation, which, in turn, opens a pathway to high-effort participation, especially when individuals perceive such participation as satisfying and able to address social and political issues. These findings provide a nuanced understanding of the pathways from incidental exposure to political participation, contributing to the broader literature on political communication and behavior in the digital age.
With the exponential growth in social media usage, the rapid spread of misinformation has become a critical global challenge. Recent advances in large language models (LLMs) have shown promising potential in automated misinformation detection. This survey provides a comprehensive review of LLM-based approaches for detecting misinformation in textual data on social media platforms. In this work, we analyze 70+ recent papers; to examine the evolution, implementation, and effectiveness of various LLM architectures in this domain. Our analysis reveals that BERT-based models dominate the field, appearing in approximately 85% of studies, with domain-specific variants like CT-BERT demonstrating superior performance in specialized contexts such as COVID-19 misinformation detection. We provide detailed comparisons of model architectures, implementation strategies, and performance metrics across different domains. Additionally, seven major datasets commonly used in this field were analyzed, examining their characteristics, limitations, and suitability for different detection tasks. The survey also addresses key challenges, including linguistic nuances, model interpretability, and ethical considerations. Our findings indicate that while LLM-based approaches achieve impressive accuracy metrics, significant challenges remain in cross-domain generalization and real-time detection. This survey concludes by identifying promising research directions and providing recommendations for robust model evaluation frameworks.
While research has shown how different political actors adapt engagement-triggering strategies in organic election campaigns on social media platforms, we know little about how such strategies are used in paid political advertising and how they affect algorithmic ad performance. To address this gap, this study analyzes how party attributes shape spending on, impressions of, and the cost efficiency of negative campaigning in digital political advertising. Manually content-analyzing all Facebook ads from 48 parties across 10 countries during the European Parliament election in 2019, we find that, on average, parties do not allocate more resources to negative ads, which also do not achieve higher impression counts and are even associated with higher costs per 1,000 impressions (cost efficiency). However, party characteristics matter: Whereas opposition parties invest more in negative campaigning and consequently gain more impressions than government parties, negativity only pays off in terms of cost efficiency for extreme parties. Our findings contribute to a better understanding of the interplay between advertiser characteristics and content strategies in shaping ad use and performance within commercial marketing infrastructure, emphasizing the need for greater transparency to ensure fair competition in digital political advertising.
This study investigates the dynamics of anti-Muslim hate speech within Norwegian social media during the period between 2010 and 2021. Using a dataset of more than one million comments from Twitter and Facebook, we developed a custom hate speech classifier trained on an annotated corpus of 3,277 comments in Norwegian language. We identify that despite representing a small share of the total comments, hate speech content has increased over time. In an effort to understand the social network characteristics of hate speech content, we delve deeper into Twitter conversations as we can more easily identify how this content is spread. We develop network metrics to assess the prevalence, distribution, and diffusion of hateful content. The findings reveal that regardless of the number of users or tweets in a conversation, the volume of hateful content tends to remain constant. Furthermore, a small fraction of users contribute disproportionately to the dissemination of hate speech, with most conversations being limited in participant diversity. These results contribute to the growing field of computational social science by offering a novel methodology for studying hate speech in under-resourced languages and suggesting that mitigating hate speech may be possible through targeted network interventions rather than content removal alone.
This study examines the impact of performatives and evolving social media typology in shaping political activism among Kenya’s Generation Z (Gen Z) movement during the 2024 anti-tax law protests. The study addresses the questions of the role of performatives and how social media has revolutionised their production, reproduction, and consumption in political activism in Kenya. Based on qualitative content analysis and critical discourse analysis, the study employed purposive sampling of a collection of digital artefacts, including memes, protest songs, TikTok videos, graffiti-inspired art, and Twitter threads, which were drawn from the #RejectFinanceBill2024 campaign. Analytical categories were derived from literature on performative activism, postcolonial media theory, and digital political communication. The findings suggest that Kenya’s Gen Z activists adopted a highly performative mode of social media resistance, blending entertainment with activism. The content of performatives was found to function not only as expressive tools but also as mechanisms for mobilising support, challenging state narratives, and asserting digital visibility. Social media was found to circumvent traditional media gatekeeping, amplifying the voices of the marginalised, and fostering an enlightened political culture. The study identifies a cyclic loop of production and reproduction of performatives, reinforcing African people’s communal identity formation and resistance posturing. Findings highlight how Gen Z’s social media use is reshaping civic engagement in the postcolonial public sphere. The study advances theoretical understanding of how visual and performative content is democratising political discourse, disrupting power hierarchies, and deepening participatory governance in the Global South. This study contributes to the body of literature on digital media and political communication by illuminating the intersection of social movement, culture, aesthetics, and performativities in resistance. These insights are particularly relevant for scholars and practitioners interested in digital media use, activism, political communication, and youth-led social movements.
This article examines how citizens’ perceptions of public employees in the education system change as these employees adopt artificial intelligence (AI) in their work. It argues that the use of AI matters for citizens’ perceptions of decision-making not only due to AI system characteristics, but also because AI adoption alters the perceived relationship between citizens and public organizations. Rooted in assumptions of social cognition theory, the analysis tests how information about AI use by public employees alters the perceived warmth of these employees and thereby affects the acceptability of decision-making. The analysis is based on a pre-registered vignette experiment and a sample of 4,569 participants from Norway. It finds that AI use decreases both the perceived warmth and competence of public employees, that these evaluations negatively bear on the overall acceptability of decision-making, and that the effect of AI use is stronger for public employees more directly interacting with others.
Bias in word embeddings is often measured using bipolar dimensions, constructed as the difference between two anchor centroids. This technique assumes both poles are symmetrical and equally informative. However, normativity literature shows that one category may function as the unmarked norm, with others framed as marked deviations. In race, whiteness typically holds the normative position, and embedding-based race dimensions may inherit the skew. We test this possibility using dimensions constructed from validated African-European name anchors, probed with neutral and valence words. In three embedding models (Wiki-News, South African news, Google News), we assess whether race dimensions favour whiteness as a normative anchor, whether this skew is stronger in culturally specific models (SA, Google), and whether bipolar offsets amplify one pole, given unipolar evidence. Results show that neutral and valence terms cluster nearer to the white pole (most strongly in the Wiki-News model), indicating whiteness as the semantic default. Overshoot favoured Black in Google and Wiki-News, while White overshoot only occurred in the South African model. We argue that this captures racialised variance where the pole with more spread tends to exert greater leverage on the bipolar axis. The study provides quantitative evidence of white-normative anchoring and diagnostics for asymmetric amplification in embedding-based bias measures.
Computational Public Opinion Measurement (CPOM) uses natural language processing and machine learning to infer public attitudes from social media. However, the methodological foundations and validity constraints of CPOM remain inadequately documented. This systematic literature review examines the computational methods, sentiment representations, and measurement strategies employed in CPOM research, and identifies the structural and methodological limitations that constrain CPOM as an approach to measurement. This systematic review (PRISMA, 2020) searched seven databases, identifying 56 studies from 5,108 records (2008-2025). Methods shifted substantially across eras: lexicon-based approaches declined from 77.8% (2011-2015) to 28.6% (2021-2025) while deep learning grew from 0% to 40.0%. Despite this technical evolution, empirical validation remains rare. Only 11 of 56 studies (19.6%) computed a quantitative statistic against an external benchmark such as a survey or poll. In other words, methodological sophistication and validation rigour moved in opposite directions: none of the 14 deep learning studies computed a quantitative validation statistic, while all 11 that did used lexicon-based or classical machine learning. The broader field shows an even larger gap: 49.4% of all full-text-assessed papers were excluded for providing no external validation or representativeness analysis at all. The recognition-action gap widened over time: TSE awareness grew from 44.4% to 80.0% across eras while quantitative validation rates fell from 33% to 11%, showing that growing awareness has not translated into action. CPOM must move beyond technical sophistication toward systematic criterion validation, demographic adjustment, and transparent reporting.
Early modern writers had much to say about how humans and other animals compared to machines. Ren & eacute; Descartes (1596-1650) argued that an automaton could never convincingly impersonate human intelligence, however lifelike it seemed. To set the stage, I will briefly discuss Descartes' imitation game, along with its philosophical and theological motivations. I will then move on to another imitation game conceived of by Nicolas Filleau de la Chaise (1631-1688), who asked whether a human could successfully imitate a machine. Finally, I will reflect on how Thomas Hobbes (1588-1679) portrayed the mechanical intelligence of the state as a safeguard against the deceptions of its enemies.
Topic discovery and integration are vital for maintaining vocabularies that categorize textual corpora. Automated approaches are often computationally expensive and lack domain-specific conceptual nuance; manual approaches are costly in terms of time and potential bias. To address this dilemma, we introduce the segments-as-topic (SAT) methodology, a four-stage process that combines automation and human expertise to assess candidate topics for vocabulary inclusion. In the SAT generation stage, a topic is formulated and refined through collaboration with domain experts, and then a sentence-level semantic similarity model retrieves corpus segments semantically aligned with the topic. The SAT expansion stage uses this seed set to find additional semantically similar segments, which are iteratively accepted or rejected to build a final segment set. During the review stage, a panel of scholars evaluates the topic for inclusion. In the integration stage, all segments in the final segment set are automatically tagged with the new topic. We apply this methodology to the Comparative Constitutions Project vocabulary that tracks over 330 topics in national constitutions, and demonstrate the addition of three new topics to the vocabulary. The SAT approach balances computational efficiency with expert judgment, offering a systematic, user-friendly, and replicable framework for social scientists to expand domain-specific vocabularies.
This essay conceives artificial intelligence as a chapter in the history of writing through reconsidering the eighteenth-century automaton writer created by Jaquet Droz, a Swiss clock making workshop. As much an innovation in writing technology as it was an early example of artificial intelligence, the Jaquet Droz automaton writer reveals how artificial intelligence is a historical idea and material artifact deeply entangled with the history of writing, an embodied as well as deeply emotional form of cognitive activity and one of the oldest human technologies.
The ecological inference canonical problem in the social sciences consists of estimating the unobserved internal counts of a global RxC table from the known margins of a set of units. This paper proposes a new, computation-based strategy designed to better exploit the information contained in the unit margins. This approach can be integrated into any ecological inference method that explicitly estimates unit tables and accounts for differences in unit size. We evaluate its performance using as a baseline the fastest ecological inference linear programming method, relying on real electoral data from over 550 datasets where true contingency tables are known. In this extensive assessment, the proposed strategy reduces average global errors by more than 21% relative to the baseline, outperforming it in 95% of cases. It also improves upon nslphom-identified in the literature as the most accurate algorithm for this dataset-reducing average global errors by over 5% and outperforming it in 60% of cases. The versatility of the approach is further illustrated by also integrating it into three more computationally intensive methods, including the two main statistical ecological inference models-ei.MD.bayes and BPF-and nslphom, yielding consistent improvements over their respective baselines in a small set of examples.
Drawing on information ecology theory, this study investigates how online toxicity spreads and drives emotional group polarization in Reddit discussions surrounding the Israel-Palestine conflict. Based on a dataset comprising 8,725 posts and 1,628,366 comments, we employ Google's Perspective API to detect toxic content, BERTopic to extract discussion topics, and a dictionary-based method to measure affective features. The results show that, from the information content perspective, content related to conflict and political topics is more prone to generating toxicity. From the information user perspective, post toxicity reduces the scale of user interaction while simultaneously increasing the level of comment toxicity. Furthermore, the analysis shows that post toxicity provokes comment toxicity by triggering users' negative or high-arousal emotions, with discrete negative emotions such as anger, sadness, and fear enhancing the positive impact of post toxicity on comment toxicity. From the information environment perspective, the results indicate that post toxicity is not significantly associated with emotional group polarization. Accordingly, post toxicity alone plays a very limited role in explaining emotional group polarization in online discussions. These findings advance our understanding of toxicity dynamics in online environments and offer evidence-based strategies for moderators, platform designers, and policymakers to mitigate harmful discourse and foster healthier online communities.
The use of algorithmic hiring is on the rise, becoming more and more common due to its capacity to process a large number of applications with efficiency. Despite the advantages of this approach, there have been numerous cases of discriminatory hiring outcomes. A significant contributing factor to such outcomes is the lack of representativeness in the data used to develop these systems. This results in a significant decline in performance for underrepresented groups, disproportionately impacting marginalized communities. Addressing bias in algorithmic hiring requires access to comprehensive datasets that include curriculum vitae (CV) and demographic information reflecting diverse backgrounds. Unfortunately, there is a lack of datasets that serve such a purpose. This paper introduces a data donation campaign designed to collect real-world CVs, including demographic and sensitive information, from a representative sample of individuals in the EU workforce. The paper discusses the design decisions underpinning the campaign, along with the challenges encountered during its deployment and execution. Finally, it offers lessons learned and practical solutions to overcome these challenges, thereby contributing valuable insights for future efforts in this domain.
This essay argues that the black box-both as cryptic device and as critique of illegibility-is not unique to modern technology and has deep roots in the medieval Mediterranean world. Technical opacity was frequently addressed in Latin and Arabic sources, often with a critical undertone. Then as now, technoskeptical writers saw the self-acting device as treacherous, due to its reliance on hidden labor and mechanisms. This critique arose especially in relation to unfamiliar or foreign devices, like animated idols; as such, it was often racializing, attributing opacity as well as deceit to the object and its makers. Modern critiques of technology that focus on invisible labor may reproduce similar biases by enforcing a privileged, first-world perspective. A transhistorical approach thus not only shows the enduring history of the black box; it also illuminates the religious genealogy of techno-skepticism, as well as the biases that inhere in the black box, especially when deployed as a critical discourse.
Centuries before the advent of computers, the German philosopher and mathematician Gottfried Wilhelm von Leibniz (1646-1716) sketched out a "computational ontology" whereby information operates as an organic principle imposing order, molding and driving it, in such a way that the world gains a form of consciousness, and thought produces its being at the same time as it thinks itself.