
The 2022 FIFA World Cup in Qatar was as much a political event as a sporting one, marked by controversies over the death of construction workers, concerns over LGBTQ+ rights, and allegations of corruption. Drawing on 573,927 tweets from UK-based accounts and 15,812 excerpts from 66 UK newspapers, we examine how political and apolitical attention to the competition was distributed across mainstream and social media, how it varied across various Twitter sub-networks, and how UK journalists navigated both arenas. Combining Topic Modelling, LLM-based classification, Social Network Analysis, and close reading, we find that political content accounted for 29.0% of tweets and 18.5% of news excerpts—a gap driven largely by Twitter discussion of issues only loosely tied to the tournament, particularly the Mahsa Amini protests in Iran. Consistent with Wright et al. (2015), we find that political talk routinely surfaced within seemingly apolitical conversations about teams, players, and matches—from debates over footballers’ working conditions to readings of Morocco’s run as a symbolic challenge to European footballing power. At the same time, political conversations on Twitter were quite segmented, with different sub-networks focusing on different political topics. Moreover, while the UK press emphasized LGBTQ+ rights and migrant worker deaths, Twitter hosted more sustained debate about perceived bias in Western media coverage of Qatar, especially within a community of users with ties to Arab countries and Islam. These insights offer unusually direct empirical support for counterpublic theory. Journalists, meanwhile, mostly used Twitter to push content but also engaged in disputes with one another over bias, accuracy, and the propriety of attending the event, reflecting a shift toward more assertive personal positioning consistent with broader trends in journalistic branding. We close by arguing that the framework of online “third spaces”—apolitical online fan environments into which politics occasionally intrudes—fits the World Cup poorly: the tournament’s scale, commercialization, and political salience make sustained apoliticism untenable, and a meta-debate over whether to “keep politics out of football” played out openly across both platforms. The concept may apply more productively to smaller, more bounded fan communities.
In a rapidly evolving digital media landscape, understanding how exposure to untrustworthy news changes over time is essential for evaluating its potential effects on public attitudes and behavior. However, there is limited evidence on how demand for untrustworthy news develops across longer time frames. Linking web data with surveys, we compare exposure to untrustworthy news sources across demographic and political groups and over a 7-year time frame for two samples of German adults (N = 1,212 in 2017 and N = 436 in 2024). Visits to untrustworthy news sources make up less than 1% of media diets and are associated with low satisfaction with democracy and a preference for a far-right party. Propensity score matching reveals stability in untrustworthy news exposure and a subtle decline in the average quality of news diets over 7 years. Our findings suggest that concerns about rising misinformation exposure may not be borne out in desktop browsing behavior, though the absence of in-app tracking data means that mobile and social media exposure remains a question to be explored further.
Computational social science has expanded the capacity of scientists to study connected human behavior at previously unprecedented scales. Yet from its beginning, scientists expressed concern that its reliance on private companies might produce a body of work that cannot be critiqued or replicated. Such commercial determinants of science have been observed in other fields including public health where science has implications for corporate liability. In this meta-scientific report, we analyze the population of computational social science articles about technology platforms published in three general scientific journals to investigate commercial determinants of scientific replicability in computational social science. We find that only 26% of those papers can be replicated today, and that 34% of computational social science studies published in leading general scientific journals rely on special arrangements with corporations that are impossible to replicate without special permission. We find that articles relying on API access, scraping or access to nonprofit platforms have much higher potential for replication. These findings are consistent with broader literature on the commercial forces that determine the direction and reliability of science.
Election betting platforms, which combine real money trading with public political discussion and social feedback, have seen explosive growth worldwide. Prediction market research emphasizes aggregate accuracy and price dynamics, while computational political communication research examines discourse in digital forums; however, these literatures remain analytically separate. Here, we provide a quantitative description of how financial, discursive, and social behaviors are jointly observable on a single platform. We analyze 778,634 accounts (unique wallet addresses) over 846 days, documenting 30,502,864 trades totaling $5.86 billion, 384,586 comments, and 571,523 reactions across three behavioral modalities (trading, commenting, reacting) on the Polymarket platform, spanning 959 election-related events. We find that within events, trading and commenting are largely decoupled and temporally ordered: among wallets exhibiting both behaviors, trading and commenting co-occur in only 21.8% of wallet-event combinations, are only moderately correlated within events, and 90.2% of joint pairs are trade-first. Participation is highly concentrated across all behaviors, with concentration varying systematically by behavior type: financial trading (Gini = 0.951; top 1% account for 73.7% of volume) concentrates most sharply, followed by social reactions (Gini = 0.845), then political discourse (Gini = 0.814; top 10% produce 76.4% of comments). This cross-behavior ordering persists when the analysis is restricted to higher minimum-activity thresholds. We also document the volume of single-action wallets: 12.6% of trading accounts execute exactly one trade, 41.1% of commenters post exactly one comment, and 34.3% of reactors leave exactly one reaction. While single-mode accounts (engaging in only one modality) dominate numerically (95.9% of accounts, predominantly Trader Only at 94.4%), a small multi-modal minority (engaging in two or more modalities; 4.1%) contributes disproportionately across financial, discursive, and social behaviors. These findings characterize observable wallet-level political activity when financial commitment is visible alongside discursive acts, informing comparative and explanatory research on hybrid platforms. More broadly, quantifying who participates, how activity concentrates, and when engagement occurs provides a foundation for studying participation across behavioral domains that have typically been analyzed separately.
We investigate the issue of stance representation at the platform level, and differences in experience at the individual level, in the context of two controversial, mostly two-sided issues: the legality of abortion in the U.S., and the Israel/Hamas conflict; we do so on TikTok. In manually annotated representative samples containing over 3.8k videos and spanning a combined 37 weeks, we measure the number of videos available on the platform and the views these videos receive, separated by the stance they represent (Pro Choice vs. Pro Life, and Pro Israel vs. Pro Palestine). We complement these platform-level analyses with a contemporaneous estimate of TikTok user opinions and a survey of users' content recommendation experiences. Around 34.5% of TikTok users identified as Pro Life and 53.7% as Pro Israel in Pew surveys fielded close to our topic time frames. However, at the platform level, we find that videos labeled as either Pro Life or Pro Israel are markedly underrepresented both in terms of video counts and viewership; we found only 3.2% of abortion-related videos (3.0% of views) represented a Pro Life stance, and 8.4% of Israel/Hamas-related videos (12.0% of views) represented a Pro Israel stance, respectively. Pro Choice and Pro Palestine respondents tended to be shown mostly videos in line with their opinion, users aligned with the opposing opinions were more likely to report being shown mostly videos they disagreed with, or seeing less on-topic content overall.
We study petition-related mobilization on Social Media across issues and ideologies. Using calls to sign petitions on X in the seven most spoken EU languages, we build the first multi-platform, multi-language map of the e-petitioning ecosystem. To ensure cross-language and cross-ideology comparability, we infer call for signatures’ issues using ManifestoBERTa and users' ideological orientation via Ideology Scaling methods calibrated using expert survey data. We classify active individuals into an ontology of mobilization types. Results show that e-petition activism is issue-specific and short-lived, with only a small portion of individuals engaging in sustained mobilization. We characterize differences in e-activism across the Left-Right spectrum, with Right-leaning users being most active on issues like political corruption and traditional morality, while environmental protection sees low engagement levels compared to its reach.
Using nationally representative data from the 2020 and 2024 American National Election Studies (ANES), this paper traces how the U.S. social media landscape has shifted across platforms, demographics, and politics. Overall platform use has declined, with the youngest and oldest Americans increasingly abstaining from social media altogether. Facebook, YouTube, and Twitter/X have lost ground, while TikTok and Reddit have grown modestly, reflecting a more fragmented digital public sphere. Platform audiences have aged and become slightly more educated and diverse. Politically, most platforms have moved toward Republican users while remaining, on balance, Democratic-leaning. Twitter/X has experienced the sharpest shift: posting has flipped nearly 50 percentage points from Democrats to Republicans. Across platforms, political posting remains tightly linked to affective polarization, as the most partisan users are also the most active. As casual users disengage and polarized partisans remain vocal, the online public sphere grows smaller, sharper, and more ideologically extreme.
Community Notes (formerly known as Birdwatch) is the first large-scale crowdsourced content moderation initiative launched by X (formerly Twitter) in January 2021. As the Community Notes model gains momentum across other social media platforms, there is a growing need to assess its underlying dynamics and effectiveness. This paper provides a descriptive investigation of Community Notes during its first four years, examining its linguistic diversity, sourcing practices, Contributor activity, rating behaviour, and interaction networks. In addition, we release a curated dataset and accompanying source code to support future research, along with a review of prior research on Community Notes. We parsed Notes and ratings data from the first four years of the program and conducted language detection across all Notes. For English-language Notes, we extracted embedded URLs and identified discussion topics in each Note. Additionally, we constructed monthly interaction networks among the Contributors. Together, the descriptive analysis, dataset, code, and literature review provide a foundation for advancing research on Community Notes and community-based content moderation more broadly.
City council meetings are vital sites for civic participation where the public can speak directly to their local government. By addressing city officials and calling on them to take action, public commenters can potentially influence policy decisions spanning a broad range of concerns, from housing, to sustainability, to social justice. Yet studies of these meetings have often been limited by the availability of large-scale, geographically-diverse data. Relying on local governments’ increasing use of YouTube and other technologies to archive their public meetings, we propose a framework that characterizes comments along two dimensions: local concerns (e.g., housing, election administration), and societal concerns (e.g., functional democracy, anti-racism). Based on a large record of public comments we collect from 15 cities in Michigan, we produce data-driven taxonomies of the local concerns and societal concerns that these comments cover, and employ machine learning methods to scalably apply our taxonomies across the entire dataset. We then demonstrate how our framework allows us to examine the salient local concerns and societal concerns that arise in our data, as well as how these aspects interact.
Conspiracy theories have long drawn public attention, but their explosive growth on platforms like Telegram during the COVID-19 pandemic raises pressing questions about their impact on societal trust, democracy, and public health. We provide a geographical, temporal and network analysis of the structure of of conspiracy-related German-language Telegram chats in a novel large-scale data set. We examine how information flows between regional user groups and influential broadcasting channels, revealing the interplay between decentralized discussions and content spread driven by a small number of key actors. Our findings reveal that conspiracy-related activity spikes during major COVID-19-related events, correlating with societal stressors and mirroring prior research on how crises amplify conspiratorial beliefs. By analysing the interplay between regional, national and transnational chats, we uncover how information flows from larger national or transnational discourse to localised, community-driven discussions. Furthermore, we find that the top 10
TikTok is now a massive platform, and has a deep impact on global events. Despite preliminary studies, issues remain in determining fundamental characteristics of the platform. We develop a method to extract a representative sample of >99% of posts from a given time range on TikTok, and use it to collect all posts from a full hour on the platform, alongside all posts from a single minute from each hour of a day. Through this, we obtain post metadata, video media, and comments from a close-to-complete slice of TikTok, and report the critical statistics of the platform. Notably, we estimate a total of 269 million posts produced on the day we looked at, that 18% of videos on the platform feature children, and that at least 0.5% of posts contain artificial intelligence-generated content.
How does Google Search direct people to information about their elected offcials? To answer this, we conducted daily searches for members of the US House of Representatives from all 435 US congressional districts and DC between September 1 and December 31, 2020, resulting in 20.1 million search engine results pages (SERPs) and 302 million search results. We find that these search results are dominated by a small number of mainstream sources (eg. Twitter, Wikipedia), with the top seven domains accounting for 64.2% of all results. There was no significant difference in the partisanship of search results depending on whether the member whose name was searched was a Democrat or Republican. Additionally, we found a clear prioritization of politician-controlled social media, government, and personal websites over news media, local news outlets over national ones, and reliable news over unreliable news. We observed a lack of sensitivity to search location, where searching for a given member’s name on the same day but from different locations yielded similar results.
A key focus in the study of digital and social media in politics has been to investigate the homophily of online interactions and content exposure, often framed as the extent to which these platforms operate as echo chambers. However, research in this area has yielded mixed findings. One possible explanation is that the degree of homophily in online behavior varies depending on the specific type of behavior examined. This paper contributes to this debate by measuring multiple indicators of political homophily within the same sample. We identified a near-census list of ∼44,000 politically active Twitter (now X) users in Uruguay and retrieved all their interactions—including retweets, replies, quotes, and likes—over a defined period between 2021 and 2022, along with the accounts they follow. This dataset, comprising 7,172,636 tweets, 9,711,053 likes, and 13,660,197 following links allows us to assess the extent of political homophily across four dimensions: following behavior, interactions among ordinary users, interactions with elites, and media consumption. Our study makes three key contributions. First, it fills the gap of a single study that estimates different outcomes of political homophily in online communication using the same sample. Second, it focuses on politically active users who play a central role in shaping online political discourse. Third, it extends research beyond the typical U.S. and Western European cases by examining Uruguay, a country with a stable party system, strong partisan attachments, and low affective polarization. Our findings show that while politically active users are not completely isolated within echo chambers they exhibit a strong tendency toward homophily while confirming expected variation across dimensions noted in previous research. This pattern of high homophily persists even in a setting of low affective polarization. These results highlight the nuanced ways political homophily manifests across different behaviors and settings.
This article explores the role of unrecognised labour in corporate innovation systems via an analysis of researcher coding and discursive contributions to R, one of the largest statistical software ecosystems. Studies of online platforms typically focus on how platform affordances constrain participants’ actions, and profit from their labour. We innovate by connecting the labour performed inside digital platforms to the professional employment of participants. Our case study analyses 8,924 R package repositories on GitHub, examining commits and communications. Our quantitative findings show that researchers, alongside non-affiliated contributors, are the most frequent owners of R package repositories and their most active contributors. Researchers are more likely to hold official roles compared to the average, and to engage in collaborative problem-solving and support work during package development. This means there is, underneath the ‘recognised’ category of star researchers who transition between academia and industry and secure generous funding, an ‘unrecognised’ category of researchers who not only create and maintain key statistical infrastructure, but also provide support to industry employees, for no remuneration. Our qualitative findings show how this unrecognised labour affects practitioners. Finally, our analysis of the ideology and practice of free, libre and open source software (FLOSS) shows how this ideology and practice legitimate the use of ‘university rents’ by Big Tech. In conclusion, we argue that existing mechanisms are insufficient to ensure these digital commons’ sustainability: FLOSS needs broader systemic support.
Exposure to wildfire smoke has serious health implications, highlighted by public and media attention each fire season. This study combines datasets from newspaper archives, social media, and a national survey to assess how hazards, impacts, and protective actions are discussed across different types of media. We found protective actions are underdiscussed in traditional and social media, especially in information-seeking contexts. The survey indicated the public would benefit from more wildfire smoke information. Results show how media and public data can help describe the risk communication environment surrounding hazards like wildfire smoke.
Election campaigns increasingly pursue their strategic goals online, using digital advertising not only to persuade voters but also to raise funds, mobilize supporters, and collect data. These varied objectives reflect how campaigns operate within broader networks of candidates, parties, and outside groups. The extended party network forms the strategic backbone of modern U.S. campaigns, yet its internal division of labor in paid advertising has remained largely unmeasured. By classifying the goals of over 377,000 Facebook and Instagram ads from the 2022 US elections, we provide empirical evidence about the strategic behavior of candidates, parties, and outside groups. To this end, we build on previous work to extend and refine a taxonomy of nine goals for online election ads: acquisition, contact, donate, event, learn, persuade, poll, purchase, and vote. We analyze how advertisers allocate spending across goals and show that some — notably persuasion, donation, and learn — receive far greater spending than others. We also show that different types of sponsors often pursue different goals. For example, candidate advertisers devote a much higher share of their ads to fundraising than non-candidate advertisers, highlighting the costs of running for political office in the US. We also find that ad goals vary depending on the timing of the ad, the type of sponsor, and the characteristics of the targeted audience. Together, these findings offer the first systematic multi-sponsor account of how digital advertising reflects a functional division of labor within the extended party network that structures contemporary U.S. campaigns.
State supervision of ideas and information circulated among the public has a longstanding history. While there is a substantial body of literature examining the government’s motives for censorship, scholarly assessments of evolving censorship strategies in the new era of artificial intelligence (AI) remain relatively scarce. This paper analyzes an automated censorship system, developed and commercialized by a leading Chinese internet company, to study its content review logic. Using multiple real-world datasets, we assess: (1) concordance between conventional human-led and automated censorship decisions; (2) disruption of keyword evasion on the system’s effcacy; and (3) varied responses toward collective action and other political threats. Despite a notable gap between human-led and automated censorship decisions, we demonstrate that the system’s primary capability lies not in perfectly mimicking human censors, but in conducting large-scale user profiling and information categorization, which complements other information control tactics in China.
Which topics do local administrations focus on when providing refugees and other immigrants with information about their first steps after arrival – and how well are they received? These questions are particularly important in federal systems, where regions bear significant responsibility for integration efforts, despite limited resources and a complex network of offices to navigate. This study explores “Integreat”, Germany’s largest multilingual information platform designed to facilitate integration at the district and city levels. Using the platform’s full website visit history from 2018–2024, it analyzes several million records from 100+ participating regions. Results document significant differences in the thematic focus across districts. While some topics, such as work and training, are widely available, others are much less frequently covered. Multivariate analyses comparing topic availability and the likelihood of visiting related pages reveal several key insights. Pages offering immediate practical value (opening hours; mobility) are highly visited when available, yet absent in many districts. Topics related to acquiring essential skills or knowledge (e.g., language learning) show intense engagement relative to supply and users readily navigate detailed pages to access them. Finally, engagement differs across translations (e.g., Arabic, Ukrainian) for given topic availability, reflecting the diverse needs of different immigrant groups. The work concludes by deriving implications for administrative practice and noting the potential of these data for further descriptive and inferential research on refugee and immigrant integration.
This article introduces a dataset of all posts by candidates during the 2024 General Election in the United Kingdom with a presence on the X (formerly Twitter) platform. The article relies on a crowd-sourcing innovation in the United Kingdom that, for the first time, provided researchers with early access to a regularly updated candidate list prior to the start of the election. This made it possible to collect real-time data on candidate posts for 1,604 candidates across 53 separate political parties. Additionally, we download and store 53,327 images and 15,982 videos posted within tweets. We enrich the data with the realized vote count and vote share for each candidate as well as text transcripts extracted from the audio of video posts. Overall, the dataset provides a uniquely comprehensive collection of online campaigning material for an election campaign and will be of considerable value to scholars of political communication, elections, and democratic responsiveness. We also analyze the topics and tone — focusing on negativity — across different media formats to identify patterns in the content and style of candidate communication across parties.
Understanding typical smartphone behavior increasingly relies on device-enabled fine-grained data sources that go beyond retrospective self-reports. This study contributes to this knowledge by studying how, when, and under what conditions people engage with their smartphones, using a rich mixed-method dataset that combines Android logging, iOS data donation, and mobile experience sampling. The dataset captures both the quantity and quality of smartphone use among a large, quota-targeted sample of German adults (n = 1,797). The descriptive findings indicate that smartphone engagement is characterized by frequent interactions. Most sessions last under seven minutes, and app use rarely exceeds two minutes. Usage rhythms vary throughout the day: shorter glances dominate during constrained periods, while longer sessions cluster around mornings and evenings. Younger users display more fragmented usage patterns, whereas older adults tend to engage in fewer but longer sessions. These patterns reflect the situational affordances and gratifications of different types of mobile interactions and highlight the temporal structure of smartphone use. By mapping these rhythms and use types, our findings offer a foundation for theorizing about the cognitive, behavioral, and emotional consequences of smartphone use and provide practical guidance for researchers employing intensive longitudinal and real-time measurement approaches.