Despite concern about exposure to content from untrustworthy sources on social media, little is known about the frequency or effects of exposure to their content. We examine 2020 data from all active US adults on Facebook and Instagram to measure exposure to content from Pages, groups, and web domains on Facebook and public accounts on Instagram that repeatedly publish misinformation. We find that average users saw relatively little content in their feeds from these untrustworthy sources; exposure was highly concentrated. A multimonth field experiment during the 2020 election among consenting users reduced feed-based exposure to content from untrustworthy sources by approximately 70% on both platforms but had no measurable effects on numerous preregistered outcomes, even among participants with high pretreatment exposure. Our results demonstrate that a feasible platform intervention can successfully reduce exposure to content from untrustworthy sources but suggest that these changes are unlikely to have immediate effects on attitudes and beliefs.
Democratic challenges are often attributed to the spread of misleading, untrustworthy, or biased information, leading scholars to focus on minimizing exposure to such “bad” content online. Instead, we introduce a scalable intervention to put factual and verified public affairs information in users’ social media feeds to make them better informed and more resilient to various online threats. We conducted 48 field quasi-experiments using Instagram ads targeting news non-users to enhance their belief accuracy, democratic attitudes, and behavioral intentions related to climate change, COVID-19 vaccines, media literacy, and election integrity. The treatment videos reached 2,496,878 Instagram accounts, 690,470 users watched at least 50% of the video, and 40,584 of those users completed post-test assessment. The intervention was effective: 46 out of 48 of the quasi-experiments had positive effect sizes and 40 out of 48 achieved statistical significance. The intervention predicted not only belief accuracy but also attitudes, media literacy, and — to some extent — behavioral intentions related to vaccination. These patterns emerged across topics, did not dissipate with time (two of three climate change quasi-experiments show continued effects), and were not contingent on persuasive appeals and format features presented in the ads.
Deceptive online networks are coordinated efforts that use identity deception to pursue strategic political or financial goals. During the US 2020 elections, these networks reached at least 37 million Facebook and 3 million Instagram users, representing 15% and 2% of the platforms' active US adult users, respectively. Only 3 networks out of 49-1 network with explicitly political aims and 2 that appeared to use politics as a lure for profit-were responsible for over 70% of users reached. Notably, accounts unaffiliated with the networks played an important role in facilitating this reach by resharing content the three networks produced. Deceptive networks, regardless of whether their goals were political or financial, reached users who were older, more conservative, more frequently exposed to content from untrustworthy sources, and spent more time on Facebook.
We study the effects of social media political advertising by randomizing subsets of 36,906 Facebook users and 25,925 Instagram users to have political ads removed from their news feeds for 6 weeks before the 2020 US presidential election. We show that most presidential ads were targeted towards parties' own supporters and that fundraising ads were the most common. On both Facebook and Instagram, we found no detectable effects of removing political ads on political knowledge, polarization, perceived legitimacy of the election, political participation (including campaign contributions), candidate favourability and turnout. This was true overall and for both Democrats and Republicans separately.
Recommendation algorithms on social media platforms optimize for user engagement, which can inadvertently amplify exposure to harmful content such as violence, sexual material, and hate speech. Platform-level moderation is often delayed, opaque, and uniform, motivating the need for complementary user-side interventions that allow individuals to reduce unwanted content in their feeds without relying on platform cooperation. Prior work largely relies on single-session or human-subject studies, limiting the ability to capture recursive recommendation feedback loops, control for users’ baseline preferences for harmful content, or systematically compare intervention strategies across harm types. To address these gaps, we propose a sock puppet simulation framework that models 30 rounds of iterative recommendation and interaction. We evaluate two user-side interventions: Downranking and Replacement, on YouTube’s Homepage and Up-Next interfaces, controlling for users’ baseline harm exposure levels (0%, 25%, 50%), yielding 18 experimental conditions with 1,000 puppets each. Our results show that user-side interventions are effective relative to the baseline, with effects concentrated on the Homepage interface. In particular, Downranking emerges as the most robust and consistent strategy, producing statistically significant improvements in both final-state outcomes (net change) and cumulative harmful exposure across baseline preference levels. For example, Downranking reverses baseline increases in harm into significant decreases (e.g., from a +0.6 percentage-point increase to a -0.9 percentage-point reduction), and yields durable reductions in cumulative exposure over time. Replacement shows weaker and less consistent effects on final-state outcomes. We further find no strong evidence of heterogeneous intervention effects across harm types, and observe that the overall reduction in harmful recommendations is largely driven by declines in Physical harm. Our work establishes the long-term efficacy of user-side interventions and provides guidance for their design.
Despite growing evidence on the popularity, the coverage, and the effects of partisan news media in the U.S., missing is a more fundamental understanding of how potential political bias manifests in partisan media coverage. We propose a comprehensive framework examining (a) the prevalence of political news that is positive toward the in-party (in-party affinity) versus negative toward the out-party (out-party hostility), accounting for key factors shaping these reporting strategies, and also attend to (b) user engagement with these different reporting strategies. Analyzing 1,011,911 Facebook posts from 50 major partisan outlets in the U.S. from 2010 to 2020 with ensemble models integrating Gemini, Llama, and DistilBERT, we show that partisan media criticized the out-party more than praised the in-party. This pattern was influenced by the changing political and media environment, party power distribution, and ideological extremity of the media outlet. Users were more likely to share and comment on posts criticizing the out-party than those praising the in-party, but less likely to react to such content (e.g., like, love, haha, angry); higher levels of reactions, but not shares or comments, in turn, predicted increased posting of out-party negative news posts.
Political polarization poses major societal risks, yet a globally relevant, interdisciplinary understanding remains lacking, as nearly half of all studies focus on the United States. We argue for renewed effort to bring global equity to polarization research, highlighting insights robust across countries, those unique to specific contexts, and key literature gaps. Closing these gaps means overcoming cultural and systemic barriers, including data-access restrictions and misaligned incentives. It also demands interdisciplinarity bridging traditional approaches and those addressing the polarizing role of the internet, social media, and AI. Otherwise, efforts to counter polarization, and its democratic harms, risk resting on unsuitable evidence.
We report the first direct comparisons of multiple alternative social media algorithms on multiple platforms on outcomes of societal interest. We used a browser extension to modify which posts were shown to desktop social media users, randomly assigning 9,386 users to a control group or one of five alternative ranking algorithms which simultaneously altered content across three platforms for six months during the US 2024 presidential election. This reduced our preregistered index of affective polarization by an average of 0.03 standard deviations (p < 0.05), including a 1.5 degree decrease in differences between the 100 point inparty and outparty feeling thermometers. We saw reductions in active use time for Facebook (-0.37 min/day) and Reddit (-0.2 min/day), but an increase of 0.32 min/day (p < 0.01) for X/Twitter. We saw an increase in reports of negative social media experiences but found no effects on well-being, news knowledge, outgroup empathy, perceptions of and support for partisan violence. This implies that bridging content can improve some societal outcomes without necessarily conflicting with the engagement-driven business model of social media.
Cheapfakes, or real images presented misleadingly or in unrelated contexts, are an increasingly prominent form of visual misinformation. While media literacy interventions can enhance individuals' ability to detect such content, motivational barriers often hinder the adoption of image verification. This study examines whether incorporating different mechanisms and types of incentives into a digital media literacy intervention improves visual misinformation discernment and image verification behavior, both immediately and over time. We conducted a pre-registered two-wave between-subjects online experiment (N = 1,421) on a professionally designed social media platform. The study used a 2 (Incentive Type: symbolic vs. monetary) x 2 (Incentive Mechanism: task- vs. result-based) factorial design with additional control groups. Results show that task-based incentives, particularly monetary ones, were most effective at initiating image verification behaviors, namely reverse image search, and boosting short-term discernment, whereas result-based incentives were more effective in sustaining discernment accuracy. These findings suggest that both the mechanism and the type of incentives play a critical role in shaping the short- and long-term effectiveness of media literacy interventions, highlighting the value of multi-phased incentive strategies for combating visual misinformation in digital environments.
Web browsers now provide AI-generated news summaries for millions of users. Despite their popularity and influence, we lack a systematic understanding of how these systems transform news before people read it. Through a large-scale audit, we investigate the factual accuracy of browser-based AI summarizers and how they alter the political bias, negative affect, and journalistic writing quality of news. Drawing on 13,777 articles from 15 U.S. news outlets, we evaluate their 41,331 summaries generated by three leading AI-powered browsers: Google Chrome (Gemini), Microsoft Edge (Copilot), and Perplexity Comet. We find that browser-based AI summarizers are broadly accurate. Furthermore, they consistently transform news by attenuating ideological bias, partisan stances, negativity, anger, and fear, while increasing clarity and reducing personal tone. With some variations, these patterns hold across browsers, outlet ideologies, and topics. Our findings identify AI-powered browsers as a new class of editorial intermediaries that systematically reshape news, with implications for democratic discourse and AI governance.
The rapid growth of social media platforms has led to concerns about radicalization, filter bubbles, and content bias. Existing approaches to classifying ideology are limited in that they require extensive human effort, the labeling of large datasets, and are not able to adapt to evolving ideological contexts. This paper explores the potential of Large Language Models (LLMs) for classifying the political ideology of online content in the context of the two-party US political spectrum through in-context learning (ICL). Our extensive experiments involving demonstration selection in label-balanced fashion, conducted on three datasets comprising news articles and YouTube videos, reveal that our approach significantly outperforms zero-shot and traditional supervised methods. Additionally, we evaluate the influence of metadata (e.g., content source and descriptions) on ideological classification and discuss its implications. Finally, we show how providing the source for political and non-political content influences the LLM's classification.
With a folk understanding that political polarization refers to socio-political divisions within a society, many have proclaimed that we are more divided than ever. In this account, polarization has been blamed for populism, the erosion of social cohesion, the loss of trust in the institutions of democracy, legislative dysfunction, and the collective failure to address existential risks such as Covid-19 or climate change. However, at a global scale there is surprisingly little academic literature which conclusively supports these claims, with half of all studies being U.S.-focused. Here, we provide an overview of the global state of research on polarization, highlighting insights that are robust across countries, those unique to specific contexts, and key gaps in the literature. We argue that addressing these gaps is urgent, but has been hindered thus far by systemic and cultural barriers, such as regionally stratified restrictions on data access and misaligned research incentives. If continued cross-disciplinary inertia means that these disparities are left unaddressed, we see a substantial risk that countries will adopt policies to tackle polarization based on inappropriate evidence, risking flawed decision-making and the weakening of democratic institutions.
We assess the phenomenon of partisan temporal selective avoidance, or individuals dynamically altering their news consumption when news is negative toward their in- and out-party. Using nine months of online behavioral data (27,648,770 visits) from 2,462 Americans paired with machine learning classifications, we examine whether changing daily news sentiment toward in- and out-party (macro-level) and exposure to articles negative toward in- or out-party during one's browsing session (micro-level) influence news use. We test if partisans change their consumption of (a) news overall, (b) partisan outlets, (c) hard versus soft news, and (d) individual articles. We find support for partisan temporal selective news avoidance; partisans alter the volume, type, and source of news because of changing news sentiment. On the macro-level, partisan asymmetries emerge, and on the micro-level negative news about either party reduce news browsing length while increasing hard news and negative news visits for both Democrats and Republicans.
Abortion has been one of the most divisive issues in the United States. Yet, missing is comprehensive longitudinal evidence on how political divides on abortion are reflected in public discourse over time, on a national scale, and in response to key events before and after the overturn of Roe v Wade. We analyze a corpus of over 3.5M tweets related to abortion over the span of one year (January 2022 to January 2023) from over 1.1M users. We estimate users' ideology and rely on state-of-the-art transformer-based classifiers to identify expressions of hostility and extract five prominent frames surrounding abortion. We use those data to examine (a) how prevalent were expressions of hostility (i.e., anger, toxic speech, insults, obscenities, and hate speech), (b) what frames liberals and conservatives used to articulate their positions on abortion, and (c) the prevalence of hostile expressions in liberals and conservative discussions of these frames. We show that liberals and conservatives largely mirrored each other's use of hostile expressions: as liberals used more hostile rhetoric, so did conservatives, especially in response to key events. In addition, the two groups used distinct frames and discussed them in vastly distinct contexts, suggesting that liberals and conservatives have differing perspectives on abortion. Lastly, frames favored by one side provoked hostile reactions from the other: liberals use more hostile expressions when addressing religion, fetal personhood, and exceptions to abortion bans, whereas conservatives use more hostile language when addressing bodily autonomy and women's health. This signals disrespect and derogation, which may further preclude understanding and exacerbate polarization.
Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users. These platforms expose users to harmful content, ranging from clickbait or physical harms to hate or misinformation. Yet, we lack a comprehensive understanding and measurement of online harm on short video platforms. Toward this end, we present two large-scale datasets of multi-modal and multi-categorical online harm: (1) 60,906 systematically selected potentially harmful YouTube videos and (2) 19,422 videos annotated by three labeling actors: trained domain experts, GPT-4-Turbo (using 14 image frames, 1 thumbnail, and text metadata), and crowdworkers (Amazon Mechanical Turk master workers). The annotated dataset includes both (a) binary classification (harmful vs. harmless) and (b) multi-label categorizations of six harm categories: Information, Hate and harassment, Addictive, Clickbait, Sexual, and Physical harms. Furthermore, the annotated dataset provides (1) ground truth data with videos annotated consistently across (a) all three actors and (b) the majority of the labeling actors, and (2) three data subsets labeled by individual actors. These datasets are expected to facilitate future work on online harm, aid in (multi-modal) classification efforts, and advance the identification and potential mitigation of harmful content on video platforms.
Although news negativity is often studied, missing is comparative evidence on the prevalence of and engagement with negative political and non-political news posts on social media. We use 6,081,134 Facebook posts published between January 1, 2020, and April 1, 2024, by 97 media organizations in six countries (U.S., UK, Ireland, Poland, France, Spain) and develop two multilingual classifiers for labeling posts as (non-)political and (non-)negative. We show that: (1) negative news posts constitute a relatively small fraction (12.6
Online panels have become an important resource for research in political science, but the compensation offered to panelists incentivizes them to become "survey professionals," raising concerns about data quality. We provide evidence on survey professionalism exploring three US samples of subjects who donated their browsing data, recruited via Lucid, YouGov, and Facebook (total $n = 3,886$ ). Survey professionalism is common, but varies across samples: by our most conservative estimate, we find 1.7% of respondents on Facebook, 7. $\color {black}6$ % on YouGov, and 34 $\color {black}.7$ % on Lucid to be professionals (under the assumption that professionals are as likely as non-professionals to donate data after conditioning on observable demographics available from all online survey takers). However, evidence that professionals lower data quality is limited: they do not systematically differ demographically or politically from non-professionals and do not exhibit more response instability. They are, however, somewhat more likely to speed, straightline, and attempt to take questionnaires repeatedly. To address potential selection issues in donating of browsing data, we present sensitivity analyses with lower bounds for survey professionalism. While concerns about professionalism are warranted, we conclude that survey professionals do not, by and large, distort inferences of research based on online panels.
The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised classifiers, and large volumes of training data, and often struggle with scalability, subjectivity, and the dynamic nature of harmful content (e.g., violent content, dangerous challenge trends, etc.). To bridge these gaps, we utilize Large Language Models (LLMs) to undertake few-shot dynamic content moderation via in-context learning. Through extensive experiments on multiple LLMs, we demonstrate that our few-shot approaches can outperform existing proprietary baselines (Perspective and OpenAI Moderation) as well as prior state-of-the-art few-shot learning methods, in identifying harm. We also incorporate visual information (video thumbnails) and assess if different multimodal techniques improve model performance. Our results underscore the significant benefits of employing LLM based methods for scalable and dynamic harmful content moderation online.
Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current moderation efforts, reliant on classifiers trained with extensive human-annotated data, struggle with scalability and adapting to new forms of harm. To address these challenges, we propose a novel re-ranking approach using Large Language Models (LLMs) in zero-shot and few-shot settings. Our method dynamically assesses and re-ranks content sequences, effectively mitigating harmful content exposure without requiring extensive labeled data. Alongside traditional ranking metrics, we also introduce two new metrics to evaluate the effectiveness of re-ranking in reducing exposure to harmful content. Through experiments on three datasets, three models and across three configurations, we demonstrate that our LLM-based approach significantly outperforms existing proprietary moderation approaches, offering a scalable and adaptable solution for harm mitigation.
Could news coverage of misinformation be harmful? Across two studies on U.S. citizens, we examine whether news coverage of misinformation generates misperceptions and decreases levels of trust in information institutions (i.e., media, professors, and scientists) and whether its effects can be comparable to those of exposure to untrustworthy content. We rely on an online experiment using mock social media posts (Study 1, N = 1,670) and also on online behavioral tracking data paired with over-time survey self-reports (Study 2, N = 804). Study 1 finds that exposure to both actual misinformation and the coverage of misinformation affect misperceptions, but does not decrease trust. Study 2 presents evidence that behaviorally tracked visits to untrustworthy sites and exposure to news coverage of misinformation—although relatively rare—do not affect misperceptions, but both predict lower levels of trust in scientists, with less consistent effects for media and university professors. These results support concerns that not only misinformation but also its coverage contribute to epistemic uncertainty by eroding confidence in credible sources of knowledge, and warrant further inquiry into the potential harms of news media’s attention to misinformation.