Misinformation poses a serious and growing challenge to democratic societies. A range of interventions have been developed to address it, yet their effects remain modest, short-lived, and prone to unintended trade-offs. We argue this is not simply a problem of weak tools but of mismatched scope: misinformation is embedded in polarized systems of identity, norms, trust, and platform incentives that no single intervention can address alone. Here, we propose coordinated intervention bundles spanning multiple levels, designed so components address each other's limitations. We outline diagnostic evidence, a case study, and evaluation principles that foreground trade-offs, durability, and legitimacy.
We examine the dynamics of citizens’ trust in public media during government-led efforts to implement major media reforms in a highly polarized context using two cross-sectional experiments. After Poland’s Law and Justice party (PiS) lost its parliamentary majority in 2023, the new government promised to “restore impartiality” to public news outlets that PiS had previously transformed into government mouthpieces. We conducted two experiments—one before and one after the reforms. In both experiments, we showed Polish respondents a representative sample of content from public and private media outlets, randomizing the inclusion of source information. We find that, pre-reform, both content preferences and source cue effects followed a highly polarized partisan pattern. Post-reform, polarization in public media attitudes was eliminated, as trust in the public television brand increased among new government supporters and decreased among PiS supporters, such that the public television cue had a null effect on trust for both pro and anti-PiS respondents.
The study of human behavior has shifted in the last fifteen years, with increasing reliance on opt-in non-probability online data sources. We offer an analysis of nine such data sources (total N = 13,053), aiming to inform researchers conducting experiments or correlational studies. We assess response validity (attentiveness, effort, honesty, speeding, and attrition), the extent to which samples represent the underlying population (observable demographics, measured attitude representativeness, and responding to experimental treatments), and professionalism (number of studies taken, frequency of taking studies, and modality of device on which the study is taken). We document substantial variation across these samples on each dimension. Samples that employ demographic quotas display relatively higher amounts of representativeness across multiple indicators (beyond demographics) but often exhibit less response validity. However, the inclusion of two attention checks early in a study enhances response validity without negatively impacting representativeness. We offer guidance for choosing opt-in samples, depending on the purpose of the research and resource constraints.
The search for effective interventions to counter misinformation has yielded disappointing results despite considerable effort among researchers. This disappointment stems from two interconnected problems: unrealistic expectations about behavioral science’s capacity to change beliefs, and an overly narrow focus on finding single “magic bullet” solutions. The replication crisis revealed that many celebrated behavioral interventions produce smaller effects than initially reported, yet these inflated expectations continue to shape both public discourse and assessments of intervention efficacy. Meanwhile, researchers in the field have pursued individual interventions in a competitive rather than collaborative framework, overlooking potential synergies between approaches. I argue that advancing the field requires both recalibrating our understanding of what constitutes success and reconceptualizing intervention design around bundled, complementary strategies. In addition, proper evaluation must focus on truth discernment – people’s ability to distinguish true from false content – rather than simply measuring reductions in false belief. By combining multiple modest interventions strategically, we can build layered systems of protection that are more robust than any single approach.
Can reading a chapter of popular nonfiction shift politically relevant attitudes among resistant partisans, and, if so, can an AI-generated summary do so just as effectively? We address these questions in a preregistered experiment where N = 555 Republicans were randomized to a treatment aimed at increasing trust in civil servants in which they read “The Cyber Sleuth” by Geraldine Brooks—a profile of IRS cybercrime investigators from Michael Lewis’s Who Is Government?—or a control condition in which they read an excerpt from Jonathan Haidt’s The Happiness Hypothesis. Participants were further randomized to read either the full original text (∼5000 words) or an AI-generated summary that was less than 1/3 of the original’s length. Reading “The Cyber Sleuth” produced large, wide-ranging attitude change relative to the control: IRS favorability increased substantially (d= 0.62), with significant spillover to civil-service favorability more generally (d= 0.37) and favorability toward Social Security Administration workers (d = 0.23), as well as opposition to DOGE-led workforce reductions (d = 0.40). Nearly all these effects remained significant two months later (ds 0.14–0.19). Strikingly, the much shorter AI-generated summaries were equivalently persuasive across nearly all outcomes. Popular nonfiction can be an effective tool for politically-relevant attitude change, even for polarized and partisan issues. The success of the AI summaries further suggests that this persuasive impact can be transmitted through concise summaries that preserve core informational content.
Both academic researchers and political pundits have generally accepted two over-time features of persuasion by partisan media: that the persuasive effects of partisan media might be temporary and decay quickly after a single exposure, and that these effects accumulate from multiple exposures. That effects decay may serve to ameliorate concerns about the broad impact of such media on partisan polarization. Yet the assumption that persuasive effects accumulate may raise larger concerns from real-world repeat exposure. To explore these possibilities, we implement a novel set of multiwave experiments that allow us to examine concerns about media effects over time. We present estimates from three studies suggesting that the persuasive effect of exposure to just a short article or video clip can persist for up to a week. In contrast to this persistence, our results suggest that an experiment adequately powered to detect the cumulative effect from multiple doses of partisan media—let alone one powered to detect cumulative effects among subgroups of the population—would require an unrealistic number of respondents. These cumulative effects are thus difficult to test in an experimental setting with limited resources.
Low confidence in the integrity of elections is a growing concern in the US, and questioning election integrity has become a core part of Republican identity in recent years. These beliefs appear to be the result of an uninformed or misinformed electorate. However, despite growing evidence that factual information can shift political beliefs even on contentious issues, election integrity beliefs have so far proven unusually resistant to information-based approaches, arguably because they are tightly linked to partisan identity and reinforced by a polarized information environment. To examine whether election integrity beliefs are indeed resistant to corrective information, we develop and test an informational treatment that provides a high volume of politically balanced accurate evidence on election integrity. Immediately prior to the 2024 general election, we randomly assigned N = 871 Republicans to either the experimental group or a control group engaging with general political information. The treatment substantially increased participants’ overall beliefs about the integrity of US elections, retrospective beliefs about the integrity of the 2020 election, and prospective beliefs about the expected integrity of the upcoming 2024 election (.6 < ds < .8). Furthermore, a follow-up shows that the effects persist two weeks later, following the 2024 election. These findings demonstrate that even beliefs closely tied to partisan identity are responsive to credible factual information.
There is great public concern about the potential use of generative artificial intelligence (AI) for political persuasion and the resulting impacts on elections and democracy1-6. We inform these concerns using pre-registered experiments to assess the ability of large language models to influence voter attitudes. In the context of the 2024 US presidential election, the 2025 Canadian federal election and the 2025 Polish presidential election, we assigned participants randomly to have a conversation with an AI model that advocated for one of the top two candidates. We observed significant treatment effects on candidate preference that are larger than typically observed from traditional video advertisements7-9. We also document large persuasion effects on Massachusetts residents' support for a ballot measure legalizing psychedelics. Examining the persuasion strategies9 used by the models indicates that they persuade with relevant facts and evidence, rather than using sophisticated psychological persuasion techniques. Not all facts and evidence presented, however, were accurate; across all three countries, the AI models advocating for candidates on the political right made more inaccurate claims. Together, these findings highlight the potential for AI to influence voters and the important role it might play in future elections.
Antisemitic conspiracy theories have been central to anti-Jewish prejudice for centuries. Given their longevity and deep ties to religious, ethnic, and ideological identities, debunking them presents a particularly difficult challenge. Here, we test whether having believers discuss their chosen antisemitic conspiracy with a large language model (LLM) prompted to debunk such conspiracies can reduce belief and improve attitudes toward Jews. In a preregistered experiment (N = 1,224 U.S. adults endorsing an antisemitic conspiracy theory), participants were randomized to a dialogue with an LLM (Claude 3.5 Sonnet) prompted to debunk their belief, or one of two control conditions. The debunking dialogue substantially reduced belief in antisemitic conspiracies relative to controls, and increased favorability toward Jews among initially unfavorable participants. These findings show that even deeply rooted, identity-linked conspiracies can be effectively debunked through factual correction, offering new insight into prejudice reduction and suggesting that LLM chatbots may help reduce antisemitism at scale.
An enormous body of literature argues that recommendation algorithms drive political polarization by creating "filter bubbles" and "rabbit holes." Using four experiments with nearly 9,000 participants, we show that manipulating algorithmic recommendations to create these conditions has limited effects on opinions. Our experiments employ a custom-built video platform with a naturalistic, YouTube-like interface presenting real YouTube videos and recommendations. We experimentally manipulate YouTube's actual recommendation algorithm to simulate filter bubbles and rabbit holes by presenting ideologically balanced and slanted choices. Our design allows us to intervene in a feedback loop that has confounded the study of algorithmic polarization-the complex interplay between supply of recommendations and user demand for content-to examine downstream effects on policy attitudes. We use over 130,000 experimentally manipulated recommendations and 31,000 platform interactions to estimate how recommendation algorithms alter users' media consumption decisions and, indirectly, their political attitudes. Our results cast doubt on widely circulating theories of algorithmic polarization by showing that even heavy-handed (although short-term) perturbations of real-world recommendations have limited causal effects on policy attitudes. Given our inability to detect consistent evidence for algorithmic effects, we argue the burden of proof for claims about algorithm-induced polarization has shifted. Our methodology, which captures and modifies the output of real-world recommendation algorithms, offers a path forward for future investigations of black-box artificial intelligence systems. Our findings reveal practical limits to effect sizes that are feasibly detectable in academic experiments.
Recent advancements in generative AI have raised widespread concern about the use of this technology to spread audio and visual misinformation. In response, there has been a major push among policymakers and technology companies to label AI-generated media appearing online. It remains unclear, however, what types of labels are most effective for this purpose. Here, we evaluate two (potentially complementary) strategies for labeling AI-generated content online: (i) a process-based approach, aimed at clarifying how content was made and (ii) a harm-based approach, aimed at highlighting content's potential to mislead. Using two preregistered survey experiments focused on misleading, AI-generated images (total n = 7,579 Americans), we assess the consequences of these different labeling strategies for viewers' beliefs and behavioral intentions. Overall, we find that all of the labels we tested significantly decreased participants' belief in the presented claims. However, in both studies, labels that simply informed participants that content was generated using AI tended to have little impact on respondents' stated likelihood of engaging with their assigned post. Together, these results shed light on the relative advantages and disadvantages of different approaches to labeling AI-generated media online.
Misinformation exposure is widely known to affect individuals’ own judgements. Here, we demonstrate that exposure to misinformation also has social consequences which have been largely overlooked. Across multiple experiments (total N=6,273 Americans), we extend the study of misinformation from studying what people believe themselves (first-order beliefs) to also examining what people believe about what others believe (second-order beliefs). Given the substantial evidence that second-order beliefs play a central role in social coordination, collective action, and trust in institutions, understanding the relationship between misinformation and second-order beliefs—while largely overlooked—is of great importance. First, we demonstrate that the illusory truth effect extends to second-order beliefs: exposure to false content increases participants' belief that others believe the claim just as much as it increases their own belief in the claim. Second, we find that individuals tend to overestimate how widely false statements—but not true claims—are believed by others, indicating miscalibrated meta-perceptions. Third, we show that common interventions designed to reduce first-order belief to misinformation— media literacy tips and descriptive norms—also reduce second-order beliefs to a similar extent. By demonstrating that misinformation alters not only private belief but also perceived social consensus, we uncover second-order belief distortion as a critical mechanism through which falsehoods influence societies; and by demonstrating that anti-misinformation interventions also reduce second-order belief in false claims, we reveal a previously underappreciated channel through which such intervention strategies exert social benefits.
Media platforms have recently introduced initiatives to label AI-generated media, aiming to increase transparency about content creation. Yet such efforts may carry unintended consequences. AI-generated media often accompany informational content that can vary in veracity. However, labeling may confound perceptions of the media's authenticity and the content's veracity, reducing belief in true information. Moreover, since it isn’t feasible to label all AI-generated media, partial labeling may lead people to assume that the absence of a label implies authenticity and/or veracity. We test for these labeling and implied effects in two survey experiments (N = 11,044), where respondents evaluated political news posts. Labeling decreased perceptions of the authenticity of AI-generated images but also lowered belief in and willingness to share posts—even when the associated claims were true. Furthermore, exposure to partial labeling increased the perceived authenticity of unlabeled content. These results highlight the need for carefully designed labeling practices online.
Content moderation is a critical aspect of platform governance on social media and of particular relevance to addressing the belief in and spread of misinformation. However, current content moderation practices have been criticized as unjust. This raises an important question-who do Americans want deciding whether online content is harmfully misleading? We conducted a nationally representative survey experiment (n = 3,000) in which US participants evaluated the legitimacy of hypothetical content moderation juries tasked with evaluating whether online content was harmfully misleading. These moderation juries varied on whether they were described as consisting of experts (e.g. domain experts), laypeople (e.g. social media users), or nonjuries (e.g. computer algorithm). We also randomized features of jury composition (size and necessary qualifications) and whether juries engaged in discussion during content evaluation. Overall, participants evaluated expert juries as more legitimate than layperson juries or a computer algorithm. However, modifying layperson jury features helped increase legitimacy perceptions-nationally representative or politically balanced composition enhanced legitimacy, as did increased size, individual juror knowledge qualifications, and enabling juror discussion. Maximally legitimate layperson juries were comparably legitimate with expert panels. Republicans perceived experts as less legitimate compared with Democrats, but still more legitimate than baseline layperson juries. Conversely, larger lay juries with news knowledge qualifications who engaged in discussion were perceived as more legitimate across the political spectrum. Our findings shed light on the foundations of institutional legitimacy in content moderation and have implications for the design of online moderation systems.
Social scientists rely heavily on data collected from human participants via surveys or experiments. To obtain these data, many social scientists recruit participants from opt-in online panels that provide access to large numbers of people willing to complete tasks for modest compensation. In a large study (total N=13,053), we explore nine opt-in non-probability samples of American respondents drawn from panels widely used in social science research, comparing them on three dimensions: response quality (attention, effort, honesty, speeding, and attrition), representativeness (observable demographics, measured attitude typicality, and responding to experimental treatments), and professionalism (number of studies taken, frequency of taking studies, and modality of device on which the study is taken). We document substantial variation across these samples on each dimension. Most notably, we observe a clear tradeoff between sample representativeness and response quality (particularly regarding attention), such that samples with more attentive respondents tend to be less representative, and vice versa. Even so, we find that for some samples, this tension can be largely eliminated by adding modest attention filters to more representative samples. This and other insights enable us to provide a guide to help researchers decide which online opt-in sample is optimal given one’s research question and constraints.
Interventions to reduce misinformation sharing have been a major focus in recent years. Developing “content-neutral” interventions that do not require specific fact-checks or warnings related to individual false claims is particularly important in developing scalable solutions. Here, we provide the first evaluations of a content-neutral intervention to reduce misinformation sharing conducted at scale in the field. Specifically, across two on-platform randomized controlled trials, one on Meta’s Facebook (N=33,043,471) and the other on Twitter (N=75,763), we find that simple messages reminding people to think about accuracy—delivered to large numbers of users using digital advertisements—reduce misinformation sharing, with effect sizes on par with what is typically observed in digital advertising experiments. On Facebook, in the hour after receiving an accuracy prompt ad, we found a 2.6% reduction in the probability of being a misinformation sharer among users who had shared misinformation the week prior to the experiment. On Twitter, over more than a week of receiving 3 accuracy prompt ads per day, we similarly found a 3.7% to 6.3% decrease in the probability of sharing low-quality content among active users who shared misinformation pre-treatment. These findings suggest that content-neutral interventions that prompt users to consider accuracy have the potential to complement existing content-specific interventions in reducing the spread of misinformation online.
The surge in online self-administered surveys has given rise to an extensive body of literature on respondent inattention, also known as careless or insufficient effort responding. This burgeoning literature has outlined the consequences of inattention and made important strides in developing effective methods to identify inattentive respondents. However, differences in terminology, as well as a multiplicity of different methods for measuring and correcting for inattention, have made this literature unwieldy. We present an overview of the current state of this literature, highlighting commonalities, emphasizing key debates, and outlining open questions deserving of future research. Additionally, we emphasize the key considerations that survey researchers should take into account when measuring attention.
The spread of misinformation through media and social networks threatens many aspects of society, including public health and the state of democracies. One approach to mitigating the effect of misinformation focuses on individual-level interventions, equipping policymakers and the public with essential tools to curb the spread and influence of falsehoods. Here we introduce a toolbox of individual-level interventions for reducing harm from online misinformation. Comprising an up-to-date account of interventions featured in 81 scientific papers from across the globe, the toolbox provides both a conceptual overview of nine main types of interventions, including their target, scope and examples, and a summary of the empirical evidence supporting the interventions, including the methods and experimental paradigms used to test them. The nine types of interventions covered are accuracy prompts, debunking and rebuttals, friction, inoculation, lateral reading and verification strategies, media-literacy tips, social norms, source-credibility labels, and warning and fact-checking labels. Kozyreva et al. review evidence from individual-level interventions for fighting online misinformation featured in 81 scientific papers. They classify the interventions in nine different types and summarize their findings in a toolbox.
Debates on how tech companies ought to oversee the circulation of content on their platforms are increasingly pressing. In the U.S., questions surrounding what, if any, action should be taken by social media companies to moderate harmfully misleading content on topics such as vaccine safety and election integrity are now being hashed out from corporate boardrooms to federal courtrooms. But where does the American public stand on these issues? Here we discuss the findings of a recent nationally representative poll of Americans’ views on content moderation of harmfully misleading content.
Labeling is a commonly proposed strategy for reducing the risks of generative artificial intelligence (AI). This approach involves applying visible content warnings to alert users to the presence of AI-generated media online (e.g., on social media, news sites, or search . . .