BackgroundHPV vaccination rates in the United States remain substantially below the 80% target of Healthy People 2030. New approaches for increasing vaccination rates are needed. We tested whether brief text conversations with the artificial intelligence large language model (LLM) GPT-4 Turbo could significantly increase intentions to vaccinate children against HPV.MethodsWe conducted a pre-registered randomized clinical trial in which participants were assigned to a text conversation with GPT-4 Turbo in which the LLM (with no specific training) tried to persuade them to vaccinate against HPV, a CDC informational brochure, or a control conversation with the LLM about an unrelated topic. In October 2025, we recruited 7368 Americans via Cint Exchange who had unvaccinated eligible children or did not yet have children, and whose intentions to vaccinate their children against HPV was under 75 on a 0-100 scale. Participants completed pre- and post-intervention measures of HPV vaccination intentions for sons and daughters. Randomization was via Qualtrics, with no masking possible. Our primary outcome was change in HPV vaccination intentions.FindingsThe AI-dialogue significantly increased vaccination intentions relative to the control (sons: 9.73 points, 95% CI [6.89, 12.56], p < .001; daughters: 7.72 points, 95% CI [4.98, 10.46], p < .001); as did the brochure (sons: 8.04 points, 95% CI [4.83, 11.25], p < .001; daughters: 5.51 points, 95% CI [2.52, 8.50], p < .001). The differences between the AI-dialogue and the brochure were not statistically significant (sons, p = .268; daughters, p = .120). Post-assignment attrition rates did not significantly differ between treatment arms (AI-dialgue, 11.9pp; brochure, 14.2pp), although both arms had significantly more attrition than the control (5.1pp); our results were robust to baseline-carried-forward imputation. Post-hoc analysis found that the AI-dialogue was significantly more likely than the CDC brochure to produce large increases in vaccination intentions (e.g. >30 points). Fact-checking of all 15,559 AI-generated claims using an independent AI model with web search capability revealed high overall accuracy (median = 100/100, mean = 95.4/100 on a scale where 0 = completely inaccurate and 100 = completely accurate), with only 0.4% of claims being medically inaccurate.InterpretationDialogues with GPT-4 Turbo offer a potential strategy for increasing HPV vaccination intentions at scale.
Research on misinformation has expanded rapidly over the past decade. This review examines areas of consensus and contention within the literature on misinformation, with particular attention to conceptual definitions, measurement approaches, and theoretical explanations. We outline views on what it means to be susceptible to misinformation and advocate for assessing susceptibility by focusing on the ability to discern true from false claims. Although exposure to fake news is less common online than many think, we argue that misinformation-more broadly construed-continues to be a major problem. We review evidence on cognitive and socioaffective mechanisms underlying susceptibility to misinformation, including repetition effects, lack of analytic reasoning, political identity, and information environments. We also evaluate interventions such as fact-checking, prebunking, psychological inoculation, and accuracy prompts. Overall, the literature reveals several robust findings but also significant methodological limitations and theoretical disagreements that suggest fruitful future directions for misinformation research.
Misinformation on social media threatens democracy, public health, and civic discourse. In fast-paced feeds, users often rely on quick cues when deciding what to attend to, believe, and share. One particularly consequential linguistic cue is confidence—widely treated as a signal of knowledge and credibility. Yet, we show that in social media news contexts, confidence is counterintuitively anti-diagnostic: it signals lower-quality information while also predicting greater diffusion. Using computational linguistics and machine learning, we analyze over 20 million posts across eight major online platforms. Posts expressed with greater confidence are more likely to link to lower-quality, more misinformative news domains, even controlling for user baselines, toxicity, and political lean. At the same time, confident posts also receive more engagement. Together, these findings show that confidence—often treated as a cue to knowledge—acts as an anti-diagnostic signal of information quality while simultaneously predicting the diffusion of misinformation.
Misinformation poses a serious and growing challenge to democratic societies. A range of interventions have been developed to address it, yet their effects remain modest, short-lived, and prone to unintended trade-offs. We argue this is not simply a problem of weak tools but of mismatched scope: misinformation is embedded in polarized systems of identity, norms, trust, and platform incentives that no single intervention can address alone. Here, we propose coordinated intervention bundles spanning multiple levels, designed so components address each other's limitations. We outline diagnostic evidence, a case study, and evaluation principles that foreground trade-offs, durability, and legitimacy.
Despite decades of research, the conditions under which punishment promotes cooperation remain unclear. Through an integrative experiment varying 14 design parameters of public goods games across 360 experimental conditions (147,618 decisions from 7100 participants), we reveal substantial heterogeneity in punishment effectiveness: Its impact on welfare ranges from 43% improvement to 44% reduction depending on the game parameters. To characterize these patterns, we developed models that outperformed human forecasters in predicting punishment effectiveness in new experiments. Communication emerges as the most important factor, followed by contribution framing (opt out versus opt in), contribution type (variable versus all-or-nothing), game length, and outcome visibility, though these factors often interact. The results reframe the debate from whether punishment works to when it does, demonstrating how integrative experiments enable discovery of generalizable patterns in social phenomena.
We examine the dynamics of citizens’ trust in public media during government-led efforts to implement major media reforms in a highly polarized context using two cross-sectional experiments. After Poland’s Law and Justice party (PiS) lost its parliamentary majority in 2023, the new government promised to “restore impartiality” to public news outlets that PiS had previously transformed into government mouthpieces. We conducted two experiments—one before and one after the reforms. In both experiments, we showed Polish respondents a representative sample of content from public and private media outlets, randomizing the inclusion of source information. We find that, pre-reform, both content preferences and source cue effects followed a highly polarized partisan pattern. Post-reform, polarization in public media attitudes was eliminated, as trust in the public television brand increased among new government supporters and decreased among PiS supporters, such that the public television cue had a null effect on trust for both pro and anti-PiS respondents.
Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against ("debunking") or for ("bunking") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.
A consistent pattern emerges from the history of psychology: Technological advances change the way that we understand ourselves. We argue that, in addition to various uses that are already common (e.g., qualitative coding), large language models can be integrated into survey software and act as a virtual research assistant that can generate tailored stimuli on the fly. This creates unprecedented flexibility in developing materials for psychological theory testing. We present an illustrative case study to show how a major lingering debate in the field-that is, whether people really change their mind according to evidence or, instead, rely on motivated reasoning-was pushed forward by using artificial intelligence (AI) to administer personalized experimental treatments. We discuss various potential uses of AI to test hypotheses in psychological science and argue that psychologists should seriously consider using AI to better understand human intelligence.
The study of human behavior has shifted in the last fifteen years, with increasing reliance on opt-in non-probability online data sources. We offer an analysis of nine such data sources (total N = 13,053), aiming to inform researchers conducting experiments or correlational studies. We assess response validity (attentiveness, effort, honesty, speeding, and attrition), the extent to which samples represent the underlying population (observable demographics, measured attitude representativeness, and responding to experimental treatments), and professionalism (number of studies taken, frequency of taking studies, and modality of device on which the study is taken). We document substantial variation across these samples on each dimension. Samples that employ demographic quotas display relatively higher amounts of representativeness across multiple indicators (beyond demographics) but often exhibit less response validity. However, the inclusion of two attention checks early in a study enhances response validity without negatively impacting representativeness. We offer guidance for choosing opt-in samples, depending on the purpose of the research and resource constraints.
Political polarization poses major societal risks, yet a globally relevant, interdisciplinary understanding remains lacking, as nearly half of all studies focus on the United States. We argue for renewed effort to bring global equity to polarization research, highlighting insights robust across countries, those unique to specific contexts, and key literature gaps. Closing these gaps means overcoming cultural and systemic barriers, including data-access restrictions and misaligned incentives. It also demands interdisciplinarity bridging traditional approaches and those addressing the polarizing role of the internet, social media, and AI. Otherwise, efforts to counter polarization, and its democratic harms, risk resting on unsuitable evidence.
Can reading a chapter of popular nonfiction shift politically relevant attitudes among resistant partisans, and, if so, can an AI-generated summary do so just as effectively? We address these questions in a preregistered experiment where N = 555 Republicans were randomized to a treatment aimed at increasing trust in civil servants in which they read “The Cyber Sleuth” by Geraldine Brooks—a profile of IRS cybercrime investigators from Michael Lewis’s Who Is Government?—or a control condition in which they read an excerpt from Jonathan Haidt’s The Happiness Hypothesis. Participants were further randomized to read either the full original text (∼5000 words) or an AI-generated summary that was less than 1/3 of the original’s length. Reading “The Cyber Sleuth” produced large, wide-ranging attitude change relative to the control: IRS favorability increased substantially (d= 0.62), with significant spillover to civil-service favorability more generally (d= 0.37) and favorability toward Social Security Administration workers (d = 0.23), as well as opposition to DOGE-led workforce reductions (d = 0.40). Nearly all these effects remained significant two months later (ds 0.14–0.19). Strikingly, the much shorter AI-generated summaries were equivalently persuasive across nearly all outcomes. Popular nonfiction can be an effective tool for politically-relevant attitude change, even for polarized and partisan issues. The success of the AI summaries further suggests that this persuasive impact can be transmitted through concise summaries that preserve core informational content.
Charitable donations frequently fail to maximize cost-effectiveness (the amount of good a donation does per dollar). This failure is often attributed to charitable motivations being affective and thus insensitive to evidence-based arguments. We challenge this perspective, hypothesizing that evidence can substantially increase effective giving—if that evidence is sufficiently compelling. We test this prediction in a pre-registered experiment (N = 1,949 Americans) by leveraging the ability of artificial intelligence large language models (LLMs) to generate persuasive content. Participants allocated $1 between their favorite charity and a highly effective charity (the Against Malaria Foundation), before and after an LLM dialogue or a static LLM message advocating for the effective charity, or a control conversation. The LLM dialogue and static LLM message both significantly increased effective donations (45.9% and 28.7%, respectively) in comparison to control, while the LLM dialogue also shifted moral attitudes. Effective giving can be meaningfully increased through evidence-based persuasion with LLMs.
Large language models (LLMs) are increasingly embedded directly into social media platforms, enabling users to request real-time fact-checks of online content. Using an exhaustive dataset of 1,671,841 English-language fact-checking requests made to Grok and Perplexity on X between February and September 2025, we provide the first large-scale empirical analysis of how LLM-based fact-checking operates in the wild. Fact-checking requests comprise 7.6% of all interactions with the LLM bots, and focus primarily on politics, economics, and current events. We document clear partisan asymmetries in usage. Users requesting fact-checks from Grok are much more likely to be Republican than Democratic, while the opposite is true for fact-check requests from Perplexity -- indicating emerging polarization in attitudes toward specific AI models. At the same time, both Democrats and Republicans are more likely to request fact-checks on posts authored by Republicans, and - consistent with prior work using professional fact-checkers and crowd judgments - posts from Republican-leaning accounts are more likely to be rated as inaccurate by both LLMs. Across posts rated by both LLM bots, evaluations from Grok and Perplexity agree 52.6% of the time and strongly disagree (one party rates a claim as true and the other as false) 13.6% of the time. For a sample of 100 fact-checked posts, 54.5% of Grok bot ratings and 57.7% of Perplexity bot ratings agreed with ratings of human fact-checkers, which is significantly lower than the inter-fact-checker agreement rate of 64.0%; but API-access versions of Grok had higher agreement with fact-checkers than did not significantly differ from inter-fact-checker agreement. Finally, in a preregistered survey experiment with 1,592 U.S. participants, exposure to LLM fact-checks meaningfully shifts belief accuracy, with effect sizes comparable to those observed in studies of professional fact-checking. However, responses to Grok fact-checks are polarized by partisanship when model identity is disclosed, whereas responses to Perplexity are not. Together, these findings show that LLM-based fact-checking is rapidly scaling, is generally informative although far from perfect, while also becoming entangled with polarization and partisanship. Our work highlights both the promise and the risks of integrating AI fact-checking into online public discourse.
Low confidence in the integrity of elections is a growing concern in the US, and questioning election integrity has become a core part of Republican identity in recent years. These beliefs appear to be the result of an uninformed or misinformed electorate. However, despite growing evidence that factual information can shift political beliefs even on contentious issues, election integrity beliefs have so far proven unusually resistant to information-based approaches, arguably because they are tightly linked to partisan identity and reinforced by a polarized information environment. To examine whether election integrity beliefs are indeed resistant to corrective information, we develop and test an informational treatment that provides a high volume of politically balanced accurate evidence on election integrity. Immediately prior to the 2024 general election, we randomly assigned N = 871 Republicans to either the experimental group or a control group engaging with general political information. The treatment substantially increased participants’ overall beliefs about the integrity of US elections, retrospective beliefs about the integrity of the 2020 election, and prospective beliefs about the expected integrity of the upcoming 2024 election (.6 < ds < .8). Furthermore, a follow-up shows that the effects persist two weeks later, following the 2024 election. These findings demonstrate that even beliefs closely tied to partisan identity are responsive to credible factual information.
This piece distills lessons for increasing the impact of social science research based on our work at the Massachusetts Institute of Technology Applied Cooperation Initiative, which facilitates collaborations between social scientists and practitioners to promote real-world prosocial behaviors. We describe an iterative model of collaboration- intervention, randomized trial, publication, and press-that has sustained a robust research pipeline, broadening the dissemination and application of findings. We highlight two promising directions for future research in organizational contexts: (i) adapting interventions to intraorganizational dynamics, in which norms and incentives differ from public settings, and (ii) developing strategies for establishing or reforming counterproductive norms that may hinder organizational performance. We conclude with recommendations for supporting translational research: introducing article types dedicated to translation, removing barriers to publishing practitioner-relevant review articles, streamlining institutional review board and legal processes, and funding cross-institutional collaborations.
With a folk understanding that political polarization refers to socio-political divisions within a society, many have proclaimed that we are more divided than ever. In this account, polarization has been blamed for populism, the erosion of social cohesion, the loss of trust in the institutions of democracy, legislative dysfunction, and the collective failure to address existential risks such as Covid-19 or climate change. However, at a global scale there is surprisingly little academic literature which conclusively supports these claims, with half of all studies being U.S.-focused. Here, we provide an overview of the global state of research on polarization, highlighting insights that are robust across countries, those unique to specific contexts, and key gaps in the literature. We argue that addressing these gaps is urgent, but has been hindered thus far by systemic and cultural barriers, such as regionally stratified restrictions on data access and misaligned research incentives. If continued cross-disciplinary inertia means that these disparities are left unaddressed, we see a substantial risk that countries will adopt policies to tackle polarization based on inappropriate evidence, risking flawed decision-making and the weakening of democratic institutions.
Persuasion is a powerful capability of large language models (LLMs) that both enables beneficial applications (e.g. helping people quit smoking) and raises significant risks (e.g. large-scale, targeted political manipulation). Prior work has found models possess a significant and growing persuasive capability, measured by belief changes in simulated or real users. However, these benchmarks overlook a crucial risk factor: the propensity of a model to attempt to persuade in harmful contexts. Understanding whether a model will blindly ``follow orders'' to persuade on harmful topics (e.g. glorifying joining a terrorist group) is key to understanding the efficacy of safety guardrails. Moreover, understanding if and when a model will engage in persuasive behavior in pursuit of some goal is essential to understanding the risks from agentic AI systems. We propose the Attempt to Persuade Eval (APE) benchmark, that shifts the focus from persuasion success to persuasion attempts, operationalized as a model's willingness to generate content aimed at shaping beliefs or behavior. Our evaluation framework probes frontier LLMs using a multi-turn conversational setup between simulated persuader and persuadee agents. APE explores a diverse spectrum of topics including conspiracies, controversial issues, and non-controversially harmful content. We introduce an automated evaluator model to identify willingness to persuade and measure the frequency and context of persuasive attempts. We find that many open and closed-weight models are frequently willing to attempt persuasion on harmful topics and that jailbreaking can increase willingness to engage in such behavior. Our results highlight gaps in current safety guardrails and underscore the importance of evaluating willingness to persuade as a key dimension of LLM risk. APE is available at github.com/AlignmentResearch/AttemptPersuadeEval
Empirical social networks are characterized by a high degree of triadic closure (i.e., transitivity, clustering), whereby network neighbors of the same individual are also likely to be directly connected. It is unknown to what degree this results from dispositions to form such ties (i.e., to close open triangles) per se or from other processes, such as homophily and more opportunities for exposure. These are difficult to disentangle in many settings, but in social media not only can they be decomposed, but platforms frequently make decisions that depend on these distinct processes. Here, using a field experiment on social media, we randomize the existing network structure that a user faces when followed by a target account that we control, and we examine whether they reciprocate this tie formation. Being randomly assigned to have an existing tie to an account that follows the target user increases tie formation by 35%. Through the use of multiple control conditions in which the relevant tie is absent (never existent or removed), we attribute this effect specifically to a minimal cue that indicates the presence of a potential mutual follower. Theory suggests that triadic closure should be especially likely in open triads of strong ties, and we find larger effects when the subject has interacted more with the existing follower. These results indicate a substantial role for tendencies toward triadic closure, but one that is substantially smaller than what might be inferred from prior observational studies. Platforms and others may rely on these tendencies in encouraging tie formation, with broader implications for network structure and information diffusion in online networks.
There are widespread fears that conversational AI could soon exert unprecedented influence over human beliefs. Here, in three large-scale experiments (N=76,977), we deployed 19 LLMs-including some post-trained explicitly for persuasion-to evaluate their persuasiveness on 707 political issues. We then checked the factual accuracy of 466,769 resulting LLM claims. Contrary to popular concerns, we show that the persuasive power of current and near-future AI is likely to stem more from post-training and prompting methods-which boosted persuasiveness by as much as 51
Although conspiracy beliefs are often viewed as resistant to correction, recent evidence shows that personalized, fact-based dialogues with a large language model (LLM) can reduce them. Is this effect driven by the debunking facts and evidence, or does it rely on the messenger being an AI? In other words, would the same message be equally effective if delivered by a human? To answer this question, we conducted a preregistered experiment (N = 955) in which participants reported either a conspiracy belief or a nonconspiratorial but epistemically unwarranted belief and interacted with a LLM that argued against that belief using facts and evidence. We randomized whether the debunking LLM was characterized as an AI tool or a human expert and whether the model used human-like conversational tone. The conversations significantly reduced participants' confidence in both conspiracies and epistemically unwarranted beliefs, with no significant differences across conditions. Thus, AI persuasion is not reliant on the messenger being an AI model: it succeeds by generating compelling messages.