Pursuing replicability - independent evidence for previous claims - is important for creating generalizable knowledge(1,2). Here we attempted replications of 274 claims of positive results from 164 quantitative papers published from 2009 to 2018 in 54 journals in the social and behavioural sciences. Replications were high powered on average to detect the original effect size (median of 99.6%), used original materials when relevant and available, and were peer reviewed in advance through a standardized internal protocol. Replications showed statistically significant results in the original pattern for 151 of 274 claims (55.1% (95% confidence interval (CI) 49.2-60.9%)) and for 80.8 of 164 papers (49.3% (95% CI 43.8-54.7%)), weighed for replicating multiple claims per paper. We observed modest variation in replication rates across disciplines (42.5-63.1%), although some estimates had high uncertainty. The median Pearson's r effect size was 0.25 (95% CI 0.21-0.27) for original studies and 0.10 (95% CI 0.09-0.13) for replication studies, an 82.4% (95% CI 67.8-88.2%) reduction in shared variance. Thirteen methods for evaluating replication success provided estimates ranging from 28.6% to 74.8% (median of 49.3%). Some decline in effect size and significance is expected based on power to detect original effects and regression to the mean because we replicated only positive results. We observe that challenges for replicability extend across social-behavioural sciences, illustrating the importance of identifying conditions that promote or inhibit replicability(3,4).
We evaluate an intervention designed to give lab-based research teams an opportunity to intentionally discuss project-related data management practices within their labs, examining how such communication might inform perceptions of the relationship between formal data management plans and lab members’ day-to-day data management practices. In an earlier study, we developed a lab-based intervention that encouraged a deliberative approach to discussions among lab members regarding the practices of data management and authorship, and an exploration of the ethical dimensions of those practices. This present study builds on this prior work, both as partial replication and extension. Here we show the significant effects of the intervention across several dimensions, but importantly and specific to this project, this deliberative communication approach enhances the likelihood that lab members share an understanding of and a commitment to its data management practices, in part because they have been actively involved in shaping those practices. Fostering shared understanding and commitment is crucial for maintaining rigor and responsibility across all aspects of a lab’s work, and is essential for cultivating legitimate, defensible, and ethical approaches to producing scientific knowledge.
The Credibility Revolution advances internally valid research designs intended to identify causal effects from quantitative data. The ensuing emphasis on internal validity, however, has enabled a neglect of construct and external validity. We show that ignoring construct and external validity within identification strategies undermines the Credibility Revolution’s own goal of understanding causality deductively. Without assumptions regarding construct validity, one cannot accurately label the cause or outcome. Without assumptions regarding external validity, one cannot label the conditions enabling the cause to have an effect. If any of the assumptions regarding internal, construct, and external validity are missing, the claim is not deductively supported. The critical role of theoretical and substantive knowledge in deductive causal inference is illuminated by making such assumptions explicit. This article critically reviews approaches to identification in causal inference while developing a framework called causal specification. Causal specification augments existing identification strategies to enable and justify deductive, generalized claims about causes and effects. In the process, we review a variety of developments in the philosophy of science and causality and interdisciplinary social science methodology.
How can we effectively model arguments communicated in diverse environments? On the one hand, there is a great opportunity with the abundance of digitized speech across different contexts including online forums, official proceedings, or transcripts of spoken debates. On the other hand, there is a great challenge in correctly detecting arguments, especially since each medium has its own set of conventions, lingo, affordances, and styles of argumentative engagement. We propose WIBA, a novel framework and suite of methods that enable the comprehensive understanding of “What Is Being Argued” across contexts. Our approach develops a comprehensive framework that detects: (a) the existence, (b) the topic, and (c) the stance of an argument, correctly accounting for the logical dependence among the three tasks. Our algorithm leverages the fine-tuning and prompt-engineering of Large Language Models. We evaluate our approach and show that it performs well in all the three capabilities. First, we develop and release an Argument Detection model that can classify a piece of text as an argument with an F_1 score between 79 F_1 score between 71
How can we capture the dynamics of deliberation in a debate? In an increasingly divided and misinformed world, understanding the relationship between who is arguing and what they are arguing about is becoming critical for fostering a meaningful exchange of ideas. Given the vast array of available platforms for people to express their viewpoints and deliberate on issues, how can we develop methods to accurately analyze these processes? Luckily, there is an abundance of debate data available, ranging from: (a) formal proceedings, such as committee hearings in legislatures, to (b) online discussion forums, such as Reddit. Here we introduce DALiSM, a data-driven argument-centric framework, to analyze discourse dynamics in diverse and multi-party spaces at scale. We develop methods to harness and extend the state-of-the-art in computational argumentation for: (a) identifying arguments from long-form raw texts, (b) calculating the intensity of deliberation, and (c) modeling the evolution of discourse over time. We deploy our framework as a comprehensive and interactive dashboard for dynamically viewing the outputs of DALiSM to clearly understand the nature of a discourse event. To showcase the importance and utility of DALiSM, we apply our framework to U.S. congressional committee hearings from 2005 to 2023 (109th to 117th Congresses), and to selected Reddit communities from 2008 to 2023. This case study reveals substantive insights into deliberative behavior in online and offline spaces.
Public deliberation grows increasingly prevalent yet remains costly in terms of money and time. Accordingly, some suggest supplanting talk-based practices with individual, "deliberation within." Yet we have little evidence either way on the additional benefits of public deliberation over its individual variant. We evaluate the benefits of public deliberation with a field experiment. With the cooperation of two sitting US Senators, we recruited several hundred of their constituents to deliberate on immigration reform. Participants were randomly assigned to either deliberate publicly in an online discussion, to deliberate individually, or to an information-only control. Across several measures, public deliberation yielded more benefits than individual deliberation. We find, moreover, little evidence to ground worries that differences in education, race, conflict avoidance, gender, or gender composition of deliberating groups will render public talk less valuable than individual deliberation.
The social media platforms of the twenty-first century have an enormous role in regulating speech in the USA and worldwide1. However, there has been little research on platform-wide interventions on speech2,3. Here we evaluate the effect of the decision by Twitter to suddenly deplatform 70,000 misinformation traffickers in response to the violence at the US Capitol on 6 January 2021 (a series of events commonly known as and referred to here as ‘January 6th’). Using a panel of more than 500,000 active Twitter users4,5 and natural experimental designs6,7, we evaluate the effects of this intervention on the circulation of misinformation on Twitter. We show that the intervention reduced circulation of misinformation by the deplatformed users as well as by those who followed the deplatformed users, though we cannot identify the magnitude of the causal estimates owing to the co-occurrence of the deplatforming intervention with the events surrounding January 6th. We also find that many of the misinformation traffickers who were not deplatformed left Twitter following the intervention. The results inform the historical record surrounding the insurrection, a momentous event in US history, and indicate the capacity of social media platforms to control the circulation of misinformation, and more generally to regulate public discourse. Difference-in-differences analysis indicates that the decision by Twitter to deplatform 70,000 users following the events at the US Capitol on 6 January 2021 had wider effects on the spread of misinformation.
This article reviews and summarizes current reproduction and replication practices in political science. We first provide definitions for reproducibility and replicability. We then review data availability policies for 28 leading political science journals and present the results from a survey of editors about their willingness to publish comments and replications. We discuss new initiatives that seek to promote and generate high-quality reproductions and replications. Finally, we make the case for standards and practices that may help increase data availability, reproducibility, and replicability in political science.
How can we model arguments and their dynamics in online forum discussions? The meteoric rise of online forums presents researchers across different disciplines with an unprecedented opportunity: we have access to texts containing discourse between groups of users generated in a voluntary and organic fashion. Most prior work so far has focused on classifying individual monological comments as either argumentative or not argumentative. However, few efforts quantify and describe the dialogical processes between users found in online forum discourse: the structure and content of interpersonal argumentation. Modeling dialogical discourse requires the ability to identify the presence of arguments, group them into clusters, and summarize the content and nature of clusters of arguments within a discussion thread in the forum. In this work, we develop ArguSense, a comprehensive and systematic framework for understanding arguments and debate in online forums. Our framework consists of methods for, among other things: (a) detecting argument topics in an unsupervised manner; (b) describing the structure of arguments within threads with powerful visualizations; and (c) quantifying the content and diversity of threads using argument similarity and clustering algorithms. We showcase our approach by analyzing the discussions of four communities on the Reddit platform over a span of 21 months. Specifically, we analyze the structure and content of threads related to GMOs in forums related to agriculture or farming to demonstrate the value of our framework.
Public deliberation grows increasingly prevalent yet remains costly in terms of money and time. Accordingly, some suggest supplanting talk-based practices with individual, “deliberation within.” Yet we have little evidence either way on the additional benefits of public deliberation over its individual variant. We evaluate the benefits of public deliberation with a field experiment. With the cooperation of two sitting US Senators, we recruited several hundred of their constituents to deliberate on immigration reform. Participants were randomly assigned to either deliberate publicly in an online discussion, to deliberate individually, or to an information-only control. Across several measures, public deliberation yielded more benefits than individual deliberation. We find, moreover, little evidence to ground worries that differences in education, race, conflict avoidance, gender, or gender composition of deliberating groups will render public talk less valuable than individual deliberation.
Theoretical expectations regarding communication patterns between legislators and outside agents, such as lobbyists, agency officials or policy experts, often depend on the relationship between legislators' and agents' preferences. However, legislators and non-elected outside agents evaluate the merits of policies using distinct criteria and considerations. We develop a measurement method that flexibly estimates the policy preferences for a class of outside agents -- witnesses in committee hearings -- separate from that of legislators’ and compute their preference distance across the two dimensions. In our application to Medicare hearings, we find that legislators in the U.S. Congress heavily condition their questioning of witnesses on preference distance, showing that legislators tend to seek policy information from like-minded experts in committee hearings. We do not find this result using a conventional measurement placing both actors on one dimension. The contrast in results lends support for the construct validity of our proposed preference measures.
Canonical theories of democratic representation envision legislators cultivating familiarity to enhance esteem among their constituents. Some scholars, however, argue that familiarity breeds contempt, which if true would undermine incentives for effective representation. Survey respondents who are unfamiliar with their legislator tend not to provide substantive answers to attitude questions, and so we are missing key evidence necessary to adjudicate this important debate. We solve this problem with a randomized field experiment that gave some constituents an opportunity to gain familiarity with their Member of Congress through an online Deliberative Town Hall. Relative to controls, respondents who interacted with their member reported higher esteem as a result of enhanced familiarity, a mediation effect supporting canonical theories of representation. This effect is statistically significant among constituents who are the same political party as the member but not among those of the opposite party, although in neither case did familiarity breed contempt.
How can we comprehensively understand the main concerns and beliefs of the GMO debate in online forums? Genetically Modified Organisms (GMOs) have historically been a hotly debated topic, both within and outside of the agriculture industry. Understanding the complexity of these beliefs can lend policy makers the knowledge necessary to counteract misinformation. In this paper we develop Forumlyze, a systematic framework to understand user beliefs in online discourse surrounding an issue. As a case study, we focus on data collected from Reddit between 2019–2020 from four sub-forums: farming, agriculture, horticulture, and vegetable gardening. In our approach we (a) illustrate the fundamental and temporal characteristics of the issue (b) extract and characterize sentiments surrounding the issue (c) uncover the dominate concepts prevalent in this discussion and the context surrounding these concepts. The comprehensive nature of this analysis led to the following results. (1) The dominant concepts surrounding GMOs are Climate Change, Monsanto and Soil Science. (2) The sentiment of discourse around GMOs and its related concepts indicates a polarized affective system. (3) Evidence that real-world events impact online forum communities' sentiment surrounding GMOs-related concepts.
Congress Overwhelmed: The Decline in Congressional Capacity and Prospects for Reform. Edited by Timothy M. LaPira, Lee Drutman, and Kevin R. Kosar. Chicago: University of Chicago Press, 2020. 352p. 35.00 paper. - Volume 20 Issue 3
How are the sentiment and stance of online users affected by real-world events? Previous studies have ignored the role of events in co-determining sentiment and stance and hence have failed to understand the relationship between these two important aspects of public opinion. In this paper, we develop SentiStance, a systematic framework to understand the intertwined change of sentiment and stance due to real-world events in online discussions. In our approach: (a) we customize state-of-the-art NLP techniques to overcome domain-specific constraints, and (b) we provide an efficient way to quantify the change of sentiment and stance in tandem. As a case study, we focus on the 2020 United States Election events and we analyze 7.5 million posts from 4chan, Reddit, and Parler over a span of three months from November 2020 to January 2021. We showcase our framework by describing the effect that the Jan 6 insurrection had on concepts "Pence" and "Trump." Parler users turn significantly against Pence with (33.1% increase in Against stance and Negative sentiment), while Reddit users' opinion improves (with a drop of 7.1% in the same combination of sentiment and stance). By contrast, the effect of the same event on the concept "Trump" shows no statistically significant change. In addition, our results suggest that conditioning on significant events strengthens the correlation between sentiment and stance, which provides a new perspective on the debate around the correlation between sentiment and stance. Overall, we see our work as a fundamental building block towards a data-driven understanding of the interplay of preferences and emotions of online forum users towards a concept.
Theoretical expectations regarding the legislative influence of outside agents, such as lobbyists, agency officials or policy experts, often depend on the relationship between legislators' and agents' preferences. Non-elected agents, however, typically have preferences defined on a dimension that is different from that of elected legislators. I develop a bridging method that accommodates a shift and rotation of agent preferences relative to the legislator roll-call preference space, and that identifies distances across the two dimensions in units necessary to test institutional hypotheses. I provide both a theoretical and a computational proof of identification, and a simulation to show the model exactly recovers benchmark parameters. In my application to Medicare hearings, I show that the agent preference space has an orthogonal rotation anchored by a quality-cost latent dimension, and that legislators heavily condition their questioning of agents on preference distance in a way consistent with ``cheap talk'' informational models of strategic lobbying.
Given an online forum, how can we quantify changes in user affect towards a person or an idea over time? We argue that online political forums constitute an untapped opportunity for understanding sentiment toward aspects under discussion. However, the analysis of such forums has received little attention from the research community. In this paper, we develop RAFFMAN, a systematic approach to quantify the impact of external events on the affect of forum users towards a concept, such as a person or an entity. First, we develop an approach to capture and quantify the observed activity: we identify related keywords, filter threads, and establish correlations between events and spikes in the activity. Second, we modify and evaluate state-of-the-art NLP techniques to achieve high accuracy (74%) in a three-class sentiment classification problem. As a case study, we deploy our method to quantify the effect of President Trump’s impeachment on several concepts including: President Trump, Speaker Pelosi, and QAnon. Our data consists of 32M posts from Reddit and 4chan over a span of 6 months from September 2019 to February 2020. This initial analysis hints at an increase in political polarization, especially for people’s affect towards the President. Overall, our work is a building block towards mining the affect of online forum user towards a concept, which constitutes a untapped, massive, and publicly-available source of information.
The causal mediation literature has developed techniques to assess the sensitivity of an inference to pretreatment confounding, but these techniques are limited to the case of a single mediator. In this article, we extend sensitivity analysis to possible violations of pretreatment confounding in the case of multiple mediators. In particular, we develop sensitivity analyses under three alternative approaches to effect decomposition: (1) jointly considered mediators, (2) identifiable direct and indirect paths, and (3) interventional analogues effects. With reasonable assumptions, each approach reduces to a single procedure to assess sensitivity in the presence of simultaneous pre- and posttreatment confounding. We demonstrate our sensitivity analysis techniques with a framing experiment that examines whether anxiety mediates respondents' attitudes toward immigration in response to an information prompt.
By advancing causal identification, the credibility revolution has facilitated tremendous progress in political science. The ensuing emphasis on internal validity however has led to the neglect of construct and external validity. This article develops a framework we call causal specification. The framework formally demonstrates the joint necessity of internal, construct and external validity for causal generalization. Indeed, the lack of any of the three types of validity undermines the credibility revolution's own goal to understand causality deductively. Without construct, one cannot accurately label the cause or outcome. Without external, one cannot understand the conditions enabling the cause to have an effect. Our framework clarifies the assumptions necessary for each of the three. We show why internal validity should not have lexical priority, and advocate for equally valuing construct and external validity. Political scientists should use causal specification, not just causal identification, as a framework for the goal of causal generalization.