We consider the problem of generating confidence sets in randomized experiments with noncompliance. We show that a refinement of a randomization-based procedure proposed by Imbens & Rosenbaum (2005) has desirable properties. Specifically, we show that using a studentized Anderson-Rubin statistic as a test statistic yields confidence sets that are finite-sample exact under treatment effect homogeneity and remain asymptotically valid for the local average treatment effect when the treatment effects are heterogeneous. We provide a uniform analysis of this procedure and efficient algorithms to construct the confidence sets.
Many traditional approaches to constructing confidence intervals for randomized experiments with binary outcomes are based on a binomial model for the outcome distribution. However, the assumptions underlying the binomial model are highly problematic in typical experimental designs [Robins 1988].
We consider design-based causal inference for spatial experiments in which treatments may have effects that bleed out and feed back in complex ways. Such spatial spillover effects violate the standard “no interference” assumption for standard causal inference methods. The complexity of spatial spillover effects also raises the risk of misspecification and bias in model-based analyses. We offer an approach for robust inference in such settings without having to specify a parametric outcome model. We define a spatial “average marginalized effect” (AME) that characterizes how, in expectation, units of observation that are a specified distance from an intervention location are affected by treatment at that location, averaging over effects emanating from other intervention nodes. We show that randomization is sufficient for non-parametric identification of the AME even if the nature of interference is unknown. Under mild restrictions on the extent of interference, we establish asymptotic distributions of estimators and provide methods for both sample-theoretic and randomization-based inference. We show conditions under which the AME recovers a structural effect. We illustrate our approach with a simulation study. Then we re-analyze a randomized field experiment and a quasi-experiment on forest conservation, showing how our approach offers robust inference on policy-relevant spillover effects.
We argue that randomized controlled trials (RCTs) are special even among settings where average treatment effects are identified by a nonparametric unconfoundedness assumption. This claim follows from two results of Robins and Ritov (1997): (1) with at least one continuous covariate control, no estimator of the average treatment effect exists which is uniformly consistent without further assumptions, (2) knowledge of the propensity score yields a uniformly consistent estimator and honest confidence intervals that shrink at parametric rates with increasing sample size, regardless of how complicated the propensity score function is. We emphasize the latter point, and note that successfully-conducted RCTs provide knowledge of the propensity score to the researcher. We discuss modern developments in covariate adjustment for RCTs, noting that statistical models and machine learning methods can be used to improve efficiency while preserving finite sample unbiasedness. We conclude that statistical inference has the potential to be fundamentally more difficult in observational settings than it is in RCTs, even when all confounders are measured.
From the social sciences to machine learning, it has been well documented that metrics to be optimized are not always aligned with social welfare. In healthcare, Dranove et al. (2003) showed that publishing surgery mortality metrics actually harmed the welfare of sicker patients by increasing provider selection behavior. We analyze the incentive misalignments that arise from such average treated outcome metrics, and show that the incentives driving treatment decisions would align with maximizing total patient welfare if the metrics (i) accounted for counterfactual untreated outcomes and (ii) considered total welfare instead of averaging over treated patients. Operationalizing this, we show how counterfactual metrics can be modified to behave reasonably in patient-facing ranking systems. Extending to realistic settings when providers observe more about patients than the regulatory agencies do, we bound the decay in performance by the degree of information asymmetry between principal and agent. In doing so, our model connects principal-agent information asymmetry with unobserved heterogeneity in causal inference.
Quantitative empirical inquiry in international relations often relies on dyadic data. Standard analytic techniques do not account for the fact that dyads are not generally independent of one another. That is, when dyads share a constituent member (e.g., a common country), they may be statistically dependent, or “clustered.” Recent work has developed dyadic clustering robust standard errors (DCRSEs) that account for this dependence. Using these DCRSEs, we reanalyzed all empirical articles published in International Organization between January 2014 and January 2020 that feature dyadic data. We find that published standard errors for key explanatory variables are, on average, approximately half as large as DCRSEs, suggesting that dyadic clustering is leading researchers to severely underestimate uncertainty. However, most (67% of) statistically significant findings remain statistically significant when using DCRSEs. We conclude that accounting for dyadic clustering is both important and feasible, and offer software in R and Stata to facilitate use of DCRSEs in future research.
Freedman (2008a,b) showed that the linear regression estimator is biased for the analysis of randomized controlled trials under the randomization model. Under Freedman's assumptions, we derive exact closed-form bias corrections for the linear regression estimator. We show that the limiting distribution of the bias corrected estimator is identical to the uncorrected estimator. Taken together with results from Lin (2013), our results show that Freedman's theoretical arguments against the use of regression adjustment can be resolved with minor modifications to practice.
Kevin Munger argues that, when an agnostic approach is applied to social scientific inquiry, the goal of prediction to new settings is generically impossible. We aim to situate Munger’s critique in a broader scientific and philosophical literature and to point to ways in which gnosis can and, in some circumstances, must be used to facilitate the accumulation of knowledge. We question some of the premises of Munger’s arguments, such as the definition of statistical agnosticism and the characterization of knowledge. We further emphasize the important role of microfoundations and particularism in the social sciences. We assert that Munger’s conclusions may be overly pessimistic as they relate to practice in the field.
The regression discontinuity (RD) design offers identification of causal effects under weak assumptions, earning it a position as a standard method in modern political science research. But identification does not necessarily imply that causal effects can be estimated accurately with limited data. In this paper, we highlight that estimation under the RD design involves serious statistical challenges and investigate how these challenges manifest themselves in the empirical literature in political science. We collect all RD-based findings published in top political science journals in the period 2009–2018. The distribution of published results exhibits pathological features; estimates tend to bunch just above the conventional level of statistical significance. A reanalysis of all studies with available data suggests that researcher discretion is not a major driver of these features. However, researchers tend to use inappropriate methods for inference, rendering standard errors artificially small. A retrospective power analysis reveals that most of these studies were underpowered to detect all but large effects. The issues we uncover, combined with well-documented selection pressures in academic publishing, cause concern that many published findings using the RD design may be exaggerated.
We consider the properties of listwise deletion when both n and the number of variables grow large. We show that when (i) all data has some idiosyncratic missingness and (ii) the number of variables grows superlogarithmically in n, then, for large n, listwise deletion will drop all rows with probability 1. Using two canonical datasets from the study of comparative politics and international relations, we provide numerical illustration that these problems may emerge in real world settings. These results suggest, in practice, using listwise deletion may mean using few of the variables available to the researcher.
This paper presents methods for analyzing spatial experiments when complex spillovers, displacement effects, and other types of “interference” are present. We present a robust, design-based approach to analyzing effects in such settings. The design-based approach derives inferential properties for causal effect estimators from known features of the experimental design, in a manner analogous to inference in sample surveys. The methods presented here target a quantity of interest called the “average marginalized response,” which is equal to the average effect of activating a treatment at an intervention node that is a given distance away, averaging ambient effects emanating from other intervention nodes. We provide a step-by-step tutorial based on the SpatialEffect package for R. We apply the methods to a randomized experiment on payments for community forest conservation in Uganda, showing how our methods reveal possibly substantial spatial spillovers that more conventional analyses cannot detect.
Advances in machine learning have made possible “deepfakes,” or realistic, computer-generated videos of public figures saying something they have not actually said. Policymakers have expressed concern that deepfakes could mislead voters, but prior research has found that such videos have minimal effects. There has nevertheless been extensive media coverage of the dangers of deepfakes, urging voters to be critical consumers of political videos. We explore whether these well-intentioned messages have an unintended consequence: if voters are warned about deepfakes, they may begin to distrust all political videos. We conducted two online survey experiments, and found that informing participants about deepfakes did not enhance participants’ ability to successfully spot manipulated videos but consistently induced them to believe the videos they watched were fake, even when they were real. Our findings suggest that even if deepfakes are not themselves persuasive, information about deepfakes can nevertheless be weaponized to dismiss real political videos.
We investigate large-sample properties of treatment effect estimators under unknown interference in randomized experiments. The inferential target is a generalization of the average treatment effect estimand that marginalizes over potential spillover effects. We show that estimators commonly used to estimate treatment effects under no interference are consistent for the generalized estimand for several common experimental designs under limited but otherwise arbitrary and unknown interference. The rates of convergence depend on the rate at which the amount of interference grows and the degree to which it aligns with dependencies in treatment assignment. Importantly for practitioners, the results imply that if one erroneously assumes that units do not interfere in a setting with limited, or even moderate, interference, standard estimators are nevertheless likely to be close to an average treatment effect if the sample is sufficiently large. Conventional confidence statements may, however, not be accurate.
Recent advances in machine learning have led to the development of the “deepfake,” a convincingly realistic, computer-generated video of a public figure saying something they have not actually said. Policymakers have expressed concern that deepfakes could mislead voters and affect election outcomes, but existing research has found minimal persuasive effects. In this paper, we explore a downstream consequence of deepfakes: if voters are repeatedly warned of the existence and dangers of deepfakes, they may simply begin to distrust all political video footage – whether real or fake. Through two online survey experiments, we found that voters were unable to discriminate between a real video and a deepfake. Statements warning about the existence of deepfakes did not enhance participants’ ability to successfully spot manipulated video content. Instead, these warnings consistently induced participants to believe that the videos they watched were fake, even when the videos were real. The warnings were not specific to the video participants were watching; simply stating that deepfakes exist increased distrust of any accompanying video. Our findings suggest that even if deepfakes are not themselves persuasive, rhetoric about deepfakes can nevertheless be weaponized by politicians and campaigns to dismiss and disown real videos.
In an influential critique of empirical practice, Freedman (2008) showed that the linear regression estimator was biased for the analysis of randomized controlled trials under the randomization model. Under Freedman's assumptions, we derive exact closed-form bias corrections for the linear regression estimator with and without treatment-by-covariate interactions. We show that the limiting distribution of the bias corrected estimator is identical to the uncorrected estimator, implying that the asymptotic gains from adjustment can be attained without introducing any risk of bias. Taken together with results from Lin (2013), our results show that Freedman's theoretical arguments against the use of regression adjustment can be completely resolved with minor modifications to practice.
Lucid has become increasingly popular as a low-cost provider of online survey responses. In this memo, we share our concerns about Lucid’s recent data quality. First, a large and increasing number of survey respondents are failing attention checks. Second, respondents who pass and respondents who fail attention checks are systematically different. Many respondents who fail standard attention checks appear to provide low-quality data. We conclude that researchers should exercise caution when analyzing data recently collected from Lucid unless respondents were subject to stringent attention checks.
Book review published as: Aronow, Peter M. and Fredrik Sävje (2020), "The Book of Why: The New Science of Cause and Effect." Journal of the American Statistical Association, 115: 482-485.
We present current methods for estimating treatment effects and spillover effects under "interference", a term which covers a broad class of situations in which a unit's outcome depends not only on treatments received by that unit, but also on treatments received by other units. To the extent that units react to each other, interact, or otherwise transmit effects of treatments, valid inference requires that we account for such interference, which is a departure from the traditional assumption that units' outcomes are affected only by their own treatment assignment. Interference and associated spillovers may be a nuisance or they may be of substantive interest to the researcher. In this chapter, we focus on interference in the context of randomized experiments. We review methods for when interference happens in a general network setting. We then consider the special case where interference is contained within a hierarchical structure. Finally, we discuss the relationship between interference and contagion. We use the interference R package and simulated data to illustrate key points. We consider efficient designs that allow for estimation of the treatment and spillover effects and discuss recent empirical studies that try to capture such effects.
We report a number of irregularities in the replication dataset posted for LaCour and Green (Science, "When contact changes minds: An experiment on transmission of support for gay equality," 2014) that jointly suggest the dataset (LaCour 2014) was not collected as described. These irregularities include baseline outcome data that is statistically indistinguishable from a national survey and over-time changes that are unusually small and indistinguishable from perfectly normally distributed noise. Other elements of the dataset are inconsistent with patterns typical in randomized experiments and survey responses and/or inconsistent with the claimed design of the study. A straightforward procedure may generate these anomalies nearly exactly: for both studies reported in the paper, a random sample of the 2012 Cooperative Campaign Analysis Project (CCAP) form the baseline data and normally distributed noise are added to simulate follow-up waves.
Dropping subjects based on the results of a manipulation check following treatment assignment is common practice across the social sciences, presumably to restrict estimates to a subpopulation of subjects who understand the experimental prompt. We show that this practice can lead to serious bias and argue for a focus on what is revealed without discarding subjects. Generalizing results developed in Zhang and Rubin (2003) and Lee (2009) to the case of multiple treatments, we provide sharp bounds for potential outcomes among those who would pass a manipulation check regardless of treatment assignment. These bounds may have large or infinite width, implying that this inferential target is often out of reach. As an application, we replicate Press, Sagan, and Valentino (2013) with a design that does not drop subjects that failed the manipulation check and show that the findings are likely stronger than originally reported. We conclude with suggestions for practice, namely alterations to the experimental design.