We review, from a practical standpoint, the evolving literature on assessing external validity (EV) of estimated treatment effects. We review existing EV measures, and focus on methods that permit multiple datasets (Hotz et al., 2005). We outline criteria for practical usage, evaluate the existing approaches, and identify a gap in potential methods. Our practical considerations motivate a novel method utilizing the Group Lasso (Yuan and Lin, 2006) to estimate a tractable regression-based model of the conditional average treatment effect (CATE). This approach can perform better when settings have differing covariate distributions and allows for easily extrapolating the average treatment effect to new settings. We apply these measures to a set of identical field experiments upgrading slum dwellings in three different countries (Galiani et al., 2017).
This paper uses synthetic controls to reevaluate the passage of Right-to-Work legislation in several states and its effect on union density levels in those states. Building upon recent work, we include data from several new legislative changes and also pool evidence across events to increase the inferential power for detecting a common effect. This adds to the literature by expanding the number of states investigated as well as allowing for more robust statistical testing on the impact of Right-to-Work. We estimate that modern Right-to-Work laws have a statistically significant effect and precipitated union density declines of about two to three percentage points.
In this article, we present statacons, an SCons-based build tool for Stata. Because of the integration of Stata and Python in recent versions of Stata, we are able to adapt SCons for Stata workflows without the use of an external shell or extensive configuration. We discuss the usefulness of build tools generally, provide examples of the use of statacons in Stata workflows, present key elements of the syntax of statacons, and discuss extensions, alternatives, and limitations. We provide recommendations for collaborative workflows and, at the end of the article, installation instructions.
Estimating heterogeneous treatment effects in domains such as healthcare or social science often involves sensitive data where protecting privacy is important. We introduce a general meta-algorithm for estimating conditional average treatment effects (CATE) with differential privacy (DP) guarantees. Our meta-algorithm can work with simple, single-stage CATE estimators such as S-learner and more complex multi-stage estimators such as DR and R-learner. We perform a tight privacy analysis by taking advantage of sample splitting in our meta-algorithm and the parallel composition property of differential privacy. In this paper, we implement our approach using DP-EBMs as the base learner. DP-EBMs are interpretable, high-accuracy models with privacy guarantees, which allow us to directly observe the impact of DP noise on the learned causal model. Our experiments show that multi-stage CATE estimators incur larger accuracy loss than single-stage CATE or ATE estimators and that most of the accuracy loss from differential privacy is due to an increase in variance, not biased estimates of treatment effects.
Restricting randomization in the design of experiments (e.g., using blocking/stratification, pair-wise matching, or rerandomization) can improve the treatment-control balance on important covariates and therefore improve the estimation of the treatment effect, particularly for small- and medium-sized experiments. Existing guidance on how to identify these variables and implement the restrictions is incomplete and conflicting. We identify that differences are mainly due to the fact that what is important in the pre-treatment data may not translate to the post-treatment data. We highlight settings where there is sufficient data to provide clear guidance and outline improved methods to mostly automate the process using modern machine learning (ML) techniques. We show in simulations using real-world data, that these methods reduce both the mean squared error of the estimate (14%-34%) and the size of the standard error (6%-16%).
The parallel package allows parallel processing of tasks that are not interdependent. This allows all flavors of Stata to take advantage of multiprocessor machines. Even Stata/MP users can benefit because many community-contributed programs are not automatically parallelized but could be under our framework.
When pre-processing observational data via matching, we seek to approximate each unit with maximally similar peers that had an alternative treatment status--essentially replicating a randomized block design. However, as one considers a growing number of continuous features, a curse of dimensionality applies making asymptotically valid inference impossible (Abadie and Imbens, 2006). The alternative of ignoring plausibly relevant features is certainly no better, and the resulting trade-off substantially limits the application of matching methods to "wide" datasets. Instead, Li and Fu (2017) recasts the problem of matching in a metric learning framework that maps features to a low-dimensional space that facilitates "closer matches" while still capturing important aspects of unit-level heterogeneity. However, that method lacks key theoretical guarantees and can produce inconsistent estimates in cases of heterogeneous treatment effects. Motivated by straightforward extension of existing results in the matching literature, we present alternative techniques that learn latent matching features through either MLPs or through siamese neural networks trained on a carefully selected loss function. We benchmark the resulting alternative methods in simulations as well as against two experimental data sets--including the canonical NSW worker training program data set--and find superior performance of the neural-net-based methods.
We introduce the Pricing Engine package to enable the use of Double ML estimation techniques in general panel data settings. Customization allows the user to specify first-stage models, first-stage featurization, second stage treatment selection and second stage causal-modeling. We also introduce a DynamicDML class that allows the user to generate dynamic treatment-aware forecasts at a range of leads and to understand how the forecasts will vary as a function of causally estimated treatment parameters. The Pricing Engine is built on Python 3.5 and can be run on an Azure ML Workbench environment with the addition of only a few Python packages. This note provides high-level discussion of the Double ML method, describes the packages intended use and includes an example Jupyter notebook demonstrating application to some publicly available data. Installation of the package and additional technical documentation is available at $\href{https://github.com/bquistorff/pricingengine}{github.com/bquistorff/pricingengine}$.
The synthetic control methodology (Abadie and Gardeazabal, 2003, American Economic Review 93: 113–132; Abadie, Diamond, and Hainmueller, 2010, Journal of the American Statistical Association 105: 493–505) allows for a data-driven approach to small-sample comparative studies. synth_runner automates the process of running multiple synthetic control estimations using synth. It conducts placebo estimates in space (estimations for the same treatment period but on all the control units). Inference ( p-values) is provided by comparing the estimated main effect with the distribution of placebo effects. It also allows several units to receive treatment, possibly at different time periods. It allows automatic generation of the outcome predictors and diagnostics by splitting the pretreatment into training and validation portions. Additionally, it provides diagnostics to assess fit and generates visualizations of results.
This paper analyzes a geographic quasi-experiment embedded in a cluster-randomized experiment in Honduras.In the experiment, average treatment effects on school enrollment and child labor were large-especially in the poorest blocks-and could be generalized to a policyrelevant population given the original sample selection criteria.In contrast, the geographic quasiexperiment yielded point estimates that, for two of three dependent variables, were attenuated.A judicious policy analyst without access to the experimental results might have provided misleading advice based on the magnitude of point estimates.We assessed two main explanations for the difference in point estimates, related to external and internal validity.
National and region capital cities are substantially larger than other cities in their countries and typically larger than would be predicted by Zipf's law. Though this empirical regularity is well established in the urban economics literature, causal estimates of the effect of capital designation on city size are hard to obtain. I use the 1960 relocation of the Brazilian capital from Rio de Janeiro to Brasília to identify the causal estimate of the effect of being a capital on city size, employment, and GDP. Using a synthetic controls strategy I find that losing the capital designation had no significant effects on Rio de Janeiro. In contrast, I find that Brasília experienced large and significant increases in population, employment and GDP after being designated the new Brazilian capital. Further, these increases are not entirely explained by the mechanical movement of government workers to Brasília; Instead, I find evidence of large spillovers in the private sector from this public sector shock. Finally, I show that some cities may violate the independence assumptions typically used when employing synthetic controls and I develop a method to deal with the resulting contamination of the donor units.
Low rates of adoption of and low willingness to pay for preventative health technologies pose an ongoing puzzle in development economics. In the case of water-borne disease, the burden is high both in terms of poor health and cost of treatment. Inexpensive preventative technologies are available, but willingness to pay (WTP) for products such as chlorine treatment or ceramic filters has been observed to be low in a number of contexts. In this paper, we investigate whether time payments (micro-loans or dedicated micro-savings) can increase WTP for a high-quality ceramic water filter among 400 households in slums of Dhaka, Bangladesh, where water quality is poor and the burden of water-borne disease high. We use a modified Becker-Degroot-Marschak mechanism to elicit WTP for the filter under a variety of payment plans. Crucially, we obtain valuations from each household across all payment plans, which (a) increases power and (b) allows us to investigate the mechanisms behind differences in WTP across plans. We find that time payments significantly increase WTP: compared to a lump-sum up-front purchase, median WTP increases 83% with a six-month loan and 115% with a 12-month loan. Similarly, coverage can be greatly increased: at an unsubsidized price (50% subsidy) coverage is 12% (27%) under a lump-sum but as high as 45% (71%) given time payments. We use our rich within-household WTP data, the design of the payment plans, and a simple structural model of time preference and credit constraints to investigate the mechanisms. We find that households are not impatient with respect to health goods and that therefore time-preferences do not contribute to low baseline WTP. We find strong evidence for the presence of credit constraints.