We consider inference for parameters of the form θ_0 = E[F_Y^-1∘ F_Z(X)] for some variables X, Y and Z. Such parameters appear, in particular, in the “changes-in-changes” model of . We first establish that , a plug-in estimator of θ_0, is root-n consistent and asymptotically normal under weaker conditions than those previously available, allowing in particular for unbounded variables. Next, we propose a new estimator of the asymptotic variance of and show its consistency, also allowing for unbounded variables. Monte Carlo simulations suggest that the conditions for root-n consistency and asymptotic normality are, in some sense, minimal. These simulations highlight that our variance estimator also leads to more accurate inference than some alternative approaches.
This paper studies analytic inference along two dimensions of clustering. In such setups, the commonly used approach has two drawbacks. First, the corresponding variance estimator is not necessarily positive. Second, inference is invalid in non-Gaussian regimes, namely when the estimator of the parameter of interest is not asymptotically Gaussian. We consider a simple fix that addresses both issues. In Gaussian regimes, the corresponding tests are asymptotically exact and equivalent to usual ones. Otherwise, the new tests are asymptotically conservative. We also establish their uniform validity over a certain class of data generating processes. Independently of our tests, we highlight potential issues with multiple testing and nonlinear estimators under two-way clustering. Finally, we compare our approach with existing ones through simulations.
The command did_multiplegt_dyn can be used to estimate event-study effects in complex designs with a potentially non-binary and/or non-absorbing treatment. This paper starts by providing an overview of the estimators computed by the command. Then, simulations based on three real datasets are used to demonstrate the estimators' properties. Finally, the command is used on four real datasets to estimate event-study effects in complex designs. The first example has a binary treatment that can turn on an off. The second example has a continuous absorbing treatment. The third example has a discrete multivalued treatment that can increase or decrease multiple times over time. The fourth example has two, binary and absorbing treatments, where the second treatment always happens after the first.
We study linear regressions in a context where the outcome of interest and some of the covariates are observed in two different datasets that cannot be matched. Traditional approaches obtain point identification by relying, often implicitly, on exclusion restrictions. We show that without such restrictions, coefficients of interest can still be partially identified, with the sharp bounds taking a simple form. We obtain tighter bounds when variables observed in both datasets, but not included in the regression of interest, are available, even if these variables are not subject to specific restrictions. We develop computationally simple and asymptotically normal estimators of the bounds. Finally, we apply our methodology to estimate racial disparities in patent approval rates and to evaluate the effect of patience and risk-taking on educational performance.
This paper considers the identification of dynamic treatment effects with panel data, in complex designs where the treatment may not be binary and may not be absorbing. We first show that under no-anticipation and parallel-trends assumptions, we can identify event-study effects comparing outcomes under the actual treatment path and under the status-quo path where all units would have kept their period-one treatment throughout the panel. Those effects can be helpful to evaluate ex-post the policies that effectively took place, and once properly normalized they estimate weighted averages of marginal effects of the current and lagged treatments on the outcome. Yet, they may still be hard to interpret, and they cannot be used to evaluate the effects of other policies than the ones that were conducted. To make progress, we impose another restriction, namely a random coefficients distributed-lag linear model, where effects remain constant over time. Under this model, the usual distributed-lag two-way-fixed-effects regression may be misleading. Instead, we show that this random coefficients model can be estimated simply. We illustrate our findings by revisiting Gentzkow et al. (2011).
Assume that an estimator is asymptotically normal for a target parameter under some conditions. Suppose also that one can test these conditions, and one conducts inference for the target only if the pre-test is not rejected. Does such pre-testing undermine inference? We show that if the tested conditions and mild regularity restrictions hold, conditional inference is still valid, albeit typically conservative. Validity holds regardless of the asymptotic dependence between the estimator and the pre-test. If the tested conditions do not hold, we exhibit conditions under which confidence intervals have larger conditional than unconditional coverage.
We consider treatment-effect estimation with a two-periods panel, where units are untreated at period one, and receive strictly positive doses at period two. First, we consider designs with some quasi-untreated units, with a period-two dose local to zero. We show that under a parallel-trends assumption, a weighted average of slopes of units' potential outcomes is identified by a difference-in-difference estimand using quasi-untreated units as the control group. We leverage results from the regression-discontinuity-design literature to propose a nonparametric estimator. Then, we propose estimators for designs without quasi-untreated units. Finally, we propose a test of the homogeneous-effect assumption underlying two-way-fixed-effects regressions.
Many treatments or policy interventions are continuous in nature. Examples include prices, taxes or temperatures. Empirical researchers have usually relied on two-way fixed effect regressions to estimate treatment effects in such cases. However, such estimators are not robust to heterogeneous treatment effects in general; they also rely on the linearity of treatment effects. We propose estimators for continuous treatments that do not impose those restrictions, and that can be used when there are no stayers: the treatment of all units changes from one period to the next. We start by extending the nonparametric results of de Chaisemartin et al. (2023) to cases without stayers. We also present a parametric estimator, and use it to revisit Deschênes and Greenstone (2012).
We study partially linear models when the outcome of interest and some of the covariates are observed in two different datasets that cannot be linked. This type of data combination problem arises very frequently in empirical microeconomics. Using recent tools from optimal transport theory, we derive a constructive characterization of the sharp identified set. We then build on this result and develop a novel inference method that exploits the specific geometric properties of the identified set. Our method exhibits good performances in finite samples, while remaining very tractable. We apply our approach to study intergenerational income mobility over the period 1850-1930 in the United States. Our method allows us to relax the exclusion restrictions used in earlier work, while delivering confidence regions that are informative.
We develop a new permutation test for inference on a subvector of coefficients in linear models. The test is exact when the regressors and the error terms are independent. Then, we show that the test is asymptotically of correct level, consistent and has power against local alternatives when the independence condition is relaxed, under two main conditions. The first is a slight reinforcement of the usual absence of correlation between the regressors and the error term. The second is that the number of strata, defined by values of the regressors not involved in the subvector test, is small compared to the sample size. The latter implies that the vector of nuisance regressors is discrete. Simulations and empirical illustrations suggest that the test has good power in practice if, indeed, the number of strata is small compared to the sample size.
Many real-world data sets can be presented in the form of a matrix whose entries correspond to the interaction between two entities of different natures (number of times a web user visits a web page, a student's grade in a subject, a patient's rating of a doctor, etc.). We assume in this paper that the mentioned interaction is determined by unobservable latent variables describing each entity. Our objective is to estimate the conditional expectation of the data matrix given the unobservable variables. This is presented as a problem of estimation of a bivariate function referred to as graphon. We study the cases of piecewise constant and H\"older-continuous graphons. We establish finite sample risk bounds for the least squares estimator and the exponentially weighted aggregate. These bounds highlight the dependence of the estimation error on the size of the data set, the maximum intensity of the interactions, and the level of noise. As the analyzed least-squares estimator is intractable, we propose an adaptation of Lloyd's alternating minimization algorithm to compute an approximation of the least-squares estimator. Finally, we present numerical experiments in order to illustrate the empirical performance of the graphon estimator on synthetic data sets.
Two-way fixed effects (TWFE) regressions with period and group fixed effects are widely used to estimate policies' effects: 26 of the 100 most cited papers published by the American Economic Review from 2015 to 2019 estimate such regressions. Researchers have long thought that TWFE estimators are equivalent to differences-in-differences (DID) estimators, that rely on a partly testable parallel trends assumption. In two-groups two-periods designs where a treatment group is untreated at both dates and a treatment group becomes treated at the second period, the treatment coefficient in a TWFE is indeed equivalent to a DID. Motivated by this fact, researchers have also estimated TWFE regressions in more complicated designs with many groups and periods, variation in treatment timing, treatments switching on and off, and/or non-binary treatments, confident that there as well, TWFE was giving them an estimation method that only relied on a partly testable parallel trends assumption. Two recent strands of literature have shattered that confidence. First, it has recently been shown that even if parallel trends holds, TWFE may produce misleading estimates, if the policy's effect is heterogeneous between groups or over time, as is often the case. The realization that one of the most commonly used empirical methods in the quantitative social sciences relies on an often-implausible assumption has spurred a flurry of methodological papers. Some of them have diagnosed this issue and analyzed its origins. Other papers have proposed alternative estimators relying on parallel trends conditions, like TWFE estimators, but robust to heterogeneous effects, unlike TWFE estimators. Hereafter, those alternative estimators are referred to as heterogeneity-robust DID estimators. Second, in a recent paper, Roth (2022) has shown that tests of the parallel trends assumption often lack statistical power, and may fail to detect differential trends between treated and control locations that are often large enough to account for a significant share of the policy's estimated effect. This realization has spurred a growing interest among practitioners for a second strand of literature, that has proposed alternative estimation methods relying on weaker assumptions than parallel trends. Examples include estimators relying on a conditional parallel trends assumption (see, e.g., Abadie, 2005), estimators assuming bounded differential trends (see, e.g., Manski and Pepper, 2018; Rambachan and Roth, 2023), estimators assuming a factor model with interactive fixed effects (see, e.g., Bai, 2003) and synthetic control estimators (see, e.g., Abadie et al., 2010), and estimators assuming grouped patterns of heterogeneity (see,e.g., Bonhomme and Manresa, 2015).This textbook aims to provide an overview of these two strands of literature, as well as other panel data methods routinely used for causal inference by practitionners.
We study two-way-fixed-effects regressions (TWFE) with several treatment variables. Under a parallel trends assumption, we show that the coefficient on each treatment identifies a weighted sum of that treatment's effect, with possibly negative weights, plus a weighted sum of the effects of the other treatments. Thus, those estimators are not robust to heterogeneous effects and may be contaminated by other treatments' effects. We further show that omitting a treatment from the regression can actually reduce the estimator's bias, unlike what would happen under constant treatment effects. We propose an alternative difference-in-differences estimator, robust to heterogeneous effects and immune to the contamination problem. In the application we consider, the TWFE regression identifies a highly non-convex combination of effects, with large contamination weights, and one of its coefficients significantly differs from our heterogeneity-robust estimator.
The synthetic control method is a an econometric tool to evaluate causal effects when only one unit is treated. While initially aimed at evaluating the effect of large-scale macroeconomic changes with very few available control units, it has increasingly been used in place of more well-known microeconometric tools in a broad range of applications, but its properties in this context are unknown. This paper introduces an alternative to the synthetic control method, which is developed both in the usual asymptotic framework and in the high-dimensional scenario. We propose an estimator of average treatment effect that is doubly robust, consistent and asymptotically normal. It is also immunized against first-step selection mistakes. We illustrate these properties using Monte Carlo simulations and applications to both standard and potentially high-dimensional settings, and offer a comparison with the synthetic control method.