When assessing the causal effect of an exposure on two or more outcomes in an observational study, a linear combination of outcomes may lessen the sensitivity of a test of the global null hypothesis to potential unmeasured biases. While all linear combinations of scored outcomes can be considered using Scheffe projections or constrained variants thereof, finding the combination that minimizes sensitivity to unmeasured biases requires corrections for multiple testing, which can erode power, especially when many outcomes are of interest. To mitigate this issue, we propose splitting the sample into a planning sample to identify an optimal linear combination and an analysis sample to conduct inference. We provide a novel characterization of the set of linear combinations for which this approach is guaranteed to achieve the same asymptotic power as full-sample alternatives and conduct extensive simulation studies that demonstrate enhanced power in finite samples. Finally, we apply our method to investigate the effects of poverty on the emergence of cardiovascular disease risk factors in children and adolescents. We discover adverse consequences on outcomes related to body composition, physical activity, and tobacco exposure. Although the impact of poverty on elevated tobacco exposure shows some robustness to unmeasured confounding, the other findings remain sensitive to potential biases.
Sensitivity analysis asks how strong unmeasured confounding needs to be to explain away an observational study's conclusion. The conventional approach in matched studies conducts inference conditional upon the potential outcomes as well as both observed and unobserved confounders, and then finds the worst-case distribution for the conditional treatment assignments across all possible realizations of the unobserved confounder. The resulting worst-case allocation imagines strong, near perfect, correlations between the potential outcomes and hidden bias. We propose a stochastic sensitivity analysis that instead targets inference conditional upon potential outcomes and observed confounders while treating the hidden confounders as random with unknown conditional laws. Rather than finding the worst-case realizations for the hidden confounders, we instead determine the worst-case conditional law over a broad class of distributions. This preserves the adversarial spirit of sensitivity analysis while allowing for imperfect alignment between hidden bias and potential outcomes to a degree controlled by a scalar sensitivity parameter. We consider restrictions to both an interpretable class with no parametric assumptions and a Bernoulli class of conditional laws. Design sensitivity calculations and real-data demonstrations illustrate that allowing for even a small degree of stochasticity can materially increase reported robustness to hidden bias relative to the conventional approach.
We present a new procedure for conducting a sensitivity analysis in matched observational studies. For any candidate test statistic, the approach defines tilted modifications dependent upon the proposed strength of unmeasured confounding. The framework subsumes both (i) existing approaches to sensitivity analysis for sign-score statistics; and (ii) sensitivity analyses using conditional inverse probability weighting, wherein one weights the observed test statistic based upon the worst-case assignment probabilities for a proposed strength of hidden bias. Unlike the prevailing approach to sensitivity analysis after matching, there is a closed form expression for the limiting worst-case distribution when matching with multiple controls. Moreover, the approach admits a closed form for its design sensitivity, a measure used to compare competing test statistics and research designs, for matching with multiple controls, whereas the conventional approach generally only does so for pair matching. The tilted sensitivity analysis improves design sensitivity under a host of generative models. The proposal may also be adaptively combined with the conventional approach to attain a design sensitivity no smaller than the maximum of the individual design sensitivities. Data illustrations indicate that tilting can provide meaningful improvements in the reported robustness of matched observational studies.
In network settings, interference between units makes causal inference more challenging as outcomes may depend on the treatments received by others in the network. Typical estimands in network settings focus on treatment effects aggregated across individuals in the population. We propose a framework for estimating node-wise counterfactual means, allowing for more granular insights into the impact of network structure on treatment effect heterogeneity. We develop a doubly robust and non-parametric estimation procedure, KECENI (Kernel Estimator of Causal Effect under Network Interference), which offers consistency and asymptotic normality under network dependence. The utility of this method is demonstrated through an application to microfinance data, revealing the impact of network characteristics on treatment effects.
In randomized experiments, adjusting for observed features when estimating treatment effects has been proposed as a way to improve asymptotic efficiency. However, only linear regression has been proven to form an estimate of the average treatment effect that is asymptotically no less efficient than the treated-minus-control difference in means regardless of the true data generating process. Randomized treatment assignment provides this "do-no-harm" property, with neither truth of a linear model nor a generative model for the outcomes being required. We present a general calibration method which confers the same no-harm property onto estimators leveraging a broad class of nonlinear models. This recovers the usual regression-adjusted estimator when ordinary least squares is used, and further provides non-inferior treatment effect estimators using methods such as logistic and Poisson regression. The resulting estimators are non-inferior to both the difference in means estimator and to treatment effect estimators that have not undergone calibration. We show that our estimator is asymptotically equivalent to an inverse probability weighted estimator using a logit link with predicted potential outcomes as covariates. In a simulation study, we demonstrate that common nonlinear estimators without our calibration procedure may perform markedly worse than both the calibrated estimator and the unadjusted difference in means.
The debate surrounding an observational study rarely centers around the K covariates that have been accounted for through matching, weighting, or some other form of adjustment; rather, a critic's counterclaim typically rests upon the existence of a nettlesome https://www.w3.org/1998/Math/MathML"> K + 1 https://s3-euw1-ap-pe-df-pch-content-public-p.s3.eu-west-1.amazonaws.com/9781003102670/60801d0b-e7e6-4d6c-aa75-7056e29cb275/content/math25_1.tif" xmlns:xlink="https://www.w3.org/1999/xlink"/> st covariate for which the researcher did not account, an unobserved covariate whose relationship with treatment assignment and the outcome of interest may explain away the study's purported finding. Should the totality of evidence assume the absence of hidden bias, as is common with studies proceeding under the usual expedient of strong ignorability, the critic need merely suggest the existence of bias to cast doubt upon the posited causal mechanism. It is thus incumbent upon researchers not only to anticipate such criticism, but also to arm themselves with a suitable rejoinder.
We develop sensitivity analyses for the sample average treatment effect in matched observational studies while allowing unit-level treatment effects to vary. The methods may be applied to studies using any optimal without-replacement matching algorithm. In contrast to randomized experiments and to paired observational studies, we show for general matched designs that over a large class of test statistics, any procedure bounding the worst-case expectation while allowing for arbitrary effect heterogeneity must be unnecessarily conservative if treatment effects are actually constant across individuals. We present a sensitivity analysis which bounds the worst-case expectation while allowing for effect heterogeneity, and illustrate why it is generally conservative if effects are constant. An alternative procedure is presented that is asymptotically sharp if treatment effects are constant, and that is valid for testing the sample average effect under additional restrictions which may be deemed benign by practitioners. Simulations demonstrate that this alternative procedure results in a valid sensitivity analysis for the weak null hypothesis under a host of reasonable data-generating processes. The procedures allow practitioners to assess robustness of estimated sample average treatment effects to hidden bias while allowing for effect heterogeneity in matched observational studies.
In finite population causal inference exact randomization tests can be constructed for sharp null hypotheses, hypotheses which impute the missing potential outcomes. Oftentimes inference is instead desired for the weak null that the sample average of the treatment effects takes on a particular value while leaving the subject-specific treatment effects unspecified. Tests valid for sharp null hypotheses can be anti-conservative should only the weak null hold. We develop a general framework for unifying modes of inference for sharp and weak nulls, wherein a single procedure simultaneously delivers exact inference for sharp nulls and asymptotically valid inference for weak nulls. We employ randomization tests based upon prepivoted test statistics, wherein a test statistic is first transformed by a suitably constructed cumulative distribution function and its randomization distribution assuming the sharp null is then enumerated. For a large class of test statistics, we show that prepivoting may be accomplished by employing the push-forward of a sample-based Gaussian measure based upon a suitable covariance estimator. The approach enumerates the randomization distribution (assuming the sharp null) of a p-value for a large-sample test known to be valid under the weak null, and uses the resulting randomization distribution for inference. The versatility of the method is demonstrated through many examples, including rerandomized designs and regression-adjusted estimators in completely randomized designs.
We investigate the efficacy of surgical versus nonsurgical management for two gastrointestinal conditions, colitis and diverticulitis, using observational data. We deploy an instrumental variable design with surgeons’ tendencies to operate as an instrument. Assuming instrument validity, we find that nonsurgical alternatives can reduce both hospital length of stay and the risk of complications, with estimated effects larger for septic patients than for nonseptic patients. The validity of our instrument is plausible but not ironclad, necessitating a sensitivity analysis. Existing sensitivity analyses for IV designs assume effect homogeneity, unlikely to hold here because of patient-specific physiology. We develop a new sensitivity analysis that accommodates arbitrary effect heterogeneity and exploits components explainable by observed features. We find that the results for nonseptic patients prove more robust to hidden bias despite having smaller estimated effects. For nonseptic patients, two individuals with identical observed characteristics would have to differ in their odds of assignment to a high tendency to operate surgeon by a factor of 2.34 to overturn our finding of a benefit for nonsurgical management in reducing length of stay. For septic patients, this value is only 1.64. Simulations illustrate that this phenomenon may be explained by differences in within-group heterogeneity. Supplementary materials for this article, including a standardized description of the materials available for reproducing the work, are available as an online supplement.
In many observational studies, the interest is in the effect of treatment on bad, aberrant outcomes rather than the average outcome. For such settings, the traditional approach is to define a dichotomous outcome indicating aberration from a continuous score and use the Mantel-Haenszel test with matched data. For example, studies of determinants of poor child growth use the World Health Organization's definition of child stunting being height-for-age z-score <= - 2. The traditional approach may lose power because it discards potentially useful information about the severity of aberration. We develop an adaptive approach that makes use of this information and asymptotically dominates the traditional approach. We develop our approach in two parts. First, we develop an aberrant rank approach in matched observational studies and prove a novel design sensitivity formula enabling its asymptotic comparison with the Mantel-Haenszel test under various settings. Second, we develop a new, general adaptive approach, the two-stage programming method, and use it to adaptively combine the aberrant rank test and the Mantel-Haenszel test. We apply our approach to a study of the effect of teenage pregnancy on stunting.
We present a general approach to constructing permutation tests that are both exact for the null hypothesis of equality of distributions and asymptotically correct for testing equality of parameters of distributions while allowing the distributions themselves to differ. These robust permutation tests transform a given test statistic by a consistent estimator of its limiting distribution function before enumerating its permutation distribution. This transformation, known as prepivoting, aligns the unconditional limiting distribution for the test statistic with the probability limit of its permutation distribution. Through prepivoting, the tests permute one minus an asymptotically valid $p$-value for testing the null of equality of parameters. We describe two approaches for prepivoting within permutation tests, one directly using asymptotic normality and the other using the bootstrap. We further illustrate that permutation tests using bootstrap prepivoting can provide improvements to the order of the error in rejection probability relative to competing transformations when testing equality of parameters, while maintaining exactness under equality of distributions. Simulation studies highlight the versatility of the proposal, illustrating the restoration of asymptotic validity to a wide range of permutation tests conducted when only the parameters of distributions are equal.
We present a multivariate one-sided sensitivity analysis for matched observational studies, appropriate when the researcher has specified that a given causal mechanism should manifest itself in effects on multiple outcome variables in a known direction. The test statistic can be thought of as the solution to an adversarial game, where the researcher determines the best linear combination of test statistics to combat nature's presentation of the worst-case pattern of hidden bias. The corresponding optimization problem is convex, and can be solved efficiently even for reasonably sized observational studies. Asymptotically, the test statistic converges to a chi-bar-squared distribution under the null, a common distribution in order-restricted statistical inference. The test attains the largest possible design sensitivity over a class of coherent test statistics, and facilitates one-sided sensitivity analyses for individual outcome variables while maintaining familywise error control through its incorporation into closed testing procedures.
A fundamental limitation of causal inference in observational studies is that perceived evidence for an effect might instead be explained by factors not accounted for in the primary analysis. Methods for assessing the sensitivity of a study's conclusions to unmeasured confounding have been established under the assumption that the treatment effect is constant across all individuals. In the potential presence of unmeasured confounding, it has been argued that certain patterns of effect heterogeneity may conspire with unobserved covariates to render the performed sensitivity analysis inadequate. We present a new method for conducting a sensitivity analysis for the sample average treatment effect in the presence of effect heterogeneity in paired observational studies. Our recommended procedure, called the studentized sensitivity analysis, represents an extension of recent work on studentized permutation tests to the case of observational studies, where randomizations are no longer drawn uniformly. The method naturally extends conventional tests for the sample average treatment effect in paired experiments to the case of unknown, but bounded, probabilities of assignment to treatment. In so doing, we illustrate that concerns about certain sensitivity analyses operating under the presumption of constant effects are largely unwarranted.
BACKGROUND:Emerging resistance to anti-malarial drugs has led malaria researchers to investigate what covariates (parasite and host factors) are associated with resistance. In this regard, investigation of how covariates impact malaria parasites clearance is often performed using a two-stage approach in which the WWARN Parasite Clearance Estimator or PCE is used to estimate parasite clearance rates and then the estimated parasite clearance is regressed on the covariates. However, the recently developed Bayesian Clearance Estimator instead leads to more accurate results for hierarchial regression modelling which motivated the authors to implement the method as an R package, called "bhrcr".METHODS:Given malaria parasite clearance profiles of a set of patients, the "bhrcr" package performs Bayesian hierarchical regression to estimate malaria parasite clearance rates along with the effect of covariates on them in the presence of "lag" and "tail" phases. In particular, the model performs a linear regression of the log clearance rates on covariates to estimate the effects within a Bayesian hierarchical framework. All posterior inferences are obtained by a "Markov Chain Monte Carlo" based sampling scheme which forms the core of the package.RESULTS:The "bhrcr" package can be utilized to study malaria parasite clearance data, and specifically, how covariates affect parasite clearance rates. In addition to estimating the clearance rates and the impact of covariates on them, the "bhrcr" package provides tools to calculate the WWARN PCE estimates of the parasite clearance rates as well. The fitted Bayesian model to the clearance profile of each individual, as well as the WWARN PCE estimates, can also be plotted by this package.CONCLUSIONS:This paper explains the Bayesian Clearance Estimator for malaria researchers including describing the freely available software, thus making these methods accessible and practical for modelling covariates' effects on parasite clearance rates.
The conventional model for assessing insensitivity to hidden bias in paired observational studies constructs a worst-case distribution for treatment assignments subject to bounds on the maximal bias to which any given pair is subjected. In studies where rare cases of extreme hidden bias are suspected, the maximal bias may be substantially larger than the typical bias across pairs, such that a correctly specified bound on the maximal bias would yield an unduly pessimistic perception of the study's robustness to hidden bias. We present an extended sensitivity analysis which allows researchers to simultaneously bound the maximal and typical bias perturbing the pairs under investigation while maintaining the desired Type I error rate. We motivate and illustrate our method with two sibling studies on the impact of schooling on earnings, one containing information of cognitive ability of siblings and the other not. Cognitive ability, clearly influential of both earnings and degree of schooling, is likely similar between members of most sibling pairs yet could, conceivably, vary drastically for some siblings. The method is straightforward to implement, simply requiring the solution to a quadratic program.
Applied analysts often use the differences-in-differences (DID) method to estimate the causal effect of policy interventions with observational data. The method is widely used, as the required before and after comparison of a treated and control group is commonly encountered in practice. DID removes bias from unobserved time-invariant confounders. While DID removes bias from time-invariant confounders, bias from time-varying confounders may be present. Hence, like any observational comparison, DID studies remain susceptible to bias from hidden confounders. Here, we develop a method of sensitivity analysis that allows investigators to quantify the amount of bias necessary to change a study's conclusions. Our method operates within a matched design that removes bias from observed baseline covariates. We develop methods for both binary and continuous outcomes. We then apply our methods to two different empirical examples from the social sciences. In the first application, we study the effect of changes to disability payments in Germany. In the second, we re-examine whether election day registration increased turnout in Wisconsin.
In paired randomized experiments individuals in a given matched pair may differ on prognostically important covariates despite the best efforts of practitioners. We examine the use of regression adjustment as a way to correct for persistent covariate imbalances after randomization, and present two regression assisted estimators for the sample average treatment effect in paired experiments. Using the potential outcomes framework, we prove that these estimators are consistent for the sample average treatment effect under mild regularity conditions even if the regression model is improperly specified. Further, we describe how asymptotically conservative confidence intervals can be constructed. We demonstrate that the variances of the regression assisted estimators are at least as small as that of the standard difference-in-means estimator asymptotically. Through a simulation study, we illustrate the appropriateness of the proposed methods in small and moderate samples. The analysis does not require a superpopulation model, a constant treatment effect, or the truth of the regression model, and hence provides a mode of inference for the sample average treatment effect with the potential to yield improvements in the power of the resulting analysis over the classical analysis without imposing potentially unrealistic assumptions.
While attractive from a theoretical perspective, finely stratified experiments such as paired designs suffer from certain analytical limitations not present in block-randomized experiments with multiple treated and control individuals in each block. In short, when using an appropriately weighted difference-in-means to estimated the sample average treatment effect, the traditional variance estimator in a paired experiment is conservative unless the pairwise average treatment effects are constant across pairs; however, in more coarsely stratified experiments, the corresponding variance estimator is unbiased if treatment effects are constant within blocks, even if they vary across blocks. Using insights from classical least squares theory, we present an improved variance estimator appropriate in finely stratified experiments. The variance estimator is still conservative in expectation for the true variance of the difference-in-means estimator, but is asymptotically no larger than the classical variance estimator under mild conditions. The improvements stem from the exploitation of effect modification, and thus the magnitude of the improvement depends upon on the extent to which effect heterogeneity can be explained by observed covariates. Aided by these estimators, a new test for the null hypothesis of a constant treatment effect is proposed. These findings extend to some, but not all, super-population models, depending on whether or not the covariates are viewed as fixed across samples in the super-population formulation under consideration.
The conventional model for assessing insensitivity to hidden bias in paired observational studies constructs a worst-case distribution for treatment assignments subject to bounds on the maximal bias to which any given pair is subjected. In studies where rare cases of extreme hidden bias are suspected, the maximal bias may be substantially larger than the typical bias across pairs, such that a correctly specified bound on the maximal bias would yield an unduly pessimistic perception of the study's robustness to hidden bias. We present an extended sensitivity analysis which allows researchers to simultaneously bound the maximal and typical bias perturbing the pairs under investigation while maintaining the desired Type I error rate. We motivate and illustrate our method with two sibling studies on the impact of schooling on earnings, one containing information of cognitive ability of siblings and the other not. Cognitive ability, clearly influential of both earnings and degree of schooling, is likely similar between members of most sibling pairs yet could, conceivably, vary drastically for some siblings. The method is straightforward to implement, simply requiring the solution to a quadratic program. $\texttt{R}$ code is provided in the supplementary materials.
We present methods for conducting hypothesis testing and sensitivity analyses for composite null hypotheses in matched observational studies when outcomes are binary. Causal estimands discussed include the causal risk difference, causal risk ratio, and the effect ratio. We show that inference under the assumption of no unmeasured confounding can be performed by solving an integer linear program, while inference allowing for unmeasured confounding of a given strength requires solving an integer quadratic program. Through simulation studies and data examples, we demonstrate that our formulation allows these problems to be solved in an expedient manner even for large datasets and for large strata. We further exhibit that through our formulation, one can assess the impact of various assumptions about the potential outcomes on the performed inference. R scripts are provided that implement our methods. Supplementary materials for this article are available online.