
Many fairness criteria constrain the policy or choice of predictors, which can have unwanted consequences, in particular, when optimizing the policy under such constraints. Here, we instead suggest that fairness can be directly analyzed as a property of the utility function. Instead of imposing fairness constraints on the policy, we suggest to simply maximize a utility function satisfying certain fairness properties. Concretely, we define value of information fairness, which prescribes that there must not be an incentive to infer the protected attribute. This principle suggests modifying utility functions such that they satisfy value of information fairness. We describe how such modifications can be achieved and discuss consequences for the corresponding optimal policies. We apply our framework to thought experiments and the COMPAS data, demonstrating that focusing on utility functions sometimes provides answers that better align with intuitive judgments about what is fair. Moreover, we are not aware of any intuitively fair policy that violates value of information fairness; and when we find that value of information fairness recommends an intuitively unfair policy, no realizable policy is intuitively fair.
Time-varying treatments can be confounded by time-varying prognostic factors influenced by earlier treatments. Structural nested mean models (SNMMs) adjust for time-varying confounders by parameterizing causal effects conditioned on variable history, and their parameters can be estimated using doubly robust g-estimation. For binary outcomes, log-linear SNMM parameters may be variation dependent on nuisance parameters due to probability range constraints. If the fitted model ignores this restriction, the predicted probabilities may exceed one, introducing variability into g-estimation functions. For time-fixed treatments, a nuisance parameterization for the confounder-conditional log odds product, which is variation independent of the causal parameters, provides a bivariate mapping to outcome regressions under binary treatments. The existing extension for time-varying treatments ensures all parameters constituting fully conditional regressions are variation independent of each other but requires a monotonicity assumption on the regressions when handling continuous time-varying covariates. When the interest lies in SNMM parameters and the objective is reducing variability of g-estimators, SNMM parameters may not necessarily be variation independent of all nuisance parameters across time points. We propose specifying the log odds product of SNMMs under two hypothetical treatments, avoiding the monotonicity assumption. Simulations with unbounded confounders and an application study demonstrate the improved efficiency of the g-estimators.
When making causal inferences from observational data, researchers must consider the effects of confounding. In a regression discontinuity design (RDD), individuals receive a treatment based on whether they score below or above a threshold value measured on a continuous variable. By assuming continuous regression lines for the potential outcomes at the threshold, RDD methods remove the confounding bias in estimating the treatment effect at the threshold. This effect is estimated by the jump in the regression line for the observed outcome at the threshold. Although RDD methods have received deserved attention in economics, the social sciences, and epidemiology, we show that inferences from RDDs using local and global linear regression estimators are prone to regression to the mean bias in certain situations. A common situation where the bias occurs is when a running variable has a normal distribution and the cutoff is relatively far from the mean of this distribution. We derive the expression for the limiting bias in this case. In general, the bias occurs when some units receive (or do not receive) treatment when their running variable values are extreme relative to the typical value of the running variable. Through simulations, we show that the regression to the mean bias can lead to inflated type I error rates and bias toward the null in typical settings. Simulations show that the RTM effect can be different for different estimators. We develop a novel method to correct this bias and provide valid inferences. We verify our correction method in simulations and apply it to a real-life example of the incumbency advantage in U.S. House elections.
Vaccine safety surveillance programs that monitor possible short-term rare adverse events following vaccination usually only have access to data on vaccinated individuals who experienced the event of interest. The Self-Controlled Case Series design employs such data and compares the risk of the event in an “risky” period immediately after vaccination to that in a “baseline risk” period where the transient risk should be gone. To ensure valid analysis, some assumptions have been given in the literature while others are made implicitly through parametric modeling. In this work, we provide a complete formal causal framework for the Self-Controlled Cases Series design. We provide sufficient conditions for a causal interpretation of the contrasts estimated from the data on exposed individuals who experienced the event. These conditions are intuitive but often cannot be tested from the available data. We describe practical settings where these conditions may be violated. When the conditions are not met, the contrasts estimated in practice do not clearly relate to any quantities of causal interest, and should therefore be interpreted with caution.
Understanding the complex interactions among multiple environmental exposures is critical for assessing their combined impact on health outcomes. This study introduces InterXshift, a novel semiparametric method that provides a nonparametric definition of interaction and facilitates both the discovery and efficient estimation of interaction effects in mixed exposures. Leveraging stochastic shift interventions and ensemble machine learning, InterXshift identifies and quantifies interactions through a model-independent target parameter, estimated using targeted maximum likelihood estimation (TMLE) and cross-validation. The approach contrasts expected outcomes from joint interventions against those from individual exposures, enabling the detection of synergistic and antagonistic interactions. Validation through simulations and application to the National Institute of Environmental Health Sciences (NIEHS) Mixtures Workshop data demonstrate InterXshift’s efficacy in accurately identifying true interaction directions and consistently highlighting significant impacts. We apply our methodology to National Health and Nutrition Examination Survey (NHANES) data to understand the interaction effect (if any) of furan exposure on leukocyte telomere length. This method enhances the analysis of multi-exposure interactions within high-dimensional datasets, offering robust methodological improvements for elucidating complex exposure dynamics in environmental health research. Additionally, we provide an open-source implementation of InterXshift in the InterXshift R package, facilitating its adoption and application by the research community.
Regulations of chemical exposures often focus on individual substances, neglecting the amplified toxicity that can arise from multiple concurrent exposures. We propose a novel methodology to identify critical thresholds in multivariate exposure spaces and estimate the effects of policy interventions that limit exposures within these thresholds. Our approach employs a recursive partitioning algorithm integrated with targeted maximum likelihood estimation (TMLE) to discover regions in the exposure space where the expected outcome is minimized or maximized. To address potential overfitting bias from using the same data for threshold discovery and effect estimation, we utilize cross-validated TMLE (CV-TMLE), which ensures asymptotic unbiasedness and efficiency. Simulation studies demonstrate convergence to the optimal exposure region and accurate estimation of intervention effects. We apply our method to synthetic mixture data, successfully identifying true interactions, and to NHANES data, discovering harmful metal exposures affecting telomere length. Our approach provides a flexible and interpretable framework for policy-makers to assess the impact of exposure regulations, and we offer an open-source implementation in the CVtreeMLE R package.
Standard methods for estimating average causal effects require complete observations of the exposure and confounders. In observational studies, however, missing data are ubiquitous. Motivated by a study on the effect of prescription opioids on mortality, we propose methods for estimating average causal effects when exposures and potential confounders may be missing. We consider missingness at random and additionally propose several specific missing not at random (MNAR) assumptions. Under our proposed MNAR assumptions, we show that the average causal effects are identified from the observed data and derive corresponding influence functions, which form the basis of our proposed estimators. Our simulations show that standard multiple imputation techniques paired with a complete data estimator is unbiased when data are missing at random (MAR) but can be biased otherwise. For each of the MNAR assumptions, we instead propose doubly robust targeted maximum likelihood estimators (TMLE), allowing misspecification of either (i) the outcome models or (ii) the exposure and missingness models. The proposed methods are suitable for any outcome types, and we apply them to a motivating study that examines the effect of prescription opioid usage on all-cause mortality using data from the National Health and Nutrition Examination Survey (NHANES).
We study the properties of the score confidence set for the local average treatment effect in non and semiparametric instrumental variable models. This confidence set is constructed by inverting a score test based on an estimate of the nonparametric influence function for the estimand, and is known to be uniformly valid in models that allow for arbitrarily weak instruments; because of this, the confidence set can have infinite diameter at some laws. We characterize the six possible forms the score confidence set can take: a finite interval, an infinite interval (or a union of them), the whole real line, an empty set, or a single point. Moreover, we show that, at any fixed law, the score confidence set asymptotically coincides, up to a term of order 1/n, with the Wald confidence interval based on the doubly robust estimator which solves the estimating equation associated with the nonparametric influence function. This result implies that, in models where the efficient influence function coincides with the nonparametric influence function, the score confidence set is, in a sense, optimal in terms of its diameter. We also show that under weak instrument asymptotics, where the strength of the instrument is modelled as local to zero, the doubly robust estimator is asymptotically biased and does not follow a normal distribution. A simulation study confirms that, as expected, the doubly robust estimator performs poorly when instruments are weak, whereas the score confidence set retains good finite-sample properties in both strong and weak instrument settings. Finally, we provide an algorithm to compute the score confidence set, which is now available in the DoubleML package for double machine learning.
Surrogate markers are most commonly studied within the context of randomized clinical trials. However, the need for alternative outcomes also extends to real-world public health and social science research, where randomized trials are often impractical. While standard methods for evaluating surrogate markers largely rely on the assumption of randomized treatment, there is a significant gap in applying these techniques to observational data, where the central challenge shifts to managing confounding. The few methods that do allow for non-randomized treatment/exposure do not offer a way to examine surrogate heterogeneity with respect to patient characteristics. In this paper, we propose a framework to assess surrogate heterogeneity in non-randomized data and implement this framework using meta-learners. Our approach allows us to quantify heterogeneity in surrogate strength with respect to patient characteristics while accommodating confounders through the use of flexible, off-the-shelf machine learning methods. In addition, we use our framework to identify covariate profiles where the surrogate is a valid replacement of the primary outcome. We examine the performance of our methods via a simulation study and application to examine heterogeneity in the surrogacy of hemoglobin A1c as a surrogate for fasting plasma glucose.
Large language models (LLMs) offer the potential to automate a large number of tasks that previously have not been possible to automate, including some in science. There is considerable interest in whether LLMs can automate the process of causal inference by providing the information about causal links necessary to build a structural model. We use the case of confounding in the Coronary Drug Project (CDP), for which there are several studies listing expert-selected confounders that can serve as a ground truth. LLMs exhibit mediocre performance in identifying confounders in this setting, even though text about the ground truth is in their training data. Variables that experts identify as confounders are only slightly more likely to be labeled as confounders by LLMs compared to variables that experts consider non-confounders. Further, LLM judgment on confounder status is highly inconsistent across models, prompts, and, for some prompts, irrelevant concerns like multiple-choice option ordering, reducing the chance that a prompt shown to be successful in a setting with ground truth would continue to work in a setting where correct answers are not predetermined. LLMs do not yet have the ability to automate the reporting of causal links.
The positivity assumption is central in the identification of a causal effect. Especially its stochastic variant is an issue many applied researchers face. Yet positivity is rarely discussed, especially in conjunction with continuous treatments or Modified Treatment Policies. One common recommendation for dealing with a violation is to change the estimand. However, an applied researcher is faced with two problems: First, how can she tell whether there is a stochastic positivity violation given her estimand of interest, preferably without having to estimate a model first? Second, if she finds a problem with stochastic positivity, how should she change her estimand in order to arrive at an estimand which does not face the same issues? We suggest a novel diagnostic which allows the researcher to answer both questions by providing insights into how well an estimation for a certain estimand can be made for each observation using the data at hand. We provide a simulation study on the general behaviour of different Modified Treatment Policies (MTPs) at different levels of stochastic positivity violations and show how the diagnostic helps understand where bias is to be expected. We illustrate the application of our proposed diagnostic in a pharmacoepidemiological study based on data from CHAPAS-3, a trial comparing different treatment regimens for children living with HIV.
Difference-in-Differences and Event-Study regression estimators are commonly used to estimate treatment effects. However, when effects are heterogeneous, the standard two-way fixed effects (TWFE) regression can provide biased estimates. Recent literature proposes alternative estimators to overcome the so-called “bad comparisons” problem, allowing unbiased estimates of the treatment effect to be recovered. To date, attention has primarily focused on linear models despite the prevalence of other outcome types in applied research. In this paper, we address this gap by extending five of these alternative estimators for estimating treatment effects on count and binary outcomes and exploring their relative performance within a simulation exercise where true effects are known. Our simulation results indicate that while these estimators can produce unbiased estimates for linear outcomes, they may fail to recover the true treatment effect for nonlinear outcomes if used ‘straight out of the box’. We show that Interaction-weighted, IPW, and Extended-TWFE estimators can be readily adjusted to provide unbiased estimates. Finally, we apply these estimators to prior published work examining the effect of star coauthorship on peers’ productivity using citation data, revealing moderate differences in the magnitude of the effects across the estimators.
Understanding causal relationships in the presence of complex, structured data remains a central challenge in modern statistics and science in general. While traditional causal inference methods are well-suited for scalar outcomes, many scientific applications demand tools capable of handling functional data – outcomes observed as functions over continuous domains such as time or space. Motivated by this need, we propose DR-FoS, a novel method for estimating the Functional Average Treatment Effect (FATE) in observational studies with functional outcomes. DR-FoS exhibits double robustness properties, ensuring consistent estimation of FATE even if either the outcome or the treatment assignment model is misspecified. By leveraging recent advances in functional data analysis and causal inference, we establish the asymptotic properties of the estimator, proving its convergence to a Gaussian process. This guarantees valid inference with simultaneous confidence bands across the entire functional domain. Through extensive simulations, we show that DR-FoS achieves robust performance under a wide range of model specifications. Finally, we illustrate the utility of DR-FoS in a real-world application, analyzing functional outcomes to uncover meaningful causal insights in the SHARE (Survey of Health, Aging and Retirement in Europe) dataset.
There has been widespread use of causal inference methods for the rigorous analysis of observational studies and to identify policy evaluations. In this article, we consider a class of generalized coarsened procedures for confounding. At a high level, these procedures can be viewed as performing a clustering of confounding variables, followed by treatment effect and attendant variance estimation using the confounder strata. In addition, we propose two new algorithms for generalized coarsened confounding. While previous authors have developed some statistical properties for one special case in our class of procedures, we instead develop a general asymptotic framework. We provide asymptotic results for the average causal effect estimator as well as providing conditions for consistency. In addition, we provide an asymptotic justification for the variance formulae for coarsened exact matching. A bias correction technique is proposed, and we apply the proposed methodology to data from two well-known observational studies.
Inverse probability (IP) weighting of marginal structural models (MSMs) can provide consistent estimators of time-varying treatment effects under correct model specifications and identifiability assumptions, even in the presence of time-varying confounding. However, this method has two problems: (i) inefficiency due to IP-weights cumulating all time points and (ii) bias and inefficiency due to the MSM misspecification. To address these problems, we propose (i) new IP-weights for estimating parameters of the MSM that depends on partial treatment history and (ii) closed testing procedures for selecting partial treatment history (how far back in time the MSM depends on past treatments). We derive the theoretical properties of our proposed methods under known IP-weights and discuss their extension to estimated IP-weights. Although some of our theoretical results are derived under additional assumptions beyond standard identifiability assumptions, some of which can be checked empirically from the data. In simulation studies, our proposed methods outperformed existing methods both in terms of performance in estimating time-varying treatment effects and in selecting partial treatment history. Our proposed methods have also been applied to real data of hemodialysis patients with reasonable results.
In experimental and observational data settings, researchers often have limited knowledge of the reasons for missing outcomes. To address this uncertainty, we propose bounds on causal effects for missing outcomes, accommodating the scenario where missingness is an unobserved mixture of informative and non-informative components. Within this mixed missingness framework, we explore several assumptions to derive bounds on causal effects, including bounds expressed as a function of user-specified sensitivity parameters. We develop influence-function based estimators of these bounds to enable flexible, non-parametric, and machine learning based estimation, achieving root-n convergence rates and asymptotic normality under relatively mild conditions. We further consider the identification and estimation of bounds for other causal quantities that remain meaningful when informative missingness reflects a competing outcome, such as death. We conduct simulation studies and illustrate our methodology with a study on the causal effect of antipsychotic drugs on diabetes risk using a health insurance dataset.
Prediction invariance of causal models under heterogeneous settings has been exploited by a number of recent methods for causal discovery, typically focussing on recovering the causal parents of a target variable of interest. Existing methods require observational data from a number of sufficiently different environments, which is rarely available. In this paper, we consider a structural equation model where the target variable is described by a generalized linear model conditional on its parents. Besides having finite moments, no modelling assumptions are made on the conditional distributions of the other variables in the system, and nonlinear effects on the target variable can naturally be accommodated by a generalized additive structure. Under this setting, we characterize the causal model uniquely by means of two key properties: the Pearson risk invariant under the causal model and, conditional on the causal parents, the causal parameters maximize the expected likelihood. These two properties form the basis of a computational strategy for searching the causal model among all possible models. A stepwise greedy search is proposed for systems with a large number of variables. Crucially, for generalized linear models with a known dispersion parameter, such as Poisson and logistic regression, the causal model can be identified from a single data environment. The method is implemented in the R package causalreg.
The dishonest casino is a well-known hidden Markov model (HMM) often used in education to introduce HMMs and graphical models. A sequence of die rolls is observed with the casino switching between a fair and a loaded die. Instead of recovering the latent regime through filtering, smoothing, or the Viterbi algorithm, we ask a counterfactual question: how much of the gambler’s winnings are caused by the casino’s cheating? We introduce a class of structural causal models (SCMs) consistent with the HMM and define the expected winnings attributable to cheating (EWAC). Because EWAC is only partially identifiable, we bound it via linear programs (LPs). Numerical experiments help to develop intuition using benchmark SCMs based on independence, comonotonic, and countermonotonic copulas. Imposing a time homogeneity condition on the SCM yields tighter bounds, whereas relaxing it produces looser bounds that admit an explicit LP solution. Domain knowledge such as pathwise monotonicity or counterfactual stability can be incorporated through additional linear constraints. Finally, we show the time averaged EWAC becomes fully identifiable as the number of time periods tends to infinity. Our work is the first to develop LP bounds for counterfactuals in a HMM setting, benefiting educational contexts where counterfactual inference is taught.
Estimating and obtaining reliable inference for the marginally adjusted causal dose-response curve for continuous treatments without relying on parametric assumptions is a well-known statistical challenge. Parametric models risk introducing significant bias through model misspecification, compromising the accurate representation of the underlying data and dose-response relationship. On the other hand, nonparametric models face difficulties as the dose-response curve is not pathwise differentiable, preventing consistent estimation at standard rates. The Highly Adaptive Lasso (HAL) maximum likelihood estimator offers a promising approach to this issue. In this paper, we introduce a HAL-based plug-in estimator for the causal dose-response curve, bridge theoretical development and empirical application, and assess its empirical performance against other estimators. This work emphasizes not just theoretical proofs, but also demonstrates their application through comprehensive simulations, thereby filling an essential gap between theory and practice. Our comprehensive simulations demonstrate that the HAL-based estimator achieves pointwise asymptotic normality with valid inference and consistently outperforms existing approaches for estimating the causal dose-response curve.
The average treatment effect (ATE) is a common parameter estimated in causal inference literature, but it is only defined for binary exposures. Thus, despite concerns raised by some researchers, many studies seeking to estimate the causal effect of a continuous exposure create a new binary exposure variable by dichotomizing the continuous values into two categories. In this paper, we affirm binarization as a statistically valid method for answering causal questions about continuous exposures by showing the equivalence between the binarized ATE and the difference in the average outcomes of two specific modified treatment policies. These policies impose cut-offs corresponding to the binarized exposure variable and assume preservation of relative self-selection. Relative self-selection is the ratio of the probability density of an individual having an exposure equal to one value of the continuous exposure variable versus another. The policies assume that, for any two values of the exposure variable with non-zero probability density after the cut-off, this ratio will remain unchanged. Through this equivalence, we clarify the assumptions underlying binarization and discuss how to properly interpret the resulting estimator. Additionally, we introduce a new target parameter that can be computed after binarization that considers the observed world as a benchmark. We argue that this parameter addresses more relevant causal questions than the traditional binarized ATE parameter. We present a simulation study to illustrate the implications of these assumptions when analyzing data and to demonstrate how to correctly implement estimators of the parameters discussed. Finally, we present an application of this method to evaluate the effect of a law in the state of California which seeks to limit exposures to oil and gas wells on birth outcomes to further illustrate the underlying assumptions.