Causal graphs may inform covariate adjustment for estimating causal effects and improve estimation efficiency by exploiting the graphical structure. In many applications, however, the target causal parameter may not be point-identified due to the presence of unmeasured confounding. Sensitivity analysis methods address this challenge by characterizing bounds on the causal parameter under varying assumptions about the magnitude or form of unmeasured confounding. We focus on semiparametric efficient estimation of causal effects in non-identifiable settings, assuming a known (or hypothesized) causal graph. We propose an influence function projection approach that exploits the conditional independence constraints implied by the graph to improve the efficiency of semiparametric estimators of upper and lower bounds on the average causal effect under a given sensitivity analysis model. Our approach applies across multiple sensitivity analysis frameworks and causal estimands, thereby connecting knowledge of graphical structure with the sensitivity analysis literature. We illustrate our approach through simulations and real data examples thought to be affected by unmeasured confounding, including the effect of labor training program on post-intervention earnings, and the effect of low ejection fraction on heart failure death.
Chronic diseases, including Alzheimer's disease (AD) and related dementia (ADRD), do not exist solely as isolated entities. Instead, they weave concomitant trajectories of multiple diseases, conditions, behaviors, and risks, mutually influencing each other's course and natural history, in ways yet unexplored. Electronic health records (EHRs) provide us with a unique opportunity to look at related and unrelated clinical trajectories over time, thus potentially providing insight into unrecognized prodromes, while incorporating the complexities of patients' lives. We harmonize and federate a three-city EHR metaplatform of nearly 10 million patients (∼60,000 with AD/ADRD), which we further embed within census tracts, to contextualize these health trajectories. Our multidisciplinary approach ambitions a unique dynamic platform to inform strategies to tailor risk prediction, complex clinical management, and real-world evaluation of future treatments of AD/ADRD. We present the rationale for and design of the Multimorbidity Three-City Alzheimer's Disease EHR (M3AD) Study and real-world data metaplatform, progress and demonstration of feasibility, its expected singular and complementary contributions to the field. HIGHLIGHTS: Our success in living longer lives often brings chronic conditions and multimorbidity. Alzheimer's research should comprise life trajectories' complexity in multimorbidity. New real-world analytical approaches allow integrated prediction of Alzheimer's disease. We are building a three-city electronic health record (EHR) metaplatform for prediction, prevention, and impact We further embed EHR within census tracts to contextualize Alzheimer's trajectories.
Associations between exposure to ambient air pollution and progression of emphysema have been identified in longitudinal observational studies. However, previous work has not used statistical causal inference methods tailored to address bias from time-varying confounding. The objective of this study is to propose an analytical approach for estimating longitudinal health effects of air pollution while accounting for time-varying confounding using marginal structural models and to re-analyze data on air pollution and emphysema progression from the Multi-Ethnic Study of Atherosclerosis using this analytical approach. We estimate weights for continuous exposure levels using two techniques: quantile binning of the exposure and a semiparametric model for the requisite conditional densities. The latter approach incorporates flexible machine learning methods. We find evidence for the harmful effects of ambient ozone pollution during study follow-up on the progression of emphysema, consistent with previously reported results. We find no evidence of effects of NOx during study follow-up. This investigation demonstrates that analyses based on marginal structural models are feasible in studies of the health effects of air pollution and may address possible sources of bias that traditional regression-based methods fail to address. Further investigation is warranted to understand differences between our findings and previously published results.
Longitudinal data often contains outcomes measured at multiple visits, and scientific interest may lie in quantifying the effect of an intervention on an outcome's rate of change. For example, one may wish to study the progression (or trajectory) of a disease over time under different hypothetical interventions. We extend the longitudinal modified treatment policy (LMTP) methodology to estimate effects of complex, exposure-dependent interventions on rates of change in an outcome over time. We exploit the theoretical properties of a nonparametric efficient influence function (EIF)-based estimator to introduce a novel inference framework that can be used to construct simultaneous confidence intervals for a variety of causal effects of interest and to formally test relevant global and local hypotheses about rates of change. We demonstrate the utility of our framework in investigating whether a longitudinal shift intervention affects an outcome's counterfactual trajectory, as compared with no intervention. We present results from a simulation study to illustrate the performance of our inference framework in a longitudinal setting with time-varying confounding and a continuous exposure. We also apply our inference framework to the Columbia Brain Health DataBank (CBDB) to examine the effect of shifting blood pressure on the progression of dementia.
Automated decision systems (ADS) leverage predictions about individual future outcomes to inform consequential decision-making in organizational settings. Across various settings - including criminal pretrial release, clinical triage, student support, and more - it is often assumed that improved predictive accuracy is the priority consideration in determining better downstream outcomes upon the deployment of ADS. In practice, real-world case studies reveal that this is far from the case: introducing individual predictions into decision-making modifies organizational workflows, assessment, and decision-making processes in ways that require a complete re-consideration of our approach to the design, evaluation, and deployment of ADS. As a result, this Perspective develops an integrated framework for studying ADS in social systems, shifting current priorities from a purely prediction-based paradigm towards an intervention-oriented view that accounts for real-world conditions. Our aim is to improve our understanding of ADS and more meaningfully anticipate its downstream societal and organizational consequences.
Background Medical treatment decisions are often based on estimated global risk scores. When heterogeneity in treatment effects exists, assigning treatment according to estimated individualized treatment rules (ITRs) instead has the potential to improve mean outcomes. This article aims to investigate racial and ethnic group differences in treatment rates when comparing antihypertensive medication recommendations from an estimated ITR with a risk score approach. Methods Data were simulated to emulate observational data with underlying treatment effect heterogeneity in survival times. An ITR and risk score approach were compared to illustrate how the resulting recommendations may disagree. An ITR for prescribing antihypertensives was estimated from 3281 adults from MESA (Multi-Ethnic Study of Atherosclerosis), an observational longitudinal cohort study, and compared with the risk-based approach recommended by cardiovascular care guidelines. Hypothetical treatment rates under each "rule" were computed. In the simulation study, the proportion of individuals treated optimally under each rule was calculated. Using MESA, a Chi-square test of independence was performed to determine whether treatment rates differed across racial and ethnic groups. Results Two benefits of ITRs were shown: they (1) maximize expected survival times and (2) may mitigate racial disparities when treatment effect heterogeneity is expected. Using MESA, the ITR recommended treatment to more participants than the risk score approach across all racial and ethnic groups. A Chi-square test suggested that treatment rates for different "rules" differed significantly across racial and ethnic groups (P<0.001). Conclusions Treatment recommendations varied substantially when assigning treatment using an ITR versus a risk-based approach.
Algorithms for constraint-based causal discovery select graphical causal models among a space of possible candidates (e.g., all directed acyclic graphs) by executing a sequence of conditional independence tests. These may be used to inform the estimation of causal effects (e.g., average treatment effects) when there is uncertainty about which covariates ought to be adjusted for, or which variables act as confounders versus mediators. However, naively using the data twice, for model selection and estimation, would lead to invalid confidence intervals. Moreover, if the selected graph is incorrect, the inferential claims may apply to a selected functional that is distinct from the actual causal effect. We propose an approach to post-selection inference that is based on a resampling and screening procedure, which essentially performs causal discovery multiple times with randomly varying intermediate test statistics. Then, an estimate of the target causal effect and corresponding confidence sets are constructed from a union of individual graph-based estimates and intervals. We show that this construction has asymptotically correct coverage for the true causal effect parameter. Importantly, the guarantee holds for a fixed population-level effect, not a data-dependent or selection-dependent quantity. Most of our exposition focuses on the PC-algorithm for learning directed acyclic graphs and the multivariate Gaussian case for simplicity, but the approach is general and modular, so it may be used with other conditional independence based discovery algorithms and distributional families.
Many automated decision systems (ADS) are designed to solve prediction problems – where the goal is to learn patterns from a sample of the population and apply them to individuals from the same population. In reality, these prediction systems operationalize holistic policy interventions in deployment. Once deployed, ADS can shape impacted population outcomes through an effective policy change in how decision-makers operate, while also being defined by past and present interactions between stakeholders and the limitations of existing organizational, as well as societal, infrastructure and context. In this work, we consider the ways in which we must shift from a prediction-focused paradigm to an interventionist paradigm when considering the impact of ADS within social systems. We argue this requires a new default problem setup for ADS beyond prediction, to instead consider predictions as decision support, final decisions, and outcomes. We highlight how this perspective unifies modern statistical frameworks and other tools to study the design, implementation, and evaluation of ADS systems, and point to the research directions necessary to operationalize this paradigm shift. Using these tools, we characterize the limitations of focusing on isolated prediction tasks, and lay the foundation for a more intervention-oriented approach to developing and deploying ADS.
The health sciences largely focus on disease. However, the interconnected determinants of diseases suggest that we need a science of health, a framework to examine the biology of homeodynamics in a changing environment and how this affects the health we value. We build on first principles and recent discoveries on biological system dynamics to develop the concept of intrinsic health, a field-like state emerging from the dynamic interplay of energy, communication, and structure within the organism, giving rise to robustness/resilience, plasticity, performance, and sustainability. Intrinsic health is a quantifiable property of individuals that declines with age and interacts with context. We propose a measurement framework and describe how it will contribute to achieving the shared goals of medicine and public health.
The primary practice of healthcare artificial intelligence (AI) starts with model development, often using state-of-the-art AI, retrospectively evaluated using metrics lifted from the AI literature like AUROC and DICE score. However, good performance on these metrics may not translate to improved clinical outcomes. Instead, we argue for a better development pipeline constructed by working backward from the end goal of positively impacting clinically relevant outcomes using AI, leading to considerations of causality in model development and validation, and subsequently a better development pipeline. Healthcare AI should be "actionable," and the change in actions induced by AI should improve outcomes. Quantifying the effect of changes in actions on outcomes is causal inference. The development, evaluation, and validation of healthcare AI should therefore account for the causal effect of intervening with the AI on clinically relevant outcomes. Using a causal lens, we make recommendations for key stakeholders at various stages of the healthcare AI pipeline. Our recommendations aim to increase the positive impact of AI on clinical outcomes.
Early-life growth adversity is important to later-life health, but precision assessment in adulthood is challenging. We evaluated whether the difference between attained and genotype-predicted adult height (“height-GaP”) would associate with prospectively ascertained early-life growth adversity and later-life all-cause and cardiovascular mortality. Data were first analyzed from the Avon Longitudinal Study of Parents and Children (ALSPAC; n = 4582; 56/43
[This corrects the article PMC5711475.].
The study of disparities in the liver transplantation process may focus on quantifying causal effects, particularly the average, direct, or indirect effects of various social determinants of health on being listed as a candidate for transplant. Selection bias arises when the data sample does not represent the target population, defined here as all individuals referred to the transplant clinic. Listing decisions are made for the subset of patients who complete the evaluation process, who may differ systematically from the referred population. There is evidence that selection is associated with patient characteristics that also impact outcomes. Using data only from the selected population may yield biased causal effect estimates. However, incorporating data from the referred population allows for analytic correction. This correction leverages hypothesized causal relationships among selection, the outcome (getting listed), exposures, and mediators. Using directed acyclic graphs (DAGs), we establish graphical conditions under which a reweighted mediation formula identifies effect of interest - direct, indirect, and path-specific effects - in the presence of sample selection. In a clinical case study, we investigate mediated and direct effects of a patient's socioeconomic position on being listed for transplant, allowing selection to depend on race, gender, age, and other social determinants.
We propose a set of causal estimands that we call “the mediated probabilities of causation.” These estimands quantify the probabilities that an observed negative outcome was induced via a mediating pathway versus a direct pathway in a stylized setting involving a binary exposure or intervention, a single binary mediator, and a binary outcome. We outline a set of conditions sufficient to identify these effects given observed data and propose a doubly robust projection-based estimation strategy that allows for the use of flexible nonparametric and machine learning methods for estimation. We argue that these effects may be more relevant than the probability of causation, particularly in settings where we observe both some negative outcome and negative mediating event, and we wish to distinguish between settings where the outcome was induced via the exposure inducing the mediator versus the exposure inducing the outcome directly. We motivate these estimands by discussing applications to legal and medical questions of causal attribution.
We estimated the effect of community-level natural hazard exposure during prior developmental stages on later anxiety and depression symptoms among young adults and potential differences stratified by gender. We analyzed longitudinal data (2002–2020) on 5585 young adults between 19 and 26 years in Ethiopia, India, Peru, and Vietnam. A binary question identified community-level exposure, and psychometrically validated scales measured recent anxiety and depression symptoms. Young adults with three exposure histories (“time point 1,” “time point 2,” and “both time points”) were contrasted with their unexposed peers. We applied a longitudinal targeted minimum loss-based estimator with an ensemble of machine learning algorithms for estimation. Young adults living in exposed communities did not exhibit substantially different anxiety or depression symptoms from their unexposed peers, except for young women in Ethiopia who exhibited less anxiety symptoms (average causal effect [ACE] estimate = − 8.86 [95% CI: − 17.04, − 0.68] anxiety score). In this study, singular and repeated natural hazard exposures generally were not associated with later anxiety and depression symptoms. Further examination is needed to understand how distal natural hazard exposures affect lifelong mental health, which aspects of natural hazards are most salient, how disaster relief may modify symptoms, and gendered, age-specific, and contextual differences.