
Abstract: Cornfield et al.'s 1959 paper introduced a seminal framework for evaluating the impact of unmeasured confounding by quantifying the strength an unobserved factor would need to overturn an observed association. This commentary revisits Cornfield's original inequalities, clarifying their interpretation, and highlighting the reverse-reasoning principle that underlies modern sensitivity analysis. We summarize key methodological extensions developed by later researchers–ranging from conditional independence formulations and multiplicative bounds to the sharpened limits that led to the E-value, and illustrate how these refinements broadened the applicability of Cornfield's logic. Finally, we discuss the continued influence of Cornfield-type reasoning across epidemiology and related fields, where sensitivity analysis has become an essential component of study design and causal interpretation. More than six decades later, Cornfield's insight remains a foundational pillar of causal inference, guiding researchers in assessing the robustness of empirical findings to unmeasured confounding.
Abstract: Cornfield et al.'s contributions included documenting epidemiological and policy debates regarding the effects of smoking on lung cancer. It was in the context of those debates that Cornfield et al. expressed sensitivity of the estimated effect of smoking to the dual components of confounding—relationship of a potential confounder (e.g., genes) to smoking (the focal predictor) and to lung cancer (the outcome). Current advancements have used various graphical and tabular representations to similarly express sensitivity to the dual components of confounding. As a representative of these advancements, we present Frank's Impact Threshold for a Confounding (ITCV), which resolves the dual components of confounding into a single term based on the product of the confounder's correlations with the focal predictor and with the outcome. We calculate the ITCV for an inference of an effect of kindergarten retention on achievement and provide benchmarks based on observed covariates to interpret the ITCV. Ultimately, any sensitivity analysis should only be applied after one has estimated the best possible model. The value of the sensitivity analysis is then to inform policy-oriented discourse among as wide a range of stakeholders as possible.
Abstract: Two procedures are proposed to assess sensitivity to a binary unobserved confounder for a binary outcome when the causal effect is expressed as a (log) odds ratio, as commonly arises in standard logistic modelling, particularly in case-control studies. The methods are based on graphical tools that visualize the extent to which an unobserved confounder could attenuate, nullify, or even reverse the estimated causal effect. The second procedure relies on a single sensitivity parameter, yielding a naturally bounded assessment that can be translated into an objective measure. Connections with Cornfield's conditions on relative risks are presented, thereby enlarging the circumstances where the proposed procedures can be applied and establishing a link that allows extensions to polytomous confounders.
Abstract: Cornfield et al.'s (1959) analysis of smoking and lung cancer showed how explicit assumptions about unobserved confounding and selection mechanisms shape the boundary between statistical association and causal conclusions. Today, causal inference spans diverse fields including economics, statistics, psychology, epidemiology, and evaluation, with distinct traditions and tools. Without a shared framework, it can be difficult to accumulate evidence across studies or to diagnose why results differ when similar questions are studied independently, particularly when differences arise from upstream design and data-generation choices rather than analytic methods alone. While tools for causal identification such as Directed Acyclic Graphs (DAGs) have become widely used across disciplines, they do not naturally foreground threats arising from representation and measurement errors. In contrast, the Total Survey Error (TSE) framework enumerates representation and measurement errors across the survey life cycle and is routinely used to evaluate tradeoffs in survey design and data collection. We propose a Total Causal Error (TCE) framework that integrates causal inference with the TSE perspective to organize risks to causal conclusions along parallel paths of causal identification, representation, and measurement, across the stages of the theoretical estimand, empirical estimand, data collection, data processing, and estimation and interpretation. We propose a version of this framework, provide explicit definitions for its components, illustrate a small number of use cases, and invite further development and coordination across fields.
Abstract: Cornfield, Haenszel, Hammond, Lilienfeld, Shimkin, and Wynder (1959)'s seminal work made two major contributions that shape causal inference and epidemiology to this day. First, Cornfield et al. provided the first example of quantitative sensitivity analysis for unmeasured confounding; a result that became subsequently known as the Cornfield condition. Second, they advocated for the use of relative effect measures for causal inference and absolute effect measures in public health. This recommendation is also referred to as the Cornfield principle. This commentary critically re-examines Cornfield et al.'s landmark contributions. With the benefit of hindsight, we formalize their arguments, identify shortcomings, and contextualize their claims in the light of recent research.
Abstract: Jerome Cornfield's (1959) analysis of smoking and lung cancer introduced a foundational principle for causal inference from observational data: causal conclusions should be evaluated not by the absence of assumption violations, but by the implausibility of the violations required to overturn them. In the decades since, causal inference research has made substantial progress in explicitly defining identification assumptions and improving estimation under those assumptions. As estimation theory has advanced, concerns about model misspecification have receded, and flexible semi-parametric methods now exist for increasingly complex causal questions. Despite this progress, there has been comparatively little success in extending Cornfield-style sensitivity analysis to these modern settings, limiting both the interpretability and practical adoption of contemporary causal methods. In this commentary, we revisit Cornfield's original framework and consider its amenability for four identification assumptions that have received growing attention beyond standard unmeasured confounding: positivity, interference, parallel trends, and sequential exchangeability. For each, we evaluate its accessibility to Cornfield-style reasoning, discussing challenges and opportunities for developing interpretable, calibrated sensitivity analyses. We argue that Cornfield-style reasoning is more naturally applicable to some assumptions than others, clarifying where such approaches are promising and where fundamental obstacles remain.
Abstract: Modeling the effect of smoking on the risk of lung cancer in cohort studies requires specification of time-dependent exposure variables and appropriate handling of confounding. This can be challenging as data on both exposure and confounders are typically incompletely observed. We discuss population exposure and disease processes and review modeling challenges when exposures are dynamic and challenging to summarize. Biases arising from failure to address time-dependent confounding are highlighted, with particular reference to a paradox related to the effect of quitting on risk of lung cancer. We stress the utility of comprehensive joint models.
Abstract: Jerome Cornfield was a leading biostatistician of the mid-20th century, held in extremely high regard both by fellow statisticians and medical researchers. He was also beloved for his warmth, charm and wonderful sense of humor. I was fortunate to begin my career working with him; in this piece, I share some memorable experiences.
Abstract: Since Cornfield, Haenszel, Hammond, Lilienfeld, Shimkin, and Wynder (1959)'s seminal contribution, many sensitivity analyses have been introduced. Modern-day approaches to sensitivity analysis aim to impose relatively little structure on the unobserved confounder. However, while a sensitivity analysis may not rely on many assumptions on the unobserved confounder, the worst-case bounds that are solved for within a sensitivity analysis are often only realized when a hypothetical confounder takes on a specific form, which is not always made explicit. In the following commentary, we consider several leading sensitivity analysis approaches and review what is known about the structure of the confounder at which the worst-case bounds are realized. We show that the type of sensitivity model used can induce different data generating processes for the worst-case confounders. We argue that understanding the structure of the omitted confounder implied by the worst-case bounds is crucial for more transparent sensitivity analysis, and we suggest directions for future work.
Abstract: It is challenging to infer a causal relationship with observational data. Untestable assumptions of ignorability or exchangeability are often utilized to facilitate causal effect estimation, which opens the door to arguments for and against the validity of a causal conclusion. In their 1959 seminal paper, Cornfield and colleagues provided sophisticated statistical reasoning for causal inference with observational data and made a key technical contribution to sensitivity analysis. Cornfield's idea paved the way for the development of modern tools for assessment of causality if ignorability assumptions fail. Compared with classical/frequentist approaches, Bayesian methods for sensitivity analysis are less fully developed. In this commentary, we first introduce a Bayesian semiparametric model for causal inference, then present a sensitivity analysis strategy for the Gaussian process model. Our method is easily interpretable and avoids restrictive parametric outcome assumptions. It can also be applied to both population-level and conditional causal effects.
Abstract: Smoking and Lung Cancer (1959) is a compelling case study of the nature of proof and causation in medicine. It should be understood as the culmination of years of work to find rigorous methods for making statistically and rhetorically convincing accounts of how associations might be deemed to be causal. In this comment, we highlight several papers written during the 1950s by Jerome Cornfield that illustrate the role he played in laying the groundwork for this important and influential work.
Sensitivity analysis methods such as the Cornfield inequality and the E-value were developed to assess the robustness of observed associations against unmeasured confounding – a major challenge in observational studies. However, the calculation and interpretation of these methods can be difficult for clinicians and interdisciplinary researchers. Recent advances in large language models (LLMs) offer accessible tools that could assist sensitivity analyses, but their reliability in this context has not been studied. We assess four widely used LLMs, ChatGPT, Claude, DeepSeek, and Gemini, on their ability to conduct sensitivity analyses using Cornfield inequalities and E-values. We first extract study-specific information (exposures, outcomes, measured confounders, and effect estimates) from four published observational studies in different fields. Using such information, we develop structured prompts to assess the performance of the LLMs in three aspects: (1) accuracy of E-value calculation, (2) qualitative interpretation of robustness to unmeasured confounding, and (3) suggestion of possible unmeasured confounders. To our knowledge, there has been little prior work on using LLMs for sensitivity analysis, and this study is an early investigation in this area. The results show that ChatGPT, Claude, and Gemini accurately reproduce the E-values, whereas DeepSeek shows small biases. Qualitative conclusions from all the LLMs align with the magnitude of the E-values and the reported effect sizes, and all models identify biologically and epidemiologically plausible unmeasured confounders. These findings suggest that, when guided by structured prompts, LLMs can effectively assist in evaluating unmeasured confounding, and thereby can support study design and decision-making in observational studies.
Sensitivity analysis is widely used to assess the robustness of causal conclusions in observational studies, yet its interaction with the structure of measured covariates is often overlooked. When latent confounders cannot be directly adjusted for and are instead controlled using proxy variables, strong associations between exposure and measured proxies can amplify sensitivity to residual confounding. We formalize this phenomenon in linear regression settings by showing that a simple ratio involving the exposure model coefficient and residual exposure variance provides an observable measure of this increased sensitivity. Applying our framework to smoking and lung cancer, we document how growing socioeconomic stratification in smoking behavior over time leads to heightened sensitivity to unmeasured confounding in more recent data. These results highlight the importance of multicollinearity when interpreting sensitivity analyses based on proxy adjustment.
Sensitivity analysis informs causal inference by assessing the sensitivity of conclusions to departures from assumptions. The consistency assumption states that there are no hidden versions of treatment and that the outcome arising naturally equals the outcome arising from intervention. When reasoning about the possibility of consistency violations, it can be helpful to distinguish between covariates and versions of treatment. In the context of surgery, for example, genomic variables are covariates and the skill of a particular surgeon is a version of treatment. There may be hidden versions of treatment, and this paper addresses that concern with a new kind of sensitivity analysis. Whereas many methods for sensitivity analysis are focused on confounding by unmeasured covariates, the methodology of this paper is focused on confounding by hidden versions of treatment. In this paper, new mathematical notation is introduced to support the novel method, and example applications are described.
Experimental and observational studies often lack validity due to untestable assumptions. We propose a double machine learning approach to combine experimental and observational studies, allowing practitioners to test for assumption violations and estimate treatment effects consistently. Our framework proposes a falsification test for external validity and ignorability under milder assumptions. We provide consistent treatment effect estimators even when one of the assumptions is violated. However, our no-free-lunch theorem highlights the necessity of accurately identifying the violated assumption for consistent treatment effect estimation. Through comparative analyses, we show our framework's superiority over existing data fusion methods. The practical utility of our approach is further exemplified by three real-world case studies, underscoring its potential for widespread application in empirical research.
Matching is a popular nonparametric covariate adjustment strategy in empirical health services research. Matching helps construct two groups comparable in many baseline covariates but different in some key aspects under investigation. In health disparities research, it is desirable to understand the contributions of various modifiable factors, like income and insurance type, to the observed disparity in access to health services between different groups. To single out the contributions from the factors of interest, we propose a statistical matching methodology that constructs nested matched comparison groups from, for instance, White men, that resemble the target group, for instance, black men, in some selected covariates while remaining identical to the white men population before matching in the remaining covariates. Using the proposed method, we investigated the disparity gaps between white men and black men in the US in prostate-specific antigen (PSA) screening based on the 2020 Behavioral Risk Factor Surveillance System (BFRSS) database. We found a widening PSA screening rate as the white matched comparison group increasingly resembles the black men group and quantified the contribution of modifiable factors like socioeconomic status. Finally, we provide code that replicates the case study and a tutorial that enables users to design customized matched comparison groups satisfying multiple criteria.
Abstract: We respond to Aronow et al. (2025)’s paper arguing that randomized controlled trials (RCTs) are “enough,” while nonparametric identification in observational studies is not. We agree with their position with respect to experimental versus observational research, but question what it would mean to extend this logic to the scientific enterprise more broadly. We first investigate what is meant by “enough,” arguing that this is a fundamentally a sociological claim about the relationship between statistical work and larger social and institutional processes, rather than something that can be decided from within the logic of statistics. For a more complete conception of “enough,” we outline all that would need to be known – not just knowledge of propensity scores, but knowledge of many other spatial and temporal characteristics of the social world. Even granting the logic of the critique in Aronow et al. (2025), its practical importance is a question of the contexts under study. We argue that we should not be satisfied by appeals to intuition about the complexity of “naturally occurring” propensity score functions. Instead, we call for more empirical metascience to begin to characterize this complexity. We apply this logic to the example of recommender systems developed by Aronow et al. (2025) as a demonstration of the weakness of allowing statisticians’ intuitions to serve in place of metascientific data. Rather than implicitly deciding what is “enough” based on statistical applications the social world has determined to be most profitable, we argue that practicing statisticians should explicitly engage with questions like “for what?” and “for whom?” in order to adequately answer the question of “enough?”
The interventionist approach to causal inference recommends that observational studies be framed as randomized trials with well-defined interventions to more precisely define the population of interest, exposure comparisons, assignment procedures, follow-up period, and outcomes. However, others suggest causes are not restricted to interventions, and the approach is too restrictive and will limit science. The described examples usually refer to questions about causes of effects/outcomes (rather than effects of interventions), where the 'cause' of interest often represents a mediator variable between an existing or hypothesized intervention. These questions are important because they represent the foundation for improving existing interventions and developing new interventions. In this article, I show how the interventionist approach can be used to try and answer these broader 'causes of effects/outcomes' questions. I use the sufficient casual set framework popularized by Rothman, which considers causes as states rather than interventions. The method is fully consistent with the potential outcomes approach and the need for well-defined counterfactuals. Whereas the effect of an intervention can theoretically be evaluated with an idealized single randomized trial, evaluating the causes of effects/outcomes requires evidence synthesis across multiple studies that each emulate a different target randomized trial. In other words, questions related to the effects of interventions are a special case where all the evidence that needs to be synthesized is available within a single idealized trial. Finally, considering states as causes also provides a transparent way to think about well-defined counterfactuals, and how factors that are commonly referred to as "non-manipulable" such as sex can also be studied as causes.
This is a review of Peng Ding's textbook "A First Course in Causal Inference." The book builds causal inference topics up from basics in experiments to complex observational studies. This review discusses the book's style and content as well as who should use this book.
While randomized trials are the gold standard in causal inference, observational studies are often conducted when trials are infeasible either due to logistical or ethical constraints. Since treatments are not randomized in observational studies, techniques from causal inference are required to adjust for confounding. Bayesian approaches to causal estimation are desirable because they provide: 1) prior smoothing that offers useful regularization of causal effect estimates, 2) flexible models that are robust to misspecification, and 3) full inference (i.e., both point estimates and uncertainty quantification) for causal estimands. However, Bayesian causal inference is difficult to implement manually and there is a lack of user-friendly software, presenting a significant barrier to wide-spread use. Moreover, there is a lack of manuscripts aimed at explicitly connecting statistical/causal formula with implementation code. We address this gap by developing and describing causalBETA (BayesianEventTimeAnalysis) - an open-source R package for estimating causal effects on event-time outcomes using Bayesian semiparametric models. The package provides a familiar front-end to users, with syntax identical to existing survival analysis R packages, while leveraging Stan - a popular platform for high performance Bayesian computing - for efficient posterior computation. To improve user experience, the package provides custom S3 class objects and methods to facilitate visualizations and summaries of results using familiar generic functions like plot() and summary(). In this paper, we provide the methodological details of the package, a demonstration using publicly available data, and computational guidance.