When individuals participating in a randomized trial differ with respect to the distribution of effect modifiers compared compared with the target population where the trial results will be used, treatment effect estimates from the trial may not directly apply to target population. Methods for extending -- generalizing or transporting -- causal inferences from the trial to the target population rely on conditional exchangeability assumptions between randomized and non-randomized individuals. The validity of these assumptions is often uncertain or controversial and investigators need to examine how violation of the assumptions would impact study conclusions. We describe methods for global sensitivity analysis that directly parameterize violations of the assumptions in terms of potential (counterfactual) outcome distributions. Our approach does not require detailed knowledge about the distribution of specific unmeasured effect modifiers or their relationship with the observed variables. We illustrate the methods using data from a trial nested within a cohort of trial-eligible individuals to compare coronary artery surgery plus medical therapy versus medical therapy alone for stable ischemic heart disease.
We discuss generalizability analyses under a partially nested trial design, where part of the trial is nested within a cohort of trial-eligible individuals, while the rest of the trial is not nested. This design arises, for example, when only some centers participating in a trial are able to collect data on nonrandomized individuals, or when data on nonrandomized individuals cannot be collected for the full duration of the trial. Our work is motivated by the Necrotizing Enterocolitis Surgery Trial, which compared initial laparotomy versus peritoneal drain for infants with necrotizing enterocolitis or spontaneous intestinal perforation. During the first phase of the study, data were collected from randomized individuals as well as consenting nonrandomized individuals; during the second phase of the study, however, data were only collected from randomized individuals, resulting in a partially nested trial design. We propose methods for generalizability analyses with partially nested trial designs. We describe identification conditions and propose estimators for causal estimands in the target population of all trial-eligible individuals, both randomized and nonrandomized, in the part of the data where the trial is nested while using trial information spanning both parts. We evaluate the estimators in a simulation study and provide an illustration using the Necrotizing Enterocolitis Surgery Trial study.
In multicenter randomized trials, when effect modifiers have a different distribution across centers, comparisons between treatment groups that average over centers may not apply to any of the populations underlying the individual centers. Here, we describe methods for reinterpreting the evidence produced by a multicenter trial in the context of the population underlying each center. We describe how to identify center-specific effects under identifiability conditions that are largely supported by the study design and when associations between center membership and the outcome may be present, given baseline covariates and treatment ("center-outcome associations"). We then consider an additional condition of no center-outcome associations given baseline covariates and treatment. We show that this condition can be assessed using the trial data; when it holds, center-specific treatment effects can be estimated using analyses that completely pool information across centers. We propose methods for estimating center-specific average treatment effects, when center-outcome associations may be present and when they are absent, and describe approaches for assessing whether center-specific treatment effects are homogeneous. We evaluate the performance of the methods in a simulation study and illustrate their implementation using data from the Hepatitis C Antiviral Long-Term Treatment Against Cirrhosis trial.
Analyses of multi-source data, such as data from multi-center randomized trials, individual participant data meta-analyses, or pooled analyses of observational studies, combine information to estimate an overall average treatment effect. However, if average treatment effects vary across data sources, commonly used approaches for multi-source analyses may not have a clear causal interpretation with respect to a target population of interest. In this paper, we provide identification and estimation of average treatment effects in a target population underlying one of the data sources in a point treatment setting for failure time outcomes potentially subject to right-censoring. We do not assume the absence of effect heterogeneity and hence our results are valid, under certain assumptions, when average treatment effects vary across data sources. We derive the efficient influence functions for source-specific average treatment effects using multi-source data under two different sets of assumptions, and propose a novel doubly robust estimator for our estimand. We evaluate the finite-sample performance of our estimator in simulation studies, and apply our methods to data from the HALT-C multi-center trials.
Importance The National Lung Screening Trial (NLST) found that screening for lung cancer with low-dose computed tomography (CT) reduced lung cancer–specific and all-cause mortality compared with chest radiography. It is uncertain whether these results apply to a nationally representative target population. Objective To extend inferences about the effects of lung cancer screening strategies from the NLST to a nationally representative target population of NLST-eligible US adults. Design, Setting, and Participants This comparative effectiveness study included NLST data from US adults at 33 participating centers enrolled between August 2002 and April 2004 with follow-up through 2009 along with National Health Interview Survey (NHIS) cross-sectional household interview survey data from 2010. Eligible participants were adults aged 55 to 74 years, and were current or former smokers with at least 30 pack-years of smoking (former smokers were required to have quit within the last 15 years). Transportability analyses combined baseline covariate, treatment, and outcome data from the NLST with covariate data from the NHIS and reweighted the trial data to the target population. Data were analyzed from March 2020 to May 2023. Interventions Low-dose CT or chest radiography screening with a screening assessment at baseline, then yearly for 2 more years. Main Outcomes and Measures For the outcomes of lung-cancer specific and all-cause death, mortality rates, rate differences, and ratios were calculated at a median (25th percentile and 75th percentile) follow-up of 5.5 (5.2-5.9) years for lung cancer–specific mortality and 6.5 (6.1-6.9) years for all-cause mortality. Results The transportability analysis included 51 274 NLST participants and 685 NHIS participants representing the target population (of approximately 5 700 000 individuals after survey-weighting). Compared with the target population, NLST participants were younger (median [25th percentile and 75th percentile] age, 60 [57 to 65] years vs 63 [58 to 67] years), had fewer comorbidities (eg, heart disease, 6551 of 51 274 [12.8%] vs 1 025 951 of 5 739 532 [17.9%]), and were more educated (bachelor’s degree or higher, 16 349 of 51 274 [31.9%] vs 859 812 of 5 739 532 [15.0%]). In the target population, for lung cancer–specific mortality, the estimated relative rate reduction was 18% (95% CI, 1% to 33%) and the estimated absolute rate reduction with low-dose CT vs chest radiography was 71 deaths per 100 000 person-years (95% CI, 4 to 138 deaths per 100 000 person-years); for all-cause mortality the estimated relative rate reduction was 6% (95% CI, −2% to 12%). In the NLST, for lung cancer–specific mortality, the estimated relative rate reduction was 21% (95% CI, 9% to 32%) and the estimated absolute rate reduction was 67 deaths per 100 000 person-years (95% CI, 27 to 106 deaths per 100 000 person-years); for all-cause mortality, the estimated relative rate reduction was 7% (95% CI, 0% to 12%). Conclusions and Relevance Estimates of the comparative effectiveness of low-dose CT screening compared with chest radiography in a nationally representative target population were similar to those from unweighted NLST analyses, particularly on the relative scale. Increased uncertainty around effect estimates for the target population reflects large differences in the observed characteristics of trial participants and the target population.
Investigators often believe that relative effect measures conditional on covariates, such as risk ratios and mean ratios, are “transportable” across populations. Here, we examine the identification of causal effects in a target population using an assumption that conditional relative effect measures are transportable from a trial to the target population. We show that transportability for relative effect measures is largely incompatible with transportability for difference effect measures, unless the treatment has no effect on average or one is willing to make even stronger transportability assumptions that imply the transportability of both relative and difference effect measures. We then describe how marginal (population-averaged) causal estimands in a target population can be identified under the assumption of transportability of relative effect measures, when we are interested in the effectiveness of a new experimental treatment in a target population where the only treatment in use is the control treatment evaluated in the trial. We extend these results to consider cases where the control treatment evaluated in the trial is only one of the treatments in use in the target population, under an additional partial exchangeability assumption in the target population (i.e., an assumption of no unmeasured confounding in the target population with respect to potential outcomes under the control treatment in the trial). We also develop identification results that allow for the covariates needed for transportability of relative effect measures to be only a small subset of the covariates needed to control confounding in the target population. Last, we propose estimators that can be easily implemented in standard statistical software and illustrate their use using data from a comprehensive cohort study of stable ischemic heart disease.
We consider the estimation of measures of model performance in a target population when covariate and outcome data are available on a sample from some source population and covariate data, but not outcome data, are available on a simple random sample from the target population. When outcome data are not available from the target population, identification of measures of model performance is possible under an untestable assumption that the outcome and population (source or target population) are independent conditional on covariates. In practice, this assumption is uncertain and, in some cases, controversial. Therefore, sensitivity analysis may be useful for examining the impact of assumption violations on inferences about model performance. Here, we propose an exponential tilt sensitivity analysis model and develop statistical methods to determine how sensitive measures of model performance are to violations of the assumption of conditional independence between outcome and population. We provide identification results and estimators for the risk in the target population, examine the large-sample properties of the estimators, and apply the estimators to data on individuals with stable ischemic heart disease.
When planning a cluster randomized trial, evaluators often have access to an enumerated cohort representing the target population of clusters. Practicalities of conducting the trial, such as the need to oversample clusters with certain characteristics in order to improve trial economy or support inferences about subgroups of clusters, may preclude simple random sampling from the cohort into the trial, and thus interfere with the goal of producing generalizable inferences about the target population. We describe a nested trial design where the randomized clusters are embedded within a cohort of trial-eligible clusters from the target population and where clusters are selected for inclusion in the trial with known sampling probabilities that may depend on cluster characteristics (e.g., allowing clusters to be chosen to facilitate trial conduct or to examine hypotheses related to their characteristics). We develop and evaluate methods for analyzing data from this design to generalize causal inferences to the target population underlying the cohort. We present identification and estimation results for the expectation of the average potential outcome and for the average treatment effect, in the entire target population of clusters and in its non-randomized subset. In simulation studies, we show that all the estimators have low bias but markedly different precision. Cluster randomized trials where clusters are selected for inclusion with known sampling probabilities that depend on cluster characteristics, combined with efficient estimation methods, can precisely quantify treatment effects in the target population, while addressing objectives of trial conduct that require oversampling clusters on the basis of their characteristics.
Methods for extending-generalizing or transporting-inferences from a randomized trial to a target population involve conditioning on a large set of covariates that is sufficient for rendering the randomized and nonrandomized groups exchangeable. Yet, decision makers are often interested in examining treatment effects in subgroups of the target population defined in terms of only a few discrete covariates. Here, we propose methods for estimating subgroup-specific potential outcome means and average treatment effects in generalizability and transportability analyses, using outcome model--based (g-formula), weighting, and augmented weighting estimators. We consider estimating subgroup-specific average treatment effects in the target population and its nonrandomized subset, and we provide methods that are appropriate both for nested and non-nested trial designs. As an illustration, we apply the methods to data from the Coronary Artery Surgery Study (North America, 1975-1996) to compare the effect of surgery plus medical therapy versus medical therapy alone for chronic coronary artery disease in subgroups defined by history of myocardial infarction.
Extending (generalizing or transporting) causal inferences from a randomized trial to a target population requires “generalizability” or “transportability” assumptions, which state that randomized and non-randomized individuals are exchangeable conditional on baseline covariates. These assumptions are made on the basis of background knowledge, which is often uncertain or controversial, and need to be subjected to sensitivity analysis. We present simple methods for sensitivity analyses that do not require detailed background knowledge about specific unknown or unmeasured determinants of the outcome or modifiers of the treatment effect. Instead, our methods directly parameterize violations of the assumptions using bias functions. We show how the methods can be applied to non-nested trial designs, where the trial data are combined with a separately obtained sample of non-randomized individuals, as well as to nested trial designs, where a clinical trial is embedded within a cohort sampled from the target population. We illustrate the methods using data from a clinical trial comparing treatments for chronic hepatitis C infection.
Most work on extending (generalizing or transporting) inferences from a randomized trial to a target population has focused on estimating average treatment effects (i.e., averaged over the target population's covariate distribution). Yet, in the presence of strong effect modification by baseline covariates, the average treatment effect in the target population may be less relevant for guiding treatment decisions. Instead, the conditional average treatment effect (CATE) as a function of key effect modifiers may be a more useful estimand. Recent work on estimating target population CATEs using baseline covariate, treatment, and outcome data from the trial and covariate data from the target population only allows for the examination of heterogeneity over distinct subgroups. We describe flexible pseudo-outcome regression modeling methods for estimating target population CATEs conditional on discrete or continuous baseline covariates when the trial is embedded in a sample from the target population (i.e., in nested trial designs). We construct pointwise confidence intervals for the CATE at a specific value of the effect modifiers and uniform confidence bands for the CATE function. Last, we illustrate the methods using data from the Coronary Artery Surgery Study (CASS) to estimate CATEs given history of myocardial infarction and baseline ejection fraction value in the target population of all trial-eligible patients with stable ischemic heart disease.
We present methods for causally interpretable meta-analyses that combine information from multiple randomized trials to estimate potential (counterfactual) outcome means and average treatment effects in a target population. We consider identifiability conditions, derive implications of the conditions for the law of the observed data, and obtain identification results for transporting causal inferences from a collection of independent randomized trials to a new target population in which experimental data may not be available. We propose an estimator for the potential (counterfactual) outcome mean in the target population under each treatment studied in the trials. The estimator uses covariate, treatment, and outcome data from the collection of trials, but only covariate data from the target population sample. We show that it is doubly robust, in the sense that it is consistent and asymptotically normal when at least one of the models it relies on is correctly specified. We study the finite sample properties of the estimator in simulation studies and demonstrate its implementation using data from a multi-center randomized trial.
Background/Aims When the randomized clusters in a cluster randomized trial are selected based on characteristics that influence treatment effectiveness, results from the trial may not be directly applicable to the target population. We used data from two large nursing home–based pragmatic cluster randomized trials to compare nursing home and resident characteristics in randomized facilities to eligible non-randomized and ineligible facilities. Methods We linked data from the high-dose influenza vaccine trial and the Music & Memory Pragmatic TRIal for Nursing Home Residents with ALzheimer’s Disease (METRICaL) to nursing home assessments and Medicare fee-for-service claims. The target population for the high-dose trial comprised Medicare-certified nursing homes; the target population for the METRICaL trial comprised nursing homes in one of four US-based nursing home chains. We used standardized mean differences to compare facility and individual characteristics across the three groups and logistic regression to model the probability of nursing home trial participation. Results In the high-dose trial, 4476 (29%) of the 15,502 nursing homes in the target population were eligible for the trial, of which 818 (18%) were randomized. Of the 1,361,122 residents, 91,179 (6.7%) were residents of randomized facilities, 463,703 (34.0%) of eligible non-randomized facilities, and 806,205 (59.3%) of ineligible facilities. In the METRICaL trial, 160 (59%) of the 270 nursing homes in the target population were eligible for the trial, of which 80 (50%) were randomized. Of the 20,262 residents, 973 (34.4%) were residents of randomized facilities, 7431 (36.7%) of eligible non-randomized facilities, and 5858 (28.9%) of ineligible facilities. In the high-dose trial, randomized facilities differed from eligible non-randomized and ineligible facilities by the number of beds (132.5 vs 145.9 and 91.9, respectively), for-profit status (91.8% vs 66.8% and 68.8%), belonging to a nursing home chain (85.8% vs 49.9% and 54.7%), and presence of a special care unit (19.8% vs 25.9% and 14.4%). In the METRICaL trial randomized facilities differed from eligible non-randomized and ineligible facilities by the number of beds (103.7 vs 110.5 and 67.0), resource-poor status (4.6% vs 10.0% and 18.8%), and presence of a special care unit (26.3% vs 33.8% and 10.9%). In both trials, the characteristics of residents in randomized facilities were similar across the three groups. Conclusion In both trials, facility-level characteristics of randomized nursing homes differed considerably from those of eligible non-randomized and ineligible facilities, while there was little difference in resident-level characteristics across the three groups. Investigators should assess the characteristics of clusters that participate in cluster randomized trials, not just the individuals within the clusters, when examining the applicability of trial results beyond participating clusters.
We discuss the identifiability of causal estimands for generalizability and transportability analyses, both under perfect and imperfect adherence to treatment assignment. We consider a setting where the trial data contain information on baseline covariates, assignment at baseline, intervention at baseline (point treatment), and outcomes; and where the data from non-randomized individuals only contain information on baseline covariates. In this setting, we review identification results under perfect adherence and study two examples in which non-adherence severely limits the ability to transport inferences about the effects of treatment assignment to the target population. In the first example, trial participation has a direct effect on treatment receipt and, through treatment receipt, on the outcome (a "trial engagement effect" via adherence). In the second example, participation in the trial has unmeasured common causes with treatment receipt. In both examples, the effect of assignment on the outcome in the target population is not identifiable. In the first example, however, the effect of joint interventions to scale-up trial activities that affect adherence and assign treatment is identifiable. We conclude that generalizability and transportability analyses should consider trial engagement effects via adherence and selection for participation on the basis of unmeasured factors that influence adherence.
Background US policymakers are debating whether to expand the Medicare program by lowering the age of eligibility. The goal of this study was to determine the association of Medicare eligibility and enrollment with healthcare access, affordability, and financial strain from medical bills in a contemporary population of low- and higher-income adults in the US. Methods and findings We used cross-sectional data from the National Health Interview Survey (2019) to examine the association of Medicare eligibility and enrollment with outcomes by income status using a local randomization-based regression discontinuity approach. After weighting to account for survey sampling, the low-income group consisted of 1,660,188 adults age 64 years and 1,488,875 adults age 66 years, with similar baseline characteristics, including distribution of sex (59.2% versus 59.7% female) and education (10.8% versus 12.5% with bachelor’s degree or higher). The higher-income group consisted of 2,110,995 adults age 64 years and 2,167,676 adults age 66 years, with similar distribution of baseline characteristics, including sex (40.0% versus 49.4% female) and education (41.0% versus 41.6%). The share of adults age 64 versus 66 years enrolled in Medicare differed within low-income (27.6% versus 87.8%, p < 0.001) and higher-income groups (8.0% versus 85.9%, p < 0.001). Medicare eligibility at 65 years was associated with a decreases in the percentage of low-income adults who delayed (14.7% to 6.2%; −8.5% [95% CI, −14.7%, −2.4%], P = 0.007) or avoided medical care (15.5% to 5.9%; −9.6% [−15.9%, −3.2%], P = 0.003) due to costs, and a larger decrease in the percentage who were worried about (66.5% to 51.1%; −15.4% [−25.4%, −5.4%], P = 0.003) or had problems (33.9% to 20.6%; −13.3% [−23.0%, −3.6%], P = 0.007) paying medical bills. In contrast, there were no significant associations between Medicare eligibility and measures of cost-related barriers to medication use. For higher-income adults, there was a large decrease in worrying about paying medical bills (40.5% to 27.5%; −13.0% [−21.4%, −4.5%], P = 0.003), a more modest decrease in avoiding medical care due to cost (3.5% to 0.6%; −2.9% [−5.3%, −0.5%], P = 0.02), and no significant association between eligibility and other measures of healthcare access and affordability. All estimates were stronger when examining the association of Medicare enrollment with outcomes for low and higher-income adults. Additional analyses that adjusted for clinical comorbidities and employment status were largely consistent with the main findings, as were analyses stratified by levels of educational attainment. Study limitations include the assumption adults age 64 and 66 would have similar outcomes if both groups were eligible for Medicare or if eligibility were withheld from both. Conclusions Medicare eligibility and enrollment at age 65 years were associated with improvements in healthcare access, affordability, and financial strain in low-income adults and, to a lesser extent, in higher-income adults. Our findings provide evidence that lowering the age of eligibility for Medicare may improve health inequities in the US.
In this article, we examine study designs for extending (generalizing or transporting) causal inferences from a randomized trial to a target population. Specifically, we consider nested trial designs, where randomized individuals are nested within a sample from the target population, and nonnested trial designs, including composite data-set designs, where observations from a randomized trial are combined with those from a separately obtained sample of nonrandomized individuals from the target population. We show that the counterfactual quantities that can be identified in each study design depend on what is known about the probability of sampling nonrandomized individuals. For each study design, we examine identification of counterfactual outcome means via the g-formula and inverse probability weighting. Last, we explore the implications of the sampling properties underlying the designs for the identification and estimation of the probability of trial participation.
Here we describe methods for assessing heterogeneity of treatment effects over prespecified subgroups in observational studies, using outcome-model-based (g-formula), inverse probability weighting, doubly robust, and matching estimators of subgroup-specific potential outcome means, conditional average treatment effects, and measures of heterogeneity of treatment effects. We compare the finite-sample performance of different estimators in simulation studies where we vary the total sample size, the relative frequency of each subgroup, the magnitude of treatment effect in each subgroup, and the distribution of baseline covariates, for both continuous and binary outcomes. We find that the estimators' bias and variance vary substantially in finite samples, even when there is no unobserved confounding and no model misspecification. As an illustration, we apply the methods to data from the Coronary Artery Surgery Study (August 1975-December 1996) to compare the effect of surgery plus medical therapy with that of medical therapy alone for chronic coronary artery disease in subgroups defined by previous myocardial infarction or left ventricular ejection fraction.
We describe methods that extend (generalize or transport) causal inferences from cluster randomized trials to a target population of clusters, under a general nonparametric model that allows for arbitrary within-cluster dependence. We propose doubly robust estimators of potential outcome means in the target population that exploit individual-level data on covariates and outcomes to improve efficiency and are appropriate for use with machine learning methods. We illustrate the methods using a cluster randomized trial of influenza vaccination strategies conducted in 818 nursing homes nested in a cohort of 4,475 trial-eligible Medicare-certified nursing homes.
We take steps toward causally interpretable meta-analysis by describing methods for transporting causal inferences from a collection of randomized trials to a new target population, one trial at a time and pooling all trials. We discuss identifiability conditions for average treatment effects in the target population and provide identification results. We show that the assumptions that allow inferences to be transported from all trials in the collection to the same target population have implications for the law underlying the observed data. We propose average treatment effect estimators that rely on different working models and provide code for their implementation in statistical software. We discuss how to use the data to examine whether transported inferences are homogeneous across the collection of trials, sketch approaches for sensitivity analysis to violations of the identifiability conditions, and describe extensions to address nonadherence in the trials. Last, we illustrate the proposed methods using data from the Hepatitis C Antiviral Long-Term Treatment Against Cirrhosis Trial.
We consider methods for causal inference in randomized trials nested within cohorts of trial-eligible individuals, including those who are not randomized. We show how baseline covariate data from the entire cohort, and treatment and outcome data only from randomized individuals, can be used to identify potential (counterfactual) outcome means and average treatment effects in the target population of all eligible individuals. We review identifiability conditions, propose estimators, and assess the estimators' finite-sample performance in simulation studies. As an illustration, we apply the estimators in a trial nested within a cohort of trial-eligible individuals to compare coronary artery bypass grafting surgery plus medical therapy vs. medical therapy alone for chronic coronary artery disease.