The Lung Health Study was a randomized clinical trial of smoking cessation that took place between 1986 and 1988. Special intervention participants received the smoking intervention program and were compared with usual care participants. Vital status was followed up to 32.5 years. Previous work sought to assess the effect of the cessation program on all-cause and cause-specific mortality at 14.5 years. Our objective was to do so at 32.5 years. We analyzed data from 5279 participants from the United States from the Lung Health Study. The three arms were smoking intervention plus bronchodilator, smoking intervention plus placebo, or no intervention. We compared all-cause and cause-specific mortality at 32.5 years between the combined special intervention groups and usual care group. The hazard ratio for the usual care group compared with the special intervention group was 1.05 (95% CI, 0.97-1.18) at 32.5 years. The cause-specific hazard ratio for the usual care group compared with the special intervention group for death due to respiratory disease that was not lung cancer was 1.21 (95% CI, 1.04-1.42) at 32.5 years, indicating that the smoking cessation program had a protective effect against death due to non-lung cancer respiratory disease, even after a long period.
The long term consequences of unwanted pregnancies carried to term on mothers have not been much explored. We use data from the Wisconsin Longitudinal Study (WLS) and propose a novel approach, namely two team cross-screening, to study the possible effects of unwanted pregnancies carried to term on various aspects of mothers' later-life mental health, physical health, economic well-being and life satisfaction. Our method, unlike existing approaches to observational studies, enables the investigators to perform exploratory data analysis, confirmatory data analysis and replication in the same study. This is a valuable property when there is only a single data set available with unique strengths to perform exploratory, confirmatory and replication analysis. In two team cross-screening, the investigators split themselves into two teams and the data is split as well according to a meaningful covariate. Each team then performs exploratory data analysis on its part of the data to design an analysis plan for the other part of the data. The complete freedom of the teams in designing the analysis has the potential to generate new unanticipated hypotheses in addition to a prefixed set of hypotheses. Moreover, only the hypotheses that looked promising in the data each team explored are forwarded for analysis (thus alleviating the multiple testing problem). These advantages are demonstrated in our study of the effects of unwanted pregnancies on mothers' later life outcomes.
In the analyses of cluster-randomized trials, mixed-model analysis of covariance (ANCOVA) is a standard approach for covariate adjustment and handling within-cluster correlations. However, when the normality, linearity, or the random-intercept assumption is violated, the validity and efficiency of the mixed-model ANCOVA estimators for estimating the average treatment effect remain unclear. Under the potential outcomes framework, we prove that the mixed-model ANCOVA estimators for the average treatment effect are consistent and asymptotically normal under arbitrary misspecification of its working model. If the probability of receiving treatment is 0.5 for each cluster, we further show that the model-based variance estimator under mixed-model ANCOVA1 (ANCOVA without treatment-covariate interactions) remains consistent, clarifying that the confidence interval given by standard software is asymptotically valid even under model misspecification. Beyond robustness, we discuss several insights on precision among classical methods for analyzing cluster-randomized trials, including the mixed-model ANCOVA, individual-level ANCOVA, and cluster-level ANCOVA estimators. These insights may inform the choice of methods in practice. Our analytical results and insights are illustrated via simulation studies and analyses of three cluster-randomized trials.
The log-rank test and Kaplan–Meier plot are standard tools for analyzing time-to-event data in randomized clinical trials, yet neither provides a summary of the magnitude of the treatment effect. Practitioners typically fill this gap by reporting a hazard ratio from a Cox proportional-hazards model or an acceleration factor from an accelerated failure time (AFT) model, but both require assumptions beyond those needed for the log-rank test or Kaplan–Meier estimator. We propose two nonparametric confidence intervals for scalar effect-size summaries, an additive shift c and a multiplicative factor ρ, obtained by inverting the log-rank test under sharp null hypotheses of constant treatment effects. Building on the randomization-inference framework of Li and Small (2023), both intervals are valid under the randomization distribution alone, requiring no assumptions for the event-time distribution. We evaluate the proposed multiplicative interval via simulation, finding that it maintains nominal coverage across a range of censoring rates and sample sizes, including under data-generating processes that misspecify a parametric AFT model, while incurring only a modest efficiency loss compared to parametric AFT inference under correct specification. We illustrate the approach using data from a randomized trial of rhDNase for cystic fibrosis and provide R code and a Shiny application for ease of implementation.
Observational studies are valuable tools for inferring causal effects in the absence of controlled experiments. However, these studies may be biased due to the presence of some relevant, unmeasured set of covariates. One approach to mitigate this concern is to identify hypotheses likely to be more resilient to hidden biases by splitting the data into a planning sample for designing the study and an analysis sample for making inferences. We devise a powerful and flexible method for selecting hypotheses in the planning sample when an unknown number of outcomes are affected by the treatment, allowing researchers to gain the benefits of exploratory analysis and still conduct powerful inference under concerns of unmeasured confounding. We investigate the theoretical properties of our method and conduct extensive simulations that demonstrate pronounced benefits, especially at higher levels of allowance for unmeasured confounding. Finally, we demonstrate our method in an observational study of the multi-dimensional impacts of a devastating flood in Bangladesh.
Adverse childhood experiences (ACEs) have been linked to a wide range of negative health outcomes in adulthood. However, few studies have investigated what specific combinations of ACEs most substantially impact mental health. In this article, we provide the protocol for our observational study of the effects of combinations of ACEs on adult depression. We use data from the 2023 Behavioral Risk Factor Surveillance System (BRFSS) to assess these effects. We will evaluate the replicability of our findings by splitting the sample into two discrete subpopulations of individuals. We employ data turnover for this analysis, enabling a single team of statisticians and domain experts to collaboratively evaluate the strength of evidence, and also integrating both qualitative and quantitative insights from exploratory data analysis. We outline our analysis plan using this method and conclude with a brief discussion of several specifics for our study.
BACKGROUND:The BE ACTIVE trial (Behavioral Economic Approaches to Increase Physical Activity Among Patients with Elevated Risk for Cardiovascular Disease) documented the effectiveness, compared with an attention control arm that received daily text messages, of gamification, financial incentives, or gamification+financial incentives to increase steps/day. Increases in daily step count are associated with longer life expectancy, but understanding the cost-effectiveness of these interventions is essential for payers and other stakeholders seeking to implement findings. METHODS:We built a probabilistic Markov model to compare intervention costs with lifetime estimates of life-years and quality-adjusted life-years for 2 sets of comparisons: (1) each behavioral intervention versus attention control, and (2) each trial arm, including attention control, versus no intervention. Since the durability of changes in steps/day post-intervention is unknown, we modeled optimistic, intermediate, and pessimistic scenarios. RESULTS:Over the 12-month intervention, per-participant cost to deliver attention control was $878, gamification $938, financial incentives $1534, and gamification+financial incentives $1712. Compared with attention control, gamification's cost-effectiveness ranged from $261 (95% CI, 259-263) per life-year gained if mean steps/day during the last 18 weeks of follow-up are maintained (optimistic), to $30 550 (95% CI, 30 503-30 597) per life-year if steps/day continue to decline at the rate observed during the full 26-week follow-up (pessimistic). Gamification+financial incentives cost <$50 000/life-year only under the optimistic and intermediate scenarios. Financial incentives was dominated by gamification and gamification+financial incentives. When all 4 trial arms, including attention control, were compared with no intervention, gamification again cost <$50 000/life-year across all durability scenarios. CONCLUSIONS:Across a range of scenarios about the durability of increases in steps/day post-intervention, gamification consistently cost <$50 000 per life-year gained, the threshold for high value interventions set by American College of Cardiology/American Heart Association guidelines. Gamification+financial incentives was high-value except in the pessimistic scenario. Financial incentives was dominated. REGISTRATION:URL: https://www.clinicaltrials.gov; Unique identifier: NCT03911141.
The case$^2$ study, also referred to as the case-case study design, is a valuable approach for conducting inference for treatment effects. Unlike traditional case-control studies, the case$^2$ design compares treatment in cases of concern (the first type of case) to other cases (the second type of case). One of the quantities of interest is the attributable effect for the first type of case-that is, the number of the first type of case that would not have occurred had the treatment been withheld from all units. In some case$^2$ studies, a key quantity of interest is the attributable effect for the first type of case. Two key assumptions that are usually made for making inferences about this attributable effect in case$^2$ studies are (1) treatment does not cause the second type of case, and (2) the treatment does not alter an individual's case type. However, these assumptions are not realistic in many real-data applications. In this article, we present a sensitivity analysis framework to scrutinize the impact of deviations from these assumptions on inferences for the attributable effect. We also include sensitivity analyses related to the assumption of unmeasured confounding, recognizing the potential bias introduced by unobserved covariates. The proposed methodology is exemplified through an investigation into whether having violent behavior in the last year of life increases suicide risk using the 1993 National Mortality Followback Survey dataset.
Test-negative designs (TNDs) are widely used for postmarket evaluation of vaccine effectiveness (VE), particularly in cases when randomized trials are not feasible. Unlike classical TNDs, which only include healthcare seekers with symptoms, recent TNDs have involved individuals with various reasons for testing, especially in an outbreak setting. While including these data can increase sample size and hence improve precision, concerns have been raised about whether they introduce bias into the current framework of TNDs, thereby demanding a formal statistical examination of this modified design. In this article, using statistical derivations, causal graphs, and numerical demonstrations, we show that the standard odds ratio estimator may be biased if various reasons for testing are not taken into account. To eliminate this bias, we identify three categories of reasons for testing, namely symptoms, mandatory screening, and case contact tracing, and characterize associated statistical properties and estimands. Based on our characterization, we show how to consistently estimate each estimand via stratification. Furthermore, we describe when these estimands correspond to the same VE parameter and, when appropriate, propose a stratified estimator that can incorporate multiple reasons for testing and improve precision. We demonstrate the performance of our proposed method through simulation studies and a real-data analysis.
Does having firearms in the home increase suicide risk? To test this hypothesis, a matched case-control study can be performed, in which suicide case subjects are compared to living controls who are similar in observed covariates in terms of their retrospective exposure to firearms at home. In this application, cases can be defined using a broad case definition (suicide) or a narrow case definition (suicide occurred at home). The broad case definition offers a larger number of cases but the narrow case definition may offer a larger effect size. Moreover, restricting to the narrow case definition may introduce selection bias (i.e., bias due to selecting samples based on characteristics affected by the treatment) because exposure to firearms in the home may affect the location of suicide and thus the type of a case a subject is. We propose a new sensitivity analysis framework for combining broad and narrow case definitions in matched case-control studies, that considers the unmeasured confounding bias and selection bias simultaneously. We develop a valid randomization-based testing procedure using only the narrow case matched sets when the effect of the unmeasured confounder on receiving treatment and the effect of the treatment on case definition among the always-cases are controlled by sensitivity parameters. We then use the Bonferroni method to combine the testing procedures using the broad and narrow case definitions. With the proposed methods, we find robust evidence that having firearms at home increases suicide risk.
The attributable fraction among the exposed (AFe) is the proportion of disease cases among the exposed that could be avoided by eliminating the exposure. In this article, we propose a new approach to reduce sensitivity to hidden bias for conducting statistical inference on the AFe by leveraging case description information such as subtype of cancer. The proposed method is examined through an asymptotic tool, design sensitivity, simulation studies, and case studies of alcohol consumption and the risk of postmenopausal invasive breast cancer utilizing information on the subtype of cancer using data from the Women’s Health Initiative Observational Study allowing the possibility that leveraging case definition information may introduce selection bias through an additional sensitivity parameter.
This paper presents a causal inference estimation method for longitudinal observational studies with multiple outcomes. The method uses marginal structural models with inverse probability treatment weights (MSM-IPTWs). In developing the proposed method, we re-define the weights as a product of inverse weights at each time point, accounting for time-varying confounders and treatment exposures and possible correlation between and within (serial) the multiple outcomes. The proposed method is evaluated by simulation studies and with an application to estimate the effect of HIV positivity awareness on condom use and multiple sexual partners using the Malawi Longitudinal Study of Families and Health (MLSFH) data. The simulation study shows that the joint MSM-IPTW performs well with coverage within the expected 95% level for a large sample size (n = 1000) and moderate to strong between and within outcome correlation strength ( ρ j = 0.3 , 0.75, ρ k = 0.4 , 0.8) when the effects are similar. The joint MSM-IPTW performed relatively the same as the adjusted standard joint model when the treatment effect estimate was the same for the outcomes. In the application, HIV positivity awareness increased the usage of condoms and did not affect the number of sexual partners. We recommend using the proposed MSM-IPTWs to correctly control for time-varying treatment and confounders when estimating causal effects for longitudinal observational studies with multiple outcomes.
Objectives. To test low-cost, scalable interventions designed to encourage seat belt use (primary outcome) and discourage handheld phone use while driving. Methods. A randomized controlled trial assigned 1139 consenting General Motors‒connected vehicle customers in the United States to 1 of 4 groups for a 10-week intervention: (1) control, (2) behavioral engagement, (3) behavioral engagement plus raffle, and (4) behavioral engagement plus shared pot. Behavioral engagement involved education, personalized tips, a "wish outcome obstacle plan" exercise, and weekly feedback about buckling and handheld-free streaks. Participants in the behavioral engagement plus raffle group also earned a chance at a $125 prize each week they had a buckling or handheld-free streak. Those in the behavioral engagement plus shared pot group earned an equal share of this prize for each streak. The intervention was delivered virtually in spring 2023. Results. Participants in the behavioral engagement plus shared pot group had a higher buckling rate (91.3%) than those in the behavioral engagement plus raffle (89.5%), behavioral engagement (89.4%), or control (88.3%) groups-differences that remained significant at follow-up. Handheld phone use did not differ significantly. Conclusions. A behavioral intervention with a shared pot incentive could be delivered at scale to reduce injuries and deaths associated with vehicular crashes. Trial Registration. ClinicalTrials.gov identifier: NCT05469477. (Am J Public Health. 2025;115(5):758-768. https://doi.org/10.2105/AJPH.2024.307980).
Importance:Guidelines recommend that intensive care unit (ICU) clinicians consider prognosis and offer a comfort-focused treatment alternative to patients with limited prognoses to promote preference-sensitive treatment decisions. Objective:To determine whether nudging ICU clinicians to adhere to communication guidelines improves outcomes among critically ill patients at high risk of death or severe functional impairment. Design, Setting, and Participants:This 4-arm pragmatic, stepped-wedge, cluster randomized trial (conducted February 1, 2018-October 31, 2020, follow-up through April 29, 2021, and analyses December 2023-January 2024) involved 3500 encounters of adults with chronic serious illness receiving mechanical ventilation for at least 48 hours at 10 hospitals comprising 17 medical, surgical, specialty, or mixed ICUs in community, rural, and urban settings. Interventions:Two clinician-directed electronic health record nudge interventions were each compared with usual care alone and combined: document of 6-month functional prognosis and whether a comfort-focused treatment alternative was offered or a reason why not. Main Outcomes and Measures:The primary outcome was hospital length of stay, with death coded at the 99th percentile. Secondary end points included 22 measures of acute care utilization, end-of-life care processes, and mortality. Results:Of 3500 patient encounters among 3250 patients (mean [SD] age, 63.2 [13.5] years; 46.1% female), 3384 encounters (96.7%) had complete baseline data and were included in risk-adjusted analyses. The overall intervention document completion rate for all patients was 75.0% (n = 1714) and similar across groups. Among the 3500 encounters, observed hospital mortality was 35.7% (n = 1249), and the median observed length of stay was 8.93 days (IQR, 4.64-16.23). The median length of stay with deaths coded as the 99th percentile did not differ between any intervention and usual care groups (for length of stay, all adjusted median difference 95% CIs include 0; for hospital mortality, all adjusted risk difference [RD] 95% CIs include 0). Results were similar in sensitivity analyses with death coded as low at the fifth percentile and without ranking deaths. Compared with usual care, a higher percentage of patients were discharged to hospice in the treatment alternative group (10.9% vs 7.3%; adjusted RD, 6% [95% CI, 1%-10%]) and the combined group (8.9% vs 7.3%; adjusted RD, 6% [95% CI, 0%-12%]). The treatment alternative intervention led to earlier comfort-care orders (3.6 vs 4.5 days; adjusted hazard ratio, 1.42 [95% CI, 1.06-1.92]). The 20 other secondary end points were unaffected by the interventions. Conclusions and Relevance:This cluster randomized clinical trial found that electronically nudging ICU clinicians to adhere to communication guidelines was feasible but did not reduce hospital length of stay. Trial Registration:ClinicalTrials.gov Identifier: NCT03139838.
This article discusses a sensitivity analysis for an instrumental variable (IV) estimate in the presence of many instruments that are weakly associated with the endogenous variable. We study the effect of imprisonment on earnings using data on individuals sentenced for felony in Michigan in the years 2003-2006. Motivated by the random assignment of judges to cases, we construct a vector of instruments based on judges' ID. Our data has two important features that cannot be handled using standard IV approaches. First, while some judges exhibit strong tendencies toward a prison or nonprison sentence, many judges do not have strong tendencies toward a particular sentence type. Second, our data includes only cases that result in sentencing, and thus the standard analyses are subject to selection bias. We develop a sensitivity analysis procedure that is robust to the presence of many weak instruments and quantifies the effect of the selection bias on the parameter of interest. A power formula for the sensitivity analysis is also provided. Analyses show that being sentenced to prison significantly reduces the offenders' earnings. Our simulation studies highlight the value of the proposed method in terms of statistical power and also confirm the validity of our power formula.
While palliative care is increasingly commonly delivered to hospitalized patients with serious illnesses, few studies have estimated its causal effects. Courtright et al. (2016) adopted a stepped-wedge cluster-randomized design to assess the effect of palliative care on a patient-centered outcome. The randomized intervention was a nudge to administer palliative care but did not guarantee receipt of palliative care, resulting in noncompliance. A subsequent analysis using methods suited for standard trial designs produced statistically anomalous results, as an intention-to-treat analysis found no effect while an instrumental variable analysis did (Courtright et al. 2024). This highlights the need for a more principled approach to address noncompliance in stepped-wedge designs. We provide a formal causal inference framework for the stepped-wedge design with noncompliance by introducing a relevant causal estimand and corresponding estimators and inferential procedures. Through numerical studies, we compare an array of estimators and provide practical guidance in choosing an analysis method. Finally, we apply our recommended methods to reanalyze the palliative care pragmatic trial, producing point estimates suggesting a larger effect than the original analysis, but intervals that did not reach statistical significance.
Objective: Gun violence is a serious public health problem in the United States. The Gun Violence Archive (GVA) provides detailed geographic information, while the National Violent Death Reporting System (NVDRS) offers demographic, socioeconomic, and narrative data on gun homicides. We developed and tested a method for merging datasets to inform analysis and strategies to reduce gun violence rates in the United States.Materials and Methods: After preprocessing the data, we used a probabilistic record linkage program to link records from the GVA (n = 36 245) with records from the NVDRS (n = 30 592). We evaluated sensitivity (the false-match rate) by using a manual approach.Results: The linkage returned 27 420 matches of gun violence incidents from the GVA and NVDRS datasets. Because of restricted details accessible from GVA online records, only 942 of these matched records could be manually evaluated. Our framework achieved a 90.1% (849 of 942) accuracy rate in linking GVA incidents with corresponding NVDRS records.Practice Implications: Electronic linkage of gun violence data from 2 sources is feasible and can be used to increase the utility of the datasets.
Objective: Gun violence is a serious public health issue in the United States. The Gun Violence Archive (GVA) provides detailed geographic information, while The National Violent Death Reporting System (NVDRS) offers demographic, socioeconomic, and narrative data about gun homicides. We develop and test a method for merging data sets, each with its own strengths, to overcome their individual limitations. This merged data set can inform analysis and strategies to reduce high gun violence rates in the US. Methods: After preprocessing the data, we used a probabilistic record linkage program to link records from the Gun Violence Archive (GVA) (n=36,245) with records from The National Violent Death Reporting System (NVDRS) (n=30,592). Sensitivity (the false match rate) was evaluated using a manual approach. Results: The linkage returned 27,420 matches of gun violence incidents from the GVA and NVDRS data sets. Of these cases, 942 records were able to be manually evaluated due to the restricted details accessible from GVA records. Our framework achieves a 90.12 corresponding NVDRS records. Conclusion: Electronic linkage of gun violence data from two different sources is feasible, and can be used to increase the utility of the data sets.
BACKGROUND:The majority of people in the United States do not achieve recommended levels of physical activity. Even small, daily increases can have health benefits. Wearable devices paired with social incentives increased daily steps in pilot studies but have not been tested for long-term effectiveness in community settings. This paper describes the study design and baseline participant characteristics of a trial testing these approaches to increase physical activity among families in the Philadelphia area. METHODS:The trial, called STEP Together, is a Hybrid Type 1 effectiveness-implementation study. Participants enroll on family teams of 2-10 people, including at least one person 60 years old or older. Each participant receives a Fitbit device, establishes a baseline daily step count, and selects a daily step goal 1500 to 3000 steps greater than their baseline. Family teams are stratified based on family size and randomized to Control, Social Incentive Gamification, or Social Goals through Incentives to Charity. Participation is 18-months: a 12-month intervention and 6-month follow up. RESULTS:779 participants on 285 family teams were randomized. Recruitment was more difficult than anticipated due to the COVID-19 pandemic and higher-than expected numbers of participants who were already physically active and therefore ineligible. Changes to the eligibility criteria that did not impact the underlying intent or conceptual basis for the trial improved recruitment feasibility. CONCLUSION:The results from this study will contribute to the growing body of evidence about scalable, effective strategies to motivate individuals and families to increase their daily physical activity. CLINICAL TRIAL REGISTRATION NUMBER:NCT04942535.
ImportanceIncreasing inpatient palliative care delivery is prioritized, but large-scale, experimental evidence of its effectiveness is lacking.ObjectiveTo determine whether ordering palliative care consultation by default for seriously ill hospitalized patients without requiring greater palliative care staffing increased consultations and improved outcomes.Design, Setting, and ParticipantsA pragmatic, stepped-wedge, cluster randomized trial was conducted among patients 65 years or older with advanced chronic obstructive pulmonary disease, dementia, or kidney failure admitted from March 21, 2016, through November 14, 2018, to 11 US hospitals. Outcome data collection ended on January 31, 2019.InterventionOrdering palliative care consultation by default for eligible patients, while allowing clinicians to opt-out, was compared with usual care, in which clinicians could choose to order palliative care.Main Outcomes and MeasuresThe primary outcome was hospital length of stay, with deaths coded as the longest length of stay, and secondary end points included palliative care consult rate, discharge to hospice, do-not-resuscitate orders, and in-hospital mortality.ResultsOf 34 239 patients enrolled, 24 065 had lengths of stay of at least 72 hours and were included in the primary analytic sample (10 313 in the default order group and 13 752 in the usual care group; 13 338 [55.4%] women; mean age, 77.9 years). A higher percentage of patients in the default order group received palliative care consultation than in the standard care group (43.9% vs 16.6%; adjusted odds ratio [aOR], 5.17 [95% CI, 4.59-5.81]) and received consultation earlier (mean [SD] of 3.4 [2.6] days after admission vs 4.6 [4.8] days; P < .001). Length of stay did not differ between the default order and usual care groups (percent difference in median length of stay, −0.53% [95% CI, −3.51% to 2.53%]). Patients in the default order group had higher rates of do-not-resuscitate orders at discharge (aOR, 1.40 [95% CI, 1.21-1.63]) and discharge to hospice (aOR, 1.30 [95% CI, 1.07-1.57]) than the usual care group, and similar in-hospital mortality (4.7% vs 4.2%; aOR, 0.86 [95% CI, 0.68-1.08]).Conclusions and RelevanceDefault palliative care consult orders did not reduce length of stay for older, hospitalized patients with advanced chronic illnesses, but did improve the rate and timing of consultation and some end-of-life care processes.Trial RegistrationClinicalTrials.gov Identifier: NCT02505035
S. Hennessy合作论文数20