
Accurate prediction of enrollment and interim analysis (IA) timing is critical for the effective management of clinical trials employing group sequential designs. Traditional forecasting methods often focus on trials with independent time-to-event outcomes but no effective statistical methods have been developed for trials involving recurrent count data, such as vaso-occlusive crisis counts in sickle cell disease or number of bleeding events in hemophilia. This manuscript introduces a novel integrated framework designed to predict both future enrollment and recurrent count outcomes in group sequential trials in the setting of a randomized two-arm comparative trial. The framework integrates existing methodologies for count data and enrollment predictions with new approaches tailored to the recurrent nature of outcomes, aiming to enhance prediction accuracy and reliability. Specifically, we develop a prediction model based on a negative binomial distribution to handle the overdispersed count data and incorporate enrollment projections. We validate our framework through extensive numerical simulations and apply it to a real phase 3 study in sickle cell disease. The results demonstrate that the proposed method has desirable performance in the prediction for future enrollment time and timing of IA defined by the information fraction (IF) with high accuracy and reduced bias. The findings underscore the framework's potential to improve trial management by providing more accurate and reliable IA timing prediction and facilitating better decision-making in clinical trials with recurrent count outcomes.
To estimate the effect of treatment on the frequency of recurrent events-such as asthma or chronic obstructive pulmonary disease exacerbation events, repeat infections, or falls-there are two analytical approaches frequently used in practice: the Andersen-Gill model with robust variance and the negative binomial model. We evaluate these as well as an Andersen-Gill model with Gaussian random effects. This study aims to enhance our understanding of the relative operating characteristics of these methods using simulations. First, we examined constant and proportional intensity functions with variable average number of events and degrees of between-subject heterogeneity. Second, we allowed intensity to change as a function of time since randomization, time since previous event, or number of previous events. Our findings provide evidence that the negative binomial model is not as robust as the Andersen-Gill model with robust variance in settings with deviations from its distributional assumptions. The Andersen-Gill model with random effects performs well in most settings but is less sensitive due to overestimation of the variance when the intensity increases as a function of time since the previous event. The simulations are contextualized by an analysis of a randomized trial for dupilumab in patients with uncontrolled moderate-to-severe asthma.
Network meta-analysis (NMA) has become a cornerstone of evidence synthesis, enabling simultaneous comparison of multiple interventions using both direct and indirect evidence. However, recent developments in the estimand framework, formalized in ICH E9 (R1), have highlighted conceptual ambiguities in how treatment effects are defined, interpreted, and synthesized in complex evidence networks. We review the core components of the estimand framework and evaluate how each is handled in standard NMA practice. Particular attention is given to heterogeneity in strategies for handing intercurrent events across trials, including treatment switching and non-adherence, and to the reliance of NMA on transitivity and consistency assumptions when underlying estimands may not be aligned. We identify several key challenges in the estimation for NMA. First, individual trials within a network may target different estimands, despite nominally similar comparisons, leading to incoherent synthesis targets. Second, NMA models may implicitly combine estimands aligned with different intercurrent event strategies, complicating causal interpretation of NMA estimates. Failure to align estimands across trials and comparisons can compromise transitivity, distort treatment rankings, and mislead decision-making. We argue that incorporating the estimand framework formally into NMA design, analysis, and reporting is critical for improving the transparency, interpretability, and credibility of evidence synthesized by NMA.
In clinical trial conduct, it is critical to monitor some important parameters based on accumulating data to ensure quality, inform next-stage planning and protect patients' safety. The typical approach based on observed outcomes relies on a strong extrapolation assumption, which may be violated in practical settings due to heterogeneity of enrolled patients and mechanisms of missingness. Multiple imputation (MI) is an appealing alternative to retrieve unobserved outcomes based on available covariate information. In this article, we propose to leverage B-splines as the imputation model in MI to ensure robustness. Simulation studies show that the proposed MI-B can accurately and efficiently estimate the targeted parameter under varying and unknown functional forms. We apply this approach to a hypothetical study to monitor the premature study discontinuation rate and ensure a successful trial execution.
Health technology assessment (HTA) is a transdisciplinary and collaborative process evaluating the clinical, economic, and societal impact of new medical interventions, thereby informing payer decisions. The European Union HTA Regulation (HTAR) and its Joint Clinical Assessment (JCA) procedure mandate the early submission of comprehensive evidence packages by health technology developers, amplifying the need for coordinated statistical expertise throughout the development lifecycle. This review, prepared by a scientific working group of the American Statistical Association's Biopharmaceutical Section, delineates how statisticians can enhance their impact across the HTA lifecycle. It examines the evolving evidence-generation landscape, emphasizing the benefits of early collaboration between clinical development and market access teams. The article provides guidance on the conduct of robust indirect treatment comparisons and discusses best practices for the collection and analysis of patient-reported outcome (PRO) data. Within the domain of economic modeling, we detail the role of statisticians across a range of activities, and in the realm of artificial intelligence (AI), the potential applications of automation and generative and agentic AI are explored. Through technical expertise and leadership, statisticians can collaborate with cross-functional teams, ensuring that HTA decisions are grounded in rigorous, patient-centered, and timely evidence-facilitating equitable and efficient access to innovative therapies.
Biomarkers, which may offer insights into the potential benefit in the clinical outcome for an experimental treatment, play a critical role in early phase drug development before large, time-consuming, and expensive phase 3 trials are conducted. Evaluating the impact of a biomarker on a time-to-event clinical outcome could be challenging as it is difficult to model the temporal relationship between the biomarker and the clinical outcome. We demonstrate through simulation that directly incorporating the biomarker as a covariate into a Cox proportional hazards model may lead to an attenuated relationship between the biomarker and the clinical outcome when there is a delayed biomarker effect on the clinical outcome. Although we do not provide a sophisticated way to model the delayed temporal relationship between the biomarker and the clinical outcome, understanding the limitations of conventional methods should help guide decision-making. Future research is needed on how to improve statistical models in evaluating the impact of biomarkers on time-to-event clinical outcomes.
Post-randomization treatment switching is common in randomized trials and can bias treatment-effect estimates when the estimand targets the hypothetical outcome under no switching. Existing adjustments either rely on strong structural assumptions about treatment effects (e.g., on the failure-time scale) or sacrifice effective sample size by censoring switching patients. We propose a causal machine-learning framework that (i) learns an arm-specific outcome model using non-switching patients, (ii) corrects prognosis-driven selection bias via transfer-learning-based reweighting to align the covariate distribution of non-switching patients with that of switching patients, and (iii) predicts counterfactual outcomes for switching patients as if they had remained on their randomized treatment. Treatment effects for the no-switch estimand are then computed on the reconstructed no-switch population using standard estimators. The approach extends naturally to multi-directional switching and multi-arm trials. Comprehensive simulation studies and a real-data application demonstrate the practical performance of the proposed framework.
The E-value has been proposed as a tool for sensitivity analysis for unmeasured confounding in observational studies (VanderWeele and Ding). As it was developed in the setting of risk ratio comparisons, using E-values for proportional hazards models requires approximating a risk ratio from a hazard ratio. We consider the impact of the bias induced when using the suggested approximation method, as well as how doing so may exacerbate recognized E-value interpretation issues. A running example helps illustrate how challenges of biased E-values may arise in time-to-event studies and also ways in which alternative approaches can avoid the bias.
We propose an enhanced Doubly Robust Estimator (eDRE) for binary outcomes in causal inference, aiming to improve robustness by relaxing assumptions imposed by the original doubly robust estimator. The algorithm combines a semiparametric method with propensity score weighting and outcome regression model, providing consistent estimates even when any one of the model specifications is incorrect. We discuss the properties of the eDRE, highlighting its robustness to model misspecification. To confirm the feasibility and effectiveness of the eDRE, we conduct extensive simulations under various conditions. The results demonstrate that the eDRE outperforms traditional Inverse Probability Score Weighted Estimator (IPSWE) and naive estimation methods, showing reduced bias when estimating treatment effect. To further demonstrate the applicability of our method, we apply the eDRE to a clinical dataset from a knee osteoarthritis study. The analysis uncovered a larger effect than indicated by unadjusted methods, but it was insignificant. We discuss the method's limitations and future improvements in the end.
Prognostic covariate adjustment (PROCOVA) incorporates a prognostic score, trained from historical control data, into the analysis of randomized clinical trials (RCTs) to reduce variance, when randomization is appropriate and alignment is strong between the historical model and the current trial control population. However, the performance of PROCOVA under non-ideal conditions, such as improper randomization including non-randomized settings, prognostic misalignment, or multicollinearity, remains less understood. In this work, we develop a projection representation illustrating how these non-ideal conditions can induce bias or inflate variance in PROCOVA.Motivated by this, we evaluate doubly robust estimators, augmented inverse probability weighting (AIPTW) and targeted maximum likelihood estimation (TMLE), to mitigate PROCOVA-induced bias and variance inflation. Simulations demonstrate that PROCOVA regression performs well under ideal conditions, but can induce bias, inflate variance, or deliver minimal precision gain in non-ideal conditions. In contrast, AIPTW and TMLE remain valid and achieve improved efficiency across all scenarios.These findings provide practical guidance for applying PROCOVA, particularly in pre-specified trials at the design stage. Our results support incorporating prognostic scores within doubly robust estimators to enhance robustness while preserving efficiency. TMLE is especially attractive in settings that use regularized or high-dimensional nuisance modeling, and it can offer more stable performance than AIPTW.
In comparative studies with time-to-event outcomes, stratified analyses are routinely used to adjust for imbalances in baseline characteristics between treatment groups, with the stratified Cox proportional hazards model being the most common choice. However, its validity depends on strong assumptions-proportional hazards within each stratum and a common hazard ratio across strata-that are rarely satisfied in practice. To address these limitations, model-free, nonparametric alternatives have been proposed within the frequentist framework. We briefly review these methods and then introduce a Bayesian counterpart that offers greater flexibility and a more intuitive interpretation of treatment effects. We illustrate both the frequentist and Bayesian approaches using data from the beta-Blocker Evaluation of Survival Trial, evaluating treatment effects on an event-free survival endpoint. When only minimal prior information about treatment efficacy is available, Bayesian estimates typically resemble their frequentist counterparts. However, the Bayesian approach provides enhanced validity under weaker assumptions and more transparent interpretation. Moreover, it can directly quantify the probability that a treatment effect exceeds a prespecified, clinically meaningful threshold, a capability not readily available with standard frequentist methods. Open-source software for implementing the proposed Bayesian procedure is also provided.
Accurately estimating disease prevalence in finite populations is essential in epidemiology and ecological studies. Standard methods based on the binomial model often rely on assumptions of infinite populations and perfect diagnostic tests, which are frequently violated in practice. When sampling is without replacement from a finite population, the hypergeometric distribution provides the appropriate framework, but diagnostic misclassification must also be accounted for. In this paper, we present a framework for prevalence estimation in finite populations under imperfect diagnostic testing. The framework incorporates misclassification through diagnostic sensitivity and specificity, which may be treated either as fixed or as parameters estimated from independent studies. We evaluate multiple inferential approaches, including exact hypergeometric confidence intervals, a computationally efficient simulation-based approximation, a Bayesian formulation yielding posterior distributions, and a profile likelihood method that jointly accounts for uncertainty in sensitivity and specificity. Using extensive simulations we assess coverage, interval precision, and point estimate accuracy. Results show that hypergeometric-based methods consistently yield narrower and better-calibrated intervals than binomial-based alternatives. The profile likelihood approach achieves near-nominal coverage while appropriately propagating diagnostic uncertainty. An application to epidemiological surveillance data illustrates the practical relevance of the proposed methods.
The statistical evaluation of time-to-event endpoints like overall survival demands careful consideration as commonly used methods rely on the assumption of proportional hazards. This assumption is, however, violated if the treatment effect is delayed and even though usual methods may control the specified type I error rate they will have reduced power to detect a difference. To provide a basis for selecting appropriate methods we investigated the performance of various alternatives to the standard logrank test in terms of type I error and power in this setting. A broad variety of methods, including weighted logrank tests, combinations thereof or regression-based methods, were identified within a literature search and then compared in an extensive simulation study. We found that most of the methods controlled type I error. With respect to power, the methods that come close to the data-generating process performed best in the simulation study and were not sensitive to parameter misspecification. The interpretation of the results of these methods is, however, often complicated due to the absence of interpretable summary measures. Overall, the results of our simulation study together with the proposed ranking system can help to make a substantiated choice of an appropriate statistical method to analyze future studies.