
The selection of biomarker-specific patient populations is essential in targeted cancer therapies to enhance precision and efficacy. To ensure a successful launch, it is vital to promote awareness and adoption of biomarker testing at diagnosis, tailor implementation strategies to accommodate local variations, and ensure testing is accessible and reimbursed. An integrated evidence generation plan should address critical questions, including the prevalence of biomarker expression and agreement between local and central labs, across different platforms, antibodies and pathologists. Furthermore, understanding prognostic effects and associations with other biomarkers is of considerable interest. Real-world studies (RWS) play a pivotal role in addressing these questions but are inherently challenged by confounding factors, biases (e.g. immortal time bias), missing data, agreement assessment complexities, and low biomarker expression prevalence. This manuscript provides an integrated methodological roadmap for designing and analyzing RWS. We address key challenges by integrating robust statistical methodologies with advanced machine learning (ML) methods. Core methods discussed include the use of time-dependent Cox models to mitigate immortal time bias, inverse probability of biomarker weighting to adjust for confounding, and ML-based tree ensemble approaches to model complex relationships between covariates and outcomes and handle missing data. By systematically applying this integrated roadmap, we demonstrate how to enhance the validity and robustness of RW biomarker research. This approach overcomes common analytical pitfalls, enabling more reliable evidence generation for clinical decision-making. Ultimately, this roadmap helps drive precision oncology forward by ensuring that biomarker-driven therapeutic strategies are based on sound, high-quality RW evidence.
This paper introduces randomization-based analysis of covariance (RB-ANCOVA) for hypothesis testing of restricted mean survival time (RMST) differences between two randomized treatments in trials with categorized time-to-event data. RMST treatment differences over a prespecified time period are clinically meaningful and avoid assumptions like proportional hazards. The proposed method tests the strong null hypothesis that each participant's time-to-event would remain unchanged regardless of treatment assignment. Covariate adjustment is achieved by constraining baseline covariate means to be equal across arms, which reduces variance. The follow-up period is partitioned into mutually exclusive intervals, and RMST is approximated as the area under the survival curve. Under the strong null hypothesis, the joint asymptotic covariance structure of RMST and covariate means is known and used to construct a chi-squared test statistic for the covariate-adjusted RMST difference. The method accommodates stratified trial designs and supports hypothesis testing over single or several time intervals, enabling testing within internal intervals. The difference in RMSTs over an interval can be divided by its length to yield an average survival rate difference. Additionally, the approach allows the computation of essentially exact p-values via re-randomization. We illustrate the method with data from a randomized, placebo-controlled trial evaluating a test treatment for amyotrophic lateral sclerosis. This method provides a hypothesis testing approach for RMST treatment differences that leverages RB-ANCOVA to adjust for their correlations with baseline covariate differences. Under the strong null hypothesis, it enables hypothesis testing with clear control of Type I error and reduced variance without relying on model-based assumptions.
Oncology dose optimization has moved beyond the maximum tolerated dose paradigm, yet many programs still implicitly target a threshold-style minimum effective dose (MED). In serious cancers, deliberately sub-therapeutic comparators are rarely ethical or approvable, so any sharp MED threshold is typically a fragile and only partially identifiable target. We propose a statistical and operational framework that reframes dose-finding as region-based optimization. First, we define a constraint-based Dose Optimization Region (DOR) where clinically meaningful benefit (potentially multi-endpoint and PD-informed) is credible, unacceptable toxicity is bounded, exposure targets are attainable within a prespecified window, and implementation is feasible. Second, within the DOR, we identify an Operational Optimal Dose (OOD) - a label-ready dosing strategy that integrates dose, schedule, exposure targets, and adjustmentrules - using comparative evidence across binary and time-to-eventendpoints, exposure - response (PK/PD) analyses, and time-to-event methods that account for delayed effects where follow-up allows. This approach turns regulatory evidence domains into region-definingconstraints and shifts the target from a single-point estimate to a strategy-level deliverable. It is intended to sit on top of conventional dose-finding designs and PK/PD modeling as a region-defining and reporting standard, rather than to replace existing methods. We illustrate its use with small-samplevisualizations based on pairwise superiority probabilities across in-region arms and with a real-world-inspired example to show how DOR→OOD can summarize the totality of evidence in practice. The framework provides statisticians and clinicians with practical design patterns and reporting checklists, aligning with contemporary regulatory expectations and supporting dose optimization in oncology, with potential adaptation to other therapeutic areas.
HERALD (Holistic Evolving ReAssessment-Leveraged Decision-making) is a Bayesian decision-making framework anchored in the prediction of phase 3 efficacy success. At the proof-of-commercial-concept (POCC) study design stage, HERALD links available phase 3 design assumptions and success criteria with potential POCC treatment effects and yields decision boundaries that guarantee sufficient probability of success in phase 3. Since its commencement, HERALD has been widely adopted at Sanofi and received endorsement from the governance and key stakeholders across therapeutic areas. In this manuscript, we introduce how HERALD at the POCC design stage is industrialized. Through development of statistical software, standardization of governance presentation, consolidation of implementation examples, and periodical engagement programs of role-tailored trainings, we streamline the decision-making criteria discussion and empower both statisticians and nonstatisticians to communicate design options more efficiently within a cross-functional team.
Accurate sample size determination (SSD) is critical in clinical trial design, yet traditional methods that assume a fixed effect size overlooks estimation uncertainty, risking under- or over-powered studies. The assurance-based SSD approach, developed in earlier research, addresses this limitation by integrating statistical power over possible values of the effect size, weighted by their likelihood. Building on this foundation, we extend the framework to multi-regional clinical trials (MRCTs) using a hierarchical linear model (HLM) that can capture both between- and within-region variability in treatment effects. Analytical derivations and numerical studies show that assurance-based sample sizes are consistently larger than those from traditional approaches. Also, under high uncertainty, assurance-based sample sizes may not be defined, which serves as a valuable signal that the available evidence is insufficient to justify proceeding to confirmatory trials. These results demonstrate that the assurance approach provides a more reliable and transparent criterion for determining feasible sample sizes, highlighting its importance as a principled methodology, particularly in MRCTs where between-region variability further accentuates uncertainty in effect sizes.
By early 2023, there were over 100 gene, cell and RNA products approved globally (Chancellor et al. 2023). The way in which autologous cell and cell-based gene therapy products are administered can include multiple treatment stages, at times including leukapheresis, optional bridging therapy during manufacturing, conditioning regimen, and one or more doses of product. This brings important considerations in the development of the statistical analysis plan for clinical trials of autologous cell and cell-based gene therapies, including application of the estimand framework for key facets. In this article, we review all of the FDA currently approved autologous cell and cell-based gene therapy products as examples as they apply to general considerations for the statistical analysis plan as well as more in-depth discussion of treatment exposure, safety and efficacy analyses, and other aspects.
Conventional non-inferiority (NI) clinical trials often rely on large sample sizes and parametric methods with a single experimental drug. However, rare diseases pose challenges in obtaining large datasets, and data normality cannot always be assured. Nonparametric methods are a viable alternative, particularly when testing multiple experimental drugs with limited sample sizes. This paper reviews existing parametric and nonparametric NI testing methods and proposes a novel nonparametric NI method based on relative effect, a rank-based measure. Through simulations and real clinical trial data, we demonstrate that our proposed method outperforms existing ones in scenarios with small to moderate sample sizes and non-normal data. These findings highlight the potential of nonparametric approaches in NI clinical trials.
In randomized controlled trials with survival time as the primary endpoint, it can be difficult to evaluate treatment effects using hazard ratios, particularly when the proportional hazards (PH) assumption is violated. The difference in restricted mean survival time (RMST) up to a pre-specified time point, τ, is an alternative to the hazard ratio. However, the relative advantages of parametric (flexible parametric modeling [FPM]) and nonparametric approaches (direct integration of the Kaplan-Meier curve [DI], pseudo-observation [PO], and inverse probability of censoring weighting [IPCW]) for the estimation of RMST and its difference between groups are not clearly established. In this study, we performed comparative simulation studies to evaluate the performance of these methods for unadjusted and adjusted analyses in both PH and non-PH scenarios. For PO, IPCW, and FPM, we also considered a scenario where important covariates were adjusted via regression models. We obtained several key findings. For scenarios where the total number of events was >35, FPM tended to be unbiased, with higher power in PH scenarios. In non-PH scenarios, FPM exhibited slight bias with comparable power. For scenarios where the total number of events was ≤35, FPM showed unstable results. Among nonparametric methods, the bias and differences in power were negligible, except when the number of patients at risk at τ was ≤15; in this setting, IPCW showed conservative power and, under regression adjustment, a slight bias. These results provide a basis for the selection of methods to estimate RMST based on the expected prognosis.
Adapting treatments to individual patient attributes has become essential in personalized healthcare, especially for managing complex and variable chronic diseases. Despite their important role in precision medicine, traditional N-of-1 trials face design challenges related to carryover effects, washout periods, and blinding. To address these limitations, we propose a novel design, which we call SMART-of-1, that incorporates elements from a small sample Sequential, Multiple Assignment Randomized Trials (snSMART) into the N-of-1 trial. This innovative approach is tailored for comparative effectiveness studies in chronic diseases with acute presentations, such as asthma and sleep disorders. We propose a Bayesian joint stage individual model and a Bayesian hierarchical model for SMART-of-1 trials to estimate both individual treatment effects and population-level effects in a combined final analysis. In addition, the within-cycle tailored sequences of treatments can be assessed in the binary outcome setting. Two case studies, a fatigue study and a sleep study, are presented to demonstrate the improved accuracy and efficiency of the SMART-of-1 design for different types of endpoints. In the fatigue study with a continuous endpoint, we demonstrate that both the estimated individual and aggregated treatment effects from the SMART-of-1 trials have improved efficiency in terms of bias and rMSE compared to a standard N-of-1 trial. Simulations from the sleep study indicate the improved estimation of cycle-wise dynamic treatment regimen (DTR) effects for the SMART-of-1 model compared to a conventional snSMART approach, particularly in the presence of heterogeneity among individual treatment effects.
Randomized Controlled Trials (RCTs) are the gold standard in clinical trials. If Historical Data (HD) on the standard of care is available, hybrid control clinical trials can provide more evidence than a standalone RCT with unequal allocation. HD for the control group can often be derived from real-world data, which frequently includes missing covariates data. However, such missingness may introduce bias depending on missing data mechanisms and analytical methods. In this study, we propose addressing covariate missingness under the missing at random assumption by multiple imputation. In the analysis stage, we utilize a combination of propensity score matching and modified power prior. The simulation showed that complete case analysis caused bias under outcome and covariate-dependent covariate missingness, while multiple imputation provided nearly unbiased estimates and improved precision when HD was similar to the current trial data under the missing at random assumption. HD was dynamically borrowed based on outcome similarity: improved estimation accuracy with reduced bias and increased power when outcomes were similar, while reasonably controlling the type I error when outcomes were dissimilar. The proposed method was also applied to a real clinical trial data to illustrate its practical performance. These results suggest that the approach may be useful in hybrid control clinical trials that utilize real world data with missing covariates.
We describe a Bayesian approach for sample size estimation for multi-arm randomized controlled trials utilizing response adaptive randomization (RAR) for continuous outcomes. Assuming normally distributed treatment effects and unknown but common treatment variance, this design incorporates outcome data to estimate posterior distributions of treatment parameters, modifies allocation proportions to favor more effective treatments, and re-estimates the number of participants needed. The sample size should be sufficient to show that at least one group difference is greater than 0 (success), or that all effect sizes are smaller than a desired threshold (futility) at prespecified probabilities. Using simulations for a 4-arm trial, we compare sample sizes based on hypothesis testing to those using a Bayesian approach: [1] without interim analysis; [2] with interim analyses and non-adaptive randomization; [3] with interim analyses and RAR. We demonstrate that two interim analyses, conducted when outcomes are available among 25% and 50% of participants, could result in fewer enrolled participants. Especially for distinctly large treatment effects, RAR-based sample size estimates increase due to imbalanced allocation of participants. Application of the approach to two completed randomized controlled trials confirms these observations. We conclude that early and frequent interim analyses may reduce the number of participants needed for a conclusive trial where participant allocation does not change throughout the trial. For trial designs incorporating RAR, the within-trial patient benefits of allocating more patients to favorable arms with larger sample size requirements should be considered against the efficiency of equal group allocation.
Motivated by the growing use of unanchored indirect treatment comparisons (ITCs) in the absence of head-to-head trials, this paper evaluates and compares methods for estimating treatment effects across two data sources for continuous and binary outcomes. Within the Neyman-Rubin causal framework, we target the population average treatment effect for the treated, which in our setting quantifies the counterfactual difference between receiving treatment 0 and treatment 1 among individuals in Study 1, the population for which individual patient data (IPD) may not be accessible. We review six existing methods based on propensity score weighting and outcome regression and introduce a novel doubly robust (DR) estimator tailored to settings in which IPD is unavailable for the target population. Through a simulation study and a clinical case study, we demonstrate that doubly robust estimators consistently perform well across a range of practical settings. In particular, we recommend the proposed DR estimator for unavailable IPD settings.
The past decade has seen significant advancements for cell and gene therapies (CGTs). FDA lists more than 30 approved CGTs on their Approved Cellular and Gene Therapy Products website as of August 2024. With the promising treatment effect brought by the currently approved CGTs, there are noticeable limitations, such as, manufacturing delays and durability of response, that encourage continued development in this field. New CGTs can potentially benefit patients by enhancing efficacy or addressing the limitations of currently available therapies. However, the development of new CGTs faces unique challenges as highly efficacious approved first-generation products could be potentially established as the current standard of care (SOC). When the approved CGT products are recommended as the comparator, questions arise, especially, is it feasible to set up a head-to-head comparison between the new and the approved CGT products? What could be the feasible study design options to evaluate the effectiveness of new CGT products? In this article, we introduce the challenges of developing new CGTs in conventional randomized control and single arm trials, as well as discuss some possible strategies for pivotal trials such as hybrid study designs. We believe that more dialogue about this topic among all stakeholders involved in drug development is crucial and hope that this article will contribute towards facilitating such a dialogue.
Vaginal systems (VSs) or intravaginal rings are flexible, polymer-based, drug-device delivery systems used for release of one or more drugs directly to the vaginal-cervical tissues for an extended period. These complex products can undergo chemistry, manufacturing, and controls (CMC) changes. Comparative in-vitro release testing between prechange and postchange product can be used to support minor/moderate CMC changes. Due to long drug release period, the cumulative profile may be insensitive to visualize release rates' trend over time. Any summary statistic, e.g. the similarity factor f2 for dissolution profiles' comparison, may hide a large deviation at some time points since a summary statistic is equivalent to comparing the mean difference between two profiles averaged over time. This is problematic since the amount released per day is critical to ensure maintaining product quality and product performance consistently. Therefore, a summary statistic does not fit for the use. Another way to compare two profiles is to compare the estimated parameters after fitting the data to a curve which can be described by a few parameters. However, due to irregularity of these curves with limited number of time points, modeling approach, resulting in a large deviation at some points, isn't practical. In addition, specifications on in-vitro release rates of drug product at several pre-selected time points are generally recommended when a batch is released to the market and through the stability. We propose an equivalence testing approach for comparing in-vitro release rates between the prechange and the postchange VS at each of preselected time points.
This study introduces a non-inferiority test, which is based on the method of variance estimates recovery (MOVER), for the odds ratio in two independent binomial populations. In order to address the limitations of existing methods for single binomial proportions, we propose modified asymptotic confidence intervals within the MOVER framework, incorporating asymmetric parameters. We systematically evaluate the type I error control and statistical power of seven MOVER variants (M1-M7) and the Gart adjusted logit test (GAL) in non-inferiority settings. Comprehensive evaluations demonstrate that the M6 achieves optimal performance, maintaining the highest proportion of type I error rates below the significance level α, while retaining competitive power. When stricter type I error control is prioritized, M5 or M4 provides viable alternatives. Although M7 yields the highest power, its type I error is significantly inflated. The M1-M3 exhibit adequate error control but critically low power, while GAL shows moderate performance. The MOVER-based approach offers enhanced flexibility to accommodate diverse research requirements, with empirical results supporting its practical utility in non-inferiority testing. The proposed tests were illustrated with a real-world example.
Various test statistics have been proposed for comparing the predictive values of two binary diagnostic tests. However, these existing test statistic were derived based on a paired random sample where both diagnostic tests are administered on each subject in a random sample. It is not immediately clear whether or not they can be applied to a paired case-control sample where both diagnostic tests are administered on each subject in a case-control sample. In this paper, we propose a Wald test statistic and a weighted generalized score test statistic for comparing two predictive values based on a paired case-control sample and show that both test statistics can be calculated by applying their respective counterparts based on the paired random sample as if the paired case-control sample were a paired random sample. We further derive an optimal case-control ratio and demonstrate that when the disease is rare, the paired case-control sampling design with optimal case-control allocation achieves a substantial variance reduction, compared with a paired random sampling design.
Treatment effect heterogeneity has often been observed in clinical settings, which can be explained, at least partially, by predictive biomarkers. Although a variety of continuous biomarkers are measured in clinical practice, using biomarkers to guide treatment decisions remains challenging. We consider specifying a statistically and clinically suitable cutpoint of a continuous biomarker by identifying a sensitive subpopulation. We use a Bayesian posterior probability to analyze a time-to-event outcome and biomarkers observed in a randomized clinical trial. The proposed method aims not only to control the false-positive rate under a desired level, thereby deriving a statistically suitable decision rule, but also to incorporate a minimum clinically important difference, thereby providing a clinically reasonable interpretation. Simulation studies are conducted to evaluate the operating characteristics of the proposed method under a wide range of clinical scenarios. For illustration, we apply the proposed method to data from a phase III randomized clinical trial involving patients with advanced breast cancer.
The shelf-life of a drug product is determined conventionally based on up to 24 months of observed values for critical quality attributes from available batches prior to the regulatory approval of a new drug application. The data typically consist of three bio batches and are analyzed using the statistical method recommended in Guidance for Industry: Q1E Evaluation of Stability Data, International Council for Harmonisation (ICH). Many statistical approaches for shelf-life determination have been proposed, especially when the number of batches used is large. We compared two approaches for shelf-life determination using various number of batches. One approach is an Analysis of Covariance (ANCOVA) method by which the shelf-life is determined by the worst batch, if batches cannot be pooled and is otherwise determined by the pooled data of all batches. The other approach uses a tolerance interval based on a linear mixed-effects model. In this approach, the shelf-life is determined by the lower confidence limit of the 5th percentile of shelf lives of all batches. We compared these two approaches using simulated datasets of 3, 6, 20, and 40 batches. The ANCOVA approach is appropriate when fewer than six batches have available stability data. When the number of batches increases, the batches are less likely to be pooled, which yields an underestimation of shelf-life. The estimated shelf-life using the tolerance interval-based approach alleviates such underestimation when the number of batches is six or more.