
Stability assessment of vaccine drug products typically requires years of data collection. Accelerated stability studies offer a practical alternative by leveraging high-temperature data collected over shorter time periods. We model degradation using arbitrary nth order kinetics and derive a closed-form solution that enables extrapolation across time and temperature. We compare interval estimation methods including the delta method, bootstrap resampling, and Bayesian inference and propose a frequentist-Bayesian hybrid (FBH) posterior approximation framework based on a multivariate Student's t-distribution. FBH is closely related to objective Bayesian inference under noninformative priors in simple Gaussian settings and fully propagates uncertainty in both model parameters and residual variance. Simulation results show that FBH outperforms standard approaches by consistently achieving nominal coverage, particularly in small-sample settings where competing methods exhibit undercoverage. The approach provides well-calibrated confidence and prediction intervals while maintaining computational efficiency comparable to analytical methods and avoiding the cost of full posterior sampling. The methodology is illustrated using two vaccine development case studies and implemented in the AccelStab R package.
In many studies, multiple longitudinal outcomes are collected, and interest lies in studying the association between these outcomes. Joint modeling is then required, but full likelihood estimation becomes infeasible as the number of outcomes increases. To address this, the pairwise-fitting approach was developed. However, the robustness of this pseudo-likelihood-based approach under missing at random (MAR) remains unclear. We investigate the impact of MAR dropout on the pairwise-fitting approach through a case and simulation study and compare the results to full likelihood estimation. In the simulation study, we simulate three continuous longitudinal outcomes so that full likelihood estimation remains computationally feasible, allowing a comparison with the pairwise fitting approach. Various settings are examined, including random intercept and random intercept-and-slope models, in which we vary the standard deviation of the error terms and the degree of correlation between random effects. Our results show that bias remains limited in random intercept models and in most random intercept-and-slope models. However, when the standard deviation of the error terms becomes large compared to that of the random effects, some bias appears in the covariances between the random effects of the outcomes not driving dropout. This bias is mitigated using multiple imputation. As a case study, we analyzed data from a schizophrenia study using both full likelihood and pseudo-likelihood approaches and compared the results.
Several study designs for identifying optimal biological dose (OBD) have been proposed for phase I, II, and I/II clinical trials, considering toxicity and efficacy of anticancer drugs, especially molecular-targeted therapies and immune checkpoint inhibitors. Among these, the multiple-dose randomized phase II trial (MERIT) design selects OBD candidates using hypothesis testing for toxicity and efficacy outcomes. However, it does not consistently control the type I error rate below the significance level, as the null hypothesis is defined only at specific points in a two-dimensional null space. To address this limitation, we developed a design treating toxicity and efficacy as co-primary endpoints, ensuring strict dose-level type I error control across the entire null space. A Bonferroni correction addressed multiplicity in dose selection, providing overall type I error control. Unlike the existing method, the rejection region is determined analytically rather than by simulation, reducing computational costs. The type I error rate of the existing method exceeds the significance level in regions outside points considered in its sample size calculation. By contrast, the proposed method maintains the type I error rate below the significance level across the null space, though conservatively due to the co-primary endpoint framework and Bonferroni adjustment. Required sample sizes of the proposed method tend to be larger than those of the existing one.
In multiregional clinical trials (MRCTs), regional consistency probabilities (RCPs) must be evaluated at the trial design stage in addition to the power for the entire trial population. The Japanese Ministry of Health, Labour and Welfare guidance proposes two widely used criteria for this purpose, referred to as Method 1 and Method 2. While multiplicity adjustment for familywise error rate control is standard practice when multiple primary endpoints are employed, its implications for regional consistency evaluation in MRCTs have not been addressed in the existing literature. In this article, we demonstrate that applying unadjusted consistency thresholds under the Bonferroni procedure leads to monotone inflation of the null RCP with increasing number of endpoints for both Method 1 and Method 2, because the Bonferroni-adjusted critical value progressively selects more extreme overall treatment effect estimates, which in turn inflates the probability of satisfying the regional consistency criterion even when the true treatment effect is absent. To correct for this inflation, we propose adjusted consistency thresholds that restore the null RCP to its single-endpoint reference level, and show that these thresholds can be efficiently obtained via a standard root-finding algorithm. Through numerical studies, we demonstrate that the proposed adjusted thresholds result in only a modest reduction in RCP under the alternative hypothesis, and that thresholds derived under the Bonferroni procedure also prevent null RCP inflation when step-down procedures such as Holm, Hochberg, and Hommel are applied. The methodology is illustrated using a hypothetical MRCT with multiple primary endpoints, and an R package "rcpmcp" implementing all proposed methods is publicly available on GitHub.
Quantitative dose optimization in early phase clinical trials for investigational new drugs has been expanding across drug modalities and disease indications in response to the limitations of non-quantitative or algorithmic methods of dose progression. In the context of the FDA's Project Optimus Initiative, which emphasizes dose optimization rather than reliance on the maximum tolerated dose, these challenges motivate the development of alternative quantitative frameworks tailored to gene therapies. In this manuscript we describe the state of the art for quantitative dose optimization with an eye to applications in cell and gene therapy modalities. The application of quantitative dose progression methods in this setting poses unique challenges, but the regulatory, scientific, medical, and statistical environment has advanced to the point where new approaches can be employed to characterize the safety and efficacy profile of these powerful biopharmaceutical products. We discuss the current use of quantitative dose optimization, the limitations and pitfalls of these options for gene therapy indications, and provide a tutorial example of how an established Bayesian logistic regression framework can be adapted to distinguish clinically distinct toxicity classes. The example is intended to illustrate safety-constrained decision logic in a sparse early-phase setting, rather than to establish comparative superiority over existing dose-finding designs. The proposed framework should be viewed as one component of broader dose optimization that may also incorporate pharmacodynamic, exposure-response, efficacy, durability, and practical administration considerations.
We conducted a comprehensive comparative analysis of causal machine learning (ML) methods to assess their utility in improving the efficiency of clinical trial designs with or without historical data. Specifically, we compared standard ANCOVA analysis used in a Randomized Controlled Trial (RCT) with several causal ML methods, including PROCOVA, TMLE, DML, and GRF. PROCOVA is gaining popularity in RCT design and requires a historical data for prognostic scores, but other methods can be applied with or without such data. Our primary focus was on strict RCT setting without borrowing historical control data, though we also explored the impact of borrowing data. The historical data used consisted of placebo data from two Phase 3 Ophthalmology studies with a continuous primary endpoint. We employed a generative AI approach, specifically Generative Adversarial Networks (GANs), to simulate RCT data from the historical data under various scenarios, varying treatment effects with and without treatment effect heterogeneity, RCT sizes, and bias. Results showed that causal ML methods can increase power even without borrowing historical data. For example, TMLE increased effective sample size by 21% in one scenario. In scenarios with borrowing of controls, PROCOVA increased power while controlling type 1 error, showing robustness to model misspecification.
Platform trials evaluate multiple treatments within a single trial infrastructure. Such designs have gained a lot of attraction in clinical research. If information gained from platform trials should provide confirmatory evidence for regulatory decisions, control of the Type I error rate is key. One critical issue is information leakage, for example, if any information of the ongoing trial is available, especially if it may impact the further conduct of treatments still in the platform trial and bias their results. This paper evaluates the potential impact of information leakage on the control of the Type I error rate in platform trials with time-to-event endpoints such as overall survival. We explore different strategies how information on the treatment effect of an still ongoing treatment could be obtained if a pre-planned analysis for another arm is conducted. This (leaked) information might be used to decide whether to continue the other arm as planned or conduct its final analysis immediately. By means of clinical trial simulations we evaluate the impact of different levels of information leakage on the Type I error rate. We show how the conditional error principle can be applied to estimate worst case Type I error rate inflation for the different forms of information leakage. We do not aim to quantify the exact maximum Type I error rate inflation but rather to raise awareness of the potential risk for estimation of comparative results. Finally, we discuss the regulatory implications of information leakage and propose strategies to mitigate these risks.
Accurately distinguishing between healthy and diseased states is fundamental to clinical diagnostics. This paper introduces the Harmonic Fowlkes-Mallows (HFM) index, a novel and robust metric for assessing diagnostic accuracy and identifying optimal cut-off points. The proposed HFM index integrates performance across both positive and negative classes by combining the traditional Fowlkes-Mallows Index (FM) with the proposed Negative Fowlkes-Mallows Index (NFM), using a weighted harmonic mean. Unlike conventional measures such as the F1-score or Youden Index, HFM provides a more comprehensive evaluation of classification performance by simultaneously addressing sensitivity and specificity. Additionally, it incorporates a tunable β parameter to adjust for asymmetries in class importance. Through simulation studies, the HFM index demonstrates strong performance in binary classification tasks and proves effective in selecting optimal decision thresholds. To further demonstrate its practical utility, we apply the HFM index to real-world breast cancer data and compare its performance with other diagnostic accuracy measures.
Chimeric Antigen Receptor (CAR)-T cell is an immunotherapy which revolutionised the treatment of relapsed/refractory lymphoma and leukaemia. It is shown to have a higher response rate, higher mid-to-long term overall survival, and lower toxicity than standard treatments. However, due to a lack of dose-limiting toxicity (DLT) and unclear dose-effect relationship, traditional phase I designs of clinical trials cannot lead to accurate selections of the optimal dose (OD). Beside clinical outcomes, the CAR-T cell expansion from serial blood samples is measured at various time points. We propose a novel early phase dose-finding design for CAR-T cells, using both toxicity and activity endpoints to locate the OD. The number of CAR-T cells measured in the peripheral blood is used to indicate activity, which is more sensitive than the short-term clinical responses traditionally used. A Bi-Exponential model is used for the repeated measures of the number of cells for each patient, and is estimated under a Bayesian framework. The model is motivated by biological concerns and is flexible enough to accommodate different shapes of the cell-expansion curve. Three criteria for activity are considered: (1) the number of cells at specific time points, (2) the duration before all cells are eliminated, (3) the area under the cell-expansion curve. Simulation studies show that the OD can be selected with high accuracy even under small sample sizes.
An equal randomization ratio (1:1) is the most commonly used allocation strategy in confirmatory clinical trials. Recently, there has been growing discussion around the use of unequal randomization, which may be favored for several reasons, including encouraging trial recruitment, reducing costs, and improving the robustness of estimates in the treatment arm. However, despite these potential benefits, unequal randomization is still rarely applied in trial design. A key barrier is the lack of a method to determine the optimal randomization ratio that balances the various considerations of a trial. To address this challenge, we developed an optimization framework that identifies the optimal randomization ratio to maximize a trial's probability of success and profit-two major considerations in trial design. The framework incorporates trial parameters such as prior knowledge of treatment efficacy, sample size, budget, and per-subject cost. We evaluate the proposed method through simulations and a hypothetical trial. The results show that the optimal randomization ratio is highly influenced by sample size and cost differences between treatment arms. Our framework demonstrates the ability to reduce costs while maintaining a high probability of success.
Dose optimization in oncology clinical trials has shifted from solely seeking the maximum tolerated dose to identifying the Optimal Biological Dose (OBD) that balances therapeutic benefits and risks across multiple clinical attributes. Existing advanced dose-finding methods can integrate multiple endpoints but may not be suitable to compare dose levels from randomized dose optimization studies with small sample sizes. To address these challenges, we propose a clinical utility index (CUI) based analysis and decision framework for dose optimization in multiple-dose, multiple-outcome randomized trials (CUI-MET). This framework integrates multiple binary endpoints into a combined CUI for each dose level by weighting together multiple endpoints. Marginal summaries of individual endpoints can be estimated either empirically or via one or more choices of different parametric dose-response models. These estimated probabilities are then combined using endpoint-specific weights to compute a utility score for each dose. The dose with the highest score within a prespecified clinically acceptable dose set is selected as optimal. Bootstrap analysis provides confidence intervals for the CUI and estimates the probability that each dose is selected as optimal, thereby evaluating the robustness of dose selection. To enhance usability, we implemented these methods in an interactive R Shiny application and demonstrated functionality through case examples. The framework's flexibility allows for different model selections and endpoint weighting schemes to reflect specific clinical priorities and sensitivity analyses. By integrating multiple endpoints into a single utility index and incorporating user-friendly visualizations, CUI-MET offers a flexible and accessible solution for dose optimization in early-phase oncology trials, supporting informed decision-making and the integration of patient-relevant outcomes.
Artificial Intelligence (AI) is rapidly becoming more visible in all aspects of our lives. This Viewpoint discusses the impact AI will have on statisticians working within the pharmaceutical industry. While aspects of the statistician's role will change, I propose that AI will make the core components of a statistician's skillset even more critical in the future.
Adaptive clinical dose-finding trials aim to identify an optimal drug dose for use in subsequent phase II and III trials. In adaptive dose-finding trials, dose levels for newly included patients are informed by outcomes of patients that received the drug earlier in the trial. The body of methodological research on adaptive dose-finding trials is extensive, but a clear overview is lacking. The goal of this paper is to provide a knowledge base of designs and statistical methods for adaptive clinical dose-finding trials by means of a literature review. We identified 315 adaptive dose-finding trial methodology articles of which the majority was inspired by oncology. Recent methods focused on identification of an optimal dose considering both toxicity and efficacy endpoints, and addressing challenges related to subgroup-specific dose-finding, dose-finding for combination therapies, and incorporation of delayed outcomes in dose-finding. These developments are driven by the emergence of newer classes of cancer drugs, such as targeted therapies and immunotherapies, and by initiatives like Project Optimus. Most articles focused on model-based designs, like the Continual Reassessment Method (CRM), but recent years have seen a strong increase in model-assisted or interval-based designs, including Toxicity Probability Interval (TPI) and Bayesian Optimal Interval (BOIN) designs and expansions thereof. Considering the increasing availability and large variety of adaptive dose-finding trial designs, it is challenging for researchers to find relevant designs tailored to their needs. Therefore, we provide an interactive map summarizing the classification of our results to facilitate the identification of relevant designs/methods, which we will update regularly.
Biological measurements such as percent monomer from size exclusion chromatography (SEC) or cell-based percentages from flow cytometry are continuous quantities bounded between 0 and 1. These continuous bounded data are frequently skewed and therefore construction of normal-based confidence, prediction, and tolerance intervals may not be appropriate. Transformation based approaches (e.g., logit or arcsine-square-root) are often recommended to analyze bounded data; however, for highly skewed data with sample size, they may yield intervals that are too wide for practical applications. Although the beta distribution is a natural model for continuous bounded data, routine interval procedures remain limited in Chemical, Manufacturing, and Controls (CMC) practice. The Kumaraswamy distribution is a flexible and mathematically tractable alternative. In this manuscript, we develop two new pivotal quantities for the Kumaraswamy shape parameter α $$ \alpha $$ . Combined with an existing generalized pivotal quantity for β $$ \beta $$ , these pivot quantities support fiducial-based construction of statistical intervals for functions of the shape parameters α $$ \alpha $$ and β $$ \beta $$ , such as confidence intervals for the mean and fiducial-type tolerance and prediction intervals. We investigate finite-sample performance of these intervals through a simulation study and compare the proposed methods with a published pivotal approach and commonly used transformations. We also illustrate the methods with a simulated percent monomer data representative of CMC bioanalytical practice. Practical recommendations for routine use are provided.
In medical diagnostic testing, classifications are commonly divided into two primary types: binary tests, which distinguish between diseased and nondiseased cases, and ordinal states, which categorize cases into nondiseased states and disease stages ranging from 1 to k. Additionally, there exists a multiclass classification scheme, referred to as tree or umbrella ordering, in which a classifier determines whether a biomarker measurement for one class is higher or lower than that of the other classes. We introduce a novel concept called Nested Disease Subtypes Within the Disease Ordinal States, which extends the summary measures of receiver operating characteristic (ROC) curves to accommodate this unique classification approach. Our method aims to refine current diagnostic categorization by accounting for the complexity of certain diseases that do not conform to traditional classifications. To validate the effectiveness of our approach, we conducted simulation studies and applied it to real data on tuberculosis. These findings highlight the potential of our methods to enhance both the precision and comprehension of medical diagnostic testing, particularly, for complex diseases.
The COVID-19 pandemic spurred the adoption of platform trials for efficient drug development. These trials evaluated multiple treatments simultaneously but faced challenges due to temporal shifts in patient characteristics, trial conduct, and standard of care. A key concern was the use of nonconcurrent control (NCC) data, which can enhance statistical power but may introduce bias if the effects of temporal shifts are not properly adjusted. In this study, we empirically evaluated temporal shifts in clinical outcomes (i.e., time to recovery through Day 28, clinical status on Day 14, and mortality through Day 28) using real-world platform trial data from the ACTIV-1 Immune Modulator trial, accounting for the evolving landscape of SARS-CoV-2 variants in the CoV-Spectrum database. We have investigated the operating characteristics of the existing methods using concurrent control and NCC data based on a resampling-based simulation study using ACTIV-1 IM trial data, as well as simulations under hypothetical variant-driven temporal-shift scenarios that explicitly modeled changes in circulating SARS-CoV-2 variants. Our empirical analyses revealed that assumptions in simulation studies-such as specific patterns of temporal shifts and constant treatment effects over time-might not hold in real-world settings. Simulation studies further showed that in such complex settings, even adjustment methods for leveraging NCC could introduce an unacceptably large bias. Our findings provided an empirical foundation for improving the integration of NCC data and ensuring reliability and efficiency. Real-world evidence underscored the need for robust statistical methods that account for the temporal shifts in platform trials.
Analysis of genomics data for predicting disease outcomes is a fast-growing field in medical research. There often exist categorical, specifically, ordinal outcomes that need to be predicted based on genomic profiles. This has led to recent development of some high-dimensional ordinal classification methods that can address the large dimensionality of the genomic covariate set. These high-dimensional ordinal models tend to vary widely in their performance depending on the data they are applied to and the evaluation criteria used. In this article, we outline an ensemble ordinal classifier that integrates different ordinal modeling approaches through bootstrap-based model evaluation, multi-metric performance assessment, and rank aggregation to produce a final prediction that can alleviate the uncertainty of relying on a single model. Through multiple simulated studies and real genomic data analyses, we show that the ensemble method consistently ranks among the top-performing models. These findings underscore the potential of ensemble learning to improve the robustness and predictive accuracy of high-dimensional ordinal classification in genomic research.
As Phase III trial costs and durations rise, pharmaceutical companies increasingly use quantitative methods to decide if a drug should progress beyond Phase II. A key method is the probability of success (PoS) for Phase III, calculated using the power function averaged across a treatment effect distribution estimated from Phase II. This paper explores PoS's role, particularly in moving from Phase II trials with putative surrogate endpoints to Phase III trials with clinical endpoints. Since the relationship between these endpoints is often unknown, expert input is necessary (prior elicitation). We propose the bivariate meta-analysis and a copula-based extension to characterize their relationship, using visual tools to simplify parameter elicitation. Specifically, we begin by eliciting the marginal distributions of the two quantities of interest. Then, to assist in eliciting the concordance parameter, we use the distribution of the treatment effect on the clinical endpoint conditional on the treatment effect on the putative surrogate. Our approach is illustrated in prophylactic vaccine development, linking immunological and clinical endpoints.
Precision medicine is a dynamic and evolving field that relies on biomarkers to guide decisions about patient population enrichment, thereby shaping the direction of drug development. Determining an accurate cutoff for a biomarker is important to identify patients who are more likely to benefit from a given therapy. In conventional methods, a cutoff is derived from the minimum p-value of a dichotomized biomarker and the treatment of interest. In contrast, we have developed a predictive biomarker graphical approach, PRIME, that can be used to incorporate clinical significance and evaluate biomarkers on a continuous scale. We adapted a treatment selection approach and extended it to account for covariates via G-computation. With this proposed approach, we aim to evaluate biomarkers in a continuous manner from a predicted risk standpoint, using a model that includes the interaction between a biomarker and a treatment. The predicted risk can be graphically displayed to delineate the relationship between the outcome and the biomarker, thereby indicating a cutoff for biomarker-positive and -negative groups of patients. Other features of PRIME include the ability to compare biomarkers by using net gain summary measures and calibration to assess the model fit. The PRIME approach can be used to consider a variety of outcomes and covariates. We developed an R package, performed simulation studies, and used examples to demonstrate the PRIME approach.
Dose optimization (DO) is a significant paradigm in clinical trials, with the goal of identifying an optimal dose that preserves maximal efficacy while minimizing toxicity. In this paper, we propose a Randomized Adaptive Bayesian Optimization Two-stage Seamless (RABOTS) design that integrates DO with proof-of-concept (PoC) evaluation. In Stage I, an adaptive dose selection process is implemented to drop sub-optimal dose levels based on a pre-specified decision algorithm. In Stage II, a Bayesian Go/No-Go criterion is applied to the selected dose, and a dynamic borrowing approach is recommended to enhance statistical efficiency by leveraging information from the dropped dose. We evaluate the performance of RABOTS by applying different information borrowing methods across multiple dose levels through a range of simulation scenarios. Results show that the RABOTS dose selection stage reliably identifies the dose with superior efficacy and favors the lower dose when both doses exhibit similar efficacy, thereby emphasizing safety. In the PoC stage, the SAM prior borrowing method adaptively addresses prior-data conflict, effectively balancing Bayesian power and Bayesian type I error rate. We further examine how the RABOTS key tuning parameters, including the clinically meaningful effect threshold ( c θ $$ {c}_{\theta } $$ ) and Stage I sample size ( N 1 $$ {N}_1 $$ ), influence its operating characteristics. The RABOTS offers a flexible and efficient framework for dose optimization and early efficacy assessment in oncology drug development. Future extensions may include more complex dose structures, the incorporation of safety and alternative endpoint types.