Target trial emulation prompts investigators to frame their analysis question in terms of a hypothetical clinical trial. Although this does not solve the problem of confounding, the framework can protect against other sources of bias. A natural question is this: what kinds of trials can be emulated? In a late-phase trial (Phase 3 or 4), the goal is to obtain a well-defined causal estimate that closely approximates the impact of a proposed intervention. In an early-phase trial (Phase 2 or earlier), the estimate is a means to an end rather than an end in itself. An early-phase trial provides proof-of-concept evidence on the impact of an intervention in the exposure in terms of efficacy and safety, but the estimand may not correspond to the intervention to be implemented in practice. In a natural experiment where causal inferences rely on a plausibly random (or quasi-random) comparison, the estimand may not be directly translatable to applied practice. In this case, the analysis may be conceptualized as an early-phase target trial. This provides less specific evidence than a late-stage target trial, but in many cases, a more valid but less applicable comparison is preferable to a more applicable comparison that is more susceptible to bias.
BACKGROUND:Because confirmatory clinical trials are costly, large-scale endeavors, the choice of their design carries significant weight. While the current methodological landscape offers tools to address residual pre-trial uncertainty through prespecified adaptations to design elements, design choices remain frequently constrained by prevailing orthodoxies. METHODS:We examine how unacknowledged design uncertainties (e.g., sample size calculations based on incorrect effect size or variance estimates) impact the statistical, ethical, and resource-stewardship goals of confirmatory trials. We contrast how fixed and adaptive design strategies handle these uncertainties, detailing the operational safeguards, firewalls, and trade-offs (e.g., complexity, operational bias) required to preserve trial integrity. RESULTS:Strict adherence to default templates prevents stakeholders from matching the design strategy to the specific complexities of the research question. We observe that the gold standard status of fixed designs often obscures their limitations in handling design uncertainties. A rigid adherence to these conventions can ethically hinder a trial's ability to deliver conclusive results. Conversely, while adaptive designs can offer tools to efficiently reduce uncertainty, we acknowledge that adaptive elements are not without cost; their implementation requires rigorous safeguards to manage specific risks to trial integrity and potential operational biases. CONCLUSION:To better align clinical research with its ethical and scientific mandates, we issue two calls to action. First, trial designers, including statisticians, must clearly articulate how target and design uncertainties impact ethical obligations, promoting an open-minded evaluation of both fixed and adaptive methods. Second, stakeholders must foreground ethical considerations during design selection, requiring explicit justification for the use-or non-use-of adaptive elements. The ethical path forward is not to default to the old or blindly adopt the new, but to explicitly justify the chosen design based on its responsiveness to the specific uncertainties of the trial. We posit that trial design must transition from a habit-based process to one where pre-trial uncertainties are openly discussed and addressed.
Response-adaptive randomization (RAR) can increase participant benefit in clinical trials, but also complicates statistical analysis. The burn-in period (a non-adaptive initial stage) is commonly used to mitigate this disadvantage, yet guidance on its optimal duration is scarce. To address this critical gap, this paper introduces an exact evaluation approach to investigate how the burn-in length impacts statistical operating characteristics of two-arm binary Bayesian RAR (BRAR) designs. We show that (1) commonly used calibration and asymptotic tests show substantial type I error rate inflation for BRAR designs without a burn-in period, and increasing the total burn-in length to more than half the trial size reduces but does not fully mitigate type I error rate inflation, necessitating exact tests; (2) exact tests conditioning on total successes show the highest average and minimum power up to large burn-in lengths; (3) the burn-in length substantially influences power and participant benefit, which are often not maximized at the maximum or minimum possible burn-in length; (4) the test statistic influences the type I error rate and power; (5) estimation bias decreases quicker in the burn-in length for larger treatment effects and increases for larger trial sizes under the same burn-in length. Our approach is illustrated by re-designing the ARREST trial.
Bayesian adaptive designs enable flexible clinical trials by adapting features based on accumulating data. Among these, Bayesian Response-Adaptive Randomisation (BRAR) skews patient allocation towards more promising treatments based on interim data. Implementing BRAR requires the relatively quick evaluation of posterior probabilities. However, the limitations of existing closed-form solutions mean trials often rely on computationally intensive approximations which can impact accuracy and the scope of scenarios explored. While faster Gaussian approximations exist, their reliability is not guaranteed. Critically, the approximation method used is often poorly reported, and the literature lacks practical guidance for selecting and comparing these methods, particularly regarding the trade-offs between computational speed, inferential accuracy, and their implications for patient benefit.Focusing on BRAR trials with binary endpoints, a novel algorithm is developed that efficiently and exactly computes these posterior probabilities, enabling a robust assessment of existing approximation methods in use. Leveraging these exact computations, a comprehensive benchmark is established for evaluating approximation methods based on their computational speed, patient benefit, and inferential accuracy. The comprehensive analysis, conducted through a range of simulations and a re-analysis of a real-life multi-arm trial, reveals that the exact calculation algorithm is highly efficient and often the fastest approach for trials with a small to moderate number of arms. Furthermore, it is demonstrated that commonly used approximation methods can lead to significant power loss and type I error rate inflation, with the Gaussian approximation emerging as a particularly unreliable option. To address the exponential scaling inherent in exact multi-arm computations, a formal, quantitative framework is provided. This practical guidance equips practitioners with precise thresholds to seamlessly transition between exact calculations and safe approximation methods across various clinical trial settings.
Despite extensive research, the use of response-adaptive randomization (RAR) in clinical trials has remained controversial. Korn and Freidlin's 2011 article reignited this debate back then, prompting numerous responses, including one by Zhu, Rosenberger, and Hu that remained unpublished until now. This article features the original response by Zhu, Rosenberger, and Hu, providing a valuable opportunity to revisit the original arguments, examine subsequent developments in RAR methodology, and offer a more complete historical perspective on this enduring debate. The piece also includes a concluding section by one of the special issue's co-editor that explores the nuances and complexities of RAR implementation in the context of contemporary clinical trial design.
Response-adaptive clinical trial designs allow targeting a given objective by skewing the allocation of participants to treatments based on observed outcomes. Response-adaptive designs face greater regulatory scrutiny due to potential type I error rate inflation, which limits their uptake in practice. Existing approaches for type I error control either only work for specific designs, have a risk of Monte Carlo/approximation error, are conservative, or computationally intractable. To this end, a general and computationally tractable approach is developed for exact analysis in two-arm response-adaptive designs with binary outcomes. This approach can construct exact tests for designs using either a randomized or deterministic response-adaptive procedure. The constructed conditional and unconditional exact tests generalize Fisher's and Barnard's exact tests, respectively. Furthermore, the approach allows for complexities such as delayed outcomes, early stopping, or allocation of participants in blocks. The efficient implementation of forward recursion allows for testing of two-arm trials with 1,000 participants on a standard computer. Through an illustrative computational study of trials using randomized dynamic programming it is shown that, contrary to what is known for equal allocation, the conditional exact Wald test based on total successes has, almost uniformly, higher power than the unconditional exact Wald test. Two real-world trials with the above-mentioned complexities are re-analyzed to demonstrate the value of the new approach in controlling type I errors and/or improving the statistical power.
Although response-adaptive randomisation (RAR) has gained substantial attention in the literature, it still has limited use in clinical trials. Amongst other reasons, the implementation of RAR in real world trials raises important practical questions, often neglected in the technical literature. Motivated by an innovative phase-II stratified RAR rare-disease trial, this paper addresses two challenges: (1) How to ensure that RAR allocations are desirable, that is, both acceptable and faithful to the intended probabilities, particularly in small samples? and (2) What adaptations to trigger after interim analyses in the presence of missing data? To answer (1), we propose a Mapping strategy that discretises the randomisation probabilities into a vector of allocation ratios, resulting in improved frequentist errors. Under the implementation of Mapping, we answer (2) by analysing the impact of missing data on operating characteristics in selected scenarios. Finally, we discuss additional concerns including: pooling data across trial strata, analysing the level of blinding in the trial, and reporting safety results.
Rationale: Imatinib, 400 mg daily, reduces pulmonary vascular resistance and improves exercise capacity in patients with pulmonary arterial hypertension. Concerns about safety and tolerability limit its use. Objectives: We sought to identify a safe and tolerated dose of oral imatinib between 100 mg and 400 mg daily and evaluate its efficacy. Methods: Oral imatinib was added to the background therapy of 17 patients with pulmonary arterial hypertension, including 13 who were implanted with devices that provide daily measurements of cardiopulmonary hemodynamics and physical activity. The first patient was started on 100 mg daily. The next 12 patients, recruited serially, were started on 200 mg, 300 mg, or 400 mg daily, following a continuous reassessment dose-finding model. An extension cohort (Patients 14-17) received 100 mg or 200 mg daily. Measurements and Main Results: The continuous reassessment model recommended starting dose was 200 mg daily. The most common side effect was nausea. Imatinib reduced mean pulmonary artery pressure (-6.5 mm Hg; 95% confidence interval [CI] = -2.4 to -10.6; P < 0.01) and total pulmonary resistance (-2.8 Wood units; 95% CI = -1.5 to -4.2; P < 0.001), with no significant change in cardiac output. The reduction in total pulmonary resistance was dose and exposure dependent; the reduction from baseline with imatinib, at 200 mg daily, was -20.3% (95% CI = -14.3 to -26.3%). Total pulmonary resistance and night heart rate declined steadily over the first 28 days of treatment and remained below baseline up to 40 days after imatinib withdrawal. Conclusions: Oral imatinib, 200 mg daily, is well tolerated as an add-on treatment for pulmonary arterial hypertension. A delay in the return of cardiopulmonary hemodynamics to baseline was observed after imatinib was stopped.
The majority of response-adaptive randomisation (RAR) designs in the literature use efficacy data to dynamically allocate patients. Their applicability in settings where the efficacy measure is observable with a random delay, such as overall survival, remains challenging. This paper introduces a RAR design referred to as SAFER (Safety-Aware Flexible Elastic Randomisation) design, which uses early-emerging safety data to inform treatment allocation decisions in oncology trials. However, the design is applicable to a range of settings where it may be desirable to favour the arm demonstrating a superior safety profile. This is particularly relevant in non-inferiority trials, which aim to demonstrate an experimental treatment is not inferior to the standard of care, while offering advantages in terms of safety and tolerability. Consequently, an unavoidable and well-established trade-off arises for such designs: to balance the goals of preserving inferential efficiency for the primary non-inferiority outcome while incorporating safety considerations into the randomisation process through RAR. Our method, defines a randomisation procedure which prioritises the assignment of patients to better-tolerated arms and adjusts the allocation proportion according to the observed association between safety and efficacy endpoints. We illustrate our procedure through a comprehensive simulation study, inspired by the CAPP-IT Phase III oncology trial. Our results demonstrate that SAFER preserves statistical power even when efficacy and safety endpoints are weakly associated and offers power gains when a strong positive association is present. Moreover, the approach enables a faster/slower adaptation when efficacy and safety endpoints are temporally aligned/misaligned, respectively.
Maximizing statistical power in experimental design often involves imbalanced treatment allocation, but several challenges hinder its practical adoption: (1) the misconception that equal allocation always maximizes power, (2) when only targeting maximum power, more than half the participants may be expected to obtain inferior treatment, and (3) response-adaptive randomization (RAR) targeting maximum statistical power may inflate type I error rates substantially. Recent work identified issue (3) and proposed a novel allocation procedure combined with the asymptotic score test. Instead, the current research focuses on finite-sample guarantees. First, we analyze the power for traditional power-maximizing RAR procedures under exact tests, including a novel generalization of Boschloo's test. Second, we evaluate constrained Markov decision process (CMDP) RAR procedures under exact tests. These procedures target maximum average power under constraints on pointwise and average type I error rates, with averages taken across the parametric space. A combination of the unconditional exact test and the CMDP procedure protecting allocations to the superior arm gives the best performance, providing substantial power gains over equal allocation while allocating more participants in expectation to the superior treatment. Future research could focus on the randomization test, in which CMDP procedures exhibited lower power compared to other examined RAR procedures.
The principle of allocating an equal number of patients to each arm in a randomized controlled trial remains widely believed to be optimal for maximising statistical power. However, this long-held belief only holds true if the treatment groups have equal outcome variances, a condition that is often not met or, is simply not assessed in practice. This paper reasserts the fact that a departure from a 1:1 ratio can maintain or improve statistical power while increasing the benefits to participants. The benefit is particularly self-evident for binary and time-to-event endpoints, where variances are determined by the assumed success or event rates. To illustrate this, we present two case studies: a small-scale metastatic melanoma trial with a binary endpoint and a larger trial evaluating virtual reality for pain reduction with a continuous endpoint. Our simulations compare equal randomisation, preplanned fixed unequal randomisation, and response-adaptive randomisation targeting Neyman allocation. Results show that unequal allocation can increase the proportion of patients receiving the superior treatment without reducing power, with modest power gains observed in both binary and continuous settings, highlighting the practical relevance of optimised allocation strategies across trial types and sizes.
Traditional hypothesis tests for differences between binomial proportions are at risk of being too liberal (Wald test) or overly conservative (Fisher's exact test). This problem is exacerbated in small samples. Regulators favour exact tests, which provide robust type I error control, even though they may have lower power than non-exact tests. To target an exact test with high power, we extend and evaluate an overlooked approach, proposed in 1969, which determines the rejection region through a binary decision for each outcome vector and uses integer programming to, in line with the Neyman-Pearson paradigm, find an optimal decision boundary that maximizes a power objective subject to type I error constraints. Despite only evaluating the type I error rate for a finite parameter set, our approach guarantees type I error control over the full parameter space. Our results show that the test maximizing average power exhibits remarkable robustness, often showing highest power among comparators while maintaining exact type I error control. The method can be further tailored to prior beliefs by using a weighted average. The findings highlight both the method's practical utility and how techniques from combinatorial optimization can improve statistical methodology.
Missing data is a widespread issue in clinical trials, but is particularly problematic for digital health interventions where disengagement is common and outcomes are likely to be missing not at random (MNAR). Trials that use response-adaptive designs need to handle missingness online and not simply at the end of the trial. We propose a novel online imputation strategy which allows previous imputations to be re-imputed given updated estimates of success probabilities. We additionally consider: (i) truncation of deterministic algorithms to prevent extreme realised treatment imbalance and (ii) changing the random component of semi-randomised algorithms. Through a simulation study based on a trial for a digital smoking cessation intervention, we illustrate how the strategy for handling missing responses can affect the exploration-exploitation tradeoff and the bias of the estimated success probabilities at the end of the trial. In the settings explored, we found that the exploration-exploitation tradeoff is affected particularly when arms have very different rates of missingness and we identified combinations of response-adaptive designs and missingness strategies that are particularly problematic. Further, we show that estimated success probabilities at the end of the trial can be biased not only due to optimistic sampling, but potentially also due to an MNAR missingness mechanism.
It is now commonly known that using response-adaptive designs for data collection offers great potential in terms of optimizing expected outcomes, but poses multiple challenges for inferential goals. In many settings, such as phase-II or confirmatory clinical trials, a main barrier to their practical use is the lack of type-I error guarantees and/or power efficiency, especially in finite samples. This work addresses this gap. Specifically, focusing on a novel test statistic defined on the randomization probabilities of the (randomized) adaptive design, we derive its finite-sample and asymptotic guarantees. Further theoretical properties are evaluated for Thompson sampling, a Bayesian response-adaptive design that is commonly used both in clinical applications and beyond (eg, recommendation systems or mobile health). The frequentist error control advantages of the proposed approach-also able to preserve expected outcome optimalities-are illustrated in a real-world phase-II oncology trial and in simulation experiments.
Consistent physical inactivity poses a major global health challenge. Mobile health (mHealth) interventions, particularly Just-in-Time Adaptive Interventions (JITAIs), offer a promising avenue for scalable, personalized physical activity (PA) promotion. However, developing and evaluating such interventions at scale, while integrating robust behavioral science, presents methodological hurdles. The PEARL study was the first large-scale, four-arm randomized controlled trial to assess a reinforcement learning (RL) algorithm, informed by health behavior change theory, to personalize the content and timing of PA nudges via a Fitbit app. We enrolled and randomized 13,463 Fitbit users into four study arms: control, random, fixed, and RL. The control arm received no nudges. The other three arms received nudges from a bank of 155 nudges based on behavioral science principles. The random arm received nudges selected at random. The fixed arm received nudges based on a pre-set logic from survey responses about PA barriers. The RL group received nudges selected by an adaptive RL algorithm. We included 7,711 participants in primary analyses (mean age 42.1, 86.3 We observed an increase in PA for the RL group compared to all other groups from baseline to 1 and 2 months. The RL group had significantly increased average daily step count at 1 month compared to all other groups: control (+296 steps, p=0.0002), random (+218 steps, p=0.005), and fixed (+238 steps, p=0.002). At 2 months, the RL group sustained a significant increase compared to the control group (+210 steps, p=0.0122). Generalized estimating equation models also revealed a sustained increase in daily steps in the RL group vs. control (+208 steps, p=0.002). These findings demonstrate the potential of a scalable, behaviorally-informed RL approach to personalize digital health interventions for PA.
This work revisits optimal response-adaptive designs from a type-I error rate perspective, highlighting when and how much these allocations exacerbate type-I error rate inflation - an issue previously undocumented. We explore a range of approaches from the literature that can be applied to reduce type-I error rate inflation. However, we found that all of these approaches fail to give a robust solution to the problem. To address this, we derive two optimal proportions, incorporating the more robust score test (instead of the Wald test) with finite sample estimators (instead of the unknown true values) in the formulation of the optimization problem. One proportion optimizes statistical power and the other minimizes the total number failures in a trial while maintaining a predefined power level. Through simulations based on an early-phase and a confirmatory trial we provide crucial practical insight into how these new optimal proportion designs can offer substantial patient outcomes advantages while controlling type-I error rate. While we focused on binary outcomes, the framework offers valuable insights that naturally extend to other outcome types, multi-armed trials and alternative measures of interest.