BACKGROUND:In randomized trials where some standard-treatment arm patients cross to the experimental treatment, it is frequently of interest to estimate the between-arm survival difference as if no patients on the standard-treatment arm had crossed over to the experimental treatment. Rank-preserving structural failure time models, an extension of semiparametric accelerated-failure-time models, are a popular method for accomplishing this because they do not require modeling which patients will crossover. METHODS:In trying to apply the rank-preserving structural failure time model in practice, we noted some unusual behavior of the estimated acceleration parameter (differential treatment effect). Simple examples and limited simulations are provided to examine and understand this behavior. RESULTS:The simulations show that rank-preserving structural failure time model estimator of the acceleration parameter can take on extreme values, especially when the intent-to-treat analysis favors the standard-treatment arm. Furthermore, the addition of censoring is paradoxically shown to reduce the estimator's variability compared to the uncensored data when the underlying observations are exponentially distributed. Use of a Weibull distribution with short tails for the survival times eliminates this unusual behavior. CONCLUSION:The rank-preserving structural failure time model estimators of the acceleration parameter are not based on the joint ranks of the original data, and it is suggested that this makes acceleration-parameter estimator unstable with long-tailed survival distributions.
Improvements in cancer treatment are guided by randomized clinical trials (RCTs) that reliably assess the therapeutic impact of adding experimental treatments into clinical practice. The fundamental value of RCTs is that the difference in outcomes between the experimental and control arms provides an unbiased estimate of the risk-benefit ratio for the experimental treatment. However, the clinical relevance of trial results depends on the control arm accurately reflecting the appropriate standard of care (SOC) in the intended-use population. Due to the accelerating pace of cancer treatment development, it is becoming more common for the SOC to change during an ongoing randomized trial. This increasingly dynamic SOC landscape may have a nontrivial impact on the conduct and interpretation of future RCTs, and thus it is important to appreciate these issues and, when possible, prospectively address them when designing and conducting clinical trials. We review common scenarios where the SOC changes and provide recommendations on how to design a trial to minimize the impact of such changes when the trial is ongoing.
Background/aimsRandomized clinical trials often use stratification to ensure balance between arms. Analysis of primary endpoints of these trials typically uses a "stratified analysis," in which analyses are performed separately in each subgroup defined by the stratification factors, and those separate analyses are weighted and combined. In the phase 3 setting, stratified analyses based on a small number of stratification factors can provide a small increase in power. The impact on power and type-1 error of stratification in the setting of smaller sample sizes as in randomized phase 2 trials has not been well characterized.MethodsWe performed computational studies to characterize the power and cross-arm balance of modestly sized clinical trials (less than 170 patients) with varying numbers of stratification factors (0-6), sample sizes, randomization ratios (1:1 vs 2:1), and randomization methods (dynamic balancing vs stratified block).ResultsWe found that the power of unstratified analyses was minimally impacted by the number of stratification factors used in randomization. Analyses stratified by 1-3 factors maintained power over 80%, while power dropped below 80% when four or more stratification factors were used. These trends held regardless of sample size, randomization ratio, and randomization method. For a given randomization ratio and sample size, increasing the number of factors used in randomization had an adverse impact on cross-arm balance. Stratified block randomization performed worse than dynamic balancing with respect to cross-arm balance when three or more stratification factors were used.ConclusionStratified analyses can decrease power in the setting of phase 2 trials when the number of patients in a stratification subgroup is small.
Abstract In recent years, there has been increased interest in incorporation of backfilling into dose-escalation clinical trials, which involves concurrently assigning patients to doses that have been previously cleared for safety by the dose-escalation design. Backfilling generates additional information on safety, tolerability, and preliminary activity on a range of doses below the maximum tolerated dose (MTD), which is relevant for selection of the recommended phase II dose and dose optimization. However, in practice, backfilling may not be rigorously defined in trial protocols and implemented consistently. Furthermore, backfilling designs require careful planning to minimize the probability of treating additional patients with potentially inactive agents (and/or subtherapeutic doses). In this paper, we propose a simple and principled approach to incorporate backfilling into the Bayesian optimal interval design (BOIN). The design integrates data from the dose-escalation and backfilling components of the design and ensures that the additional patients are treated at doses where some activity has been seen. Simulation studies demonstrated that the proposed backfilling BOIN design (BF-BOIN) generates additional data for future dose optimization, maintains the accuracy of the MTD identification, and improves patient safety without prolonging the trial duration.
Phase III trials that randomly assign patients to a control treatment (C), an experimental treatment (A), or a combination treatment (AB) should be designed with the goal to recommend the best treatment: AB (if it is better than A and C), A (if it is better than C, and AB is not better than A), or C (if neither AB nor A is better than C). However, this goal can be challenging to achieve with statistical confidence. We performed a survey of cancer trials published in five journals from January 2018 to May 2024 to assess the trial designs being used in this setting and found that three quarters of them did not have a provision for a formal comparison of the AB treatment arm with the A treatment arm, a possible shortcoming. A limited simulation evaluates two analysis strategies that incorporate an AB versus A comparison and is used to formulate some recommendations for designing these types of trials.
In a randomized clinical trial, instead of allocating patients equally between the treatment arms, some trials in oncology assign a higher proportion of patients to receive the experimental treatment arm (eg, a two-to-one randomization). In this commentary, we first briefly review the common reasons given for the use of a two-to-one randomization and provide some examples of trials using these designs. We then explain why the risk-benefit ratio of this approach may not be favorable as is commonly assumed.
New oncology therapies that extend patients' lives beyond initial expectations and improving later-line treatments can lead to complications in clinical trial design and conduct. In particular, for trials with event-based analyses, the time to observe all the protocol-specified events can exceed the finite follow-up of a clinical trial or can lead to much delayed release of outcome data. With the advent of multiple classes of oncology therapies leading to much longer survival than in the past, this issue in clinical trial design and conduct has become increasingly important in recent years. We propose a straightforward prespecified backstop rule for trials with a time-to-event analysis and evaluate the impact of the rule with both simulated and real-world trial data. We then provide recommendations for implementing the rule across a range of oncology clinical trial settings.
PURPOSE A phase II/III trial is a type of phase III trial that has embedded in it an intermediate phase II go/no-go decision as to whether to continue the accrual to the phase III sample size. We examine the design characteristics and experience of a well-defined set of National Cancer Institute phase II/III trials, with special emphasis on designed accrual suspensions while awaiting the data to become mature enough for the phase II analysis. This experience is used to highlight the potential of using a calendar backstop to avoid an inordinately long accrual suspension. METHODS We identified all phase II/III trials conducted by NRG Oncology or its precursor National Cancer Institute Cooperative Groups (Radiation Therapy Oncology Group, Gynecologic Oncology Group, and National Surgical Adjuvant Breast and Bowel Project). The design characteristics were recorded, and, for completed trials, the trial results in terms of sample sizes and timing of analyses were tabulated. RESULTS Twenty-two trials were identified, 14 of which had a time-to-event end point for their phase II component. Thirteen of these 14 trials had designed accrual suspensions. Seven of the eight completed trials had designed accrual suspensions, all of which went on longer than their projected suspension times (3-20 months longer than planned). The trade-offs for using a backstop are discussed using one of these trials as an example. CONCLUSION Phase II/III trials with an accrual suspension and a predefined backstop for the phase II analysis can be a useful tool for minimizing patient exposure to ineffective experimental treatments while still obtaining the trial results in a timely fashion.
When designing a randomized clinical trial to compare two treatments, the sample size required to have desired power with a specified type 1 error depends on the hypothesis testing procedure. With a binary endpoint (e.g., response), the trial results can be displayed in a 2 × 2 table. If one does the analysis conditional on the number of positive responses, then using Fisher's exact test has an actual type 1 error less than or equal to the specified nominal type 1 error. Alternatively, one can use one of many unconditional “exact” tests that also preserve the type 1 error and are less conservative than Fisher's exact test. In particular, the unconditional test of Boschloo is always at least as powerful as Fisher's exact test, leading to smaller required sample sizes for clinical trials. However, many statisticians have argued over the years that the conditional analysis with Fisher's exact test is the only appropriate procedure. Since having smaller clinical trials is an extremely important consideration, we review the general arguments given for the conditional analysis of a 2 × 2 table in the context of a randomized clinical trial. We find the arguments not relevant in this context, or, if relevant, not completely convincing, suggesting the sample‐size advantage of the unconditional tests should lead to their recommended use. We also briefly suggest that since designers of clinical trials practically always have target null and alternative response rates, there is the possibility of using this information to improve the power of the unconditional tests.
This article provides a summary of discussions from the American Statistical Association (ASA) Biopharmaceutical (BIOP) Section Open Forum organized by the ASA BIOP Statistical Methods in Oncology Scientific Working Group in coordination with the US FDA Oncology Center of Excellence and LUNGevity Foundation on January 14, 2021, and February 8, 2021. Diverse stakeholders including oncologists, patient advocates, experts from international regulatory agencies, academicians, and representatives of the pharmaceutical industry engaged in a discussion on how best to incorporate lessons learned during the COVID-19 pandemic into the design of future oncology trials. While recognizing that decentralized or hybrid cancer trials may increase variability associated with measurement error and potentially increase bias in treatment effect estimation, panel discussions highlighted the importance of flexibility for decreasing patient burden, which has the potential to increase access to and retention in cancer clinical trials and may broaden the representation of real-world patients in the trial setting.
Recent therapeutic advances have led to improved patient survival in many cancer settings. Although prolongation of survival remains the ultimate goal of cancer treatment, the availability of effective salvage therapies could make definitive phase III trials with primary overall survival (OS) end points difficult to complete in a timely manner. Therefore, to accelerate development of new therapies, many phase III trials of new cancer therapies are now designed with intermediate primary end points (eg, progression-free survival in the metastatic setting) with OS designated as a secondary end point. We review recently published phase III trials and assess contemporary practices for designing and reporting OS as a secondary end point. We then provide design and reporting recommendations for trials with OS as a secondary end point to safeguard OS data integrity and optimize access to the OS data for patient, clinician, and public-health stakeholders.
This article provides a summary of discussions from the American Statistical Association (ASA) Biopharmaceutical (BIOP) Section Open Forum organized by the ASA BIOP Statistical Methods in Oncology Scientific Working Group in coordination with the US FDA Oncology Center of Excellence and LUNGevity Foundation on June 24, 2021, and January 13, 2022. Diverse stakeholders engaged in a discussion on how best to use various innovative clinical trial designs in designing future pediatric oncology trials. While standard randomized controlled trials are preferred to evaluate treatment effect in an unbiased manner, given the rarity of pediatric cancers, innovative strategies are needed to promote and assure timely cancer drug development in pediatric populations. The discussions highlighted the need to consider innovative designs with less stringent Type I error specification and Bayesian designs borrowing from external control data, or borrowing treatment effect information from adult data, or both. Such designs are available in the literature and some examples are summarized under the FDA Complex Innovative Trials Design Pilot Program (https://www.fda.gov/drugs/development-resources/complex-innovative-trial-design-meeting-program). Early consultation with global regulatory agencies for pediatric clinical trials can provide a better understanding of different features of the clinical trial design options for successful pediatric cancer drug development.
AbstractOver the past decade, multiple trials, including the precision medicine trial National Cancer Institute-Molecular Analysis for Therapy Choice (NCI-MATCH, EAY131, NCT02465060) have sought to determine if treating cancer based on specific genomic alterations is effective, irrespective of the cancer histology. Although many therapies are now approved for the treatment of cancers harboring specific genomic alterations, most patients do not respond to therapies targeting a single alteration. Further, when antitumor responses do occur, they are often not durable due to the development of drug resistance. Therefore, there is a great need to identify rational combination therapies that may be more effective. To address this need, the NCI and National Clinical Trials Network have developed NCI-ComboMATCH, the successor to NCI-MATCH. Like the original trial, NCI-ComboMATCH is a signal-seeking study. The goal of ComboMATCH is to overcome drug resistance to single-agent therapy and/or utilize novel synergies to increase efficacy by developing genomically-directed combination therapies, supported by strong preclinical in vivo evidence. Although NCI-MATCH was mainly comprised of multiple single-arm studies, NCI-ComboMATCH tests combination therapy, evaluating both combination of targeted agents as well as combinations of targeted therapy with chemotherapy. Although NCI-MATCH was histology agnostic with selected tumor exclusions, ComboMATCH has histology-specific and histology-agnostic arms. Although NCI-MATCH consisted of single-arm studies, ComboMATCH utilizes single-arm as well as randomized designs. NCI-MATCH had a separate, parallel Pediatric MATCH trial, whereas ComboMATCH will include children within the same trial. We present rationale, scientific principles, study design, and logistics supporting the ComboMATCH study.
The goal of dose optimization during drug development is to identify a dose that preserves clinical benefit with optimal tolerability. Traditionally, the maximum tolerated dose in a small phase I dose escalation study is used in the phase II trial assessing clinical activity of the agent. Although it is possible that this dose level could be altered in the phase II trial if an unexpected level of toxicity is seen, no formal dose optimization has routinely been incorporated into later stages of drug development. Recently it has been suggested that formal dose optimization (involving randomly assigning patients between 2 or more dose levels) be routinely performed early in drug development, even before it is known that the experimental therapy has any clinical activity at any dose level. We consider the relative merits of performing dose optimization earlier vs later in the drug development process and demonstrate that a considerable number of patients may be exposed to ineffective therapies unless dose optimization is delayed until after clinical activity or benefit of the new agent has been established. We conclude that patient and public health interests may be better served by conducting dose optimization after (or during) phase III evaluation, with some exceptions when dose optimization should be performed after activity shown in phase II evaluation.