BACKGROUND:Response-adaptive randomization is controversial even in the best circumstances when based on a quickly determined primary outcome. In disease settings in which the primary outcome requires long follow-up, an intermediate endpoint may be chosen to update randomization allocations. The aim of our study is to evaluate the impact of response-adaptive randomization applied to an imperfect intermediate endpoint. We use tuberculosis trials as the motivating example. METHODS:We simulated a response-adaptive randomization design, adapting randomization allocations using an imperfect intermediate endpoint, in a superiority trial of two experimental regimens and one control arm. The primary study outcome was treatment success after 73 weeks from randomization; the intermediate endpoint was culture conversion at 8 weeks. We compared different sensitivity (Se) and specificity (Spe) scenarios for the intermediate endpoint, while varying the true treatment efficacy. We evaluated the performance of response-adaptive randomization to achieve its primary goal of allocating more participants to the better arm and the impact of time-trends on type I error rate. RESULTS:Even in an ideal state of perfect accuracy (i.e. intermediate endpoint with Se = 100% and Spe = 100%), response-adaptive randomization did not always live up to its main purpose of allocating more patients to the better arm. Lower accuracy of the intermediate endpoint leads to greater divergence from the goal of more allocations to the better arm. The larger the difference in treatment efficacy between the arms, the more striking the impact of an intermediate endpoint with poor diagnostic accuracy. Time-trends inflate the type I error rate, and while stratified tests can correct this, they do so at the cost of a power loss. Allocating more patients to the worst arm increases power for comparisons with this arm but reduces power for comparisons of the best arm to control. CONCLUSION:Given the objective of evaluating several new therapeutic regimens in a timely manner, response-adaptive randomization is tempting. However, it requires at least reliance on highly accurate intermediate endpoints, which are still no guarantee of response-adaptive randomization's trustworthiness.
Background:Some people with HIV (PWH) have a poor immunologic response to antiretroviral therapy (ART). Such immunologic nonresponders are at increased risk for HIV and non-HIV related complications. Immune exhaustion may contribute to poor immune reconstitution, and blockade of PD-1, an immune checkpoint molecule, may thus be beneficial. Method:We undertook a phase 1, 3:1 randomized, placebo-controlled trial of a single 200 mg dose of pembrolizumab in PWH with CD4+ T-cell counts between 100 and 350 cells/mm3, who had been on ART and were virally suppressed for at least 12 months. The primary endpoint was the frequency of either Grade 3 or higher adverse events or Grade 2 or higher autoimmune events requiring corticosteroid therapy. Additional endpoints included levels of PD-1 expression, and changes in immune and virologic parameters. Results:Of the 7 enrolled participants, 6 received pembrolizumab and 1 received placebo. The only grade 3 event, ophthalmic zoster, developed in the placebo recipient. There were no autoimmune events requiring corticosteroid therapy. CD4+ and CD8+ T-cell counts were stable over time, though both showed increased CD38 and HLA-DR expression. PD-1 expression declined from a baseline of 60.4% (SD, 7.4%) to a nadir of 0.8%-23.5% for CD4+ T cells, and from 45.2% (SD, 9.4%) to 0.1%-23.2% for CD8+ T cells. There were no significant changes in viral killing, plasma viral loads or proviral DNA levels. Conclusions:A single dose of pembrolizumab may be safely administered to PWH who are immunologic nonresponders.
Background/Aims: In randomized two-armed clinical trials with binary endpoints, there may be uncertainty about the event probability, which is needed for sample size calculation. Survival trials are powered based on number of events rather than people, and this is advantageous because the number of events needed to achieve a desired power is less sensitive to an unknown parameter than is the number of people needed. We investigate and quantify this relative stability of number of events compared to number of people in the context of a randomized two-armed trial with equal sample sizes and a binary endpoint. In binary endpoint settings with such relative stability, we consider (1) enhancement of traditional adaptive trial design and (2) potential benefits of a simple event-driven strategy.Methods: Using sample size formulas, we determine the relative stability of the expected number of events compared to the sample size for binary outcome trials using the relative risk, odds ratio, or risk difference. Simulations consider a simple event-driven design when there is relative stability; we evaluate type I error rate and power under various analysis methods and approaches to halting the trial.Results: We find that the number of events is at least three times more stable than the sample size to achieve a specified power for the relative risk when the overall event probability is less than 1/3, and for the odds ratio when the overall event probability is less than 0.20. We show that this relative stability is independent of the type 1 and type 2 error rates and magnitude of the treatment effect. In a setting where the overall event probability is consistent with relative stability, simulations of an event-driven design show that asymptotic methods may have modestly high type I error rates, but that other approaches appear to have good operating characteristics.Conclusion: In settings with moderately low event probabilities, thinking in terms of the number of events instead of sample size may (1) facilitate the planning of clinical trials and help determine whether a trial is futile, and (2) lead to a simple event-driven design for binary endpoints that may be feasible and appealing.
At the beginning of a phase III clinical trial, there is great optimism. After all, the phase II trial results were encouraging. Then, early data from the phase III trial trend in the wrong way, but there is still an opportunity for the trend to reverse and become statistically significant by the end. At what point does optimism become denial of reality? How do we decide when a clinical trial is futile? What does futility even mean? This tutorial reviews different concepts and tools for evaluating futility, including conditional and predictive power, reverse conditional power, predicted interval plots, revised unconditional power, and beta spending functions.
There are many connections between probability, other mathematics courses, and statistics. Understanding these connections provides insights that might not be fully appreciated when considering each discipline in isolation. While the typical instruction of statistics courses relies on elucidating its foundational principles from mathematical and probability theory, it is generally less appreciated that statistics can in turn provide a deeper understanding of results in mathematics and probability. We offer several examples for which knowledge of statistics can shed new light on probability and other mathematics results. Examples span both undergraduate and graduate level material. In today's data driven-world, many students are naturally curious about statistics and are exposed to this field early in their undergraduate curriculum. Leveraging connections between statistics and mathematics and probability makes theoretical concepts more intuitive and relevant, fostering a better understanding.
The first Adaptive COVID-19 Treatment Trial (ACTT-1) showed that remdesivir improved COVID-19 recovery time compared with placebo in hospitalized adults. The secondary outcome of mortality was almost significant overall (p = 0.07) and highly significant for people receiving supplemental oxygen at enrollment (p = 0.002), suggesting a mortality benefit concentrated in this group. We explore analysis methods that are helpful when a single subgroup benefits from treatment and apply them to ACTT-1, using baseline oxygen use to define subgroups. We consider two questions: (1) is the remdesivir effect for people receiving supplemental oxygen real, and (2) does this effect differ from the overall effect? For Question 1, we apply a Bonferroni adjustment to subgroup-specific hypothesis tests and the Westfall and Young permutation test, which is valid when small cell counts preclude normally distributed test statistics (a frequently unexamined condition in subgroup analyses). For Question 2, we introduce Qmax, the largest standardized difference between subgroup-specific effects and the overall effect. Qmax simultaneously tests whether any subgroup effect differs from the overall effect and identifies the subgroup benefitting most. We demonstrate that Qmax strongly controls the familywise error rate (FWER) when test statistics are normally distributed with no mean-variance relationship. We compare Qmax to a related permutation test, SEAMOS, which was previously proposed but not extensively applied or tested. We show that SEAMOS can have inflated Type 1 error under the global null when control arm event rates differ between subgroups. Our results support a mortality benefit from remdesivir in people receiving supplemental oxygen.
BACKGROUND:The coronavirus disease 2019 pandemic highlighted the need to conduct efficient randomized clinical trials with interim monitoring guidelines for efficacy and futility. Several randomized coronavirus disease 2019 trials, including the Multiplatform Randomized Clinical Trial (mpRCT), used Bayesian guidelines with the belief that they would lead to quicker efficacy or futility decisions than traditional "frequentist" guidelines, such as spending functions and conditional power. We explore this belief using an intuitive interpretation of Bayesian methods as translating prior opinion about the treatment effect into imaginary prior data. These imaginary observations are then combined with actual observations from the trial to make conclusions. Using this approach, we show that the Bayesian efficacy boundary used in mpRCT is actually quite similar to the frequentist Pocock boundary. METHODS:The mpRCT's efficacy monitoring guideline considered stopping if, given the observed data, there was greater than 99% probability that the treatment was effective (odds ratio greater than 1). The mpRCT's futility monitoring guideline considered stopping if, given the observed data, there was greater than 95% probability that the treatment was less than 20% effective (odds ratio less than 1.2). The mpRCT used a normal prior distribution that can be thought of as supplementing the actual patients' data with imaginary patients' data. We explore the effects of varying probability thresholds and the prior-to-actual patient ratio in the mpRCT and compare the resulting Bayesian efficacy monitoring guidelines to the well-known frequentist Pocock and O'Brien-Fleming efficacy guidelines. We also contrast Bayesian futility guidelines with a more traditional 20% conditional power futility guideline. RESULTS:A Bayesian efficacy and futility monitoring boundary using a neutral, weakly informative prior distribution and a fixed probability threshold at all interim analyses is more aggressive than the commonly used O'Brien-Fleming efficacy boundary coupled with a 20% conditional power threshold for futility. The trade-off is that more aggressive boundaries tend to stop trials earlier, but incur a loss of power. Interestingly, the Bayesian efficacy boundary with 99% probability threshold is very similar to the classic Pocock efficacy boundary. CONCLUSIONS:In a pandemic where quickly weeding out ineffective treatments and identifying effective treatments is paramount, aggressive monitoring may be preferred to conservative approaches, such as the O'Brien-Fleming boundary. This can be accomplished with either Bayesian or frequentist methods.
Abstract Public health emergencies present special challenges in the monitoring of clinical trials. There is tremendous pressure to find effective treatments as quickly as possible without jeopardizing the integrity of results. Frequent early monitoring from an experienced data and safety monitoring board (DSMB) is required to ensure patient safety, especially when novel or repurposed interventions may be tested earlier than they would be in a nonemergency setting. This chapter begins by reviewing general monitoring, including the purpose, composition, and operation of DSMBs, efficacy monitoring boundaries, tools such as conditional power and predicted interval plots for judging whether continuation of a trial is futile, and practical aspects of monitoring clinical trials. We offer important historical examples, like the Cardiac Arrhythmia Suppression Trial (CAST), as well as more recent examples specific to infectious disease, such as the Adaptive COVID-19 Treatment Trial (ACTT-1). We highlight the special challenges posed by monitoring clinical trials in public health emergencies and illustrate these with a case study of the Partnership for Research on Ebola Vaccines in Liberia (PREVAIL-I) vaccine trial in West Africa. We conclude with important lessons learned.
Recent examples for unplanned external events are the global COVID-19 pandemic, the war in Ukraine, or most recently Hurricane Ian in Puerto Rico. Disruptions due to unplanned external events can lead to violation of assumptions in clinical trials. In certain situations, randomization tests can provide nonparametric inference that is robust to violation of the assumptions usually made in clinical trials. The ICH E9 (R1) Addendum on estimands and sensitivity analyses provides a guideline for aligning the trial objectives with strategies to address disruptions in clinical trials. In this article, we embed randomization tests within the estimand framework to allow for inference following disruptions in clinical trials in a way that reflects recent literature. A stylized clinical trial is presented to illustrate the method, and a simulation study highlights situations when a randomization test that is conducted under the intention-to-treat principle can provide unbiased results.
ABSTRACT Designing clinical trials for emerging infectious diseases such as COVID-19 is challenging because information needed for proper planning may be lacking. Pre-specified adaptive designs can be attractive options, but what happens if a trial with no such design needs to be modified? For example, unexpectedly high efficacy (approximately 95%) in two COVID-19 vaccine trials might cause investigators in other COVID-19 vaccine trials to increase the number of interim analyses to allow earlier stopping for efficacy. If such a decision is based solely on external data, there are no issues, but what if internal trial data by arm are also examined? Fortunately, the conditional error principle of Müller and Schäfer (2004) can be used to ensure no inflation of the type 1 error rate, even if no interim analyses were planned. We study the properties, including limitations, of this method. We provide a shiny app to evaluate changes in timing of interim analyses in response to outcome data by arm in clinical trials.
The Accelerating COVID-19 Therapeutic Interventions and Vaccines (ACTIV) Cross-Trial Statistics Group gathered lessons learned from statisticians responsible for the design and analysis of the 11 ACTIV therapeutic master protocols to inform contemporary trial design as well as preparation for a future pandemic. The ACTIV master protocols were designed to rapidly assess what treatments might save lives, keep people out of the hospital, and help them feel better faster. Study teams initially worked without knowledge of the natural history of disease and thus without key information for design decisions. Moreover, the science of platform trial design was in its infancy. Here, we discuss the statistical design choices made and the adaptations forced by the changing pandemic context. Lessons around critical aspects of trial design are summarized, and recommendations are made for the organization of master protocols in the future.
Accelerating COVID-19 Treatment Interventions and Vaccines (ACTIV) was initiated by the US government to rapidly develop and test vaccines and therapeutics against COVID-19 in 2020. The ACTIV Therapeutics-Clinical Working Group selected ACTIV trial teams and clinical networks to expeditiously develop and launch master protocols based on therapeutic targets and patient populations. The suite of clinical trials was designed to collectively inform therapeutic care for COVID-19 outpatient, inpatient, and intensive care populations globally. In this report, we highlight challenges, strategies, and solutions around clinical protocol development and regulatory approval to document our experience and propose plans for future similar healthcare emergencies.
Abstract This chapter outlines the demanding, complex task of monitoring an expedited clinical trial in a resource-limited, insecure setting. The PALM Ebola therapeutics trial was implemented in the Democratic Republic of the Congo (DRC) despite minimal basic infrastructure, widespread suspicion and hostility among the population, and outbreaks of violence in the region. In order to keep participant enrollment numbers manageable, the PALM trial was designed to detect a large treatment effect, and its primary endpoint was mortality within 28 days after randomization. The course of an Ebola epidemic can wax and wane unpredictably, causing fluctuations in enrollment and a lag in information presented to the data and safety monitoring board (DSMB). The DSMB had to meet frequently, usually virtually, to monitor the four treatments being investigated and detect safety and efficacy signals. Since both resources and medical personnel were at their limits, collecting and verifying all data was difficult. Therefore, streamlining data collection and focusing on the most critical outcomes for safety and efficacy monitoring were paramount. Results indicated that two of the four agents under study, MAb114 and REGN-EB3, were clearly superior to the ZMapp control arm, so much so that the trial was stopped early.
Background:Toxoplasmic encephalitis (TE) is a life-threatening complication of people with human immunodeficiency virus (PWH) with severe immunodeficiency, especially those with a CD4+ T-cell count <100 cells/µL. Following a clinical response to anti-Toxoplasma therapy, and immune reconstitution after initiation of combination antiretroviral therapy (ART), anti-Toxoplasma therapy can be discontinued with a low risk of relapse. Methods:To better understand the evolution of magnetic resonance imaging (MRI)-defined TE lesions in PWH receiving ART, we undertook a retrospective study of PWH initially seen at the National Institutes of Health between 2001 and 2012, who had at least 2 serial MRI scans. Lesion size and change over time were calculated and correlated with clinical parameters. Results:Among 24 PWH with TE and serial MRI scans, only 4 had complete clearance of lesions at the last MRI (follow-up, 0.09-5.8 years). Of 10 PWH off all anti-Toxoplasma therapy (median, 3.2 years after TE diagnosis), 6 had persistent MRI enhancement. In contrast, all 5 PWH seen in a pre-ART era study who were followed for >6 months had complete clearance of lesions. TE lesion area at diagnosis was associated with the absolute change in area (P < .0001). Conclusions:Contrast enhancement can persist even when TE has been successfully treated and anti-Toxoplasma therapy has been stopped, highlighting the need to consider diagnostic alternatives in successfully treated patients with immune reconstitution presenting with new neurologic symptoms.
Rare events can sometime arise in clinical development of treatments. For example, CYPIDES was a single-arm study of the CYP11A1 inhibitor ODM-208 to treat metastatic prostate cancer.1 Preclinical testing of the compound identified elevated thyroid-stimulating hormone (TSH) and bilirubin in rats and dogs. Unusual findings in preclinical testing focus attention and magnify evidence if similar results occur in humans. By analogy, imagine a murder trial in which the only evidence against the defendant arose from a database search of DNA matching the partial profile found at the crime scene. Multiple people could match, so without other evidence, the perpetrator could be any of them.
SARS-CoV-2 mRNA booster vaccines provide protection from severe disease, eliciting strong immunity that is further boosted by previous infection. However, it is unclear whether these immune responses are affected by the interval between infection and vaccination. Over a 2-month period, we evaluated antibody and B cell responses to a third-dose mRNA vaccine in 66 individuals with different infection histories. Uninfected and post-boost but not previously infected individuals mounted robust ancestral and variant spike-binding and neutralizing antibodies and memory B cells. Spike-specific B cell responses from recent infection (<180 days) were elevated at pre-boost but comparatively less so at 60 days post-boost compared with un-infected individuals, and these differences were linked to baseline frequencies of CD27lo B cells. Day 60 to baseline ratio of BCR signaling measured by phosphorylation of Syk was inversely correlated to days be-tween infection and vaccination. Thus, B cell responses to booster vaccines are impeded by recent infection.
The PREDICT TB trial tests noninferiority of an abbreviated treatment regimen (arm A) vs a conventional treatment regimen (arm C). Treatment trials of drug-susceptible tuberculosis are expected to have low event rates (ie, relapse probabilities around 3-5%). We examine the question of what is the "best" way to test for noninferiority in a setting with low event rates. In a series of simulations supported by theoretical arguments, we examine operating characteristics of five tests, including normal approximation, exact, and simulation-based tests. Two of these tests are constructed from Kaplan-Meier based-estimators, which account for variable follow-up time (and those lost to follow-up). We evaluate the effect of loss to follow-up via simulations. We also examine the results of the five tests on a data set similar to PREDICT TB, the REMoxTB trial. We find that the normal approximation tests perform well, albeit with small type I error rate inflation. We also find that the Kaplan-Meier methods generally have larger power than the other tests, especially when there is between 10-30% loss to follow-up.
Background:Immune dysregulation contributes to poorer outcomes in severe Covid-19. Immunomodulators targeting various pathways have improved outcomes. We investigated whether infliximab provides benefit over standard of care.Methods:We conducted a master protocol investigating immunomodulators for potential benefit in treatment of participants hospitalized with Covid-19 pneumonia. We report results for infliximab (single dose infusion) versus shared placebo both with standard of care. Primary outcome was time to recovery by day 29 (28 days after randomization). Key secondary endpoints included 14-day clinical status and 28-day mortality.Results:A total of 1033 participants received study drug (517 infliximab, 516 placebo). Mean age was 54.8 years, 60.3% were male, 48.6% Hispanic or Latino, and 14% Black. No statistically significant difference in the primary endpoint was seen with infliximab compared with placebo (recovery rate ratio 1.13, 95% CI 0.99-1.29; p=0.063). Median (IQR) time to recovery was 8 days (7, 9) for infliximab and 9 days (8, 10) for placebo. Participants assigned to infliximab were more likely to have an improved clinical status at day 14 (OR 1.32, 95% CI 1.05-1.66). Twenty-eight-day mortality was 10.1% with infliximab versus 14.5% with placebo, with 41% lower odds of dying in those receiving infliximab (OR 0.59, 95% CI 0.39-0.90). No differences in risk of serious adverse events including secondary infections.Conclusions:Infliximab did not demonstrate statistically significant improvement in time to recovery. It was associated with improved 14-day clinical status and substantial reduction in 28- day mortality compared with standard of care.Trial registration:ClinicalTrials.gov ( NCT04593940 ).