Participants who are randomized to treatment but have no post-baseline data pose a unique challenge. These participants need to be included to preserve randomization. Because there is no information about the outcome or the intercurrent event(s) that led to missing data, an estimand of interest, and the focus of this study, is a hypothetical strategy to estimate what would have been observed if participants had not discontinued. Various imputation-based and likelihood-based analyses were compared in simulated and real clinical trial data. Models that used baseline as a covariate or constrained baseline values to be equal yielded similar results and had greater power than an unconstrained analysis that fit baseline as a response. Assigning change to the first post-baseline visit as 0 and applying an analysis with baseline as a covariate controlled Type I error at the nominal level and had power equal to or greater than other methods. Treatment contrasts were not biased when the reason for missing all post-baseline data was random or treatment related. Within-group bias occurred with outcome-related missingness of all post-baseline data, but the bias was nearly equal in the two arms, leading to unbiased treatment contrasts. Bias occurred when missing all post-baseline data was related to treatment and outcome. Given the idiosyncratic nature of clinical trials, no universally best analytic approach exists for dealing with participants that have a baseline but no post-baseline data. Analysts can choose among the methods to tailor an approach to the situation at hand.
IntroductionChoice of the primary outcome can be problematic in 1) diseases with heterogeneous signs and symptoms, 2) trials of disease-modifying treatments (DMTs) that are expected to affect all aspects of the disease, 3) in complex and rare diseases with minimal data on clinical trial outcomes. In such situations, a single outcome measuring a single domain of disease will rarely suffice as a primary outcome. To address this issue, regulatory bodies often suggest co-primary endpoints or require efficacy on both primary and key secondary endpoints for confirmatory trials, to ensure that at least two different domains of disease are affected by treatment. However, obtaining statistical significance on two outcomes is a much stricter requirement than on a single primary outcome. Global statistical tests (GST) combine multiple outcomes into a single score and could provide a viable alternative to the co-primary approach. Importantly for rare diseases, combining multiple assessments reduces the risk of selecting a poorly performing outcome simply because it has not been studied extensively.MethodsWe conducted simulations to compare GST to single primary and co-primary endpoint approaches with two moderately or highly correlated outcomes with various effect sizes.ResultsFor scenarios with the same true effect size on both outcomes, the GST had greater power than single primary and co-primary approaches, regardless of the correlation level between outcomes. This was also true with different effect size combinations at the same correlation level. With an effect observed on one outcome only, GST was more likely to yield statistical significance than the co-primary approach. Unlike the co-primary approach, The GST yielded lower p-values in scenarios with lower correlation between the outcomes due to the independence of the information from each endpoint.ConclusionsUsing GST as a prespecified endpoint is appropriate in trials where a clear primary endpoint has not been identified, sample sizes are insufficient to support multiple primary endpoints, and/or more comprehensive assessments across multiple endpoints are needed to fully evaluate outcomes.
Objective: Previous experience with antidepressant studies highlight the difficulties in discriminating an effective drug from placebo. In hopes of improving signal detection, three easy-to-implement methodologies were employed during the development of a recently approved antidepressant. Experimental Design: Results from alternative and traditional methods could be compared directly because most studies employed both methods. This database included 11 double-blind, placebo-controlled trials (some with multiple dose arms and/or active comparators) yielding 22 treatment arms of antidepressants at or above the minimally effective dose noted in their U.S. labels. Principal Observations: Results agreed with the previous evidence showing that the performance of a likelihood-based, mixed-effects model repeated measures (MMRM) analysis was superior to that of analysis of covariance with missing values imputed using the last observation carried forward (LOCF) approach; MMRM correctly identified drug as superior to placebo in 14/22 (63.6%) comparisons versus 11/22 (50.0%) for LOCF. In agreement with previous studies, use of subscales of the Hamilton Depression Rating scale (HAMD) improved signal detection compared to the HAMD total score. Using MMRM with HAMD subscales correctly identified drug as superior to placebo in up to 17/22 (77.3%) comparisons. Excluding double-blind, placebo lead-in responders did not increase the frequency of correctly identifying drug-versus-placebo differences. Conclusions: The 22 drug-versus-placebo comparisons in this report offer a small amount of evidence and therefore may not be convincing on their own, although results do agree with previous research. Researchers may be able to take advantage of these easy-to-implement methods while we wait for further improvements in other areas. Psychopharmacology Bulletin. 2007;40(2):101-114.
A new era of Alzheimer’s disease (AD) research is beginning now that multiple monoclonal antibodies (mABs) are on the market. Their use may not be widespread initially but will be more common in clinical sites likely to participate in clinical trials and will continue to grow. Many AD investigational treatments have been studied as add-on to acetylcholinesterase inhibitors; however, putative disease-modifying therapies (DMTs) like mABs are expected to alter the underlying rate of progression, potentially reducing our ability to detect effects of other DMTs on top of mABs. In addition, if the mechanism of action (MOA) also targets amyloid, the treatment effect may be further reduced by co-administration (antagonism). Alternatively, treatments that target complementary mechanisms may actually perform better (synergy). This new context requires careful consideration when powering clinical trials. We consider two clinical trial design scenarios. The first is a two-arm trial to test a treatment as an add-on to a mAB. The second is a four-arm combination trial to represent the scenario where the mAB is started at the same time as investigational treatment. We calculate the required sample sizes for studies in three different patient populations: secondary prevention (MSDR = 0.34, 2-year study), early AD (MSDR = 0.55, 18-month study), and mild-to-moderate AD (MSDR = 0.67, 1-year study). We consider scenarios with additivity, 30% antagonism, and 30% synergy. As shown in the table, the expected interaction with mAB treatment can have a large effect on study design decisions, with antagonistic treatment effects requiring more than three times the sample size of a synergistic treatment in most scenarios. The four-arm scenario also usually requires around a threefold increase in sample size compared to a simple add-on study. Studies evaluating investigational therapies as add-on to mABs are complex, and their cost to run will depend on how interactive the two treatments are. Treatments that are likely to work better with amyloid removed will be easier to study in this new context due to their complementary MOA. Symptomatic treatments may require fewer additional subjects than disease-modifying treatments given the same expected effect size in the absence of mABs.
INTRODUCTION:The importance of biomarkers as a primary outcome or as supportive evidence of clinical effect is rising as the field shifts toward disease-modifying treatments and earlier intervention, because they have lower variability and can indicate disease progression earlier than clinical outcomes. This study assessed the performance of plasma pTau 181 and 217 as a predictive biomarker and potential primary endpoint in early-phase Alzheimer's disease (AD) trials. METHODS:Summary data from recent monoclonal antibody (mAb) trials including plasma pTau 181 and 217 were analyzed to evaluate associations between plasma pTau 181/217 and clinical outcomes. The suitability of plasma pTau 181/217 as a surrogate endpoint for internal decision making was assessed using Prentice criteria. Simulations were conducted to explore the statistical power of using plasma pTau 181/217 as a primary outcome in dose-escalation, proof-of-concept (POC) trial designs. Additional criteria for biomarker validation were applied to simulated data. RESULTS:A strong group-level correlation (r = 0.781) was observed between treatment effects on plasma pTau 181/217 and Clinical Dementia Rating scale - Sum of Boxes (CDR-SB). Mean change in plasma pTau 181/217 significantly predicted mean change in CDR-SB (p = 0.013). The treatment effect on pTau 181/217 was ∼2.6 times greater than on CDR-SB. Prentice Criteria 1, 2, and 4 were met or reasonably met; Criterion 3 is not applicable in the POC setting. CONCLUSION:Plasma pTau 181/217 at 6 months shows future promise to reasonably likely predict clinical benefit for drugs that reduce pTau 181/217 levels, supporting its use as a primary endpoint in early-phase trials. With effect sizes similar to those seen with donanemab, adequately powered trials may require as few as 100 participants. Such trials should include prespecified analyses to evaluate individual-level Prentice criteria, and pTau 181/217 results can be used to predict potential Phase 3 clinical outcomes. Highlights:The group-level correlation between a biomarker treatment effect and clinical endpoint treatment effect is a measurement of the biomarker's ability to predict clinical outcome.The correlation of group level plasma pT217 or pT181 effect size at 6 months with clinical outcome Clinical Dementia Rating scale - Sum of Boxes (CDR-SB) effect size at 12 months was approximately 0.781 with p values of 0.013.Cohen's d effect size of plasma pTau as an outcome was 2.6 times greater than the Cohen's d of CDR-SB, leading to higher power or lower sample sizes.As a primary endpoint, plasma pTau meets or reasonably meets Prentice Criteria 1, 2, and 4, while Criterion 3 was deemed not applicable in the proof-of-concept study setting.
Alzheimer’s disease (AD) clinical trials led to the recent successes with monoclonal antibodies targeting amyloid, which opens up many new directions for research into treatment for AD in the future. Gaining greater understanding from these successes and failures will help researchers to focus their efforts on avenues that have the highest potential benefit. We performed a meta-analysis of over a hundred studies of 70+ AD treatments. A global statistical test (GST) combining ADAS-cog, ADCS-ADL and CDR-sb was used to assess the overall efficacy in these studies, in a fair way across multiple outcomes. Disease modifying treatment effects and symptomatic effects impacting all disease symptoms will perform better on this GST than treatments impacting only one symptom. In addition, a false positive effect is less likely to occur on all 3 outcomes simultaneously, making the GST a more reliable outcome for detecting true treatment effects. Composite scores, ADCOMS and iADRS, were used to target true disease progression in the lecanemab and donanemab phase 2 studies, respectively, and have similar advantages to a GST. A similar approach would have possibly averted much of the controversy surrounding the aducanumab approval. Some studies that were previously determined to be failures show some indication of positive effects. And conversely, some studies previously thought to be promising are shown to be more clearly failures. Meta-analyses combining similar treatments across programs illuminate overlooked mechanisms as well as more conclusive failures. The success and failure of Alzheimer’s treatments is partly hidden by including 3 different domains of clinical efficacy: cognition, function and global. True treatment efficacy and true lack of efficacy are much easier to detect with multiple, combined outcomes.
AbstractIn progressive diseases, like Alzheimer’s disease, treatments that slow progression should start early in the disease course to longer maintain higher levels of functioning. In corresponding clinical trials, the treatment effect is usually expressed in terms of mean differences on a clinical scale. Early in the disease course, however, treatment effects expressed on a clinical scale are often small but may nonetheless correspond to an important slowing of disease progression. This complicates the appreciation of the relevance of observed treatment effects. For example, it may be difficult to determine whether a 2-point improvement on a clinical scale is relevant for clinical practice. In this paper, we propose the meta Time-Component Tests (meta TCT). This new approach leads to estimators of treatment effects on the time scale, in terms of time saved or percentage slowing of progression, that are easy to interpret. This approach is based on estimates obtained from an arbitrary model for longitudinal data and is, therefore, very flexible. Asymptotic properties of the Meta TCT estimators are derived and evaluated in an extensive simulation study. Meta TCT is then applied to a phase 2/3 clinical trial for Alzheimer’s disease, which was first analyzed with a mixed model. In this trial, meta TCT leads to important additional insights into the treatment effect. We believe that meta TCT will facilitate the estimation of interpretable treatment effects in clinical trials for progressive diseases, and that this, in turn, will fine-tune the evaluation of the clinical relevance of new treatments.
We would like to make a correction to the paper titled "Using principal stratification in analysis of clinical trials" in Statistics in Medicine. In Section 11.2, Equations (14) and (15) for expressions T 3 $$ {T}_3 $$ and T 4 $$ {T}_4 $$ , respectively, have to be corrected as shown in the updated text below.
Background: The heavily right-skewed data seen in recently reported Alzheimer’s disease (AD) clinical trials influenced treatment contrasts when data were analyzed via the typical mixed-effects model for repeated measures (MMRM). Methods: An MMRM analysis similar to what is commonly used in AD clinical trials was compared versus robust regression (RR) and the non-parametric Hodges–Lehman estimator (HL). Results: Results in simulated data patterned after AD trials showed that imbalance across treatment arms in the number of patients in the extreme right tail (those with rapid disease progression) frequently occurred. Each analysis method controlled Type I error at or below the nominal level. The RR analysis yielded smaller standard errors and more power than MMRM and HL. In data sets with appreciable imbalance in the number of rapid progressing patients, MMRM results favored the treatment arm with fewer rapid progressors. Results from HL showed the same trend but to a lesser degree. Robust regression yielded similar results regardless of the ratio of rapid progressors. Conclusions: Although more research is needed over a wider range of scenarios, it should not be assumed that MMRM is the optimal approach for trials in early Alzheimer’s disease.
Introduction: The reliable assessment of treatment outcomes for disease-modifying therapies (DMT) in neurodegenerative disease is challenging. The objective of this paper is to describe a generalized framework for developing composite scales that can be applied in diverse, degenerative conditions, termed "GENCOMS." Composite scales optimize the sensitivity for detecting clinically meaningful effects that slow disease progression. Methods: The GENCOMS method relies on robust natural history data and/or placebo arm data from DMT trials. Validated scales that are core to the disease process have been identified, and item level data obtained to standardize the response outcomes from 0 (best possible score) to 1 (worst possible score). A partial least squares regression analysis was conducted with temporal change as the dependent variable and change scores in standardized items as the explanatory variables. The derived model coefficients constitute a weighted sum of items that most effectively measure disease progression. Results: The resultant composite scale was optimized to detect disease progression and can be examined in a range of slow or fast progressing populations. The scale can be used in studies with comparable patient populations as an endpoint optimized to measure disease progression and therefore ideally suited to assess treatment effects in DMTs. Conclusion: The methodology presented here provides a generalizable framework for developing composite scales in the assessment of neurodegenerative disease progression and evaluation of DMT effects. By objectively selecting and weighting items from previously validated measures based solely on their sensitivity to disease progression, this methodology allows for the creation of a more responsive measurement of clinical decline. This heightened sensitivity to clinical decline can be utilized to detect modest yet meaningful treatment effects in the early stages of neurogenerative diseases, when it is optimal to begin a DMT.
Gingipains, toxic protease virulence factors from the bacterial pathogen P. gingivalis ( Pg ), were discovered in postmortem brains of patients with Alzheimer’s disease. Gingipain levels correlated with tau and ubiquitin pathology, and oral infection of wild-type mice with Pg resulted in brain inflammation and neurodegeneration that was blocked by gingipain inhibitors. LHP588 is a second-generation orally bioavailable and brain-penetrant lysine gingipain inhibitor that reduces the toxicity of Pg and the bacterial load. A first-generation molecule (COR388/atuzaginstat) previously showed reduction of cognitive decline in prespecified cohorts defined by their infection load, but it was discontinued due to a hepatic safety signal. We will review new data from the LHP588 SAD/MAD, and the design of the Phase 2b study planned for initiation later this year. The Phase 1 study of LHP588 enrolled 32 individuals in the SAD component with 4 cohorts and concurrent placebo (25 mg, 50 mg, 100 mg, 200 mg) and 24 healthy subjects in the 10-day MAD portion, with 3 cohorts and concurrent placebo (50, 100 mg, and 200 mg). In the study of atuzaginstat, significance was not observed in the full intent-to-treat population, however, prespecified subgroup analyses indicated efficacy in patients with Pg -positive saliva ( Pg +), slowing cognitive decline compared with placebo on the ADAS-Cog11 by approximately 50% (p = 0.02). Changes in Pg DNA in saliva correlated significantly with changes on the ADAS-Cog, CDR-SB, and MMSE. The second-generation lysine gingipain inhibitor LHP588 was well-tolerated in the SAD and 10-day MAD study; adverse events in the active arms were mild and sporadic. PK with once-daily dosing achieved target concentrations sufficient for reduction of systemic Pg infection at doses >25 mg of LHP588, and exposures equivalent or greater than those achieved with the high dose of atuzaginstat. LHP588 was also detected in the CSF. LHP588 was well-tolerated in healthy volunteers without evidence of hepatic safety signals to date, and its PK profile was supportive of once daily dosing. The Phase 2 trial of LHP588 proposed to start in late 2023 will be similar in design to the prior atuzaginstat study but will be restricted to subjects with Pg + saliva.
Many clinical trials have open label extension (OLE) phases in which all patients from the randomized controlled study go onto active treatment for some period of time. These OLEs provide the opportunity for placebo patients to have access to active treatment. Trial sponsors also want to get meaningful information from the OLE data. We evaluate the types of hypotheses that can be tested with OLE data in degenerative diseases and propose analyses aligned with these hypotheses. Challenges with OLE data are also enumerated and discussed. For these analyses, and for simplicity of examples, we are assuming that the randomized phase and the OLE phase are each 1 year in duration. 6 key hypotheses can be addressed with OLE data. 1. Treatment effects differ during the randomized phase and the OLE, 2. Treatment groups differ at the end of the OLE, 3. The slope for placebo patients in the randomized phase differs from their slope during the OLE, 4. The slope over the entire duration of the randomized phase and the OLE phase differs between groups, 5. The slope of decline during the first 12 months of treatment combined across the original active patients and the placebo patients during the OLE phase differs from the placebo slope of decline during the original 12 month randomized phase. 6. The slope of the original placebo group differs from the original active group during the OLE phase. If a significant treatment effect is seen during the randomized phase, then hypotheses 2 and 6 can demonstrate disease modification in a pseudo staggered start analysis if they demonstrate that the treatment groups also differ at the end of the OLE phase, and are not converging during the OLE. Challenges with OLE analyses include high dropout rates, differential dropout between groups, ceiling and floor effects resulting in differing slopes (nonlinearity) during the randomized and OLE phases. OLE phases can add to the weight of evidence for efficacy in clinical trials and can identify disease modifying effects as long as issues such as dropout and non-linearity are accounted for.
Until the accelerated approvals of aducanumab and lecanemab, the only approved drugs for Alzheimer’s disease (AD) were symptomatic treatments, which showed relatively large improvements but did not alter disease progression. Multiple decades of designing and analyzing trials for symptomatic treatments have optimized the process for acetyl cholinesterase inhibitors and NMDA receptor antagonists, but this paradigm is not optimal for disease-modifying therapies. As the AD research community continues to transition towards treatments with the potential to alter disease progression, significant course corrections in the design and analysis of clinical trials are required to ensure that treatments with the largest potential for impact do not fail because they’re being evaluated out of context. Several design aspects from AD clinical trials are evaluated in the context of disease-modifying treatments, including the expected effect size, patient population, outcomes, multiplicity adjustments, trial duration, use of biomarkers, and timing of visits. We use publicly available trial results to demonstrate that both the clinical trial process and public interpretation of data will result in costly losses unless significant improvements are accepted and implemented In the context of donepezil’s expected 3-point change on ADAS-cog over 6 months, the results for lecanemab (1.5 points over 18 months) seem small, yet lecanemab improvements can be reasonably expected to match donepezil after three years of treatment and continue to increase. Symptomatic treatments were required to demonstrate significant improvement on multiple outcomes, but disease-modifying treatments should address the underlying disease process. Trials of aducanumab, donanemab, and lecanemab show consistent disease modification of 5 to 6 months less progression over 18 months when outcomes are converted to units of time and combined. This time-converted composite (which may include biomarkers) can show consistent results earlier in the disease course, allowing prodromal and early AD patients to be studied with greater success. The current message sent by recent approvals is that only sponsors with resources to field clinical trials with thousands of patients have a chance of gaining regulatory approval. By modernizing clinical trials, more sponsors will have a better chance of success, meaning better treatments available to more patients sooner.
Abstract Disease modifying therapies (DMTs) are hypothesized to be most beneficial in early disease when progression is slow and mean changes will be small. Therefore, even highly effective therapies will yield small absolute differences whose clinical relevance may be hard to interpret. Time component tests (TCTs) translate differences between treatments in mean change – the vertical distance between longitudinal trajectories, into an intuitively understood metric of time saved – the horizontal distance between trajectories. This corresponds to maintenance of independence with active treatment. DMTs are likely to impact multiple disease domains simultaneously and on the timescale these outcomes can be readily combined in a global time component test (gTCT). Use of gTCTs reflects a critical shift from emphasizing single outcomes and minimally clinically important effects to valuing true disease slowing, and incremental, but permanent benefits on an entire progressive disease. gTCTs are particularly helpful early in clinical development because combining across scales measuring multiple domains reduces noise and improves power. Clinical outcomes, such as ADAS-Cog, ADCS-ADL, and CDR-sb, reflect different aspects of disease progression and convergence of time savings results across these outcomes is evidence of an upstream effect on the cascade of events leading to neurodegeneration. By tailoring the statistical analysis to treatments with disease modifying effects, treatment effect estimates will be more precise thereby increasing statistical power when multiple endpoints are affected. Results will have less power with symptomatic treatments that primarily impact only one endpoint. The TCT was applied to a phase II clinical trial with a composite scale as the primary outcome. The AD04 2 mg group, showed some statistically significant effects compared with other study arms. It is unclear whether the observed 3.8-point difference on the composite measure is clinically meaningful; however, the TCT results show a time savings of 11 months in an 18 month study with AD04 2 mg. The relevance of 11 months saved is more universally understood than a mean difference of 3.8 points in the composite outcome. These results suggest that a combination of a composite approach and a gTCT (time savings) interpretation offers a powerful approach for detecting disease modifying effects.
Alzheimer’s disease progression causes worsening on multiple symptoms of disease which can be measured by cognitive, functional and global scales. These scales can be thought of as surrogate outcomes for the true disease progression. The degenerative and therefore progressive nature of the disease indicates that time is a measurable gold standard against which we can assess all other outcomes. Composite endpoints and global statistical tests have been used to approximate this progression outcome since it is not directly measurable. Using time as the gold standard, one can calculate the alignment of any measure against the gold standard using a signal to noise ratio. A combination of the ADAS-cog, ADCS-ADL and CDR-sb aligns more than 90% with the unmeasurable progression outcome. Converting each of the clinical scales to a time scale allows us to measure progression in “disease progression” time: i.e. months of progression per month of treatment, which should be linear in nature despite any non-linearity in the clinical scale. A recent meta-analysis was performed and data from relevant studies have been estimated using an overall progression time metric. Results for five studies of interest (three monoclonal antibodies and two other treatments) Aducanumab, Donanemab, Lecanemab, AD04, and Souvenaid show time savings range from 2 months per year to 7 months per year. Results for these five studies are presented in Table 1. Measuring time savings in a degenerative disease reflects the true goal of disease modifying treatments which is to slow down a progressive disease which will result in slowing of progression on all aspects or symptoms of disease. Progressive diseases would cease to be diseases at all if there was no progression of symptoms. Focusing on the clinical meaningfulness of each separate point on each outcome distracts us from this overall goal. Progression time savings is meaningful to patients, caregivers and all stakeholders and can be calculated from meaningful, clinically acceptable outcomes.
L’aducanumab est un AC monoclonal humain ciblant les formes agrégées d’Aβ (oligomères solubles et fibrilles insolubles). EMERGE et ENGAGE sont 2 études de phase III mondiales, randomisées, en double insu, contrôlées vs placebo, de 18 mois. Ces études, aux méthodologies identiques, ont évalué l’efficacité et la sécurité d’emploi de l’aducanumab chez des patients de 50 à 85 ans avec une maladie d’Alzheimer (MA) débutante (troubles cognitifs légers dus à la MA ou démence légère liée à la MA). Les critères d’inclusion comprenaient une TEP-amyloïde positive, un MMSE de 24 à 30, un score CDR Global de 0,5, et un score RBANS-DMI ≤ 85. Pendant la période contrôlée vs placebo de 18 mois, les patients ont été randomisés 1/1/1 pour recevoir l’aducanumab à dose faible, l’aducanumab à dose élevée ou un placebo, administré en IV toutes les 4 semaines. Le critère de jugement principal était la variation du score CDR-SB à la sem 78. Les critères secondaires étaient la variation des scores MMSE, ADAS-Cog13 et ADCS-ADL-MCI. Après l’analyse de futilité planifiée, une analyse réalisée suite au gel de base final a montré que EMERGE avait satisfait son critère de jugement principal, conformément au plan d’analyse statistique préspécifié. Les patients traités par aducanumab à dose élevée ont présenté une réduction significative du déclin clinique mesuré par les CDR-SB à la semaine 78 (22 % vs placebo, p = 0,01). ENGAGE n’a pas satisfait son critère de jugement principal. Cependant, les données des patients de l’étude ENGAGE qui avaient atteint une exposition suffisante à l’aducanumab à dose élevée ont corroboré les résultats d’EMERGE. Le profil de sécurité d’emploi et de tolérance de l’aducanumab dans les études EMERGE et ENGAGE concordait avec celui observé dans les études précédentes. EMERGE a satisfait à son critère de jugement principal, conformément au plan d’analyse statistique préspécifié. Les données d’un sous-groupe de patients d’ENGAGE corroborent les résultats d’EMERGE.
The National Research Council's report on the prevention and treatment of missing data highlighted the need to clearly specify causal estimands. This focus fundamentally changed how the missing data problem was perceived and addressed in clinical trials. The recent ICH E9(R1) addendum is another major step in promoting the use of the causal estimands framework that should further influence how clinical trial protocols and statistical analysis plans are written and implemented. The language of potential outcomes that is widely accepted in the causal inference literature is not widely recognized in the clinical trialists community and was not used in defining causal estimands in the NRC report or the ICH E9(R1). In this article, we attempt to bridge the gap between the causal inference community and clinical trialists to further advance the use of causal estimands in clinical trial settings. We illustrate how concepts from causal literature, such as potential outcomes and dynamic treatment regimens, can facilitate defining and implementing causal estimands and may provide a unifying language to describing the targets for both observational and randomized clinical trials.
AbstractBackgroundAducanumab is a human monoclonal antibody that selectively targets aggregated forms of Aβ, including soluble oligomers and insoluble fibrils. EMERGE and ENGAGE are two 18‐month, randomized, double‐blind, placebo‐controlled, global Phase 3 studies with identical design that evaluated the efficacy and safety of aducanumab in patients aged 50–85 years with early Alzheimer’s disease (MCI due to AD or mild AD dementia).MethodKey inclusion criteria included positive amyloid PET, MMSE score of 24–30, CDR Global score of 0.5, and an RBANS‐DMI score ≤85. During the 18‐month placebo‐controlled period, patients were randomized 1:1:1 to low‐dose aducanumab, high‐dose aducanumab, or placebo, administered via IV infusion every 4 weeks. The primary endpoint for EMERGE and ENGAGE was change from baseline at Week 78 on the CDR‐SB. Secondary endpoints included change from baseline on MMSE, ADAS‐Cog13, and ADCS‐ADL‐MCI.ResultFollowing pre‐planned futility analysis, analysis of the data from the final database lock showed that EMERGE met its primary endpoint, based on the pre‐specified statistical analysis plan. Patients treated with high dose aducanumab showed a significant reduction of clinical decline from baseline in CDR‐SB scores at 78 weeks (22% versus placebo, P = 0.01). ENGAGE did not meet its primary endpoint. However, data from patients in ENGAGE who achieved sufficient exposure to high dose aducanumab supported the findings of EMERGE.ConclusionEMERGE met its primary endpoint, based on the pre‐specified statistical analysis plan. Data from a subset of patients in ENGAGE support the results of EMERGE. The safety and tolerability profile of aducanumab in EMERGE and ENGAGE was consistent with previous studies of aducanumab.