In this paper we review recent advances in statistical methods for the evaluation of the heterogeneity of treatment effects (HTE), including subgroup identification and estimation of individualized treatment regimens, from randomized clinical trials and observational studies. We identify several types of approaches using the features introduced in Lipkovich, Dmitrienko and D'Agostino (2017) that distinguish the recommended principled methods from basic methods for HTE evaluation that typically rely on rules of thumb and general guidelines (the methods are often referred to as common practices). We discuss the advantages and disadvantages of various principled methods as well as common measures for evaluating their performance. We use simulated data and a case study based on a historical clinical trial to illustrate several new approaches to HTE evaluation.
There has been much interest in the evaluation of heterogeneous treatment effects (HTE) and multiple statistical methods have emerged under the heading of personalized/precision medicine combining ideas from hypothesis testing, causal inference, and machine learning over the past 10-15 years. We discuss new ideas and approaches for evaluating HTE in randomized clinical trials and observational studies using the features introduced earlier by Lipkovich, Dmitrienko, and D'Agostino that distinguish principled methods from simplistic approaches to data-driven subgroup identification and estimating individual treatment effects and use a case study to illustrate these approaches. We identified and provided a high-level overview of several classes of modern statistical approaches for personalized/precision medicine, elucidated the underlying principles and challenges, and compared findings for a case study across different methods. Different approaches to evaluating HTEs may produce (and actually produced) highly disparate results when applied to a specific data set. Evaluating HTE with machine learning methods presents special challenges since most of machine learning algorithms are optimized for prediction rather than for estimating causal effects. An additional challenge is in that the output of machine learning methods is typically a "black box" that needs to be transformed into interpretable personalized solutions in order to gain acceptance and usability.
Data-driven subgroup analysis plays an important role in clinical trials. This paper focuses on practical considerations in post-hoc subgroup investigations in the context of confirmatory clinical trials. The analysis is aimed at assessing the heterogeneity of treatment effects across the trial population and identifying patient subgroups with enhanced treatment benefit. The subgroups are defined using baseline patient characteristics, including demographic and clinical factors. Much progress has been made in the development of reliable statistical methods for subgroup investigation, including methods based on global models and recursive partitioning. The paper provides a review of principled approaches to data-driven subgroup identification and illustrates subgroup analysis strategies using a family of recursive partitioning methods known as the SIDES (subgroup identification based on differential effect search) methods. These methods are applied to a Phase III trial in patients with metastatic colorectal cancer. The paper discusses key considerations in subgroup exploration, including the role of covariate adjustment, subgroup analysis at early decision points and interpretation of subgroup search results in trials with a positive overall effect.
In this paper, we consider randomized controlled clinical trials comparing two treatments in efficacy assessment using a time to event outcome. We assume a relatively small number of candidate biomarkers available in the beginning of the trial, which may help define an efficacy subgroup which shows differential treatment effect. The efficacy subgroup is to be defined by one or two biomarkers and cut-offs that are unknown to the investigator and must be learned from the data. We propose a two-stage adaptive design with a pre-planned interim analysis and a final analysis. At the interim, several subgroup-finding algorithms are evaluated to search for a subgroup with enhanced survival for treated versus placebo. Conditional powers computed based on the subgroup and the overall population are used to make decision at the interim to terminate the study for futility, continue the study as planned, or conduct sample size recalculation for the subgroup or the overall population. At the final analysis, combination tests together with closed testing procedures are used to determine efficacy in the subgroup or the overall population. We conducted simulation studies to compare our proposed procedures with several subgroup-identification methods in terms of a novel utility function and several other measures. This research demonstrated the benefit of incorporating data-driven subgroup selection into adaptive clinical trial designs.
Inclusion body myositis (IBM) is a progressive skeletal muscle autoimmune disease with no established treatments. The development of therapies for IBM, as in other diseases, critically depends on the evaluation of efficacy, assessed through clinical outcome assessments (COAs)(1) that measure a patient's symptoms or function, and biomarkers, objective indicators of biological processes.(2) For drug approval, the Food and Drug Administration (FDA) and other regulatory agencies generally require well-controlled clinical trials that use COA endpoints demonstrating improvements in how patients feel, function, or survive.(3) However, biomarker endpoints can aid drug development, for example, by providing proof of mechanism or concept or as potential surrogate endpoints (markers known or reasonably likely to predict clinical benefit) and intermediate clinical endpoints (markers that are reasonably likely to predict irreversible morbidity), which can be the basis for accelerated approvals.(3)
OBJECTIVES:This study aimed to characterize corrected QT (QTc) prolongation in a cohort of hospitalized patients with coronavirus disease-2019 (COVID-19) who were treated with hydroxychloroquine and azithromycin (HCQ/AZM). BACKGROUND:HCQ/AZM is being widely used to treat COVID-19 despite the known risk of QT interval prolongation and the unknown risk of arrhythmogenesis in this population. METHODS:A retrospective cohort of COVID-19 hospitalized patients treated with HCQ/AZM was reviewed. The QTc interval was calculated before drug administration and for the first 5 days following initiation. The primary endpoint was the magnitude of QTc prolongation, and factors associated with QTc prolongation. Secondary endpoints were incidences of sustained ventricular tachycardia or ventricular fibrillation and all-cause mortality. RESULTS:Among 415 patients who received concomitant HCQ/AZM, the mean QTc increased from 443 ± 25 ms to a maximum of 473 ± 40 ms (87 [21%] patients had a QTc ≥500 ms). Factors associated with QTc prolongation ≥500 ms were age (p < 0.001), body mass index <30 kg/m2 (p = 0.005), heart failure (p < 0.001), elevated creatinine (p = 0.005), and peak troponin (p < 0.001). The change in QTc was not associated with death over the short period of the study in a population in which mortality was already high (hazard ratio: 0.998; p = 0.607). No primary high-grade ventricular arrhythmias were observed. CONCLUSIONS:An increase in QTc was seen in hospitalized patients with COVID-19 treated with HCQ/AZM. Several clinical factors were associated with greater QTc prolongation. Changes in QTc were not associated with increased risk of death.
The chapter discusses practical considerations arising in subgroup exploration exercises in late-stage clinical trials. Subgroup identification strategies are commonly applied to characterize the efficacy profile of an experimental treatment based on the results of a failed trial with a non-significant outcome in the overall patient population. Considering this setting, we present a comprehensive overview of relevant considerations related to the selection of clinically candidate biomarkers, choice of statistical models, including the role of covariate adjustment in subgroup investigation, and selection of subgroup search parameters. The subgroup identification methods considered in the chapter rely on the SIDES family of subgroup search algorithms. We discuss applications of this methodology to failed clinical trials and its key features such as biomarker screening, complexity control and Type I error rate control. The statistical methods and considerations discussed in the chapter will be illustrated using a Phase III clinical trial for the treatment of benign prostate hypertrophy.
An important step in the development of targeted therapies is the identification and confirmation of sub-populations where the treatment has a positive treatment effect compared to a control. These sub-populations are often based on continuous biomarkers, measured at baseline. For example, patients can be classified into biomarker low and biomarker high subgroups, which are defined via a threshold on the continuous biomarker. However, if insufficient information on the biomarker is available, the a priori choice of the threshold can be challenging and it has been proposed to consider several thresholds and to apply appropriate multiple testing procedures to test for a treatment effect in the corresponding subgroups controlling the family-wise type 1 error rate. In this manuscript we propose a framework to select optimal thresholds and corresponding optimized multiple testing procedures that maximize the expected power to identify at least one subgroup with a positive treatment effect. Optimization is performed over a prior on a family of models, modelling the relation of the biomarker with the expected outcome under treatment and under control. We find that for the considered scenarios 3 to 4 thresholds give the optimal power. If there is a prior belief on a small subgroup where the treatment has a positive effect, additional optimization of the spacing of thresholds may result in a large benefit. The procedure is illustrated with a clinical trial example in depression.
The Simes test is designed to test a single null hypothesis formed by taking the intersection of multiple null hypotheses using their p-values. In this short note, we explore a novel application of the Simes test for testing a single null hypothesis in a group sequential setting based on the p-values from sequential looks. We refer to this test as the group sequential Simes test (GSST). It turns out, however, that GSST suffers from some drawbacks. The main drawback is that GSST is uniformly less powerful than the reference group sequential test on which it is based. The reason is that the rejection decision of GSST can be based on a test statistic from an earlier stage, which is not a sufficient statistic. This is also a practical drawback. These drawbacks are discussed in this short note.
BACKGROUNDThe analysis of subgroups in clinical trials is essential to assess differences in treatment effects for distinct patient clusters, that is, to detect patients with greater treatment benefit or patients where the treatment seems to be ineffective.METHODSThe software application subscreen (R package) has been developed to analyze the population of clinical trials in minute detail. The aim was to efficiently calculate point estimates (eg, hazard ratios) for multiple subgroups to identify groups that potentially differ from the overall trial result. The approach intentionally avoids inferential statistics such as P values or confidence intervals but intends to encourage discussions enriched with external evidence (eg, from other studies) about the exploratory results, which can be accompanied by further statistical methods in subsequent analyses. The subscreen application was applied to 2 clinical study data sets and used in a simulation study to demonstrate its usefulness.RESULTSThe visualization of numerous combined subgroups illustrates the homogeneity or heterogeneity of potentially all subgroup estimates with the overall result. With this, the application leads to more targeted planning of future trials.CONCLUSIONThis described approach supports the current trend and requirements for the investigation of subgroup effects as discussed in the EMA draft guidance for subgroup analyses in confirmatory clinical trials (EMA 2014). The lack of a convenient tool to answer spontaneous questions from different perspectives can hinder an efficient discussion, especially in joint interdisciplinary study teams. With the new application, an easily executed but powerful tool is provided to fill this gap.
In this paper we enhance existing SIDES and SIDEScreen methods for biomarker discovery (Lipkovich et al., Stat. Med. 30:2601–2621, 2011; Lipkovich and Dmitrienko, J. Biopharm. Statist. 24:130–153, 2014; Lipkovich et al. Stat. Biopharm. Res. 9:368–378, 2017b) and apply it to a small Phase 2 clinical trial in patients with recurrent dysmenorrhea. We argue that incorporating stochastic elements in computing the variable importance, expected treatment effect and replicability index is particularly useful when dealing with relatively small data sets, so as to properly account for the uncertainty of the subgroup selection process. To demonstrate improved operating characteristics of the Stochastic SIDEScreen compared with the corresponding deterministic procedure, we conducted a small simulation study that mimics data from our Phase 2 trial. As analytical formulas for power calculations are not available for machine learning methods of biomarker/subgroup discovery, simulations utilizing existing early phase data should be conducted routinely for obtaining realistic estimates of power.
The general topic of subgroup identification has attracted much attention in the clinical trial literature due to its important role in the development of tailored therapies and personalized medicine. Subgroup search methods are commonly used in late-phase clinical trials to identify subsets of the trial population with certain desirable characteristics. Post-hoc or exploratory subgroup exploration has been criticized for being extremely unreliable. Principled approaches to exploratory subgroup analysis based on recent advances in machine learning and data mining have been developed to address this criticism. These approaches emphasize fundamental statistical principles, including the importance of performing multiplicity adjustments to account for selection bias inherent in subgroup search.This article provides a detailed review of multiplicity issues arising in exploratory subgroup analysis. Multiplicity corrections in the context of principled subgroup search will be illustrated using the family of SIDES (subgroup identification based on differential effect search) methods. A case study based on a Phase III oncology trial will be presented to discuss the details of subgroup search algorithms with resampling-based multiplicity adjustment procedures.
It is increasingly common to encounter complex multiplicity problems with several multiplicity components in confirmatory Phase III clinical trials. These components are often based on several endpoints (primary and secondary endpoints) and several dose-control comparisons. When constructing a multiplicity adjustment in these settings, it is important to control the Type I error rate over all multiplicity components. An important class of multiple testing procedures, known as gatekeeping procedures, was derived using the mixture method that enables clinical trial sponsors to set up efficient multiplicity adjustments that account for clinically relevant logical relationships among the hypotheses of interest. An enhanced version of this mixture method is introduced in this paper to construct more powerful gatekeeping procedures for a specific type of logical relationships that rely on transitive serial restrictions. Restrictions of this kind are very common in Phase III clinical trials and the proposed method is applicable to a broad class of multiplicity problems. Several examples are provided to illustrate the new method and results of simulation trials are presented to compare the performance of gatekeeping procedures derived using this method and other available methods.
Lumateperone (ITI-007) is a first-in-class investigational agent in development for the treatment of schizophrenia. Acting synergistically through serotonergic, dopaminergic and glutamatergic systems, lumateperone represents a new approach to the treatment of schizophrenia and other neuropsychiatric disorders. Lumateperone is a potent antagonist at 5-HT2A receptors and exhibits serotonin reuptake inhibition. Lumateperone also binds to dopamine D1 and D2 receptors acting as a mesolimbic/mesocortical dopamine phosphoprotein modulator (DPPM) with pre-synaptic partial agonism and post-synaptic antagonism at D2 receptors and as an indirect glutamatergic (GluN2B) phosphoprotein modulator with D1-dependent enhancement of both NMDA and AMPA currents via the mTOR protein pathway. Lumateperone was evaluated in 3 controlled clinical trials to evaluate efficacy in patients with acute schizophrenia. In Study ITI-007-005, 335 patients were randomized equally across 4 treatment arms: one of two doses of lumateperone, risperidone (active control) or placebo QAM for 4 weeks. In Study ITI-007-301, 450 patients were randomized equally across 3 treatment arms: one of two doses of lumateperone or placebo QAM for 4 weeks. In Study ITI-007-302, 696 patients were randomized equally across 4 treatment arms: one of two doses of lumateperone, risperidone (active control) or placebo QAM for 6 weeks. In all 3 studies, the primary endpoint was change from baseline on the Positive and Negative Syndrome Scale (PANSS) total score compared to placebo. Also, a 6-week open-label safety switching study was conducted. In this ITI-007-303 study 302 patients with stable schizophrenia were switched from standard-of-care (SOC) antipsychotics and treated for 6 weeks with lumateperone QPM outpatient and then switched back to SOC for 2 weeks. Two of the 3 randomized studies were positive. In Studies ITI-007-005 and ITI-007-301, lumateperone (60 mg ITI-007, equivalent to 42 mg active base) met the primary endpoint with statistically significant superior efficacy over placebo at Day 28 as measured by the PANSS total score (Study ITI-007-005 p=0.017; Study ITI-007-301 p=0.022). In Study ITI-007-302, neither dose of lumateperone separated from placebo on the primary endpoint in the intent-to-treat population; a high placebo response was observed in this study. Across all 3 efficacy trials, lumateperone improved symptoms of schizophrenia with the same trajectory and same magnitude of improvement from baseline on the PANSS total score. Lumateperone was well-tolerated with a favorable safety profile in all studies. In the two studies with risperidone included as an active control, lumateperone was statistically significantly better than risperidone on key safety and tolerability measures including prolactin, glucose, lipids and weight. In the open-label safety switching study statistically significant improvements from SOC were observed in body weight, cardiometabolic and endocrine parameters worsened again when switched back to SOC medication. In this open-label study, symptoms of schizophrenia generally remained stable or improved. Greater improvements were observed in subgroups of patients with elevated symptomatology such as those with comorbid symptoms of depression and those with prominent negative symptoms. Lumateperone represents a novel approach to the treatment of schizophrenia with a favorable safety profile in clinical trials. The lack of metabolic, motor and cardiovascular safety issues presents a safety profile differentiated from standard-of-care antipsychotic therapy.
BACKGROUND:The quality of data from clinical trials has received a great deal of attention in recent years. Of central importance is the need to protect the well-being of study participants and maintain the integrity of final analysis results. However, traditional approaches to assess data quality have come under increased scrutiny as providing little benefit for the substantial cost. Numerous regulatory guidance documents and industry position papers have described risk-based approaches to identify quality and safety issues. In particular, the position paper of TransCelerate BioPharma recommends defining risk thresholds to assess safety and quality risks based on past clinical experience. This exercise can be extremely time-consuming, and the resulting thresholds may only be relevant to a particular therapeutic area, patient or clinical site population. In addition, predefined thresholds cannot account for safety or quality issues where the underlying rate of observing a particular problem may change over the course of a clinical trial, and often do not consider varying patient exposure.METHODS:In this manuscript, we appropriate rules commonly utilized for funnel plots to define a traffic-light system for risk indicators based on statistical criteria that consider the duration of patient follow-up. Further, we describe how these methods can be adapted to assess changing risk over time. Finally, we illustrate numerous graphical approaches to summarize and communicate risk, and discuss hybrid clinical-statistical approaches to allow for the assessment of risk at sites with low patient enrollment.RESULTS:We illustrate the aforementioned methodologies for a clinical trial in patients with schizophrenia.CONCLUSIONS:Funnel plots are a flexible graphical technique that can form the basis for a risk-based strategy to assess data integrity, while considering site sample size, patient exposure, and changing risk across time.
It is very common for clinical trials to report multiple conclusions. The conclusions may be based on several clinical endpoints, several dosing arms, several subsets of the trial’s populations or several decision points (interim and final looks). Each of these components or sources induces multiplicity that must be accounted in the analysis to control the overall Type I error rate. As trial sponsors focus on more efficient development programs for new treatments, it is increasingly more common to employ strategies that combine several components, which gives rise to “multivariate” multiplicity problems. Numerous examples of Phase III trials with multiplicity problems of considerable complexity can be found in the literature. Problemswith twomultiplicity components typically arise in trials where the effect of an experimental treatment is evaluated at several dose levels using multiple endpoints, see for example, the lurasidone trials for the treatment of schizophrenia (Meltzer et al., 2011; Nasrallah et al., 2013). Three sources of multiplicity may be found in trials where a “bivariate” setting described above is augmented by features such as a multi-population design. An illustration of a “trivariate”multiplicity problem is provided by the Phase III trial of pomeglumetad mentionil with efficacy assessments conducted in two pre-defined patient populations, at two dose levels of a novel therapy, and formultiple endpoints (Downing et al., 2014). Another example of a “trivariate” setting is the multiplicity problem in the COMPASS trial (Bosch et al., 2017) where the sponsor needed to account for the following three components: several efficacy endpoints, two regimens of an experimental treatment tested versus control and three pre-planned decision points. In addition to a confirmatory setting with a fixed number of pre-defined null hypotheses, multiplicity-related considerations play an important role in exploratory settings. For example, in data-driven subgroup, identification problems dozens and even hundreds of null hypotheses may be formulated in a post-hoc fashion within patient subgroups discovered in a clinical trial. To underscore the importance of proactively addressing multiplicity issues in confirmatory Phase III trials, the U.S. Food and Drug Administration (FDA) and European Medicines Agency (EMA) have recently published draft guidance documents on this broad topic (EMA, 2017; FDA, 2017). These guidance documents provide a high-level review of key decision-making frameworks in trials with multiple clinical objectives, define expectations for confirmatory clinical trials and discuss appropriate methods for handling multiplicity in settings frequently encountered in clinical trials. This special issue brings together multiplicity experts and researchers with a regulatory, industry and academic background to give an overview of recent developments in the general field of multiplicity. The twelve papers included in this issue discuss various aspects of this multifaceted topic, including general scientific, statistical and practical considerations. The special issue features contributions from statisticians working in the USA, Europe and Japan. The first paper in this issue, by Norbert and Brandt, focuses on regulatory/scientific considerations in the assessment of confirmatory evidence in trials with several clinical objectives. The authors touch upon key topics such as the role of analytical procedures, interpretation of simultaneous confidence intervals, multiplicity in safety analysis, and so on. This article is followed by four papers that give a review of statistical methods commonly used in popular multiplicity settings. The article by Tamhane is an excellent survey of multiple testing procedures used in problems with a single family of null hypotheses of no effect or in more challenging problems with several families of hypotheses. The former class includes basic problems with a single source of multiplicity and the latter defines an advanced setting with multiple sources. This review covers both traditional and innovative approaches to multiplicity adjustment. Hamasaki, Evans and Asakura explore multiplicity issues in clinical trials with group-sequential designs and more general adaptive designs that support an option to increase the sample size at an interim analysis. Wiens focuses on an important class of
Lumateperone (ITI-007) is a first-in-class investigational agent in development for the treatment of schizophrenia. Acting synergistically through serotonergic, dopaminergic and glutamatergic systems, lumateperone represents a new approach to the treatment of schizophrenia and other neuropsychiatric disorders. Lumateperone is a potent antagonist at 5-HT2A receptors and exhibits serotonin reuptake inhibition. Lumateperone also binds to dopamine D1 and D2 receptors acting as a mesolimbic/mesocortical dopamine phosphoprotein modulator (DPPM) with pre-synaptic partial agonism and post-synaptic antagonism at D2 receptors and as an indirect glutamatergic (GluN2B) phosphoprotein modulator with D1-dependent enhancement of both NMDA and AMPA currents via the mTOR protein pathway. Lumateperone demonstrated antipsychotic efficacy in two well-controlled clinical trials and was found to be well tolerated with a safety profile similar to placebo in all trials conducted to date. In an open-label safety study, 302 patients with schizophrenia were switched from standard-of-care (SOC) antipsychotic therapy to 6 weeks treatment with lumateperone (ITI-007 60 mg, equivalent to 42 mg active base) QPM with no dose titration, then switched back to SOC for 2 weeks. The primary objective was to determine the safety of lumateperone, assessed by adverse events, body weight, 12-lead electrocardiograms, vital signs, clinical laboratory tests, motor assessments, and the Columbia-Suicide Severity Rating Scale. The secondary objectives were to determine the effectiveness of lumateperone to improve psychopathology as measured by the PANSS, social functioning as measured by the PANSS Pro-Social Factor and the Personal and Social Performance Scale (PSP), and depression as measured by the Calgary Depression Scale for Schizophrenia. Lumateperone was generally well-tolerated with a favorable safety profile. There was no drug related serious adverse event. In comparison to treatment with SOC antipsychotics at baseline, mean body weight decreased with lumateperone treatment. Lumateperone also demonstrated a favorable cardiometabolic and endocrine safety profile. Mean levels of cholesterol, triglycerides and prolactin improved with lumateperone treatment and worsened again when patients returned to SOC. The cardiovascular safety of lumateperone was favorable including no QTc interval prolongation. While efficacy data in an open-label study should be interpreted cautiously due to the absence of a parallel control group, improvements were observed in change from baseline of the PANSS total scores. Improvements were also seen in the Positive symptom subscale score, General Psychopathology subscale score, Marder Negative Factor score, and Prosocial Factor score as well as in the PSP scale. Greater improvements were observed in subgroups of patients with elevated symptomatology such as those with comorbid symptoms of depression and those with prominent negative symptoms. Lumateperone represents a novel approach to the treatment of schizophrenia with a favorable safety profile. The lack of metabolic, motor and cardiovascular safety issues presents a safety profile differentiated from standard-of-care antipsychotic therapy. Patients with stable symptoms on other antipsychotics may further improve when switched to lumateperone, with no dose titration needed. These data may warrant further investigation in placebo-controlled trials in patients with prominent negative symptoms and, separately, in patients with comorbid depression to demonstrate efficacy in these populations.
Where there are a limited number of patients, such as in a rare disease, clinical trials in these small populations present several challenges, including statistical issues. This led to an EU FP7 call for proposals in 2013. One of the three projects funded was the Innovative Methodology for Small Populations Research (InSPiRe) project. This paper summarizes the main results of the project, which was completed in 2017.The InSPiRe project has led to development of novel statistical methodology for clinical trials in small populations in four areas. We have explored new decision-making methods for small population clinical trials using a Bayesian decision-theoretic framework to compare costs with potential benefits, developed approaches for targeted treatment trials, enabling simultaneous identification of subgroups and confirmation of treatment effect for these patients, worked on early phase clinical trial design and on extrapolation from adult to pediatric studies, developing methods to enable use of pharmacokinetics and pharmacodynamics data, and also developed improved robust meta-analysis methods for a small number of trials to support the planning, analysis and interpretation of a trial as well as enabling extrapolation between patient groups. In addition to scientific publications, we have contributed to regulatory guidance and produced free software in order to facilitate implementation of the novel methods.