Risk differences allow decision makers to easily estimate the excess safety risk associated with a medical product relative to the potential benefits. However, in post-market observational surveillance studies that actively monitor (e.g., sequentially over time) for safety risk of new medical products, available methods target a relative measure (e.g., odds ratio and relative risk), which can be especially unstable in the rare event setting. These studies are typically conducted within distributed healthcare networks (e.g., Food and Drug Administration [FDA] Sentinel and Centers for Disease Control [CDC] Vaccine Safety Datalink) with patient-level data protected behind firewalls, but sharing of aggregate, deidentified data for centralized analyses. We propose an inverse probability of treatment weighting (IPTW) method that uses site-specific propensity scores to estimate site-specific risk differences that are combined to create an overall stratified risk difference estimate. This method is tailored to the rare event setting and requires minimal data sharing. The stratified IPTW approach is then extended to the active post-market surveillance setting by incorporating group sequential monitoring boundaries using a novel permutation approach. A simulation study is conducted to evaluate the performance of the new methods relative to two centralized analysis approaches, and the methods are applied to a safety surveillance study comparing the risk of febrile seizure between two vaccines using FDA Sentinel Data from three healthcare organizations.
Borrowing controls from external sources has become popular for augmenting the control arm in small randomized controlled trials (RCTs). Due to the difference between the external and RCT populations, bias can be introduced that may lead to invalid statistical inference based on combined data. To mitigate this risk, dynamic borrowing which adaptively determines the amount of borrowing, can be used together with pre-adjustment for prognostic factors in the external data. To take into account the variability due to the estimation of the amount of borrowing and the pre-adjustment, we propose a Bayesian bootstrap (BB)-based integrated Bayesian approach together with covariate balancing (CB) for pre-adjustment. We show that the proposed BB based approach is a valid approximate Bayesian approach with CB using different distances, particularly Euclidean or entropy distance. This justification is not trivial because CB has a different nature from the probability-based approach. We also propose a BB-algorithm for generating an approximate posterior sample, which is easy to implement and computationally efficient. Statistical inference for estimand of interest using combined external and internal data can be based on the bootstrapped posterior sample or on an approximate normal distribution with parameters estimated by BB. To examine the properties of the proposed approach, we conduct an extensive simulation study. The approach is illustrated by borrowing controls for an acute myeloid leukemia trial from another study.
In a randomized controlled trial with time-to-event endpoint, some commonly used statistical tests to test for various aspects of survival differences, such as survival probability at a fixed time point, survival function up to a specific time point, and restricted mean survival time, may not be directly applicable when external data are leveraged to augment an arm (or both arms) of an RCT. In this paper, we propose a propensity score-integrated approach to extend such tests when external data are leveraged. Simulation studies are conducted to evaluate the operating characteristics of three propensity score-integrated statistical tests, and an illustrative example is given to demonstrate how these proposed procedures can be implemented.
This paper focuses on the use of novel technologies and innovative trial designs to accelerate evidence generation and increase pharmaceutical Research and Development (R&D) productivity, at Bristol Myers Squibb. We summarize learnings with case examples, on how we prepared and continuously evolved to address the increasing cost, complexities, and external pressures in drug development, to bring innovative medicines to patients much faster. These learnings were based on review of internal efforts toward accelerating R&D focusing on four key areas: adopting innovative trial designs, optimizing trial designs, leveraging external control data, and implementing novel methods using artificial intelligence and machine learning.
Use of historical control data to augment a small internal control arm in a randomized control trial (RCT) can lead to significant improvement of the efficiency of the trial. It introduces the risk of potential bias, since the historical control population is often rather different from the RCT. Power prior approaches have been introduced to discount the historical data to mitigate the impact of the population difference. However, even with a Bayesian dynamic borrowing which can discount the historical data based on the outcome similarity of the two populations, a considerable population difference may still lead to a moderate bias. Hence, a robust adjustment for the population difference using approaches such as the inverse probability weighting or matching, can make the borrowing more efficient and robust. In this paper, we propose a novel approach integrating propensity score for the covariate adjustment and Bayesian dynamic borrowing using power prior. The proposed approach uses Bayesian bootstrap in combination with the empirical Bayes method utilizing quasi-likelihood for determining the power prior. The performance of our approach is examined by a simulation study. We apply the approach to two Acute Myeloid Leukemia (AML) studies for illustration.
The marketing authorization of a medicinal product is contingent upon demonstration of safety and efficacy in support of the product's labeled conditions of use. To demonstrate safety, one group of adverse events that requires detailed consideration is common adverse reactions (ARs). ARs and their frequency are reported in prescription drug labeling in the US and EU. The determination of these adverse reactions generally takes a simple approach-usually, inclusion is through the frequency of reporting and whether the adverse event (AE) rate for the drug exceeds the placebo rate. This standard method does not account for confounders or multiplicity. To overcome these limitations, we propose a Monte Carlo approach to detect drug safety signals in clinical trials. We fit regression models incorporating covariates to assess the drug effect on the rate of AEs. Adjustment for multiplicity is carried out through the construction of simultaneous confidence intervals accounting for arbitrary correlations. A computationally efficient multiplier bootstrap approach using the Rademacher sequences is developed to generate random samples from the joint distribution of the estimators for all the AE rates. Compared to Bonferroni-based methods, the proposed method leads to narrower simultaneous confidence intervals and is more powerful in detecting potential safety signals.
BackgroundMultiple criteria decision analysis (MCDA) and stochastic multi-criteria acceptability analysis (SMAA) in their current implementation cannot incorporate prior or external information on benefits and risks. We demonstrate how to incorporate prior data using a Bayesian mixture model approach while conducting quantitative benefit-risk assessments (qBRA) for medical products.MethodsWe implemented MCDA and SMAA in a Bayesian framework. To incorporate information from a prior study, we use mixture priors on each benefit and risk attribute that mixes information from a previous study with a vague prior distribution. The degree of borrowing is varied using a mixing proportion parameter.ResultsA demonstration case study for qBRA using the supplementary New Drug Application (sNDA) filing for Rivaroxaban for the indication of reduction in the risk of major thrombotic vascular events in patients with peripheral artery disease (PAD) was used to illustrate the method. Net utility scores, obtained from the randomized controlled trial data to support the sNDA, from the MCDA for Rivaraxoban and comparator were 0.48 and 0.56, respectively, with Rivaroxaban being the preferred alternative only 33% of the time. We show that with only 30% borrowing from a previous RCT, the MCDA and SMAA results are favorable for Rivaroxaban, accounting for the seemingly aberrant results on all-cause death in the trial data used to support the sNDA.ConclusionOur method to formally incorporate prior data in MCDA and SMAA is easy to use and interpret. Software in the form of an RShiny App is available here: https://sai-dharmarajan.shinyapps.io/BayesianMCDA_SMAA/.
For dynamic borrowing to leverage external data to augment the control arm of small RCTs, the key step is determining the amount of borrowing based on the similarity of the outcomes in the controls from the trial and the external data sources. A simple approach for this task uses the empirical Bayesian approach, which maximizes the marginal likelihood (maxML) of the amount of borrowing, while a likelihood-independent alternative minimizes the mean squared error (minMSE). We consider two minMSE approaches that differ from each other in the way of estimating the parameters in the minMSE rule. The classical one adjusts for bias due to sample variance, which in some situations is equivalent to the maxML rule. We propose a simplified alternative without the variance adjustment, which has asymptotic properties partially similar to the maxML rule, leading to no borrowing if means of control outcomes from the two data sources are different and may have less bias than that of the maxML rule. In contrast, the maxML rule may lead to full borrowing even when two datasets are moderately different, which may not be a desirable property. For inference, we propose a Bayesian bootstrap (BB) based approach taking the uncertainty of the estimated amount of borrowing and that of pre-adjustment into account. The approach can also be used with a pre-adjustment on the external controls for population difference between the two data sources using, e.g., inverse probability weighting. The proposed approach is computationally efficient and is implemented via a simple algorithm. We conducted a simulation study to examine properties of the proposed approach, including the coverage of 95 CI based on the Bayesian bootstrapped posterior samples, or asymptotic normality. The approach is illustrated by an example of borrowing controls for an AML trial from another study.
This proof-of-concept study retrospectively assessed the feasibility of applying a hybrid control arm design to a completed phase III randomized controlled trial (RCT; CheckMate-057) in advanced non-small cell lung cancer using a real-world data (RWD) source. The emulated trial consists of an experimental arm (patients from the RCT experimental cohort) and a hybrid control arm (patients from the RCT and RWD control cohorts). For the RWD control cohort, this study used a nationwide electronic health record-derived de-identified database. Three frequentist statistical borrowing methods were evaluated: a two-step Cox model, a fixed Cox model, and propensity score-integrated composite likelihood ( "Methods 1-3 "). The experimental treatment effect for hybrid control designs were evaluated using hazard ratios (HRs) with 95% confidence interval (CI) estimated from the Cox models accounting for covariate differences. The reduction in study duration compared to the RCT was also evaluated. All three statistical borrowing methods achieved comparable experimental treatment effects to that observed in the CheckMate-057 clinical trial, with HRs of 0.73 (95% CI: 0.59, 0.92), 0.74 (95% CI: 0.61, 0.91), 0.72 (95% CI: 0.59, 0.88) for Methods 1-3, respectively. Reduction in study duration time was 99-115 days when borrowing 30-38 events for Methods 1-3, respectively. This study demonstrated that it is feasible to emulate an RCT using a hybrid control arm design using three frequentist propensity-score based statistical borrowing methods. Selection of an appropriate, fit-for-use RWD cohort is critical to minimizing bias in experimental treatment effect.
To use historical controls for indirect comparison with single-arm trials, the population difference between data sources should be adjusted to reduce confounding bias. The adjustment is more difficult for time-to-event data with a cure fraction. We propose different adjustment approaches based on pseudo observations and calibration weighting by entropy balancing. We show a simple way to obtain the pseudo observations for the cure rate and propose a simple weighted estimator based on them. Estimation of the survival function in presence of a cure fraction is also considered. Simulations are conducted to examine the proposed approaches. An application to a breast cancer study is presented.
In the area of diagnostics, it is common practice to leverage external data to augment a traditional study of diagnostic accuracy consisting of prospectively enrolled subjects to potentially reduce the time and/or cost needed for the performance evaluation of an investigational diagnostic device. However, the statistical methods currently being used for such leveraging may not clearly separate study design and outcome data analysis, and they may not adequately address possible bias due to differences in clinically relevant characteristics between the subjects constituting the traditional study and those constituting the external data. This paper is intended to draw attention in the field of diagnostics to the recently developed propensity score-integrated composite likelihood approach, which originally focused on therapeutic medical products. This approach applies the outcome-free principle to separate study design and outcome data analysis and can mitigate bias due to imbalance in covariates, thereby increasing the interpretability of study results. While this approach was conceived as a statistical tool for the design and analysis of clinical studies for therapeutic medical products, here, we will show how it can also be applied to the evaluation of sensitivity and specificity of an investigational diagnostic device leveraging external data. We consider two common scenarios for the design of a traditional diagnostic device study consisting of prospectively enrolled subjects, which is to be augmented by external data. The reader will be taken through the process of implementing this approach step-by-step following the outcome-free principle that preserves study integrity.
The propensity score-integrated composite likelihood (PSCL) method is one method that can be utilized to design and analyze an application when real-world data (RWD) are leveraged to augment a prospectively designed clinical study. In the PSCL, strata are formed based on propensity scores (PS) such that similar subjects in terms of the baseline covariates from both the current study and RWD sources are placed in the same stratum, and then composite likelihood method is applied to down-weight the information from the RWD. While PSCL was originally proposed for a fixed design, it can be extended to be applied under an adaptive design framework with the purpose to either potentially claim an early success or to re-estimate the sample size. In this paper, a general strategy is proposed due to the feature of PSCL. For the possibility of claiming early success, Fisher's combination test is utilized. When the purpose is to re-estimate the sample size, the proposed procedure is based on the test proposed by Cui, Hung, and Wang. The implementation of these two procedures is demonstrated via an example.
We consider outcome adaptive phase II or phase II/III trials to identify the best treatment for further development. Different from many other multi-arm multi-stage designs, we borrow approaches for the best arm identification in multi-armed bandit (MAB) approaches developed for machine learning and adapt them for clinical trial purposes. The best arm identification in MAB focuses on the error rate of identification at the end of the trial, but we are also interested in the cumulative benefit of trial patients, for example, the frequency of patients treated with the best treatment. In particular, we consider Top-Two Thompson Sampling (TTTS) and propose an acceleration approach for better performance in drug development scenarios in which the sample size is much smaller than that considered in machine learning applications. We also propose a variant of TTTS (TTTS2) which is simpler, easier for implementation, and has comparable performance in small sample settings. An extensive simulation study was conducted to evaluate the performance of the proposed approach in multiple typical scenarios in drug development.
Necessity for finding improved intervention in many legacy therapeutic areas are of high priority. This has the potential to decrease the expense of medical care and poor outcomes for many patients. Typically, clinical efficacy is the primary evaluating criteria to measure any beneficial effect of a treatment. Albeit, there could be situations when several other factors (e.g. side-effects, cost-burden, less debilitating, less intensive, etc.) which can permit some slightly less efficacious treatment options favorable to a subgroup of patients. This often leads to non-inferiority (NI) testing. NI trials may or may not include a placebo arm due to ethical reasons. However, when included, the resulting three-arm trial is more prudent since it requires less stringent assumptions compared to a two-arm placebo-free trial. In this article, we consider both Frequentist and Bayesian procedures for testing NI in the three-arm trial with binary outcomes when the functional of interest is risk difference. An improved Frequentist approach is proposed first, which is then followed by a Bayesian counterpart. Bayesian methods have a natural advantage in many active-control trials, including NI trial, as it can seamlessly integrate substantial prior information. In addition, we discuss sample size calculation and draw an interesting connection between the two paradigms.
In many orphan diseases and pediatric indications, the randomized controlled trials may be infeasible because of their size, duration, and cost. Leveraging information on the control through a prior can potentially reduce sample size. However, unless an objective prior is used to impose complete ignorance for the parameter being estimated, it results in biased estimates and inflated type-I error. Hence, it is essential to assess both the confirmatory and supplementary knowledge available during the construction of the prior to avoid "cherry-picking" advantageous information. For this purpose, propensity score methods are employed to minimize selection bias by weighting supplemental control subjects according to their similarity in terms of pretreatment characteristics to the subjects in the current trial. The latter can be operationalized through a proposed measure of overlap in propensity-score distributions. In this paper, we consider single experimental arm in the current trial and the control arm is completely borrowed from the supplemental data. The simulation experiments show that the proposed method reduces prior and data conflict and improves the precision of the of the average treatment effect.
BACKGROUND:Meta-analysis of related trials can provide an overall measure of safety-signal accounting for variability across studies. In addition to an overall measure, researchers may often be interested in study-specific measures to assess safety of the product. Likelihood ratio tests (LRT) methods serve this purpose by identifying studies that appear to show a safety concern. In this paper, we present a Bayesian approach. Despite having good statistical properties, the LRT methods may not be suitable for the meta-analysis of randomized controlled trials (RCTs) when there are several studies with zero events in at least one arm.METHODS:In this article, we describe a Bayesian framework using a Zero-inflated binomial model with spike-and-slab parameterization for the treatment effects. In addition to providing an overall meta-analytic estimate, this method provides posterior probability of a safety-signal for each study.RESULTS:We illustrate the approach using two published data sets comprising several randomized controlled trials (RCTs) each and compare the model performance for different choices of priors for treatment effect.DISCUSSION:The proposed Bayesian methodological framework is useful to identify potential signal for single adverse event and to determine overall meta-analytic estimate of the magnitude of the signal. Practitioners may consider this approach as an alternative to the frequentist's LRT approach discussed in Jung et al. (J Biopharm Stat 31:47-54, 2020) when there are zero events in either the treatment arm or the control arm. In the future, this approach can be further extended to accommodate multiple adverse events.
Utilizing external data from the real world, including data from historical clinical trials, has received increasing interest in drug development. The use of external data to support drug evaluation in clinical trials has mainly been through using various matching methods for baseline characteristics to form external control arms in single-arm trials or to augment control arms of randomized controlled trials in hybrid approaches. However, matching the baseline characteristics between the trial and the external subjects can only guarantee comparability on the level of baseline characteristics. Differences in outcomes between the two data sources may still exist due to contemporaneous and operational characteristics. Similarity between the outcomes in the trial control and the external subjects with similar baseline characteristics can be critical in leveraging the external subjects in the clinical trials. In this paper, a resampling method for augmenting control arms in randomized controlled trials is proposed under the conditional borrowing framework. The new method establishes empirical distributions for the hazard ratio in outcomes between the external and trial control subjects. The borrowing decision is then derived from this empirical distribution using a measure of similarity. Once the borrowing decision is established, the borrowing weights for the external subjects, based on the similarity measure, are incorporated in the weighted partial likelihood to evaluate the treatment effect. The operating characteristics of the hybrid control arm, under both the conditional borrowing and unconditional borrowing frameworks, are evaluated. Simulation is conducted to evaluate Type I error, bias, and power. An illustrative example using simulated data is also presented.
Leveraging external data is a topic that have recently received much attention. The propensity score-integrated approaches are a methodological innovation for this purpose. In this paper we adapt these approaches, originally introduced to augment single-arm studies with external data, for the augmentation of both arms of a randomized controlled trial (RCT) with external data. After recapitulating the basic ideas, we provide a step-by-step tutorial of how to implement the propensity score-integrated approaches, from study design to outcome analysis, in the RCT setting in such a way that the study integrity and objectively are maintained. Both the Bayesian (power prior) approach and the frequentist (composite likelihood) approach are included. Some extensions and variations of these approaches are also outlined at the end of this paper.