Proof of Concept (POC) trials should be designed to inform a drug's targeted benefit-risk profile and ability to fulfill unmet medical needs, rather than simply achieving statistical significance. Pulkstenis et al. [1] proposed a unified Bayesian framework for decision-making in Proof of Concept (POC) trials to weigh evidence regarding a dual-level Target Product Profile (TPP). To accelerate decision-making while evidence is still emerging from a Phase 2 trial, we generalize this framework by incorporating interim monitoring leveraging predictive probability. An accelerated Go-decision can be achieved to initiate Phase 3 studies earlier if sufficiently promising results are observed, while an early No-Go or wait decision can be made when supporting evidence indictates. Practical guidance for implementing the approach is demonstrated by simulation studies of binary and continuous endpoints.
Motivated by a rheumatoid arthritis clinical trial, we propose a new Bayesian method called SPx, standing for synthetic prior with covariates, to borrow information from historical trials to reduce the control group size in a new trial. The method involves a novel use of Bayesian model averaging to balance between multiple possible relationships between the historical and new trial data, allowing the historical data to be dynamically trusted or discounted as appropriate. We require only trial-level summary statistics, which are available more often than patient-level data. Through simulations and an application to the rheumatoid arthritis trial we show that SPx can substantially reduce the control group size while maintaining Frequentist properties.
Effective decision-making plays a vital role throughout the drug development process, particularly when a proof-of-concept (POC) or phase II study has been completed. To determine whether to proceed to a larger-scale, confirmatory phase III study, assessing the uncertainty about the underlying treatment effect and the probability of success (POS) in the phase III study is of critical importance. In this paper, we proposed and investigated a Bayesian covariate-adjusted hierarchical modeling approach leveraging historical data with longitudinal outcome to quantitatively assess the POS of the confirmatory phase III trial. Although historical data borrowing methods are widely used and known for the advantages in alleviating recruitment and ethical challenges as well as improving trial operational efficiency, its application to predicting future trial POS with longitudinal outcome over multiple visits pose methodological challenges. This paper not only provided a comprehensive modeling approach but also demonstrated how the proposed model can be used in a Go/No-Go decision-making framework with a glaucoma eye care project example. For the approval of new drugs targeting glaucoma, regulatory agencies typically require a pivotal phase III trial to demonstrate noninferiority compared to a standard of care treatment. This may involve meeting both statistical and clinical margins across multiple visits simultaneously. Simulations were performed to evaluate the key factors that affect the operating characteristics, such as between-trial heterogeneity, subject-level variance and between-visit correlation. The proposed decision-making framework can also be applied to studies in other therapeutical areas with similar settings.
BackgroundBradykinesia, a primary symptom of Parkinson’s disease, significantly impacts patients’ quality of life. Traditional assessments like the MDS-UPDRS require in-person evaluation and can be subjective, resulting in patient burden and variability.MethodsThis study investigates digital health technologies and machine learning (ML) models, specifically XGBoost and Convolutional Neural Networks (CNNs), to improve objectivity and precision in bradykinesia quantification. Using data from the 12-month WATCH-PD study in early, untreated Parkinson’s disease, models were evaluated against MDS-UPDRS items (3.6 pronation-supination and 3.7 toe tapping) for cross-sectional accuracy and longitudinal sensitivity.ResultsResults show that the class-balanced XGBoost was the best-performing model on both tasks, outperforming a stratified-random baseline and all CNN variants (cross-entropy, focal loss, and self-attention) across every imbalance-aware metric: macro-F1 0.47 and 0.40, balanced accuracy 0.47 and 0.42, quadratic weighted kappa 0.58 and 0.47, and accuracy 0.60 and 0.53 for pronation-supination and toe tapping, respectively, with the clearest gains on the under-represented higher-severity classes. A secondary longitudinal analysis found limited sensitivity to 12-month change across the ML models.DiscussionOverall, machine-learning models, led by the class-balanced XGBoost, show promise in accurately predicting item-level MDS-UPDRS bradykinesia scores from wearable sensor data. Our results also highlight the importance of aligning feature selection and model architecture to intended clinical endpoints.
Recently, the U.S. Food and Drug Administration (FDA) released draft guidance signaling a paradigm shift that facilitates the use of Bayesian methodology as the primary analysis and decision framework for drug approval. The cornerstone and fundamental challenge of this framework is the specification and calibration of Bayesian success criteria to control decision errors, ensuring reliable clinical and regulatory outcomes. In this work, we systematically investigate various Bayesian decision-error metrics, their theoretical interrelationships, and their alignment with conventional Frequentist counterparts. This investigation provides critical theoretical insights and practical guidance on calibrating Bayesian success criteria and operating characteristics to ensure robust decision-making and the integrity of public health decisions. We illustrate this framework using a clinical trial evaluating revascularization strategies for cardiogenic shock. A Shiny application will be available at www.trialdesign.org to assist sponsors and regulators in evaluating calibration strategies consistent with recent regulatory perspectives.
The prudent use of covariates to enhance the efficiency and ethics of clinical trials has garnered significant attention, particularly following the FDA's 2023 guidance on adjusting for covariates. This article introduces a Bayesian covariate-adjusted response-adaptive design aimed at distinguishing between prognostic and predictive covariates during randomization and analysis. The proposed design allocates more patients to the superior treatment based on predictive covariates while maintaining balance across prognostic covariate levels, without sacrificing the power to detect treatment effects. Predictive covariates, which identify patients more likely to benefit from a treatment, and prognostic covariates, which predict overall clinical outcomes, are crucial for personalized medicine and ethical rigor in clinical trials. The Bayesian covariate-adjusted response-adaptive design leverages these covariates to enhance precision and ensure balanced comparison groups, addressing patient heterogeneity and improving treatment efficacy. Our approach builds on the foundation of response-adaptive randomization designs, incorporating Bayesian methodologies to manage the complexities of adaptive designs and control the Type I error rate. Comprehensive numerical studies demonstrate the advantages of our design in achieving ethical, efficient, and balancing goals.
The importance of covariate adjustment in clinical trials has been underscored by the U.S. FDA's guidance. Inference, with or without covariates, after implementing covariate adaptive randomization (CAR), is garnering increased interest. This paper investigates the sequential monitoring of covariate-adaptive randomized clinical trials through non-parametric methods, a critical advancement for enhancing the precision and efficiency of medical research. CAR, which incorporates baseline patient characteristics into the randomization process, aims to mitigate the risk of confounding and improve the balance of covariates across treatment groups, thereby addressing patients' heterogeneity. Although CAR is known for its benefits in reducing biases and enhancing statistical power, its integration into sequentially monitored clinical trials-a standard practice-poses methodological challenges, particularly in controlling the type I error rate. By employing a non-parametric approach, we demonstrate through theoretical proofs and numerical analyses that our methods effectively control the type I error rate and surpass traditional randomization and analysis methods. This paper not only fills a gap in the literature on sequential monitoring of CAR without model misspecification but also proposes practical solutions for enhancing trial design and analysis, thereby contributing significantly to the field of clinical research.
The standard logrank test may lose statistical power substantially when the underlying proportional hazards (PH) assumption is violated. Among non-PH patterns, delayed treatment effects are very commonly anticipated and actually observed not only in cancer immunotherapies and their related combinations with other therapeutic agents, but also in numerous situations across multiple therapeutic areas. A frequently considered scenario is that the PH pattern does not emerge until a certain period of time has elapsed. Based on the generalized piecewise weighted (GPW) logrank test, which is the asymptotically most powerful weighted logrank test detecting random delayed effects, we developed and evaluated a group sequential framework for maximum duration trials based on this test. A variance based procedure for the determination of the sequence of information fractions is proposed. Simulation studies are performed comparing various types of design and handling of information fractions. The procedure can control type I error rate in this group sequential design framework. Power gain of the GPW logrank test is demonstrated for the non-PH scenarios with delayed treatment effects. These results underscore the importance of considering GPW type of logrank test with appropriate design strategies when such delayed effect patterns are expected. Moreover, we have highlighted the economic advantage of sequential monitoring, providing early decision opportunities to accelerate drug development.
In Alzheimer’s Disease (AD) trials, clinical scales are used to assess treatment effect in patients. Minimizing statistical uncertainty of trial outcomes is an important consideration to increase statistical power. Machine learning models can leverage baseline data to create AI-generated digital twins – individualized predictions (or prognostic scores) of how each patient’s clinical outcomes may change during a trial assuming they received placebo. Incorporating prognostic scores into trial design and analysis as a covariate increases statistical power, or reduces sample size, in Phase 2 and 3 trials (Figures 1/2). We assessed these properties using data from a Phase 2 clinical trial of tilavonemab in patients diagnosed with early AD (NCT02880956) and digital twin (DT) methodology (PROCOVA TM ). In a double-blind, Phase 2 trial (AWARE), 453 patients aged 55-85 years with early AD (met NIA-AA clinical criteria for mild cognitive impairment or probable AD), were randomized to receive placebo or 1 of 3 doses of tilavonemab (1:1:1:1 ratio) over a 96-week treatment period. Prognostic scores were produced for the change from baseline (Δ) in Clinical Dementia Rating Scale Sum of Boxes (CDR-SB) and the Δ in AD Assessment Scale-Cognitive Subscale 14 (ADAS-Cog 14). Sample size savings were calculated from partial Pearson correlations (controlled for treatment) between prognostic scores and trial outcomes. Variability reductions were assessed using a covariance modelling approach that adjusts for the prognostic score (PROCOVA TM ). For Δ CDR-SB and Δ ADAS-Cog 14 at Week 96, standard deviations of the prognostic scores were lower than the trial’s outcomes. Partial correlation coefficients were moderate for both Δ CDR-SB (p = 0.360) and Δ ADAS-Cog 14 (p = 0.305) at Week 96. Total residual variance for both outcomes was reduced by ∼11% with DT methodology compared to an unadjusted model. Depending on correlations between prognostic scores and actual trial outcomes, a potential overall sample size reduction of 5-10% could be achieved using DT methodology (PROCOVA TM ) while maintaining statistical power, based on Δ CDR-SB and Δ ADAS-Cog 14 in the AWARE study. Sample size savings could enable shortening of the recruitment period and reduce the number of patients on placebo, encouraging greater patient participation.
The development of educational technology (EdTech) and artificial intelligence (AI) brings about a revolution in English learning by providing flexible, effective, and customized solutions. The purpose of this study is to assess the impact of AI and EdTec on education. In this article, we defined the multi-criteria decision-making (MCDM) procedure to manage ambiguity and awkward information by integrating the Technique for Order of Preference by Similarity to the Ideal Solution (Topsis) method with Circular q-Rung orthopair fuzzy set (Cq-ROFS), and Bonferroni mean (BM) operators to evaluate and prioritize AI-driven EdTech tools. The methodology incorporates multiple attributes, such as adaptability, learner engagement, cost-effectiveness, and scalability, within an MCDM framework. These results highlight the huge potential of intelligent teaching programs, flexible learning environments, and AI-powered language models to improve English ability. This research demonstrates how AI has advanced in education from simple computer-assisted language learning to complex AI-driven platforms like chatbots and intelligent systems for teaching. These developments, such as automated grading and feedback, have given teachers the ability to improve administrative effectiveness and instructional quality. Furthermore, customized and interactive learning experiences that are adapted to the needs and preferences of each student have been made possible by AI-based EdTech solutions. The study used the TOPSIS technique to rank important criteria for optimizing various solutions, highlighting their contribution to preservation, interaction, and overall efficacy in English language learning.
Adverse drug events (ADEs) are one of the major causes of hospital admissions and are associated with increased morbidity and mortality. Post-marketing ADE identification is one of the most important phases of drug safety surveillance. Traditionally, data sources for post-marketing surveillance mainly come from spontaneous reporting system such as the Food and Drug Administration Adverse Event Reporting System (FAERS). Social media data such as posts on X (formerly Twitter) contain rich patient and medication information and could potentially accelerate drug surveillance research. However, ADE information in social media data is usually locked in the text, making it difficult to be employed by traditional statistical approaches. In recent years, large language models (LLMs) have shown promise in many natural language processing tasks. In this study, we developed several LLMs to perform ADE classification on X data. We fine-tuned various LLMs including BERT-base, Bio_ClinicalBERT, RoBERTa, and RoBERTa-large. We also experimented ChatGPT few-shot prompting and ChatGPT fine-tuned on the whole training data. We then evaluated the model performance based on sensitivity, specificity, negative predictive value, positive predictive value, accuracy, F1-measure, and area under the ROC curve. Our results showed that RoBERTa-large achieved the best F1-measure (0.8) among all models followed by ChatGPT fine-tuned model with F1-measure of 0.75. Our feature importance analysis based on 1200 random samples and RoBERTa-Large showed the most important features are as follows: "withdrawals"/"withdrawal", "dry", "dealing", "mouth", and "paralysis". The good model performance and clinically relevant features show the potential of LLMs in augmenting ADE detection for post-marketing drug safety surveillance.
Most existing dose-ranging study designs focus on assessing the dose-efficacy relationship and identifying the minimum effective dose. There is an increasing interest in optimizing the dose based on the benefit-risk tradeoff. We propose a Bayesian quasi-likelihood dose-ranging design that jointly considers safety and efficacy to simultaneously identify the minimum effective dose and the maximum utility dose to optimize the benefit-risk tradeoff. The binary toxicity endpoint is modeled using a beta-binomial model. The efficacy endpoint is modeled using the quasi-likelihood approach to accommodate various types of data (e.g. binary, ordinal or continuous) without imposing any parametric assumptions on the dose-response curve. Our design utilizes a utility function as a measure of benefit-risk tradeoff and adaptively assign patients to doses based on the doses' likelihood of being the minimum effective dose and maximum utility dose. The design takes a group-sequential approach. At each interim, the doses that are deemed overly toxic or futile are dropped. At the end of the trial, we use posterior probability criteria to assess the strength of the dose-response relationship for establishing the proof-of-concept. If the proof-of-concept is established, we identify the minimum effective dose and maximum utility dose. Our simulation study shows that compared with some existing designs, the Bayesian quasi-likelihood dose-ranging design is robust and yields competitive performance in establishing proof-of-concept and selecting the minimum effective dose. Moreover, it includes an additional feature for further maximum utility dose selection.
Nocturnal scratching substantially impairs the quality of life in individuals with skin conditions such as atopic dermatitis (AD). Current clinical measurements of scratch rely on patient-reported outcomes (PROs) on itch over the last 24 h. Such measurements lack objectivity and sensitivity. Digital health technologies (DHTs), such as wearable sensors, have been widely used to capture behaviors in clinical and real-world settings. In this work, we develop and validate a machine learning algorithm using wrist-wearing actigraphy that could objectively quantify nocturnal scratching events, therefore facilitating accurate assessment of disease progression, treatment effectiveness, and overall quality of life in AD patients. A total of seven subjects were enrolled in a study to generate data overnight in an inpatient setting. Several machine learning models were developed, and their performance was compared. Results demonstrated that the best-performing model achieved the F1 score of 0.45 on the test set, accompanied by a precision of 0.44 and a recall of 0.46. Upon satisfactory performance with an expanded subject pool, our automatic scratch detection algorithm holds the potential for objectively assessing sleep quality and disease state in AD patients. This advancement promises to inform and refine therapeutic strategies for individuals with AD.
Clinical trials are an essential component of the drug development process, providing crucial data on the efficacy and safety of new treatments. However, traditional clinical trial designs can be inefficient and ineffective, leading to increased costs and a higher risk of failure. The FDA Oncology Center of Excellence (OCE) has recently initiated Project Optimus to reform the dose selection paradigm for oncology treatments. We propose the adaptive seamless phase II/III clinical trial designs (ASD) with the sequential estimation-adjusted urn (SEU) model for randomization to achieve efficient and ethical objectives. However, the combination of ASD and SEU poses a challenge in controlling the type I error rate: ASD exerts a dual influence of multiplicity and selection; all the responses and treatment assignments are not independent due to SEU. In this paper, we investigated how to overcome these difficulties, utilize the two adaptive approaches' advantages, and control the type I error rate. We provide a theoretical foundation for this procedure, and numerical studies demonstrate that our methods can assign more people to better treatments, leading to fewer failures while still controlling the type I error rate and preserving power.
Accurate prediction of a rare and clinically important event following study treatment has been crucial in drug development. For instance, the rarity of an adverse event is often commensurate with the seriousness of medical consequences, and delayed detection of the rare adverse event can pose significant or even life-threatening health risks to patients. In this machine learning case study, we demonstrate with an example originated from a real clinical trial setting how to define and solve the rare clinical event prediction problem using machine learning in pharmaceutical industry. The unique contributions of this work include the proposal of a six-step investigation framework that facilitates the communication with non-technical stakeholders and the interpretation of the model performance in terms of practical consequences in the context of patient screenings for conducting a future clinical trial. In terms of machine learning methodology, for data splitting into the training and test sets, we adapt the rare-event stratified split approach (from scikit-learn) to further account for group splitting for multiple records of a patient simultaneously. To handle imbalanced data due to rare events in model training, the cost-sensitive learning approach is employed to give more weights to the minor class and the metrics precision together with recall are used to capture prediction performance instead of the raw accuracy rate. Finally, we demonstrate how to apply the state-of-the-art SHAP values to identify important risk factors to improve model interpretability.
An adaptive platform trial (APT) is a multi-arm trial in the context of a single disease where treatment arms are allowed to enter or leave the trial based on some decision rule. If a treatment enters the trial later than the control arm, there exist nonconcurrent controls who were not randomized between the two arms under comparison. As APTs typically take long periods of time to conduct, temporal drift may occur, which requires the treatment comparisons to be adjusted for this temporal change. Under the causal inference framework, we propose two approaches for treatment comparisons in APTs that account for temporal drift, both based on propensity score weighting. In particular, to address unmeasured confounders, one approach is doubly robust in the sense that it remains valid so long as either the propensity score model is correctly specified or the time effect model is correctly specified. Simulation study shows that our proposed approaches have desirable operating characteristics with well controlled Type I error rates and high power with or without unmeasured confounders.
Introduction The identification of a new adverse event (AE) caused by a drug product is one of the key activities in the pharmaceutical industry to ensure the safety profile of a drug product. Machine learning (ML) has the potential to assist with signal detection and supplement traditional pharmacovigilance (PV) surveillance methods. This pilot ML modeling study was designed to detect potential safety signals for two AbbVie products and test the model's capability of detecting safety signals earlier than humans.Methods Drug X, a mature product with post-marketing data, and Drug Y, a recently approved drug in another therapeutic area, were selected. Gradient boosting-based ML approaches (e.g., XGBoost) were applied as the main modeling strategy.Results For Drug X, eight true signals were present in the test set. Among 12 potential new signals generated, four were true signals with a 50.0% sensitivity rate and a 33.3% positive predictive value (PPV) rate. Among the remaining eight potential new signals, one was confirmed as a signal and detected six months earlier than humans. For Drug Y, nine true signals were present in the test set. Among 13 potential new signals generated, five were true signals with a 55.6% sensitivity rate and a 38.5% PPV rate. Among the remaining eight potential new signals, none were confirmed as true signals upon human review.Conclusion This model demonstrated acceptable accuracy for safety signal detection and potential for earlier detection when compared to humans. Expert judgment, flexibility, and critical thinking are essential human skills required for the final, accurate assessment of adverse event cases.
Clinical trialists often face the challenge of balancing scientific questions with other design features, such as improving efficiency, minimizing exposure to inferior treatments, and simultaneously comparing multiple treatments. While Bayesian response adaptive randomization (RAR) is a popular and effective method for achieving these objectives, it is known to have large variability and a lack of explicit theoretical results, making its use in clinical trials a subject of concern. It is desirable to propose a design that targets the same allocation proportion as Bayesian RAR and achieves the above objectives but addresses the concerns over Bayesian RAR. We propose the frequentist doubly adaptive biased coin designs (DBCD) targeting ethical allocation proportions from the Bayesian framework to satisfy different objectives in clinical trials with time-to-event endpoints. We derive the theoretical properties of the proposed adaptive randomization design and show through comprehensive numerical simulations that it can achieve ethical objectives without sacrificing efficiency. Our combined theoretical and numerical results offer a strong foundation for the practical use of RAR in real clinical trials.
Conventionally, dose finding trials are based on dose-limiting toxicity (DLT) that only captures the most severe toxicities, for example, treatment related grade 3 or higher toxicity according to the NCI Common Terminology Criteria for Adverse Events. However, this approach is often problematic for certain novel targeted therapies and immunotherapies, which may not induce DLT within a clinically active dose range and are often characterized by low grade toxicities. This important issue has been highlighted and discussed in the American Statistical Association (ASA) Biopharmaceutical (BIOP) Section Open Forums, and is also an important consideration of the Project Optimus initiated by FDA to "reform the dose optimization and dose selection paradigm in oncology drug development." In this article, we propose an easy-to-implement model-assisted Bayesian design, known as multiple toxicity keyboard (MT-Keyboard) design, to incorporate toxicity grades and types into dose finding. The MT-Keyboard design is able to accommodate binary, quasi-binary and continuous toxicity endpoints that are constructed to account for toxicity grades and types. We further extend the MT-Keyboard design, referred to as TITE-MT-Keyboard, to accommodate late-onset toxicity using the approximated likelihood approach. Simulation shows that the MT-Keyboard and TITE-MT-Keyboard designs have desirable operating characteristics, comparable to or better than some existing designs. A web-based software to implement the design will be freely available at www.trialdesign.org. Supplementary materials for this article are available online.