Objective The designation of osteoarthritis (OA) as a serious disease allows for accelerated approval of a treatment based on an intermediate clinical endpoint. We evaluated proposed endpoints and assessed their feasibility for use in a randomized controlled trial (RCT). We examined associations between endpoints and subsequent total knee replacement (TKR), a proxy for long-term clinical benefit. Design We selected knees from the Osteoarthritis Initiative meeting typical RCT inclusion criteria, including pain and radiographic OA. Endpoints included TKR, end-stage knee OA (esKOA), Composite Knee Osteoarthritis Symptom Outcome (CKOASO), the FNIH OA Biomarkers Consortium endpoint, and several combinations of pain and/or functional limitations. We determined the cumulative incidence (CI) over four years and the associated sample size required for an RCT to detect a hazard ratio (HR) of 0.67 with 80% power (to assess feasibility). We assessed the association between reaching each endpoint and TKR over the subsequent 5 years. Results 1350 knees (one per participant) were included. 4-year CI of TKR was 4.3%. CIs were highest for FNIH (13.8%), esKOA (34.4%), and CKOASO (63.6%). The required sample size per arm ranged from 229 (CKOASO) to 3028 (TKR). The relative risk (RR) of subsequent TKR was highest for the esKOA endpoint (RR 5.7), followed by CKOASO (4.3) and FNIH (2.6). Conclusions Based on feasibility (sample size) and clinical relevance (association with subsequent TKR), we propose that the esKOA, CKOASO, and FNIH endpoints be prioritized for further consideration. Given feasibility limitations, we would not recommend TKR alone as a DMOAD trial endpoint.
OBJECTIVE:The Foundation for the National Institutes of Health (FNIH) OA Biomarkers Consortium aims to identify, develop, and qualify biomarkers to support drug development in knee osteoarthritis (OA). The project's second phase, the PROGRESS OA study, aims to externally validate prognostic and response biomarkers identified in the earlier phase (phase 1). Here we present results assessing external validation of prognostic imaging biomarkers. DESIGN:PROGRESS OA included data from the control arms of several completed randomized controlled trials (RCTs) for symptomatic knee OA. Radiographic progression was defined as joint space width loss (JSWL) ≥0.7 mm. Symptomatic progression was defined as increase of nine or more points in Western Ontario and McMaster Universities Arthritis Index pain (0-100 scale). Imaging biomarkers included quantitative measures of cartilage thickness and semiquantitative (SQ) assessments. Associations between baseline biomarkers and outcomes over 12 to 36 months were examined using logistic regression. RESULTS:A total of 320 participants from four RCTs were included. Forty-one participants (13%) had JSWL ≥0.7 mm and 64 (20%) had worsening symptoms. In univariable logistic regression, measures of quantitative and SQ cartilage, SQ Hoffa-synovitis, effusion-synovitis, and meniscal extrusion were consistently selected to predict JSWL ≥0.7 mm, similar to phase 1. SQ Hoffa-synovitis and lateral meniscal damage were consistently selected to predict symptomatic progression. Cross-validated areas under the curve were 0.69 (95% confidence interval [CI]: 0.53-0.85) for JSWL ≥0.7 mm and 0.77 (95% CI: 0.65-0.87) for symptomatic progression. CONCLUSION:The selected prognostic imaging biomarkers are candidates for enriching OA trials for structural and/or symptomatic progressors. Ongoing work includes pursuit of formal biomarker qualification by regulatory agencies, and the use of these biomarkers to capture structural progression with high sensitivity to change.
Quantifiable image patterns associated with disease progression and treatment response are critical tools for guiding individual treatment, and for developing novel therapies. Here, we show that unsupervised machine learning can identify a pattern vocabulary of liver tissue in magnetic resonance images that quantifies treatment response in diffuse liver disease. Deep clustering networks simultaneously encode and cluster patches of medical images into a low-dimensional latent space to establish a tissue vocabulary. The resulting tissue types capture differential tissue change and its location in the liver associated with treatment response. We demonstrate the utility of the vocabulary in a randomized controlled trial cohort of patients with nonalcoholic steatohepatitis. First, we use the vocabulary to compare longitudinal liver change in a placebo and a treatment cohort. Results show that the method identifies specific liver tissue change pathways associated with treatment and enables a better separation between treatment groups than established non-imaging measures. Moreover, we show that the vocabulary can predict biopsy derived features from non-invasive imaging data. We validate the method in a separate replication cohort to demonstrate the applicability of the proposed method.
This paper investigates the robustness of futility analyses in clinical trials when interim analysis population deviates from the target population. We demonstrate how population shifts can distort early stopping decisions and propose post-stratification strategies to mitigate these effects. Simulation studies illustrate the impact of subgroup imbalances and the effectiveness of naive, model-based, and hybrid post-stratification methods. We also introduce a permutation-based screening test for identifying variables contributing to population heterogeneity. Our findings support the integration of post-stratification adjustments using all available baseline data at the interim analysis to enhance the validity and integrity of futility decisions.
Platform trials are innovative clinical trials governed by a master protocol that allows for the evaluation of multiple investigational treatments that enter and leave the trial over time. Interest in platform trials has been steadily increasing over the last decade. Due to their highly adaptive nature, platform trials provide sufficient flexibility to customize important trial design aspects to the requirements of both the specific disease under investigation and the different stakeholders. The flexibility of platform trials, however, comes with complexities when designing such trials. In the past, we reviewed existing software for simulating clinical trials and found that none of them were suitable for simulating platform trials as they do not accommodate the design features and flexibility inherent to platform trials, such as staggered entry of treatments over time. We argued that simulation studies are crucial for the design of efficient platform trials. We developed and proposed an iterative, simulation-guided “vanilla and sprinkles” framework, i.e. from a basic to a more complex design, for designing platform trials. We addressed the functionality limitations of existing software as well as the unavailability of the coding therein by developing a suite of open-source software to use in simulating platform trials based on the R programming language. To give some examples, the newly developed software supports simulating staggered entry of treatments throughout the trial, choosing different options for control data sharing, specifying different platform stopping rules and platform-level operating characteristics. The software we developed is available through open-source licensing to enable users to access and modify the code. The separate use of two of these software packages to implement the same platform design by independent teams obtained the same results. We provide a framework, as well as open-source software for the design and simulation of platform trials. The software tools provide the flexibility necessary to capture the complexity of platform trials.
Platform trials have gained a lot of attention in recent years as a possible remedy for time-consuming classical two-arm randomized controlled trials, especially in early phase drug development. This article illustrates how to use the CohortPlat R package to simulate a cohort platform trial, where each cohort consists of a combination treatment and the respective monotherapies and standard-of-care. For all simulations, we assume a binary primary endpoint. The package offers extensive flexibility with respect to both platform trial trajectories, as well as treatment effect scenarios and decision rules. As a special feature, the package provides a designated function for running multiple such simulations efficiently in parallel and saving the results in a concise manner. Many illustrations of code usage are provided.
Histological assessment of autoimmune hepatitis (AIH) is challenging. As one of the possible results of these challenges, nonclassical features such as bile-duct injury stays understudied in AIH. We aim to develop a deep learning tool (artificial intelligence for autoimmune hepatitis [AI(H)]) that analyzes the liver biopsies and provides reproducible, quantifiable, and interpretable results directly from routine pathology slides. A total of 123 pre-treatment liver biopsies, whole-slide images with confirmed AIH diagnosis from the archives of the Institute of Pathology at University Hospital Basel, were used to train several convolutional neural network models in the Aiforia artificial intelligence (AI) platform. The performance of AI models was evaluated on independent test set slides against pathologist's manual annotations. The AI models were 99.4%, 88.0%, 83.9%, 81.7%, and 79.2% accurate (ratios of correct predictions) for tissue detection, liver microanatomy, necroinflammation features, bile duct damage detection, and portal inflammation detection, respectively, on hematoxylin and eosin-stained slides. Additionally, the immune cells model could detect and classify different immune cells (lymphocyte, plasma cell, macrophage, eosinophil, and neutrophil) with 72.4% accuracy. On Sirius red-stained slides, the test accuracies were 99.4%, 94.0%, and 87.6% for tissue detection, liver microanatomy, and fibrosis detection, respectively. Additionally, AI(H) showed bile duct injury in 81 AIH cases (68.6%). The AI models were found to be accurate and efficient in predicting various morphological components of AIH biopsies. The computational analysis of biopsy slides provides detailed spatial and density data of immune cells in AIH landscape, which is difficult by manual counting. AI(H) can aid in improving the reproducibility of AIH biopsy assessment and bring new descriptive and quantitative aspects to AIH histology.
Background Interventional clinical studies conducted in the regulated drug research environment are designed using International Council for Harmonisation (ICH) regulatory guidance documents: ICH E6 (R2) Good Clinical Practice—scientific guideline, first published in 2002 and last updated in 2016. This document provides an international ethical and scientific quality standard for designing and conducting trials that involve the participation of human subjects. Recently, there has been heightened awareness of the importance of integrated research platform trials (IRPs) designed to evaluate multiple therapies simultaneously. The use of a single master protocol as a key source document to fulfill trial conduct obligations has resulted in a re-examination of the templates used to fulfill the dynamic regulatory and modern drug development environment challenges. Methods Regulatory medical writing, biostatistical, and other members of EU Patient-cEntric clinicAl tRial pLatforms (EU-PEARL) developed the suite of templates for IRPs over a 3.5-year period. Stakeholders contributing expertise included academic hospitals, pharmaceutical companies, non-governmental organizations, patient representative groups, and small and medium-sized enterprises (SMEs). Results The suite of templates for IRPs based on TransCelerate’s Common Protocol Template (CPT) and statistical analysis plan (SAP) should help authors navigate relevant guidelines as they create study design content relevant for today’s IRP studies. It offers practical suggestions for adaptive platform designs which offer flexible features such as dropping treatments for futility or adding new treatments to be tested during a trial. The EU-PEARL suite of templates for IRPs comprises a preface, followed by the actual resource. The preface clarifies the intended use and underlying principles that inform resource utility. The preface lists references contributing to the development of the resource. The resource includes TransCelerate CPT guidance text, and EU-PEARL-derived guidance text, distinguished from one another using shading. Rationale comments are used throughout for clarification purposes. In addition, a user-friendly, functional, and informative Platform Trials Best Practices tool to support the setup, design, planning, implementation, and conduct of complex and innovative trials to support multi-sourced/multi-company platform trials is also provided. Together, the EU-PEARL suite of templates and the Platform Trials Best Practices tool constitute the reference user manual. Conclusions This publication is intended to enhance the use, understanding, and dissemination of the EU-PEARL suite of templates for designing IRPs. The reference user manual and the associated website ( http://www.eu-pearl ) should facilitate the designing of IRP trials.
Although platform trials have many benefits, the complexity of these designs may result not only in increased methodological but also regulatory and ethical challenges. These aspects were addressed as part of the IMI project EU Patient-Centric Clinical Trial Platforms (EU-PEARL). We reviewed the available guidelines on platform trials in the European Union and the United States. This is supported and complemented by feedback received from regulatory interactions with the European Medicines Agency and the US Food and Drug Administration. Throughout the project we collected the needs of all relevant stakeholders including ethics committees, regulators, and health technology assessment bodies through active dialog and dedicated stakeholder workshops. Furthermore, we focused on methodological aspects and where applicable identified the corresponding guidance. Learnings from the guideline review, regulatory interactions, and workshops are provided. Based on these, a master protocol template was developed. Issues that still need harmonization or clarification in guidelines or where further methodological research is needed are also presented. These include questions around clinical trial submissions in Europe, the need for multiplicity control across the whole master protocol, the use of non-concurrent controls, and the impact of different randomization schemes. Master protocols are an efficient and patient-centered clinical trial design that can expedite drug development. However, they can also introduce additional operational and regulatory complexities. It is important to understand the different requirements of stakeholders upfront and address them in the trial. While relevant guidance is increasing, early dialog with relevant stakeholders can help to further support such designs.
Purpose (the aim of the study): The FNIH OA Biomarkers Consortium aims to identify, develop, and qualify biomarkers to support new drug development in OA. Phase I of the project used data from the Osteoarthritis Initiative and sought to determine the predictive ability of biomarkers for the symptomatic and radiographic progression of knee OA; results have been published. Here we present results of the PROGRESS OA study, which aims to validate these biomarkers in the control arms of existing completed OA randomized controlled trials (RCTs).
AIMS:Metabolic dysfunction Associated Steatotic Liver Disease (MASLD) outcomes such as MASH (metabolic dysfunction associated steatohepatitis), fibrosis and cirrhosis are ordinarily determined by resource-intensive and invasive biopsies. We aim to show that routine clinical tests offer sufficient information to predict these endpoints.METHODS:Using the LITMUS Metacohort derived from the European NAFLD Registry, the largest MASLD dataset in Europe, we create three combinations of features which vary in degree of procurement including a 19-variable feature set that are attained through a routine clinical appointment or blood test. This data was used to train predictive models using supervised machine learning (ML) algorithm XGBoost, alongside missing imputation technique MICE and class balancing algorithm SMOTE. Shapley Additive exPlanations (SHAP) were added to determine relative importance for each clinical variable.RESULTS:Analysing nine biopsy-derived MASLD outcomes of cohort size ranging between 5385 and 6673 subjects, we were able to predict individuals at training set AUCs ranging from 0.719-0.994, including classifying individuals who are At-Risk MASH at an AUC = 0.899. Using two further feature combinations of 26-variables and 35-variables, which included composite scores known to be good indicators for MASLD endpoints and advanced specialist tests, we found predictive performance did not sufficiently improve. We are also able to present local and global explanations for each ML model, offering clinicians interpretability without the expense of worsening predictive performance.CONCLUSIONS:This study developed a series of ML models of accuracy ranging from 71.9-99.4% using only easily extractable and readily available information in predicting MASLD outcomes which are usually determined through highly invasive means.
Non-alcoholic fatty liver disease is a condition that affects 25% of the population. Non-alcoholic steatohepatitis (NASH) is a progressive form of the disease that can lead to severe complications such as cirrhosis and hepatocellular carcinoma. Despite its high prevalence, no drugs are currently approved for the treatment of NASH. The drug development pipeline in NASH is very active, yet most assets do not progress to phase III trials and those that do reach phase III often fail to achieve the endpoints necessary for approval by regulatory agencies. Amongst other reasons, the methodological and operational features of traditional clinical trials in NASH might impede optimal drug development. In this regard, platform trials might be an attractive complement or alternative to conventional clinical trials. Platform trials use a master protocol which enables evaluation of multiple investigational medicinal products concurrently or sequentially with a single, shared control arm. Through Bayesian interim analyses, these trials allow for early exit of drugs from the trial based on success or futility, while providing participants better chances of receiving active compounds through adaptive randomisation. Overall, platform trials represent an alternative for patients, pharmaceutical companies, and clinicians in the quest to accelerate the approval of pharmacologic treatments for NASH.
Non-alcoholic steatohepatitis (NASH) is the progressive form of nonalcoholic fatty liver disease (NAFLD) and a disease with high unmet medical need. Platform trials provide great benefits for sponsors and trial participants in terms of accelerating drug development programs. In this article, we describe some of the activities of the EU-PEARL consortium (EU Patient-cEntric clinicAl tRial pLatforms) regarding the use of platform trials in NASH, in particular the proposed trial design, decision rules and simulation results. For a set of assumptions, we present the results of a simulation study recently discussed with two health authorities and the learnings from these meetings from a trial design perspective. Since the proposed design uses co-primary binary endpoints, we furthermore discuss the different options and practical considerations for simulating correlated binary endpoints.
Platform trials can evaluate the efficacy of several treatments compared to a control. The number of treatments is not fixed, as arms may be added or removed as the trial progresses. Platform trials are more efficient than independent parallel-group trials because of using shared control groups. For arms entering the trial later, not all patients in the control group are randomised concurrently. The control group is then divided into concurrent and non-concurrent controls. Using non-concurrent controls (NCC) can improve the trial's efficiency, but can introduce bias due to time trends. We focus on a platform trial with two treatment arms and a common control arm. Assuming that the second treatment arm is added later, we assess the robustness of model-based approaches to adjust for time trends when using NCC. We consider approaches where time trends are modeled as linear or as a step function, with steps at times where arms enter or leave the trial. For trials with continuous or binary outcomes, we investigate the type 1 error (t1e) rate and power of testing the efficacy of the newly added arm under a range of scenarios. In addition to scenarios where time trends are equal across arms, we investigate settings with trends that are different or not additive in the model scale. A step function model fitted on data from all arms gives increased power while controlling the t1e, as long as the time trends are equal for the different arms and additive on the model scale. This holds even if the trend's shape deviates from a step function if block randomisation is used. But if trends differ between arms or are not additive on the model scale, t1e control may be lost. The efficiency gained by using step function models to incorporate NCC can outweigh potential biases. However, the specifics of the trial, plausibility of different time trends, and robustness of results should be considered
Platform trials have become increasingly popular for drug development programs, attracting interest from statisticians, clinicians and regulatory agencies. Many statistical questions related to designing platform trials-such as the impact of decision rules, sharing of information across cohorts, and allocation ratios on operating characteristics and error rates-remain unanswered. In many platform trials, the definition of error rates is not straightforward as classical error rate concepts are not applicable. For an open-entry, exploratory platform trial design comparing combination therapies to the respective monotherapies and standard-of-care, we define a set of error rates and operating characteristics and then use these to compare a set of design parameters under a range of simulation assumptions. When setting up the simulations, we aimed for realistic trial trajectories, such that for example, a priori we do not know the exact number of treatments that will be included over time in a specific simulation run as this follows a stochastic mechanism. Our results indicate that the method of data sharing, exact specification of decision rules and a priori assumptions regarding the treatment efficacy all strongly contribute to the operating characteristics of the platform trial. Furthermore, different operating characteristics might be of importance to different stakeholders. Together with the potential flexibility and complexity of a platform trial, which also impact the achieved operating characteristics via, for example, the degree of efficiency of data sharing this implies that utmost care needs to be given to evaluation of different assumptions and design parameters at the design stage.
Youden-index were 10.2 kPa (Sp 76%, SFR 80%, NNT = 18), 14.6 (Se 68%, SFR 69%, NNT = 24), and 10.4 kPa (Se 89%, Sp 77%, SFR 79%, NNT = 18) for LSM-VCTE; 1.16 (Sp 52%, SFR 89%, NNT = 18), 2.73 (Se 41%, SFR 79%, NNT = 39) and 1.35 (Se 88%, Sp 61%, SFR 87%, NNT = 18) for FIB-4; -1.637 (So 53%, SFR 89%, NNT = 18), 0.473 (Se 32%, SFR 82%, NNT = 50) and -0.878 (Se 77%, Sp 69%, SFR 86%, NNT = 21) for NFS (Figure 1).A cut-off combination of FIB4 1.16 or NFS -1.637 followed by LSM VCTE of 16 kPa yielded Se 59%, Sp 93%, SFR 62%, and NNT = 27 and only 10% of the entire patient group would have to undergo liver biopsy while also reducing the number of LSM-VCTE exams to 51%. Conclusion:Only LSM-VCTE achieved SFR of 50% and 60% which was at the cost of high NNT.Sequential combinations can achieve similar diagnostic accuracy while potentially reducing the number of LSM-VCTE exams and biopsies being performed, which could have favorable cost implications.