Relying on a single primary endpoint in randomized controlled trials (RCTs) is often infeasible, for example due to rare or heterogeneous events. Regulatory guidance therefore allows multiple endpoints, but different analytical strategies address different scientific questions and null hypotheses, even when applied to the same set of variables. We explored three approaches to consider multiple endpoints in the primary analysis of RCTs, as stated in the FDA and EMA guidelines on multiplicity: (i) a composite endpoint (CE), (ii) multiple testing and multiplicity correction (MTMC), and (iii) a hierarchical non-parametric procedure, called generalized pairwise comparisons (GPC). Using clinical trial simulations, we compared these strategies’ power in two-arm RCTs perform when testing strategy-specific hypotheses across a range of scenarios reflecting endpoint prioritization, correlation between endpoints, and opposing treatment effects. When testing time-to-event endpoints, global testing strategies (CE and GPC) generally achieved higher power than MTMC. However, we also demonstrate that global procedures may yield statistically significant results even when treatment effects are heterogeneous across endpoints, underscoring the importance of careful interpretation and component-wise assessment. As trials increasingly use multiple endpoints, understanding the trade-off between statistical efficiency and interpretability, and provide practical guidance for choosing endpoint definitions and primary analysis strategies in future trials.
The analysis of platform trials can be enhanced by utilizing non-concurrent controls. Since including this data might also introduce bias in the treatment effect estimators if time trends are present, methods for incorporating non-concurrent controls adjusting for time have been proposed. However, so far their behavior has not been systematically investigated in platform trials that include interim analyses. To evaluate the impact of an interim analysis in trials utilizing non-concurrent controls, we consider a platform trial featuring two experimental arms and a shared control, with the second experimental arm entering later. We focus on a frequentist regression model that uses non-concurrent controls to estimate the treatment effect of the second arm and adjusts for time using a step function to account for temporal changes. We show that performing an interim analysis in Arm 1 may introduce bias in the point estimation of the effect in Arm 2, if the regression model is used without adjustment, and investigate how the marginal bias and bias conditional on the first arm continuing after the interim depend on different trial design parameters. Moreover, we propose a new estimator of the treatment effect in Arm 2, aiming to eliminate the bias introduced by both the interim analysis in Arm 1 and the time trends, and evaluate its performance in a simulation study. The newly proposed estimator is shown to substantially reduce the bias and type I error rate inflation while leading to power gains compared to an analysis using only concurrent controls.
Platform trials evaluate multiple treatments within a single trial infrastructure. Such designs have gained a lot of attraction in clinical research. If information gained from platform trials should provide confirmatory evidence for regulatory decisions, control of the Type I error rate is key. One critical issue is information leakage, for example, if any information of the ongoing trial is available, especially if it may impact the further conduct of treatments still in the platform trial and bias their results. This paper evaluates the potential impact of information leakage on the control of the Type I error rate in platform trials with time-to-event endpoints such as overall survival. We explore different strategies how information on the treatment effect of an still ongoing treatment could be obtained if a pre-planned analysis for another arm is conducted. This (leaked) information might be used to decide whether to continue the other arm as planned or conduct its final analysis immediately. By means of clinical trial simulations we evaluate the impact of different levels of information leakage on the Type I error rate. We show how the conditional error principle can be applied to estimate worst case Type I error rate inflation for the different forms of information leakage. We do not aim to quantify the exact maximum Type I error rate inflation but rather to raise awareness of the potential risk for estimation of comparative results. Finally, we discuss the regulatory implications of information leakage and propose strategies to mitigate these risks.
This review provides a systematic overview of methods that combine covariate-based clustering of observational units (patients) with outcome models for clinical studies. We distinguish between informed-cluster models, where the outcome contributes to cluster formation, and agnostic-cluster models, where clustering is performed solely on covariates in a separate first step. Informed-cluster models include product partition models with covariates (PPMx), finite mixtures of regression models (FMR), and cluster-aware supervised learning (CluSL). Agnostic-cluster models encompass two-step procedures using either model-based or algorithmic clustering followed by cluster-specific regression models. Following a systematic search of Web of Science and PubMed, 55 records were identified that propose or evaluate such models. We describe the key models, summarise study characteristics, and present applications from biomedical and public health research. Clustering-based outcome models are particularly relevant for settings with high-dimensional covariates (e.g., biomarker panels and "omics") and heterogeneous patient populations. These models can support risk stratification and we discuss extensions to estimate subgroup-specific treatment effects. They are most valuable when the population is clustered in distinct regions of the covariate space that correspond to different outcome distributions. We discuss applications to rare disease research, covariate adjustment and borrowing from historical data, and subgroup-specific treatment effect estimation in clinical trials.
BACKGROUND:To determine long-term immunogenicity and reactogenicity of different SARS-CoV-2 messenger RNA (mRNA) vaccines in a population ≥75 years of age in a randomized trial. METHODS:Participants were randomized to receive either BNT162b2 30 µg or a double booster dose of mRNA-1273, i.e., 100 µg, as the third and fourth vaccinations (first and second booster). The primary endpoint was the rate of a two-fold geometric mean titer (GMT) antibody increase 14 days after vaccination targeting the receptor binding domain (RBD) region of wild-type SARS-CoV-2. Secondary endpoints included neutralizing capacity against wild-type and 25 variants at 14 days (D14) and 12 months (M12). Safety was assessed by monitoring adverse events (AEs) for 7 days after vaccination. FINDINGS:Between November 2021 and September 2022, 322 participants received a SARS-CoV-2 vaccine as a first (Part A) or second booster (Part B). Primary endpoint results have been published previously. In Part A, it was reached by 100% of participants in both vaccine arms, with a higher GMT increase in the mRNA-1273 arm (ratio, 1.64). At M12, the GMT of anti-RBD immunoglobulin G (IgG) was slightly higher than at D14 (9319.7 vs 8568.4 IU/mL) in the BNT162b2 arm, while in the mRNA-1273 arm, the GMT was equal (14,163.8 vs 14,266.7 IU/mL at D14). In Part B, the primary endpoint was reached by 78.5% of participants in the BNT162b2 and 87.2% in the mRNA-1273 arm (P = 0.056), respectively, with a higher GMT increase of anti-RBD IgG for mRNA-1273 (ratio, 1.38). At M12, GMT of anti-RBD IgG was markedly lower than at D14 (9962 vs 15,248.2 IU/mL) in the BNT162b2 arm as well as in the mRNA-1273 arm (12,024.3 vs 21,325.6 IU/mL). Higher neutralizing capacity in individuals who received a booster with mRNA-1273 was detected against wild-type and 15 of 25 tested variants. Fewer participants in the mRNA-1273 arm had vaccine-related AEs (29.6% vs 38.5%), but severity was more frequently grade 2 (n = 38, 28.1% vs n = 22, 16.3%). INTERPRETATION:Long-term serological immunogenicity and virus neutralization capacity in participants ≥75 years of age were numerically better with an mRNA-1273 100 µg booster, with a comparable safety profile.
BackgroundThe design and analysis of randomized clinical trials (RCTs) in filarial diseases such as onchocerciasis, loiasis, and mansonellosis pose unique statistical challenges, including skewed endpoints and limited sample sizes. This systematic review summarizes design and analysis approaches of RCTs conducted in these diseases with a focus on the statistical methodology.Methods and findingsA systematic search was conducted in PubMed and four trial registries to identify RCTs investigating treatments for onchocerciasis, loiasis, and mansonellosis published or registered between 2000 and 2024. We excluded studies focusing on new methods or pharmacokinetics, short reports, and Phase I trials. Forty-four studies met the inclusion/exclusion criteria (23 for onchocerciasis, 16 for loiasis, and 5 for mansonellosis), information was retrieved from the registries, the manuscripts and/or the study protocol. As primary efficacy endpoints, for onchocerciasis studies qualitative endpoints dominated, while quantitative endpoints were more frequently observed for loiasis and mansonellosis. The most frequently reported hypothesis tests for the primary endpoint were the Mann-Whitney U and the chi-squared tests. We found considerable heterogeneity between trials - not only in study-specific parameters such as the number of arms, type of blinding or control group - but also in design parameters or attributes that could be standardized within each disease across studies with similar objectives, such as the primary endpoint, length of follow-up, the analysis method and the primary analysis population.ConclusionsSeveral trials were well-planned with detailed information provided in either the manuscript or the registry. However, for some trials, information was sparse or incomplete, indicating a need for more structured and transparent reporting. Adopting established frameworks such as CONSORT and ICH E9 (R1) estimand approach would enhance transparency and better align trial objectives, analyses, and reported conclusions.
We propose a frequentist adaptive phase 2 trial design to evaluate the safety and efficacy of three treatment regimens (doses) compared to placebo for four types of helminth (worm) infections. This trial will be carried out in four Subsaharan African countries from spring 2025. Since the safety of the highest dose is not yet established, the study begins with the two lower doses and placebo. Based on safety and early efficacy results from an interim analysis, a decision will be made to either continue with the two lower doses or drop one or both and introduce the highest dose instead. This design borrows information across baskets for safety assessment, while efficacy is assessed separately for each basket. The proposed adaptive design addresses several key challenges: (1) The trial must begin with only the two lower doses because reassuring safety data from these doses is required before escalating to a higher dose. (2) Due to the expected speed of recruitment, adaptation decisions must rely on an earlier, surrogate endpoint. (3) The primary outcome is a count variable that follows a mixture distribution with an atom at 0. To control the familywise error rate in the strong sense when comparing multiple doses to the control in the adaptive design, we extend the partial conditional error approach to accommodate the inclusion of new hypotheses after the interim analysis. In a comprehensive simulation study we evaluate various design options and analysis strategies, assessing the robustness of the design under different design assumptions and parameter values. We identify scenarios where the adaptive design improves the trial's ability to identify an optimal dose. Adaptive dose selection enables resource allocation to the most promising treatment arms, increasing the likelihood of selecting the optimal dose while reducing the required overall sample size and trial duration.
BACKGROUND:This open-label randomised phase 2 study aimed to determine whether a single versus two-dose BNT162b2 primary vaccination regimen in children 5 to 11 years old with prior SARS-CoV-2 infection was non-inferior in terms of immunogenicity and superior in terms of safety and reactogenicity. METHODS:Participants were randomly assigned (1:1) to receive either one or two doses, spaced 3 to 12-weeks. The primary endpoint was geometric mean ratio (GMR) of neutralizing antibodies against wild-type SARS-CoV-2 at 28 days post-vaccination with non-inferiority margin defined as a 1.5-fold change in geometric mean titers (GMT). Secondary endpoints included safety and reactogenicity profile and immunogenicity up to 12 months against wild-type and Variants of Concern (VOCs). RESULTS:In total 31 participants from 3 European countries (median age 9, IQR7-10) were enrolled from May 2022 to January 2024, when the trial was prematurely terminated due to declining interest in COVID-19 vaccination among age-eligible children. Of these, 15 received two doses, and 16 received one. At day 28, GMT of neutralizing antibodies against wild-type SARS-CoV-2 was 1801.1 IU/mL(95%CI:1357.9-2388.9) in the two-dose arm and 1715.5 IU/mL(95%CI:1064.2-2765.4) in the single-dose arm. However, the non-inferiority of the single-dose could not be demonstrated (GMR:0.9; 95%CI:0.5-1.6). Titers remained above 100 IU/mL in both groups at 6 and 12 months. Both schedules elicited high anti-RBD IgG titers against wild-type and neutralizing titers against BA.5 variant at day 28. Eight participants (53%) in the two-dose arm and five (31%) in the single-dose reported a systemic adverse event grade ≥ 2 (P = 0.18) within 7 days of vaccination. CONCLUSIONS:Both regimens induced robust and sustained immune responses consistent with the possibility that, in children with prior infection, a single dose functions immunologically as a booster of the humoral response. However, the premature termination renders the primary non-inferiority comparison statistically underpowered. The vaccine was well tolerated in both groups. EudraCT registration: 2021-005043-71.
A simple device for balancing for a continuous covariate in clinical trials is to stratify by whether the covariate is above or below some target value, typically the predicted median. The object is to improve balance of the covariate and hence efficiency of the treatment estimate particularly if the trial is small, as may be the case if the disease in question is rare. This raises an issue as to which model should be used for modelling the effect of treatment on the outcome variable, Y. Should one fit, the stratum indicator, S, the continuous covariate, X, both or neither? When a covariate is added to a linear model there are three consequences for inference: (a) the mean square error effect, (b) the variance inflation factor and (c) second order precision. We consider that it is valuable to consider these three factors separately, even if, ultimately, it is their joint effect that matters. We present some simple theory, concentrating in particular on the variance inflation factor, that may be used to guide trialists in their choice of model. We also consider the case where the precise form of the relationship between the outcome and the covariate is not known. We conclude by recommending that the continuous covariate should always be in the model but that, depending on circumstances, there may be some justification in fitting the stratum indicator also.
While well-established methods for time-to-event data are available when the proportional hazards assumption holds, there is no consensus on the best inferential approach under non-proportional hazards (NPH). However, a wide range of parametric and non-parametric methods for testing and estimation in this scenario have been proposed. To provide recommendations on the statistical analysis of clinical trials where non-proportional hazards are expected, we conducted a simulation study under different scenarios of non-proportional hazards, including delayed onset of treatment effect, crossing hazard curves, subgroups with different treatment effects, and changing hazards after disease progression. We assessed type I error rate control, power, and confidence interval coverage, where applicable, for a wide range of methods, including weighted log-rank tests, the MaxCombo test, summary measures such as the restricted mean survival time (RMST), average hazard ratios, and milestone survival probabilities, as well as accelerated failure time regression models. We found a trade-off between interpretability and power when choosing an analysis strategy under NPH scenarios. While analysis methods based on weighted logrank tests typically were favorable in terms of power, they do not provide an easily interpretable treatment effect estimate. Also, depending on the weight function, they test a narrow null hypothesis of equal hazard functions, and rejection of this null hypothesis may not allow for a direct conclusion of treatment benefit in terms of the survival function. In contrast, non-parametric procedures based on well-interpretable measures like the RMST difference had lower power in most scenarios. Model-based methods based on specific survival distributions had larger power; however, often gave biased estimates and lower than nominal confidence interval coverage. The application of the studied methods is illustrated in a case study with reconstructed data from a phase III oncologic trial.
In the context of clinical research, computational models have received increasing attention over the past decades. In this systematic review, we aimed to provide an overview of the role of so-called in silico clinical trials (ISCTs) in medical applications. Exemplary for the broad field of clinical medicine, we focused on in silico (IS) methods applied in drug development, sometimes also referred to as model informed drug development (MIDD). We searched PubMed and ClinicalTrials.gov for published articles and registered clinical trials related to ISCTs. We identified 202 articles and 48 trials, and of these, 76 articles and 19 trials were directly linked to drug development. We extracted information from all 202 articles and 48 clinical trials and conducted a more detailed review of the methods used in the 76 articles that are connected to drug development. Regarding application, most articles and trials focused on cancer and imaging-related research while rare and pediatric diseases were only addressed in 14 articles and 5 trials, respectively. While some models were informed combining mechanistic knowledge with clinical or preclinical (in-vivo or in-vitro) data, the majority of models were fully data-driven, illustrating that clinical data is a crucial part in the process of generating synthetic data in ISCTs. Regarding reproducibility, a more detailed analysis revealed that only 24
Hybrid randomized controlled trials (hybrid RCTs) integrate external control data, such as historical or concurrent data, with data from randomized trials. While numerous frequentist and Bayesian methods, such as the test-then-pool and Meta-Analytic-Predictive prior, have been developed to account for potential disagreement between the external control and randomized data, they cannot ensure strict type I error rate control. However, these methods can reduce biases stemming from systematic differences between external controls and trial data. A critical yet underexplored issue in hybrid RCTs is the prespecification of external data to be used in analysis. The validity of statistical conclusions in hybrid RCTs depends on the assumption that external control selection is independent of historical trials outcomes. In practice, historical data may be accessible during the planning stage, potentially influencing important decisions, such as which historical datasets to include or the sample size of the prospective part of the hybrid trial, thus introducing bias. Such data-driven design choices can be an additional source of bias, which can occur even when historical and prospective controls are exchangeable. Through a simulation study, we quantify the biases introduced by outcome-dependent selection of historical controls in hybrid RCTs using both Bayesian and frequentist approaches, and discuss potential strategies to mitigate this bias. Our scenarios consider variability and time trends in the historical studies, distributional shifts between historical and prospective control groups, sample sizes and allocation ratios, as well as the number of studies included. The impact of different rules for selecting external controls is demonstrated using a clinical trial example.
BACKGROUND:Preoperative anaemia is a major risk factor for perioperative morbidity. Because iron deficiency is widely assumed to be the main cause of anaemia in surgical patients, treatment efforts have focused mostly on iron supplementation. However, the aetiology of anaemia is multifactorial. To further understand the underlying causes and consider a comprehensive approach to anaemia management, we studied the prevalence and aetiology of preoperative anaemia in patients undergoing major surgery. METHODS:This prospective, multicentre, observational cohort study was done in 79 hospitals in 20 countries on five continents; patients were aged at least 18 years, undergoing major surgery, and had a postoperative in-hospital stay of at least 24 h. Patients donating autologous blood before surgery were excluded. Data were extracted from the electronic hospital information system and from self-reported information during preoperative examination. The primary outcomes were the prevalence of anaemia, defined as haemoglobin less than 120 g/L for women and less than 130 g/L for men, analysed in all participants, and the aetiology of anaemia, analysed only in patients with anaemia for whom aetiology could be confirmed. The study was registered with ClinicalTrials.gov (NCT03978260) and is complete. FINDINGS:Between Aug 26, 2019, and Dec 26, 2021, 2830 patients undergoing major surgery were recruited and 2702 patients were included in the analysis (1417 [52·4%] were male, 1279 [47·3%] were female, and six [0·2%] had gender dysphoria). Overall, 856 (31·7%, 95% CI 31·2-32·2) patients had preoperative anaemia. Among 782 patients with preoperative anaemia, for whom the presence of at least one aetiology could be confirmed, 432 (55·2%, 48·9-61·6) had iron deficiency, 60 (7·7%, 6·6-8·7) had vitamin B12 deficiency, 113 (14·5%, 12·2-16·7) had folate deficiency, 68 (8·7%, 8·1-9·3) had chronic kidney disease, and 48 (6·1%, 4·5-7·8) had anaemia resulting from another cause; patients could be assigned to multiple aetiologies. Across male and female sex, all age groups, and all countries, iron deficiency was the aetiology with the highest prevalence. INTERPRETATION:The prevalence of preoperative anaemia in patients in this study who were undergoing major surgery is high. Iron deficiency is the primary cause of this anaemia; however, the substantial prevalence of vitamin B12 and folate deficiencies demands immediate attention and action. FUNDING:None. TRANSLATIONS:For the Afrikaans, Albanian, Arabic, French, German, Greek, Italian, Korean, Portuguese, Romanian, Slovenian, Spanish, and Turkish translations of the abstract see Supplementary Materials section.
Major depressive disorder (MDD) is one of the leading causes of disability globally. Despite its prevalence, approximately one-third of patients do not benefit sufficiently from available treatments, and few new drugs have been developed recently. Consequently, more efficient methods are needed to evaluate a broader range of treatment options quickly. Platform trials offer a promising solution, as they allow for the assessment of multiple investigational treatments simultaneously by sharing control groups and by reducing both trial activation and patient recruitment times. The objective of this simulation study was to support the design and optimisation of a phase II superiority platform trial for MDD, considering the disease-specific characteristics. In particular, we assessed the efficiency of platform trials compared to traditional two-arm trials by investigating key design elements, including allocation and randomisation strategies, as well as per-treatment arm sample sizes and interim futility analyses. Through extensive simulations, we refined these design components and evaluated their impact on trial performance. The results demonstrated that platform trials not only enhance efficiency but also achieve higher statistical power in evaluating individual treatments compared to conventional trials. The efficiency of platform trials is particularly prominent when interim futility analyses are performed to eliminate treatments that have either no or a negligible treatment effect early. Overall, this work provides valuable insights into the design of platform trials in the superiority setting and underscores their potential to accelerate therapy development in MDD and other therapeutic areas, providing a flexible and powerful alternative to traditional trial designs.
The graph based approach to multiple testing is an intuitive method that enables a study team to represent clearly, through a directed graph, its priorities for hierarchical testing of multiple hypotheses, and for propagating the available type-1 error from rejected or dropped hypotheses to hypotheses yet to be tested. Although originally developed for single stage non-adaptive designs, we show how it may be extended to two-stage designs that permit early identification of efficacious treatments, adaptive sample size re-estimation, dropping of hypotheses, and changes in the hierarchical testing strategy at the end of stage one. Two approaches are available for preserving the family wise error rate in the presence of these adaptive changes; the p-value combination method, and the conditional error rate method. In this investigation we will present the statistical methodology underlying each approach and will compare the operating characteristics of the two methods in a large simulation experiment.
Shared controls in platform trials comprise concurrent and non-concurrent controls. For a given experimental arm, non-concurrent controls refer to data from patients allocated to the control arm before the arm enters the trial. The use of non-concurrent controls in the analysis is attractive because it may increase the trial's power of testing treatment differences while decreasing the sample size. However, since arms are added sequentially in the trial, randomization occurs at different times, which can introduce bias in the estimates due to time trends. In this article, we present methods to incorporate non-concurrent control data in treatment-control comparisons allowing for time trends. We focus on frequentist approaches that model the time trend and Bayesian strategies that limit the borrowing level depending on the heterogeneity between concurrent and non-concurrent controls. We examine the impact of time trends, overlap between experimental treatment arms, and entry times of arms in the trial on the operating characteristics of treatment effect estimators for each method under different patterns for the time trends. We argue under which conditions the methods lead to Type 1 error control and discuss the gain in power compared to trials only using concurrent controls by means of a simulation study in which methods are compared. Supplementary materials for this article are available online.
Background Noncompressible truncal hemorrhage is a major contributor to preventable deaths in trauma patients and, despite advances in emergency care, still poses a big challenge. Objectives This study aimed to assess the clinical efficacy of trauma resuscitation care incorporating Resuscitative Endovascular Balloon Occlusion of the Aorta (REBOA) compared to standard care for managing uncontrolled torso or lower body hemorrhage. Methods This study utilized a target trial design with a matched case-control methodology, emulating randomized 1 : 1 allocation for patients receiving trauma resuscitation care with or without the use of REBOA. The study was conducted at a high-volume trauma center in Southern Austria, including trauma patients treated between January 2019 and October 2023, aged 16 and above, with suspected severe non-compressible torso hemorrhage. The primary outcome was 30-day in-hospital mortality. Secondary outcomes were in-hospital mortality rates at 3, 6, 24 h, and 90 days, need for damage control procedures, time to these procedures, computed tomography (CT) scan rates during resuscitation, complications, length of intensive care and in-hospital stay, and causes of death. Results Median age was 55 [interquartile range (IQR) 42-64] years. Median total injury severity, assessed by Injury Severity Score, was 46.5 (IQR: 43-57). There was no significant difference in 30-day in-hospital mortality between groups [9/22 (41%) vs. 9/22 (41%), odds ratio: 1.00, 95% confidence interval (CI): 0.3-3.36, P > 0.999]. Lower mortality rates within 3, 6, and 24 h were observed in the REBOA group; in a Cox proportional hazards model, hazard ratio (95% CI) for mortality in the REBOA group was 0.87 (0.35-2.15). Timing to damage control procedures did not significantly differ between groups, although patients in the REBOA group underwent significantly more CT scans. Bleeding was cited as the main cause of death less frequently in the REBOA group. Conclusion In severely injured patients presenting with possible major non-compressible torso hemorrhage, a systematically implemented resuscitation strategy including REBOA during the initial hospital phase, is not associated with significant changes in mortality.