AbstractReplication studies are recognized as essential to the scientific process. Numerous measures have been developed to quantify replication success. Most measures were developed for post hoc replications, in which the primary study has been conducted and sometimes assumed to show a specific result (e.g. statistical significance). Consequently, methodological studies have focused on evaluating replication success measures for those replications. However, recent work emphasizes the value of prospective replications, in which primary and replication studies are planned simultaneously. Such replications allow researchers to control study characteristics and thus investigate which characteristics cause effect heterogeneity. This study provides replication success measures for prospective replications, and guidelines for choosing between them. We present a taxonomy of measures based on research questions they address and evaluate existing frequentist and Bayesian approaches for their applicability to prospective replications. We illustrate their application using an example from social psychology. In simulations, we compare the statistical properties of measures that aim at the same research question. Results indicate that there is almost always a trade-off between error types. Thus, no single measure emerged as always clearly superior. We highlight the assumptions and strengths of each measure and offer recommendations for choosing a measure based on replication goals.
Background and objective: Incorporation of external information in clinical study designs has received increasing attention in recent years. When external data can be considered consistent with the data collected in the current trial, efficiency gains and a reduction in the required sample size can be achieved. This makes it particularly attractive when, e.g., recruitment of patients is difficult. However, unexpected inconsistency cannot be ruled out, and various robust borrowing approaches have been developed to limit potential losses. These include, for example, the power prior, meta-analytic predictive prior, robust mixture prior, and Bayesian hierarchical model. However, their use in actual clinical trials remains underexplored. In this scoping review, we investigated the use of external information in Phase II clinical trials. Methods: Publications were identified by literature search in PubMed central, the Cochrane Central Register of Controlled Trials and complemented with a citation search in some of the main methodological papers on Bayesian borrowing in clinical trials. Results: Of the 1282 articles retrieved, 37 were eventually included. We investigated, among others, whether and how the suitability of external data was assessed, the borrowing method used, and the specification and justification of tuning parameters that influence the degree of borrowing. Fixed downweighting was found to be the most commonly used method for external information borrowing, with only slightly more than a third of the studies providing details on its implementation. Conclusion: Overall, we found poor reporting of various aspects relating to Bayesian borrowing, signaling a gap between the methodological developments and the practical uptake and reporting of such methods.
OBJECTIVE:Recent studies have raised concerns arising from the exploitation ("mining") of public health databases for low-quality, mass-produced papers. However, it remains challenging to disambiguate whether such papers originate from paper mills (commercial entities that sell authorships on mass-produced papers) or from the uncoordinated action of individuals facilitated by AI tools and templated workflows. Our study aims to address this question for one particular database, the Global Burden of Disease Study (GBD). We selected this database after noticing that one of our papers on Bayesian age-period-cohort models has recently been experiencing a rapid surge in geographically clustered citations from GBD papers with Chinese affiliations. METHODS:We collected bibliometric and article-level metadata from GBD papers to search for indicators of mass-produced research. Moreover, we assessed 713 full-text articles for reported R versions, availability of code and data, and declaration of generative AI use. For 180 articles, we qualitatively screened the figures for graphical similarities. Finally, we conducted an exploratory scoping investigation of online platforms (social media sites, vendor websites) dedicated to do-it-yourself workflows for secondary analyses of public health data. RESULTS:Although we cannot rule out paper mill involvement, our findings suggest that the geographically clustered increase in GBD publications from China is at least partially driven by independent authors. The wide variety of R versions listed in 477 articles points against centralized paper production. Moreover, despite broad graphical similarities in figure styles that suggest the use of shared visualization tools, substantial variation in ancillary details suggest independent authors finalizing figures. This is corroborated by the identification of an online ecosystem of proprietary tools and services specializing in streamlined do-it-yourself workflows for conducting, writing, and publishing secondary analyses of public health data. CONCLUSION:Appropriate efforts should be directed towards evaluating the quality of the identified workflows. Stakeholders in scientific integrity should monitor online platforms dedicated to the rapid production of papers, especially in light of the increasing focus on AI-assisted workflows. Code sharing should be mandated for data-driven secondary analyses. Paywalls and proprietary software licenses hinder transparency and reusability, underscoring the importance of free and open-source software for trustworthy and reproducible research.
BACKGROUND:Reproducing published findings from clinical trials is a critical component of scientific transparency, yet it remains a challenging and under-practiced task. Despite increasing emphasis on reproducibility and data reuse in research policies, only few real-world examples exist where several teams have reproduced complex analyses using clinical trial data. In this case study, the aim was to reproduce the key findings of a high-impact clinical trial on rectal cancer treatment using shared trial data. METHOD:We organized a multi-team datathon, where each team was provided with the same dataset and supporting material, and was tasked to reproduce the results of the CAO/ARO/AIO-04 trial, with optional additional analyses. We contacted the original investigators for access and reuse of the data, as well as information on the clinical and scientific aspects of the study. RESULTS:Five teams used R or Python to reproduce the statistical results, and the corresponding scripts can be found on Gitlab. The key findings on disease-free survival (DFS) were consistently reproduced by most teams, reinforcing confidence in the main trial conclusions. Result robustness was investigated using different analytical software or statistical models. Some challenges were encountered because supplementary material of the original study was not easily found. Minor reporting issues were also identified in the reproduced paper. CONCLUSIONS:Reproduction of a major oncology clinical trial confirmed the reliability of its main conclusions. Divergences highlighted reporting gaps-such as incomplete protocols and broken links-that future trials should address. This case study demonstrates the value of systematic reproducibility checks for the transparency of clinical research and the challenges in data sharing for reproducibility.
Abstract Background Large-scale estimates of animal-to-human drug translation and the study characteristics associated with successful translation remain limited. The expanding preclinical literature also challenges manual evidence synthesis. We developed a natural language processing (NLP) pipeline to structure and link preclinical and clinical evidence at scale. Methods In this retrospective meta-research study, we analysed more than 500,000 neuroscience-related animal drug studies from PubMed and linked them to clinical trial and regulatory approval data. NLP methods extracted drug, disease, and experimental design characteristics from abstracts and full texts. Translation was defined as progression to completed phase III/IV trials or regulatory approval. Logistic regression assessed associations between preclinical study characteristics and successful translation. Findings Among 291,624 drug entities identified in animal studies, 6·7% entered clinical development and 3·1% reached phase III/IV trials or regulatory approval. At the drug–disease level, 4·4% entered clinical development and 1·9% achieved translation. Restricting analyses to successfully linked ontology entities increased estimates to 11·3% and 4·1%, respectively. Male-only animal studies predominated, whereas reporting of randomisation, blinding, and sample size calculations remained limited. Testing across multiple species and reporting blinding were associated with higher odds of successful translation. Interpretation Only a minority of interventions tested in animals progress to advanced clinical development or regulatory approval. Greater species diversity and blinding were associated with improved translational success. NLP-based evidence synthesis may support scalable evaluation of translational research and identification of potentially modifiable research practices. Funding Swiss National Science Foundation, UZH Digital Entrepreneurship Fellowship, Universities Federation for Animal Welfare. Research in context Evidence before this study We searched the literature for studies quantifying large-scale animal-to-human translation and factors associated with successful translation. Existing work was mainly limited to specific diseases, interventions, or manually curated datasets, and large-scale linkage of animal and clinical evidence remained limited. Added value of this study We developed a natural language processing pipeline linking more than 500,000 animal studies to clinical trial and regulatory approval data. The study provides large-scale estimates of translation and identifies experimental characteristics associated with successful translation. Implications of all the available evidence The findings suggest that only a minority of interventions tested in animals progress to advanced clinical development or regulatory approval. Greater species diversity and reporting of blinding were associated with improved translation. Automated evidence synthesis may support more systematic evaluation of translational research practices.
Confirmatory multi-lab preclinical trials are a powerful experimental strategy to enable decisions to transition from preclinical to clinical settings. With their complexity, such study designs pose several challenges in analysing and reporting experiments. To address these, we convened an expert group of biostatisticians and biomedical scientists currently involved in such trials to summarise the most common scenarios. Furthermore, we incorporated statistical advice from existing clinical trials’ guidelines and adapted it into recommendations for future preclinical trials. We describe strategies on key topics such as calculating sample sizes, handling of differences between centres, and selecting relevant covariates. Additionally, we give guidance on statistical methods to account for lab effects and proper reporting of analyses. We embed this in a general discussion on remaining open questions to advance the analysis of preclinical confirmatory studies. The provided general, non-case-specific guidance serves as a conversation starter between scientists and statisticians to develop robust statistical analysis strategies for confirmatory multi-lab preclinical trials
Background: Outcome reporting bias (ORB) occurs when study outcomes are selectively reported based on their results. ORB potentially undermines the credibility and validity of meta-analyses and contributes to research waste by distorting overall treatment effects. ORB can be viewed as a missing data problem in which unreported study outcomes introduce bias. Despite the serious implications ORB poses, it remains an underrecognized issue, with only a few adjustment methods available. Methods: We propose an approach that addresses unreported study outcomes in meta-analyses through multiple imputation for univariate and multivariate meta-analysis. To assess the impact of ORB in meta-analyses, we apply our proposed methodology to real clinical data affected by ORB, and conduct a simulation study to evaluate the method's performance under a range of scenarios. Results: The proposed method provides bias-adjusted estimates under assumed selective non-reporting mechanisms. In the application to clinical data, ORB-adjusted estimates were systematically shifted towards less extreme treatment effects compared with naive analyses, highlighting the potential magnitude of ORB in practice. The simulation study shows that the extent of adjustment depends on the assumed selection mechanism and the degree of heterogeneity, with stronger selection leading to larger adjustment. Conclusions: Imputing unreported study outcomes provides a promising approach to address ORB in meta-analyses. The multivariate approach extends ORB adjustment to jointly model correlated outcomes, allowing borrowing of strength across outcomes. Overall, we propose a practical and flexible approach for evaluating the sensitivity of univariate and multivariate meta-analytic conclusions to ORB.
BackgroundTransparency in randomized controlled trials (RCTs) has substantially improved in recent years, notably through trial registration and public availability of protocols and statistical analysis plans (SAPs). However, the reporting of protocol and SAPs modifications remains insufficiently standardized. As a result, even when these documents are publicly available, it is often challenging and time-consuming to identify what changes were made, why they were implemented, and whether they may affect the trustworthiness of the trial results.ArgumentsIn this paper, we advocate for the development of a consensus-based framework for protocol modifications in RCTs. This need arises from the inherent tension between the necessity and the risks of protocol modifications. On the one hand, such modifications are often essential to address unforeseen operational, scientific, or ethical challenges. On the other hand, they may introduce bias and undermine confidence in trial findings, particularly when changes are data-driven or insufficiently justified. Although major transparency initiatives have strengthened trial reporting, important gaps persist. We review empirical evidence demonstrating the prevalence and nature of such modifications and discuss their potential implications for the validity, interpretation, and credibility of trial findings. Furthermore, readers, reviewers, and decision-makers face substantial challenges in identifying, understanding, and evaluating the potential impact of protocol changes. In the absence of standardized reporting, key information remains dispersed across multiple documents, placing an unreasonable burden on stakeholders to identify, interpret, and assess protocol modifications and their implications for the credibility of trial results.ConclusionsStandardized and transparent reporting of protocol modifications is essential to ensure that their nature, timing, and rationale can be clearly understood and critically evaluated. We therefore advocate for the development of a consensus-based reporting framework, informed by a Delphi process, to improve transparency, facilitate critical appraisal, and strengthen confidence in RCT findings.
Over the past years, the concept of open research data (ORD) has gained traction as part of broader Open Science initiatives. The benefits of ORD, such as increased cost-effectiveness, transparency, and visibility, are well documented. However, researchers face barriers, which may be perceived rather than real, hindering the adoption of ORD practices. To address this challenge, we propose using ORD support services as sustainable enablers to stimulate cultural change around ORD. We engaged stakeholders across the University of Zurich and the Swiss ORD community, differentiating between researchers and ORD experts, to identify which services would best serve as sustainable enablers. After defining ORD support services and categorizing them into six key areas, we conducted surveys and interviews to gather insights on service preferences and barriers to ORD adoption. Among researchers, we identified a trend toward simpler and lower-resource services, highlighting the need for user-friendly and easily accessible support. ORD experts emphasized the importance of professional data stewardship, robust research data management (RDM) practices, and customized support to address discipline-specific needs. By combining survey and interview results, we provide a detailed overview of stakeholders’ ideas and suggestions for each proposed support area. Our study results in recommendations for academic institutions aiming to stimulate a cultural shift toward ORD. By focusing on findable, accessible, and user-friendly services, equipping researchers with fundamental RDM skills, and professionalizing data stewardship to provide customized support, institutions can foster the adoption of ORD practices. Ultimately, these measures can enhance the impact and reproducibility of scientific research.
The Bayes factor, the data-based updating factor from prior to posterior odds, is a principled measure of relative evidence for two competing hypotheses. It is naturally suited to sequential data analysis in settings such as clinical trials and animal experiments, where early stopping for efficacy or futility is desirable. However, designing such studies is challenging because computing design characteristics, such as the probability of obtaining conclusive evidence or the expected sample size, typically requires computationally intensive Monte Carlo simulations, as no closed-form or efficient numerical methods exist. To address this issue, we extend results from classical group sequential design theory to sequential Bayes factor designs. The key idea is to derive Bayes factor stopping regions in terms of the z-statistic and use the known distribution of the cumulative z-statistics to compute stopping probabilities through multivariate normal integration. The resulting method is fast, accurate, and simulation-free. We illustrate it with examples from clinical trials, animal experiments, and psychological studies. We also provide an open-source implementation in the bfpwr R package. Our method makes exploring sequential Bayes factor designs as straightforward as classical group sequential designs, enabling experiments to rapidly design informative and efficient experiments.
Reducing the number of experimental units is one of the three pillars of the 3R principles (Replace, Reduce, Refine) in animal research. At the same time, statistical error rates need to be controlled to enable reliable inferences and decisions. This paper proposes to adopt diagnostic likelihood ratios and the diagnostic odds ratio to statistical hypothesis tests and to adjust it for sample size to obtain a novel measure to quantify for the evidentiary value of one experimental unit. The experimental unit information index (EUII) is based on power, Type-I error and sample size, and has attractive interpretations both in terms of frequentist error rates and Bayesian posterior odds. We introduce the EUII in simple statistical test settings and show that its asymptotic value depends only on the assumed relative effect size under the alternative. We then extend the definition to adaptive designs where early stopping for efficacy or futility may cause reductions in sample size. Application to group-sequential designs show the usefulness of the approach when the goal is to maximize the evidentiary value of one experimental unit. A reanalysis of 2738 animal experiments with simulated results from (post-hoc) interim analyses illustrates the possible savings in sample size.
P-value functions are modern statistical tools that unify effect estimation and hypothesis testing and can provide alternative point and interval estimates compared to standard meta-analysis methods, using any of the many p-value combination procedures available (Xie et al., 2011, JASA). We provide a systematic comparison of different combination procedures, both from a theoretical perspective and through simulation. We show that many prominent p-value combination methods (e.g. Fisher's method) are not invariant to the orientation of the underlying one-sided p-values. Only Edgington's method, a lesser-known combination method based on the sum of p-values, is orientation-invariant and still provides confidence intervals not restricted to be symmetric around the point estimate. Adjustments for heterogeneity can also be made and results from a simulation study indicate that Edgington's method can compete with more standard meta-analytic methods.
Determining an appropriate sample size is a critical element of study design, and the method used to determine it should be consistent with the planned analysis. When the planned analysis involves Bayes factor hypothesis testing, the sample size is usually desired to ensure a sufficiently high probability of obtaining a Bayes factor indicating compelling evidence for a hypothesis, given that the hypothesis is true. In practice, Bayes factor sample size determination is typically performed using computationally intensive Monte Carlo simulation. Here, we summarize alternative approaches that enable sample size determination without simulation. We show how, under approximate normality assumptions, sample sizes can be determined numerically, and provide the R package bfpwr for this purpose. Additionally, we identify conditions under which sample sizes can even be determined in closed-form, resulting in novel, easy-to-use formulas that also help foster intuition, enable asymptotic analysis, and can also be used for hybrid Bayesian/likelihoodist design. Furthermore, we show how power and sample size can be computed without simulation for more complex analysis priors, such as Jeffreys-Zellner-Siow priors or non-local normal moment priors. Case studies from medicine and psychology illustrate how researchers can use our methods to design informative yet cost-efficient studies.
BACKGROUND:The standard regulatory approach to assess replication success is the two-trials rule, requiring both the original and the replication study to be significant with effect estimates in the same direction. The sceptical p-value was recently presented as an alternative method for the statistical assessment of the replicability of study results. METHODS:We review the statistical properties of the sceptical p-value and compare those to the two-trials rule. We extend the methodology to non-inferiority trials and describe how to invert the sceptical p-value to obtain confidence intervals. We illustrate the performance of the different methods using real-world evidence emulations of randomized controlled trials (RCTs) conducted within the RCT DUPLICATE initiative. RESULTS:The sceptical p-value depends not only on the two p-values, but also on sample size and effect size of the two studies. It can be calibrated to have the same Type-I error rate as the two-trials rule, but has larger power to detect an existing effect. In the application to the results from the RCT DUPLICATE initiative, the sceptical p-value leads to qualitatively similar results than the two-trials rule, but tends to show more evidence for treatment effects compared to the two-trials rule. CONCLUSION:The sceptical p-value represents a valid statistical measure to assess the replicability of study results and is useful in the context of real-world evidence emulations.
Continuous outcome measurements truncated by death present a challenge for the estimation of unbiased treatment effects in randomized controlled trials (RCTs). One way to deal with such situations is to estimate the survivor average causal effect (SACE), but this requires making nontestable assumptions. Motivated by an ongoing RCT in very preterm infants with intraventricular hemorrhage, we performed a simulation study to compare an SACE estimator with complete case analysis (CCA) and analysis after multiple imputation of missing outcomes. We set up nine scenarios combining positive, negative, and no treatment effect on the outcome (cognitive development) and on survival at 2 years of age. Treatment effect estimates from all methods were compared in terms of bias, mean squared error, and coverage with regard to two true treatment effects: the treatment effect on the outcome used in the simulation and the SACE, which was derived by simulation of both potential outcomes per patient. Despite targeting different estimands (principal stratum estimand, hypothetical estimand), the SACE-estimator and multiple imputation gave similar estimates of the treatment effect and efficiently reduced the bias compared to CCA. Also, both methods were relatively robust to omission of one covariate in the analysis, and thus violation of relevant assumptions. Although the SACE is not without controversy, we find it useful if mortality is inherent to the study population. Some degree of violation of the required assumptions is almost certain, but may be acceptable in practice.
Recent large-scale replication projects (RPs) have estimated concerningly low reproducibility rates. Further, they reported substantial degrees of shrinkage of effect size, where the replication effect size was found to be, on average, much smaller than the original effect size. Within these RPs, the included original-replication study-pairs can vary with respect to aspects of study design, outcome measures, and descriptive features of both original and replication study population and study team. This often results in between-study-pair heterogeneity, i.e., variation in effect size differences across study-pairs that goes beyond expected statistical variation. When broader claims about the reproducibility of an entire field are based on such heterogeneous data, it becomes imperative to conduct a rigorous analysis of the amount and sources of shrinkage and heterogeneity within and between included study-pairs. Methodology from the meta-analysis literature provides an approach for quantifying the heterogeneity present in RPs with an additive or multiplicative parameter. Meta-regression methodology further allows for an investigation into the sources of shrinkage and heterogeneity. We propose the use of location-scale meta-regressions as a means to directly relate the identified characteristics with shrinkage (represented by the location) and heterogeneity (represented by the scale). This provides valuable insights into drivers and factors associated with high or low reproducibility rates and therefore contextualises results of RPs. The proposed methodology is illustrated using publicly available data from the Replication Project Psychology and the Replication Project Experimental Economics. All analysis scripts and data are available online.
Meta-analysis can be formulated as combining p-values across studies into a joint p-value function, from which point estimates and confidence intervals can be derived. We extend the meta-analytic estimation framework based on combined p-value functions to incorporate uncertainty in heterogeneity estimation by employing a confidence distribution approach. Specifically, the confidence distribution of Edgington's method is adjusted according to the confidence distribution of the heterogeneity parameter constructed from the generalized heterogeneity statistic. Simulation results suggest that 95
Reproducibility is recognized as essential to scientific progress and integrity. Replication studies and large-scale replication projects, aiming to quantify different aspects of reproducibility, have become more common. Since no standardized approach to measuring reproducibility exists, a diverse set of metrics has emerged and a comprehensive overview is needed. We conducted a scoping review to identify large-scale replication projects that used metrics and methodological papers that proposed or discussed metrics. The project list was compiled by the authors. For the methodological papers, we searched Scopus, MedLine, PsycINFO and EconLit. Records were screened in duplicate against pre-defined inclusion criteria. Demographic information on included records and information on reproducibility metrics used, suggested or discussed was extracted. We identified 49 large-scale projects and 97 methodological papers and extracted 50 metrics. The metrics were characterized based on type (formulas and/or statistical models, frameworks, graphical representations, studies and questionnaires, algorithms), input required and appropriate application scenarios. Each metric addresses a distinct question. Our review provides a comprehensive resource in the form of a ‘live’, interactive table for future replication teams and meta-researchers, offering support in how to select the most appropriate metrics that are aligned with research questions and project goals.
Response-adaptive randomization (RAR) methods can be used to adapt randomization probabilities based on accumulating data, aiming to increase the probability of allocating patients to effective treatments. A popular RAR method is Thompson sampling, which randomizes patients proportionally to the Bayesian posterior probability that each treatment is the most effective. However, its high variability can also increase the risk of assigning patients to inferior treatments and lead to inferential problems such as confidence interval undercoverage. We propose a principled method based on Bayesian hypothesis testing to address these issues: We introduce a null hypothesis postulating equal effectiveness of treatments. Bayesian model averaging then induces shrinkage toward equal randomization probabilities, with the degree of shrinkage controlled by the prior probability of the null hypothesis. Equal randomization and Thompson sampling arise as special cases when the prior probability is set to one or zero, respectively. Simulated and real-world examples illustrate that the method balances highly variable Thompson sampling with static equal randomization. A simulation study demonstrates that the method can mitigate issues with Thompson sampling and has comparable statistical properties to Thompson sampling with common ad hoc modifications such as power transformation and probability capping. We implement the method in the free and open-source R package brar, enabling experimenters to easily perform null hypothesis Bayesian RAR and support more effective randomization of patients.