
Objective Scoping reviews are often conducted as a precursor to a systematic review, yet guidance on how this relationship should be planned, executed and reported is limited. We aimed to illustrate, through worked examples, the ways in which scoping reviews can function as precursors to systematic reviews and to derive practical guidance for teams planning linked or sequential syntheses. Study Design and Setting We undertook a methodological analysis of eight contemporary evidence-synthesis papers and protocols spanning veterinary medicine, animal welfare, public health, mental health, occupational health and patient-and-public involvement, identified through the authors’ knowledge of the evidence-synthesis literature. Each paired a scoping review with a systematic review. For every example we charted the structural relationship between the two reviews and the specific functions performed by the scoping stage. Results Across the eight purposively selected examples, four structural configurations were identified: (i) integrated sequential designs reported within a single project or paper; (ii) an antecedent, separately published scoping review whose search and records are inherited and extended by a later systematic review; (iii) a single “parent” scoping review that seeds several targeted systematic reviews; and (iv) a complementary scoping review run alongside a systematic review to contextualise sparse evidence. Across configurations, scoping reviews delivered seven recurring functions: feasibility and volume assessment; prioritisation of a tractable question; concept and definition clarification; eligibility and appraisal-tool decisions; search architecture and a reusable record set; gap identification; and conceptual or logic frameworks to structure synthesis. Conclusion Scoping reviews can provide important foundational work for systematic reviews. Drawing on these illustrative examples, we propose a decision framework and reporting recommendations to help teams refine the specific purpose of a scoping review used as a precursor.
Objective To describe a co-designed Participant Information and Consent Form (PICF) developed to improve accessibility for ethnically diverse communities and consider implications for research representativeness. Approach A generic PICF was redesigned with a consumer advisory group and researchers and piloted in community-based workshops.Findings: Participants and facilitators perceived the revised PICF as easier to understand and more supportive of engagement with the consent process. Inaccessible consent materials may contribute to differential self-selection during recruitment. Conclusion Consent accessibility is both an ethical and methodological consideration. Co-design may help reduce avoidable participation barriers, although effects on recruitment and representativeness require further evaluation.
Objective To determine whether methodological trade-offs used in rapid reviews result in changes in treatment effects or GRADE certainty of the evidence (COE) compared to systematic reviews. Design and Setting We used three systematic reviews on acupuncture, education, and TENS for the management of chronic low back pain (LBP) conducted to inform the WHO guideline on non-surgical management of chronic primary LBP and simulated the conduct of three rapid reviews. We assessed the impact of three commonly used trade-offs: restricted database selection, single reviewer screening and single reviewer risk of bias (ROB) assessment. We meta-analysed studies included for each rapid review and compared them to the results of the WHO systematic reviews and identified when treatment effects would change recommendations. We used evidence profiles to update any changes to GRADE criteria and resulting COE. Results Compared to the WHO systematic reviews, our rapid reviews resulted in no changes to recommendations based on the meta-analyses in the acupuncture or TENS reviews and 15.4% in the education review. Further, between 2.9% to 51.6% of descriptive syntheses were not performed because the single study in the analysis were excluded from the rapid review, 25.8% of descriptive syntheses became meta-analyses because of false positives from single reviewer screening, and 6.45% to 13.33% of meta-analyses became descriptive syntheses because the review was reduced to a single study. The COE was affected in 13.3% in the TENS review and 51.6% in the education review due to descriptive syntheses that were not performed (exclusion of the single study in the analysis). In the acupuncture review, the COE changed in 23.5% of analyses. Conclusions Rapid review methodological trade-offs led to few changes in treatment effects and COE. However, it is difficult to predict how they will impact quantitative and descriptive syntheses and COE because multiple factors are involved, including the breadth of the research question, volume and quality of studies, and review topic. When possible, review authors should consider few comparators and outcomes, well defined eligibility criteria and using fewer methodological trade-offs.
Objective This review aimed to identify and characterise tools designed, adapted, or validated to assess risk of bias (RoB) in causal real-world evidence (RWE), with a view to informing their potential use in evidence-informed decision-making. Study Design We conducted a rapid scoping review of records published from 2015 onwards. Eligible records included studies and documents describing tools explicitly developed for RWE or tools applicable to causal inference. Tool characteristics were extracted, and their items/domains were mapped across four stages of study generation and four core bias domains. Results Eleven tools oriented towards causal questions were included, most of which were designed primarily for primary studies. Qualitative or categorical judgements were more common than numerical scoring. Coverage of the stages of study generation was uneven: study design was the most frequently addressed stage, whereas results presentation was the least represented. The data quality stage was addressed by ten tools, although its coverage varied across instruments. Across core bias domains, no tool was concentrated in a single domain; most distributed items across several domains, although the relative emphasis placed on selection and confounding bias varied across instruments. A substantial number of items/domains could not be classified under the predefined core bias domains. Conclusion Available RoB tools differ substantially in their structure, methodological emphasis, and the way they operationalise bias across stages of study generation and core bias domains. These differences should be considered when interpreting and comparing assessments conducted with different tools.
AIM:To identify and prioritise opportunities to strengthen methodological and process innovation in global evidence synthesis, ensuring timely, equitable, and high-quality syntheses that can inform decision-making across sectors. BACKGROUND:The Evidence Synthesis Infrastructure Collaborative (ESIC) was established to optimise global capacity for producing and using living evidence syntheses. Working Group 4 (WG4) focused on developing scalable, equitable, and context-sensitive methods to make evidence synthesis "radically more timely, relevant, and affordable." METHODS:WG4 applied the UK Design Council's Double Diamond Framework for Innovation through a six-month, multi-phase process involving 19 members from 12 countries. Methods included virtual workshops, surveys, interviews, and public consultations to map and assess the maturity level of capabilities across the evidence synthesis lifecycle. Solutions were prioritised using benefit-impact-effort criteria. RESULTS:Nine key capability gaps were identified, including limited coordination, inconsistent quality assurance, inequitable access to data, and insufficient cross-sector integration. Seventy-six potential solutions were generated, of which eleven were prioritised under the AACC framework: Access, Agility, Context, and Connections, such as establishing evidence support units, shared quality and certainty-of-evidence standards, mechanisms to improve equitable access to data and synthesis capacity in low-resource settings, and a global academy for evidence synthesis. DISCUSSION:The process highlighted strong alignment between WG4's recommendations and other ESIC working groups, underscoring a shared global vision. Challenges included limited consultation time and linguistic and contextual diversity across regions, though overall engagement was broad and constructive. CONCLUSION:A coordinated, equitable infrastructure for evidence synthesis, grounded in methodological rigour, inclusivity, explicit assessment of certainty and applicability of evidence, and responsiveness is essential to ensure evidence remains credible, accessible, and actionable for global policy and practice.
OBJECTIVES:Target trial emulation offers a structured framework to reduce bias in observational studies by emulating the design of a hypothetical target randomized controlled trial. Despite a growing interest for this framework, the implementation of target trial emulation is often inconsistent, with design flaws frequently introducing avoidable biases. Our objective was to develop a design assistant tool to provide structured guidance to researchers for planning and conducting target trial emulations. STUDY DESIGN AND SETTING:A multidisciplinary steering committee defined the scope, content, and structure of Tool for Implementing TArget Trial emulatioNs (TITAN). First, we reviewed the literature to identify concepts according to the steps of planning and designing a target trial emulation. Then, we developed the tool through an iterative process with two rounds of pilot testing and subsequent refinement by the steering committee. RESULTS:TITAN provides guidance to emulate a pragmatic two-arm trial assessing a pharmacological intervention. The tool is organized in two main sections. First, the users define the research question and the target trial. Then, they emulate the target trial using the available observational data. The tool provides warnings, in addition to suggestions, to potentially minimize avoidable biases and major methodological errors, particularly focusing on the alignment of study time points (eligibility, treatment assignment, and start of follow up) and related biases. The tool's output is a synthesis that includes a summary of the target trial, design of its emulation, and possible biases and solutions. CONCLUSION:TITAN could help researchers ensure that observational studies follow the key principles of target trial emulation.
Background Systematic reviews (SRs) support patient-centered decision-making and should therefore synthesize evidence about outcomes that are most important to patients. Rehabilitation is a core component of health systems worldwide, yet little is known about whether SRs in this field adequately address patient-important outcomes (PIOs). Objectives To support the patient-centeredness of evidence used to inform rehabilitation decisions, this study aimed to (i) determine the frequency of patient-important outcomes (PIOs) and how they are reported, and (ii) compare outcome reporting and synthesis methods between patient-important and surrogate outcomes in SRs of rehabilitation interventions. Methods We conducted a methodological study of SRs evaluating rehabilitation interventions. We searched MEDLINE, EMBASE, PsycINFO, and the Cochrane Library up to October 31, 2024. We included SRs published within the previous five years to assess current reporting and synthesis practices. We included SRs of randomised controlled trials evaluating rehabilitation interventions. SRs explicitly focused on rehabilitation interventions or targeted impairments in body functions. We distinguished outcome constructs from outcome measures and characterised outcomes across methodological and functional dimensions. Using outcome measures as the unit of analysis, reviewers classified each as PIO or surrogate, and extracted outcome-level reporting and synthesis characteristics. We estimated the prevalence of PIOs at the outcome and SR levels and compared reporting and synthesis characteristics using descriptive statistics, clustered logistic regression, and non-parametric tests. Results We included 98 SRs that reported 432 outcome constructs and 580 outcome measures. At the outcome level, PIOs accounted for 41.7% of outcome measures and 49.2% of outcome constructs. At the SR level, 74.5% of SRs reported at least one PIO measure. In 16% of SRs, at least one outcome construct identified as patient-important was summarised using a surrogate measure. Compared with surrogate outcomes, PIOs were less likely to be synthesised using meta-analysis (59.1% vs 70.7%), less often analysed as continuous variables (69.4% vs 93.8%), and more often measured using patient-reported measures (55.4% vs 2.2%). Conclusion PIOs are underrepresented in rehabilitation SRs and are synthesised differently from surrogate outcomes. This limits the extent to which evidence syntheses can support patient-centered decision-making and highlights the need to improve outcome selection in research.
Background and Objectives Rapid reviews are widely used to produce timely evidence for health and policy decisions, yet considerable heterogeneity exists in how they are justified, conducted, and reported. This study aims to characterize, across a stratified random sample of recently published rapid reviews in health, the rationales for choosing a rapid review approach, the methodological streamlining strategies employed, and the reporting guidance or checklists referenced or reported as followed. Study Design and Setting This is a prospectively registered study within a review (SWAR), a methodological design that treats published evidence syntheses as the data source. Eligible rapid reviews will be identified from MEDLINE (Ovid), Embase, the Cochrane Database, and Epistemonikos, restricted to publications from January 2025 onward. If more than 150 eligible reviews are identified, a random sample of 150 will be selected, stratified by year and broad review type. Methods We will screen records and extract data using a piloted, structured extraction form organized around three domains: rationale for conducting a rapid review; methodological streamlining strategies mapped against key review stages; and reporting guidance cited. Findings will be summarized using descriptive statistics and a structured narrative synthesis. Exploratory analyses will examine the frequency of core transparency items by year and review type. Conclusion This SWAR will provide a focused, empirically grounded characterization of recent rapid review practice across rationale, methods, and reporting. Findings are intended to inform future guidance development, support training, and establish a benchmark for monitoring changes in the conduct of rapid reviews and transparency.
BACKGROUND:Treatment withdrawal effects represent a transient increase in clinical risk following treatment discontinuation that exceeds what would be expected from loss of treatment effect alone. They have been described across multiple therapeutic areas, yet remain insufficiently recognized across evidence generation, synthesis, and translation into practice. METHODOLOGICAL CHALLENGE:When unacknowledged, withdrawal effects may distort estimates of benefit and harm in randomized trials and observational studies, obscure treatment-covariate interactions, propagate bias through evidence synthesis, and undermine the applicability of guideline recommendations to routine practice. CONCEPTUAL FRAMEWORK:We summarize methodological problems introduced by unrecognized withdrawal effects and propose a conceptual framework highlighting how existing methodological approaches can be applied more explicitly to identify, interpret, and account for withdrawal-related risk across the evidence pathway. PLAIN LANGUAGE SUMMARY:When some long-term treatments are stopped, patients may feel worse or develop health problems, beyond what would be expected from simply losing the benefit of the treatment. These are called treatment withdrawal effects. They are not rare and can occur after stopping many treatments, but they are not considered enough when studies are designed, analyzed, or used to make recommendations about treatment. This matters because research studies may give false impressions about how well treatments work, or how safe they are, if they do not consider what happens when treatments are interrupted or stopped. For example, if patients worsen soon after stopping a treatment, this may be interpreted as proof that continuing the treatment is beneficial, when it may instead be a withdrawal effect. Evidence reviews and clinical guidelines may then carry this problem forward, leading to recommendations based on an inaccurate picture of benefit and harm. This may mean that some treatments appear more useful or safer than they really are, while others may be undervalued. This article explains how withdrawal effects can affect research at different stages, from individual studies to evidence reviews and treatment guidelines. It also suggests ways to identify and account for these risks. Addressing withdrawal effects more systematically could help research give a more accurate picture of how treatments work, how safe they are, and how they should be used in practice. Because this is a complex problem across the whole research pathway, progress is likely to require international collaboration among researchers, trial teams, pharmaceutical and other healthcare developers, evidence reviewers, guideline developers, regulators, funders, clinicians, and patients.
BACKGROUND AND OBJECTIVES:When pursuing participatory methods with youth, health care research teams may have the opportunity to create partnerships with transgender and gender diverse youth (TGDY). While these partnerships are crucial to ensure the voices of TGDY are included throughout health domains, research teams must ensure this marginalized population is safely and meaningfully included. It is therefore important for research teams to be aware of key pitfalls that contribute to the marginalization of TGDY in health care and research, and certain strategies to counteract these pitfalls. METHODS:Using our lived and learned expertise, this commentary seeks to support research teams wishing to conduct participatory research that includes TGDY. RESULTS:We present an overview of the mechanisms through which TGDY are marginalized throughout health care and research, providing context around the sociocultural ostracization, institutional erasure, and epistemic injustice this group faces. Considering these issues, we provide suggested approaches to ensure partnerships with TGDY are safe and meaningful. CONCLUSION:Including TGDY voices throughout health care research is key to counteract the marginalization this group faces. Through safe and meaningful partnership, research teams across health care fields can contribute to this work. PLAIN LANGUAGE SUMMARY:This commentary aims to provide insight and guidance to health care research teams who wish to include transgender and gender diverse youth (TGDY) in their participatory structures (eg, advisory committees), but may not feel adequately informed about key considerations surrounding this group. To ensure partnerships with TGDY are safe and meaningful, it is helpful for research teams to better understand the pitfalls that have bolstered the marginalization of TGDY in health care and research, and how teams may avoid these pitfalls. We therefore use our lived and learned experience to provide an overview of these pitfalls, as well as strategies health care research teams can use to support safe and meaningful engagement when including TGDY partners in their work. Using these strategies, research teams from across health care domains may contribute to research and health care practices that are more inclusive and considerate of the realities of TGDY.
OBJECTIVE:The reliability of confirmatory clinical trials (i.e. the post-trial probability) depends on the pre-trial probability. During the COVID-19 pandemic, many early COVID-19 confirmatory trials were launched despite limited phase II evidence and weak mechanistic or preclinical support. Whether this accelerated development strategy reduced the reliability of therapeutic evidence remains uncertain. STUDY DESIGN AND SETTING:Overall empirical pre-trial plausibility was estimated at the trial level as the proportion of early randomized trials evaluating interventions that were ultimately considered effective in the reference network meta-analysis. For each trial, counterfactual statistical power was estimated using the reported sample size and the number of events. Bayes' theorem was used to calculate the post-trial probability of efficacy, based on sensitivity (i.e. trial power), specificity (i.e. 1 minus 0.025), and empirical pre-trial plausibility. We then compared the positive predictive values of COVID-19 and non-COVID-19 development pipelines. RESULTS:Of the 463 early trials evaluating 150 candidate molecules, 51 assessed interventions ultimately supported by the reference network meta-analysis, corresponding to an empirical trial-level pre-trial plausibility of 11%. The median power of COVID-19 trials was 9% (min-max: 6-100%). Accordingly, the probability that a statistically significant result reflected true efficacy was 31% (20%-83%), compared to 94% for non-COVID19 trials characterized by the conventional phase II-to-phase III success benchmark of 50%. With an empirical pre-trial plausibility of 11%, a replication framework consisting of two independent, 80%-powered trials yielded a positive predictive value of 99.2%. This contrasts with a single 90%-95%-powered trial, which yielded a positive predictive value of 81%-82%. CONCLUSION:The early COVID-19 trial ecosystem was frequently underpowered and conducted under conditions of low pre-trial plausibility, resulting in lower post-trial probabilities of true efficacy than in conventional development pathways. These findings highlight the need for stronger pre-trial evidence, coordinated trial prioritization, and replication-oriented strategies in future pandemic responses.
BACKGROUND AND OBJECTIVE:The Oxford Hip Score (OHS) is widely used as a patient reported outcome measure in clinical studies and national quality databases. The content of the questionnaire suggests a potential two-subscale structure, a structure sparsely evaluated. Thus, the aim was to evaluate the validity of both the original unidimensional structure (ie, one total score) of OHS and a two-dimensional structure reflecting separate "Pain" and "Function" subscales. METHODS:Preoperative and postoperative OHS item scores were available from a prospective cohort of patients undergoing total hip arthroplasty (THA). Eight homogeneous subsamples, each comprising 100 randomly chosen patients with hip osteoarthritis were separately assessed using confirmatory factor analysis (CFA) exact-fit and close-fit indices, and Rasch analysis item fit statistics. Pooled analyses were added for robustness. Both unidimensional models and two-dimensional models were considered for each subsample. Item thresholds, targeting plots, internal consistency, and floor and ceiling effects were also evaluated. RESULTS:Exact CFA fit was rejected in 7 out of 8 subsamples for both the unidimensional and two-dimensional structures using pooled data (χ2; P-values <0.007). Close fit indices showed substantial variability across subsamples. Two of 8 subsamples reached acceptable thresholds for the unidimensional structure, and 3 of 8 subsamples for the two-dimensional structure (the root mean square error of approximation [0.031-0.120]; comparative fit index [0.876-0.997]; Tucker-Lewis index [0.842-0.996]; standardized root mean residual [0.059-0.106]). Rasch analysis consistently identified item misfit for all four preoperative subsamples, and minimal item misfit for the 1-year data. Pooling data increased misfit to the models. Substantial ceiling effects were evident at 1-year follow-up, inflicting disordered item threshold and poor targeting of the scale for 1-year data. CONCLUSION:Consistent fit to the statistical models was not demonstrated for the Danish OHS based on robust psychometric analyses of OHS data from a Danish prospective cohort of patients undergoing primary THA. These findings raise concerns regarding the validity of reporting the OHS as a single total score and separately as two subscales of pain and function. PLAIN LANGUAGE SUMMARY:The Oxford Hip Score (OHS) is a questionnaire commonly used to measure pain and function in patients undergoing total hip replacement. It is widely applied in clinical research and national quality databases, and results are usually reported as a single total score. Some researchers have suggested that the questionnaire may instead consist of the following two separate parts: one measuring pain and one measuring function. However, this structure has not been thoroughly tested. In this study, we evaluated how well the Danish version of the OHS measures what it is intended to measure. We analyzed responses from patients who completed the questionnaire before surgery and 1 year after surgery. To ensure robust results, we examined several independent patient groups and applied modern statistical methods designed to test whether questionnaire items work together as a valid measurement scale. We found that the OHS did not consistently meet statistical criteria for a well-functioning measurement scale, whether reported as a single total score or divided into separate pain and function subscales. Before surgery, several questions did not behave as expected statistically. One year after surgery, many patients achieved the highest possible score, leaving limited ability to distinguish between patients with very good outcomes. This reduced the precision of the questionnaire at follow-up. These findings suggest that caution is needed when interpreting OHS results, particularly if outcomes are measured and treatment modalities compared 1 year after hip replacement. Reporting pain and function separately is advised, but measurement limitations remain.
OBJECTIVES:To describe and characterize three partly coexisting evidentiary configurations in drug regulation and health technology assessment (HTA), and to examine the inferential consequences of the progressive shift from replication-based to coherence-oriented standards of proof. STUDY DESIGN AND SETTING:Conceptual and methodological commentary drawing on regulatory history, published methodological and epidemiological literature, documented empirical trends in regulatory approval patterns, and illustrative regulatory cases. Three evidentiary configurations are formally defined and compared: the two-trial paradigm, the one-trial paradigm, and a narrative regime characterized by coherence-based justification when independent experimental replication is limited or absent. RESULTS:The two-trial paradigm emerged from the post-Kefauver-Harris regulatory emphasis on adequate and well-controlled investigations and was consolidated in U.S. Food and Drug Administration (FDA) regulatory practice in the late 1980s and 1990s. Under realistic prior assumptions, this configuration may yield a posterior probability of false positive finding (PPFP) of approximately 0.3%. The one-trial paradigm was further legitimized by the FDA Modernization Act (1997), which allowed one adequate and well-controlled trial to be supported by confirmatory evidence, with estimated PPFPs of approximately 4%-9% under comparable assumptions. A third configuration, here termed the narrative regime, has expanded progressively through accelerated approval pathways, surrogate endpoint reliance, single-arm trial designs, external controls, real-world evidence, and extrapolation strategies; under this configuration, PPFP may increase further (potentially >15-20%), depending on the prior probability, endpoint validity, and study design. Published trends document a decline in FDA approvals supported by at least two pivotal randomized trials, from 80.6% in 1995-1997 to 52.8% in 2015-2017. The development pathway of subcutaneous belimumab in pediatric systemic lupus erythematosus (SLE) illustrates how replication-based evidence may be supplemented by pharmacokinetic bridging and extrapolation in selected settings. Regulatory agencies, HTA bodies, and guideline-development organizations operate under partly distinct institutional mandates and apply these configurations with different evidentiary thresholds and decision purposes. CONCLUSION:Maintaining the critical function of HTA requires flexible evidentiary frameworks to preserve independent comparative evaluation, analytical transparency, and explicit management of uncertainty. Without such safeguards, coherence-based inference may shift evidentiary standards toward plausibility and institutional acceptability rather than empirical testing. This shift should therefore be understood as a transformation in evidentiary standards, not merely as a technical adaptation of study designs.
BACKGROUND AND OBJECTIVE:External validation is essential for assessing the generalizability and transportability of a clinical prediction model. A previous review of studies published in 2010 identified substantial deficiencies in the reporting of external validations. In 2015, the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) statement was introduced, though its impact is unclear. Although external validation is widely recommended, its prevalence in recent oncology prediction model studies is uncertain, despite the large number of models developed in this field. We aimed to examine the proportion of oncology prediction model studies including an external validation and the reporting completeness of these studies, providing a cross-sectional overview of current practice. METHODS:We searched MEDLINE (via Ovid) for primary studies published between June 1, 2023 and July 31, 2023. Eligible studies evaluated multivariable prediction models in oncology using data not used for model development, including temporal or geographical split-sample approaches. Reporting completeness was assessed using the TRIPOD statement. Screening and data extraction were performed in duplicate. Findings were summarized using counts and percentages. RESULTS:Of 287 eligible oncology-based prediction model studies, 89 (31%) included an external validation component, with only 3 studies (1%) performing external validation without model development. Most validations (78/89, 88%) were conducted on newly developed models within the same study, and 32/89 (36%) used temporal or geographical split-sample approaches. Study design and participant characteristics were frequently reported, but outcome and predictor definitions were complete in only 64% and 46% of studies, respectively. Only one study reported a sample size calculation. Performance measures were reported with confidence intervals in 52/89 (58%). Calibration was assessed in 57/89 (64%), and clinical utility in 33/89 (37%). Only 22/89 (25%) explicitly described how predictions were calculated in the validation dataset. Despite validation data being clustered in 31/89 studies, heterogeneity across centers or subgroups was never examined. Open science practices were uncommon, including study registration (2.2%), protocol availability (1.1%), and code sharing (4.5%). CONCLUSION:In this 2-month cross-sectional snapshot in 2023, external validation studies were substantially less common than prediction model development studies and reporting of key methodological and performance details were frequently incomplete despite the availability of TRIPOD. External validation was often reported without clear articulation of its objectives or contextualization of model performance, suggesting it may sometimes be treated as a procedural step rather than a study designed to evaluate model transportability. Greater emphasis on transparently reporting external validation studies is required to enable reliable evaluation and comparison of prediction models.
OBJECTIVES:Adequate sample size is essential to the development of new prediction models with binary outcomes. We aim to relate recent approaches to traditional rules of thumb ('simple rules') and assess agreement and differences in judging adequacy of sample size for developed prediction models. STUDY DESIGN AND SETTING:A well-known simple rule considers the events per variable or, more specifically, events per predictor parameter p (EPP, eg, 'EPP > 10' or 'EPP > 20') to limit overfitting to small datasets. Another simple rule is to require a minimum absolute number of events (E) for reliable estimation of the overall event rate (eg, 'E > 100'). Recent criteria for regression-based prediction models consider: 1) limited overfitting in predictor effect estimates ('global shrinkage ≥0.9'); 2) small optimism in Nagelkerke's R2 ('δ(R2) ≤ 0.05'); and 3) precise estimation of the overall event rate (margin of error <0.05). We use statistical theory to compare simple rules to these three recent criteria. Furthermore, we compare sample size assessments for 299 published COVID-19 prediction models. RESULTS:At a 10% event rate, criterion 1 (limited overfitting) corresponded to the classic EPP>10 rule for an area under the receiver operating characteristic curve (c) of 0.767 and EPP>20 for c = 0.694. Criterion 2 implied EPP>4 (largely irrespective of c), and criterion 3 was equivalent to E > 14 at a 10% event rate. We classified largely the same COVID-19 prediction models as adequate or inadequate for sample size according to recent criteria (maximum of three) and a simple combination rule ('E = 100 plus 10∗p', or max['E = 100, 10∗p']). CONCLUSION:Simple rules have direct relations with recent criteria for sample size calculations to limit overfitting and guarantee reliability of predictions. Recent criteria are essential to inform prediction modeling efforts, while simple rules may often be sufficient to help appraise the quality of already developed regression-based prediction models. PLAIN LANGUAGE SUMMARY:Adequate sample size is essential for reliable empirical research. We focus on the development of prediction models to provide reliable individualized estimates of the risk of an event. We compared traditional simple rules for deciding if a dataset is large enough to build a binary outcome prediction model (such as "events per predictor >10" or "total events >100") with more recent, refined statistical criteria. The newer criteria aim to guarantee (1) that predictions are not too extreme (global shrinkage ≥0.9) (2), that model performance optimism is small (decrease in R2 ≤ 0.05), and (3) that the overall event rate is precisely estimated. Using theory and examples from 299 published COVID-19 models, we found that the simple rules often map closely to the newer criteria. Overall, we conclude that while the new criteria are better for planning prospective studies and developing a prediction model, simple rules remain useful and are usually adequate for judging whether already developed prediction models had enough data.
BACKGROUND AND OBJECTIVES:Diagnostic accuracy studies that compare two or more index tests (comparative diagnostic accuracy studies) can be at risk of generating biased results due to shortcomings in the statistical analysis. Understanding which statistical methods are used and how they are described in such studies is therefore crucial. The objective of this study is to evaluate statistical methods used in comparative diagnostic accuracy studies. STUDY DESIGN AND SETTING:Methodological review. We searched PubMed for comparative accuracy systematic reviews published in 2023. Of the primary studies included in these reviews, a subset of 200 comparative diagnostic accuracy studies was randomly selected. Seven reviewers extracted data in pairs about study design features, sample size calculation, statistical analysis methods used for comparing diagnostic accuracy, and methods for dealing with missing data and confounding. RESULTS:Of 200 comparative diagnostic accuracy studies, 184 (92%) had a fully paired design, 6 (3%) had a partially paired design, and 10 (5%) had an unpaired design. Eighty-four (42%, 95% CI: 35.1%-49.2%) studies did not statistically compare diagnostic accuracy estimates (ie, no confidence interval or statistical test for the difference was reported), while this was unclear in two (1%) studies. Of the remaining 114 studies including a statistical comparison, 44% (50/114, 95% CI: 34.6%-53.5%) failed to report the relevant statistical methods and of those that did report their methods, 17.2% (11/64, 95% CI: 8.9%-28.7%) used at least one inappropriate statistical method. Sixty-nine of 200 (35%, 95% CI: 27.9%-41.5%) studies reported the presence of missing, indeterminate, or intermediate test results and nearly all (96%, 66/69, 95% CI: 87.8%-99.1%) conducted complete case analysis. While 14 (7%, 95% CI: 3.9%-11.5%) studies used neither a fully paired nor randomized design, none of these reported adjusting for confounders. Overall, most studies (71%, 95% CI: 64.2%-77.2%) showed one or more shortcomings in the statistical analysis. CONCLUSION:The majority of comparative diagnostic accuracy studies suffer from suboptimal analyses, which may increase the risk of misleading or overconfident conclusions, while the incomplete reporting of statistical methods hampers interpretation. Researchers should carefully consider the choice of statistical methods for comparative accuracy and adhere to the Standards for Reporting of Diagnostic Accuracy Studies (STARD) 2015 reporting guidelines. PLAIN LANGUAGE SUMMARY:We assessed how statistical methods are used and reported in studies comparing diagnostic tests (comparative diagnostic accuracy studies). We examined 200 such studies published between 1995 and 2023. Overall, most studies (142/200, 71%) had at least one shortcoming in the statistical analysis. This finding suggests that statistical analyses in comparative diagnostic accuracy studies are often suboptimal.