
Prediction models for clinical conditions involving recurrent events require performance metrics tailored to repeated outcomes. Although we and others have developed methods to evaluate calibration and related aspects of performance, approaches for assessing discrimination remain limited. Recent contributions, including concordance-based methods and Brier-type accuracy measures, address parts of this gap, but no widely adopted, flexible framework for discrimination exists. As discrimination is a core TRIPOD+AI–recommended metric, accessible methodology and software for evaluating how well recurrent event models separate higher- and lower-risk individuals are still needed. We propose adapted discrimination statistics based on comparing predicted and observed cumulative event counts, addressing limitations of conventional concordance approaches. Statistical uncertainty is evaluated using several resampling strategies, including delete-d jackknifing, with user-friendly R code provided. We illustrate the approach using four tie-handling methods (c statistic, Kendall’s τₐ, Somers’ D, Goodman–Kruskal’s γ), four recurrent event models (negative binomial, zero-inflated negative binomial, Andersen–Gill, and Prentice–Williams–Peterson Total Time), and two clinical datasets: the SANAD epilepsy trial and an OPCRD asthma cohort. Discrimination varied substantially across models. In the asthma data, the Prentice–Williams–Peterson model showed the strongest discrimination (C = 0.939, 95
Abstract Background Children with acute abdominal pain pose a diagnostic challenge for general practitioners (GPs), as it can be difficult to distinguish appendicitis from self-limiting conditions due to overlapping symptoms. To support GPs, a diagnostic strategy for appendicitis was developed that integrates an externally validated prediction rule with C-reactive protein point-of-care testing (CRP POCT) and risk-based management advice. This study will compare the impact of this diagnostic strategy to usual care in children with acute abdominal pain in primary care in terms of: (i) clinical effectiveness, (ii) cost-effectiveness, and (iii) implementation potential. Methods We will conduct a pragmatic, hybrid type 1 effectiveness-implementation, cluster randomised controlled trial in Dutch general practice. Children aged 4–18 years with acute (≤ 7 days) abdominal pain will be included. GP practices will be randomly allocated to either: (i) use of the diagnostic strategy for appendicitis, or (ii) usual care. In the intervention group, children are stratified into low, medium, or high risk for appendicitis based on a seven-item prediction rule including symptoms and signs. Management advice is tailored to appendicitis risk: safety netting for low-risk, immediate referral for high-risk, or CRP POCT for medium-risk (CRP < 10 mg/L: safety netting; CRP 10–50 mg/L: reassessment or immediate referral (based on GPs’ clinical judgement and parent/child preference); CRP ≥ 50 mg/L: immediate referral). The primary outcome is referral efficiency (proportion of non-referrals amongst patients with no appendicitis during 30 days follow-up). Secondary outcomes include safety, proportion of children with CRP POCT, proportion of children with planned reassessment, child anxiety, patient satisfaction, quality of life, and costs. A parallel process evaluation will assess implementation outcomes, including the reach, adoption, implementation, and maintenance of the diagnostic strategy. We aim to include 566 children without appendicitis to determine an improvement of efficiency from 88% to 95%. Discussion We hypothesise that use of the diagnostic strategy will reduce referrals without increasing delayed diagnoses of appendicitis, compared to usual care in children presenting with acute abdominal pain in primary care. This may ultimately result in better patient outcomes, lower costs, and reduced health care resource use. Trial registration ClinicalTrials.gov: NCT06762275 (registration date: 2024/12/03).
Deep learning (DL)-assisted low-dose computed tomography (LDCT) may improve lung cancer screening, but the available evidence is heterogeneous and patient-level diagnostic accuracy remains uncertain. MEDLINE, Embase, and Web of Science were searched from January 2010 to December 2025 for studies evaluating DL-assisted LDCT in lung cancer screening or screening-relevant populations. Two reviewers independently screened studies, extracted data, and assessed risk of bias using QUADAS-2. AI-specific reporting completeness was assessed descriptively using items adapted from CLAIM and STARD-AI. Studies with complete or reconstructible patient-level 2 × 2 data at a defined threshold were pooled using bivariate random-effects and hierarchical summary receiver operating characteristic models. Studies without sufficient threshold-specific data were synthesised narratively. Eleven studies met the inclusion criteria. Five studies, comprising 2,220 participants, 232 lung cancer cases, and 1,988 non-cases, provided usable threshold-specific 2 × 2 data and were included in the meta-analysis. Six studies were synthesised narratively because patient-level TP, FP, FN, and TN values were unavailable or not reconstructible. Pooled sensitivity was 85.4
BACKGROUND:Sickle cell disease (SCD) and thalassaemia are inherited conditions causing chronic anaemia, increased infection risk and multi-organ failure. Standard of care (SoC) prenatal screening involves carrier blood testing for pregnant women and, if positive, carrier blood testing for biological fathers followed by invasive prenatal diagnosis for pregnancies at risk. Non-invasive prenatal testing (NIPT) presents an alternative pathway which may reduce diagnostic delays and improve equity for pregnant women when the biological father is unavailable by focusing invasive testing exclusively on fetuses shown to have a high risk of SCD by NIPT. This study compares the outcomes of SoC screening with a proposed NIPT pathway replacing the paternal blood testing stage. METHODS:A deterministic decision tree model is used to identify the outcomes of the screening pathways, focusing on the SCD population, from the perspective of the National Health Service (NHS) England. Sensitivity and specificity inputs for NIPT are informed by a separately published minimally acceptable criteria study. Diagnostic outcomes include the number of performed and declined tests and true and false positive/negative diagnoses in each pathway. Economic outcomes include the testing cost of the pathway, the cost per case detected and per accurate diagnosis, and an incremental cost threshold analysis for NIPT. Additional scenario analyses are conducted for the SCD and thalassaemia combined population and for the thalassaemia populations. RESULTS:When considering an overall cohort of 616,573 pregnancies, implementing the NIPT pathway for the screen-positive SCD population results in an incremental cost of £7,584,551. Of 276 prenatal diagnoses (PND) performed in the SoC arm, 76 show a true positive result for SCD, and 2 false positives are identified. In the NIPT arm, there are 6090 NIPTs and 543 PNDs performed, with 213 true positives and 235 false positives identified. The NIPT pathway costs £33,158 more per case detected, and £368 more per accurate diagnosis than SoC; to obtain no incremental cost per case detected versus the SoC, NIPT would need to cost £45.21. CONCLUSIONS:The presented exploratory analysis may gauge the potential cost-effectiveness of introducing NIPT into the screening pathway, pending further research on the technique's diagnostic efficacy.
Abstract Background The screening pathway for sickle cell disease (SCD) in England starts with a carrier status blood test for the pregnant woman. Following a positive result, the test is offered to the biological father. Where both results are positive, further invasive testing is offered to assess the risk of SCD to the fetus. When the father is unavailable, the timely offer of further testing may be delayed and even missed, leaving the pregnant woman underinformed and limiting reproductive choice. Non-invasive prenatal testing (NIPT) presents an alternative to paternal testing and diagnostic accuracy studies of NIPT in SCD screening are required to inform utility and application. We explored the minimally acceptable sensitivity and specificity to inform such an accuracy study of NIPT. Methods A decision tree model was produced to identify the minimally acceptable sensitivity and specificity of NIPT. Stakeholder engagement identified a positive predictive value (PPV) equal to the paternal carrier blood test and two sensitivity/specificity scenarios: Scenario 1 should result in < 10 false negative diagnoses per year; Scenario 2 should result in ≤ 2 false negative diagnoses per year. Subsequently, the minimally acceptable sensitivity and specificity were used to calculate the sample size for a hypothetical accuracy study of NIPT in each scenario. Results Scenario 1 led to a minimally acceptable sensitivity of 96.0% and specificity of 88.5% and corresponding negative predictive value (NPV) of 99.81%, PPV of 25.51% and accuracy of 88.80%. Utilising an expected prevalence of SCD of 3.94%, the resulting sample size would include 315,824 total pregnancies. Scenario 2 led to a minimally acceptable sensitivity of 99.0% and specificity of 88.0% and corresponding NPV of 99.95%, PPV of 25.29% and accuracy of 88.43%. The resulting sample size would include 65,509 overall pregnancies. Conclusions While the sample size calculated for each scenario is unfeasible if considering a prospective cohort study, approaches are presented to achieve realistic sample sizes. Increasing the sensitivity and specificity levels might be considered if this is likely to be achievable based on available studies, albeit limited in volume. Consideration could be given to study design mitigations that are consistent with guidance on accuracy studies in low prevalence settings.
Liquid biopsies offer a minimally invasive approach for early glioma detection. However, potential non-blood biomarkers in urine and CSF are under-researched. While urine is easier to sample and has fewer proteins/cells, CSF may yield higher detection rates due to its proximity to the glioma. This systematic review evaluates non-blood liquid biopsy biomarkers (proteins, metabolites, cell-free DNA, and miRNA) in urine and CSF for early glioma detection. Three electronic databases (MEDLINE, Embase, and BIOSIS) were searched from inception to September 2024. Studies which assessed urine- or CSF-derived biomarkers as a means of identifying glioma patients were included. The primary outcomes of interest were measures of diagnostic test accuracy (specificity, sensitivity, positive predictive value, negative predictive value and area under the ROC curve). QUADAS-2 was used to assess risk of bias in included studies. We extracted data from 20 studies (669 glioma patients and 597 controls) which met the inclusion criteria. Five of the included studies investigated urine-derived biomarkers, and 15 investigated CSF-derived biomarkers. Meta-analysis was not possible due to the heterogeneous nature of biomarkers collected, analysis methods, and differences in patient demographics. Conclusions are limited by the paucity of studies and small sample sizes. These limitations highlight the need for well-designed, adequately powered studies to evaluate the diagnostic utility of urine and CSF for glioma liquid biopsies. 1) Urine and cerebrospinal fluid are underexplored for glioma liquid biopsy, with typically small, underpowered studies. 2) Multiple biomarkers show promising AUCs but lack precision estimates, representing generally poor reporting. 3) Urine biomarkers show no overlap with CSF or blood; CSF shows overlap with blood but with limited replication.
Abstract Background During the early stages of the COVID-19 pandemic, prediction modelling was widely used to forecast infection rates while only few studies developed models to predict individual risk of infection for Omicron and subsequent variants. For such prediction models to perform well, it is important to carefully select predictors from a comprehensive set of potential factors, including prior infection and vaccination history, individual behaviors, and immunological markers. Methods This exploratory analysis aimed to develop and compare prediction models for Omicron infection to provide an evidence base for identifying key predictors of SARS-CoV-2 infections which can be used to inform future derivation and validation of predictive models. We used data from 710 participants from two ongoing, prospective population-based cohorts: the Zurich SARS-CoV-2 Cohort and the Zurich SARS-CoV-2 Vaccine Cohort (Switzerland). Participants were recruited between 2020 and 2021 and provided demographic data, vaccination history, self-reported infections, and longitudinal serological data (anti-S IgG, anti-S IgA, anti-N IgG antibodies) collected at 6-month intervals. Our main outcome was SARS-CoV-2 infection during the first Omicron wave (01.01.2022–31.03.2022) based on self-reported positive tests or a doubling in any of anti-S IgG, anti-S IgA or anti-N IgG. We used logistic regression models with backward stepwise selection based on the Akaike Information Criterion (AIC) to evaluate predictors and identify the best-fitting models. Results Only 17.3% of participants reported a positive SARS-CoV-2 test result during the Omicron wave. However, when including serological testing, 37.2% of participants had evidence of infection, indicating substantial underdiagnosis. The best-performing model had an AUC of 0.69 (95%CI 0.66, 0.73) and included the following predictors: age, sex, compliance with COVID-19 prevention guidelines, smoking status, comorbidities, prior anti-N IgG antibody levels, and the sequence of previous infections and vaccinations. We found that older age (≥ 65 years) was associated with a 50–60% lower odds of Omicron infection across all our models, while having fewer prior exposures (through infections or vaccinations) increased the odds of infection. Conclusion This explorative study highlights the importance of integrating comprehensive immunological, clinical and behavioral data to predict SARS-CoV-2 infection risk. Our study lays the foundation to develop and validate future prediction models that identify individuals at risk, particularly through the novel use of infection and prior vaccination sequence as an important predictor.
Abstract Background Women with gestational diabetes (GDM) are at increased risk of developing type 2 diabetes (T2D). Prognostic models have been developed and evaluated, but their methodological quality and applicability remain inconclusive. Recent reviews with the latest search date of March 31, 2025, were conducted with major shortcomings, including failure to adhering best-practice guides and inappropriately pooling heterogeneous model performances. This systematic review aims to synthesise the methodological characteristics of existing prognostic models for T2D following GDM. Methods Five electronic databases were searched from inception to January 24, 2026. Prognostic models predicting T2D following GDM, regardless of study setting were included. Data extraction adhered to existing expert guidelines. Quality and applicability were assessed using the updated Prediction model Risk Of Bias ASsessment tool. Two reviewers independently screened and assessed quality, resolving disagreements through consensus and involvement of a third reviewer. Results Our updated review identified six more studies than the previous review. Our review identified 19 studies with 20 models, half from prospective cohorts (n = 18) mostly in hospital settings across North America, Europe, Australia, and Asia. Logistic regression model was most common (n = 9), followed by machine learning (n = 6), and Cox regression (n = 5). Internal and external validation were done in only 13 and 1 models, respectively. Discrimination was widely reported (Area Under the Curve (AUC) 0.67–0.92); while calibration, overall performance, and clinical utility measures were underreported. Only one study reported appropriate sample size determination. Maternal age, pregnancy fasting glucose, and BMI were common predictors. Risk of bias was generally low during development and evaluation phases, but applicability concerns were high in 60% of models. Conclusions While several models demonstrated acceptable performance and low concern in selected quality domains, generalizability and clinical utility remain limited due to high concerns in applicability and inconsistent reporting. Adherence to best practice guides such as TRIPOD + AI, external validation of models, and exploration of novel prediction modelling techniques are recommended to advance the reporting and application of risk tools for personalised medicine post GDM. This review is the first to apply the PROBAST + AI, enabling a comprehensive evaluation of quality and applicability compared to previous works in the field. Protocol registration PROSPERO CRD420251034657.
Glioma is the most common primary brain tumor, demanding prompt and accurate diagnosis to guide therapy decisions. Conventional diagnostic methods such as histopathology and neuroimaging are limited by their invasiveness, subjectivity, or lack of intraoperative precision. Raman spectroscopy is an emerging, label-free optical technique that detects the unique molecular "fingerprint" of tissues, enabling real-time differentiation between tumor-infiltrated and normal brains based on their biochemical composition. Despite a growing body of experimental and translational studies, a comprehensive synthesis of Raman spectroscopy's clinical applicability in detecting glioma micro-infiltration is lacking. We will search four electronic databases (PubMed, Embase, Scopus, and Web of Science) for English-language studies published from 2015 to 2025. Eligible studies may use in vivo Raman (subjects with suspected or histologically confirmed glioma), ex vivo Raman (freshly excised human tissue), or in vitro analysis on archived human samples, provided the tissue originated from confirmed glioma cases. Primary outcomes will be diagnostic accuracy measures, including but not limited to sensitivity, specificity, accuracy, area under de curve (AUC), positive predictive value (PPV), negative predictive value (NPV), diagnostic odds ratio (DOR), and positive/negative likelihood ratios (LR+/LR−). Secondary outcomes include its role in intraoperative margin assessment, diagnostic time, and machine learning-assisted classification. Data traction and risk of bias assessment will be conducted independently by two reviewers. Meta-analyses will be performed using specific statistics where applicable. If a meta-analysis is not feasible, a structured narrative synthesis will be employed, following the SWiM (Synthesis Without Meta-analysis) guidelines. GRADE (Grading of Recommendations, Assessment, Development, and Evaluations) will be applied to evaluate the certainty of evidence. The review will follow Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. This review will synthesize the emerging diagnostic landscape of Raman spectroscopy for glioma, positioning it as a transformative adjunct to traditional histopathology and intraoperative decision-making. Our findings will pave the way for a translational roadmap to integrate real-time Raman spectroscopy into neurosurgical workflows with machine learning support. PROSPERO: CRD420251025922.
Abstract Background Less than 20% of patients diagnosed with advanced lung cancer will survive beyond five years and half of these will suffer a serious adverse event (SAE) caused by systemic anticancer therapy (SACT) that will result in a hospital attendance. As multiple different SACT treatments are available for patients, a risk score that predicts the likelihood of a SAE following each type of SACT treatment would improve both communication with the patient and shared decision making with all those involved in delivering care for patients. There are currently no risk scores available for use in those with advanced stage lung cancer. Aim The overarching aim of this research is to develop and internally validate a risk score that will calculate the individualised risk of SAEs for different SACT treatments for patients with late stage lung cancer. Methods Utilising linked cancer registry data (National Cancer Registration and Analysis Service (NCRAS), England) for over 20,000 late stage lung cancer patients, a risk score will be developed using a multivariable logistic regression model to predict the risk of an acute admission within 30 days of SACT administration. Model performance will be summarised using calibration and discrimination. Internal validation will be used to quantify the degree of optimism due to overfitting, using re-sampling bootstrapping. Heterogeneity will be assessed, and the model will be fine-tuned. Fine-tuning and interrogation will be used to evaluate differences in performance between hospitals. The clinical utility will be assessed through calculating the net benefit in preventing SAEs. Conclusion A developed risk score (under each treatment strategy) has real potential to support individualised treatment decisions and optimise management of SACT-induced SAEs for patients and reduce hospital attendances.
Adverse pregnancy outcomes (APOs), such as gestational diabetes, preeclampsia, and placental abruption, are major contributors to maternal and fetal morbidity and mortality, with implications for individual long-term health and health system performance. Existing prediction models for APOs rely primarily on clinical or biomarker data, with few incorporating social, behavioral, or environmental determinants that are critical for shaping perinatal outcomes. This study describes the development and validation protocol for the Adverse Pregnancy Outcomes Population Risk Tool (PregPoRT), a novel, population-based prediction model designed to estimate APO risk using population-based and routinely collected survey and administrative data in Canada. PregPoRT will be developed using a retrospective cohort of female-identifying individuals, aged 15–49, who participated in the Canadian Community Health Survey (CCHS) between 2000 and 2017, and had a subsequent delivery hospitalization within two years recorded in the Discharge Abstract Database (DAD). Pre-pregnancy predictors were selected according to a health equity-informed framework by Kramer and colleagues (2019), and include biomedical, behavioral, social, and environmental variables from the CCHS, the Canadian Marginalization Index (CAN-Marg), the Canadian Urban Environmental Health Research Consortium (CANUE), and the Canadian Active Living Environments (Can-ALE) dataset. The primary outcome is a composite measure of APOs (gestational diabetes, preeclampsia, or placental abruption), identified using validated ICD codes. A Weibull accelerated failure time model will be used to estimate the risk of experiencing an APO. Continuous variables will be modeled with restricted cubic splines. Variable selection will be performed using the Least Absolute Shrinkage and Selection Operator (LASSO), and model performance will be assessed via discrimination, calibration, and overall accuracy. Validation strategies include split-sample, bootstrap, and temporal validation using later CCHS cycles. Survey weights will be applied throughout to ensure national representativeness. PregPoRT will be the first Canadian prediction model for APOs that leverages nationally representative, linked survey and administrative data and explicitly integrates social, behavioral, and environmental determinants of health, domains that have been largely absent from prior models. By incorporating modifiable and socially patterned risk factors, the tool is designed to support public health planning, resource allocation, and maternal health equity monitoring.
The Transparent Reporting of a multivariable prediction model of Individual Prognosis Or Diagnosis (TRIPOD) statement was published to improve the reporting and critical appraisal of prediction model studies for diagnosis and prognosis. This paper describes the processes and methods that will be used to develop an extension to the TRIPOD statement (TRIPOD-Code) for the management of code associated with prediction model studies. TRIPOD-Code focuses specifically on the transparent reporting of analytical code used in prediction model studies, including code for data preprocessing, model development, and model evaluation. TRIPOD-Code will be developed following published guidance from the EQUATOR Network and will comprise five stages. Stage 1 will involve a methodological review of how code availability is reported in published prediction model studies. In Stage 2, we will consult a diverse group of key stakeholders using a Delphi process to identify items to be considered for inclusion in TRIPOD-Code. Stage 3 will consist of virtual consensus meetings to consolidate and prioritise the key items. Stage 4 will involve developing the TRIPOD-Code checklist. In the final stage, Stage 5, we will disseminate TRIPOD-Code via journals, conferences, blogs, websites (including TRIPOD and the EQUATOR Network) and social media. TRIPOD-Code will provide researchers working on prediction model studies with a reporting checklist and accompanying guidance to promote code completeness and availability. This study has been determined to be exempt from ongoing IRB oversight by the Massachusetts Institute of Technology Committee on the Use of Humans as Experimental Subjects (COUHES) under Exempt ID: E-6675. Findings from this study will be disseminated through peer-reviewed publications.
Multiple long-term conditions (MLTCs) are being made a priority by funding bodies as prevalence rates increase. Improving early detection of individuals at high risk of developing MLTCs may delay or prevent complications and poor health outcomes. Predicting MLTCs remains a challenge, and methods for singular outcomes have been proven to be inappropriate for MLTC research. The aim of this paper is to present the protocol for a systematic review to identify all published models for prediction of MLTCs, and to summarise methods used for model development. MEDLINE (Ovid), Embase (Ovid), CINAHL (EBSCOHost) and CENTRAL (Cochrane Library) will be searched from September 2015 to identify relevant clinical prediction models which predict the development of MLTCs. Screening, data extraction and the risk of bias will be undertaken by two reviewers independently. Data extraction will include primary items for methodology and model outcomes and secondary items including study descriptors, population information, measured outcomes, methodology, model performance measures, clinical usefulness measures and risk of bias. A narrative synthesis will be conducted to summarise current methodological practice and to identify areas for improvement to inform future methodological and model development. Ethical approval is not required for this systematic review as it will use published literature only. The findings of the review will be submitted for publication in a peer reviewed journal. • A comprehensive review of all methodologies used in clinical prediction modelling for MLTCs will be undertaken. • Methodologies will be included regardless of the outcome conditions predicted. • The protocol is reported according to the Preferred Reporting Items for Systematic Reviews Protocol (PRISMA – P). • Screening, data extraction and risk of bias will be conducted independently by two reviewers. • Only English language papers will be included.
Reporting of COVID-19 prognostic models frequently falls short of established standards. The TRIPOD checklist and its 2024 AI extension (TRIPOD + AI) provide a comprehensive framework for assessing reporting quality. We therefore evaluated and compared reporting completeness in conventional versus machine-learning models. Studies reporting the development, and internal and external validation of prognostic prediction models for COVID-19 using either conventional or machine learning-based algorithms were included. Literature searches were conducted in MEDLINE, Epistemonikos.org, and Scopus (up to July 31, 2024). Studies using conventional statistical methods were evaluated under TRIPOD, while machine learning-based studies were assessed using TRIPOD + AI. Data extraction followed TRIPOD and TRIPOD + AI checklists, measuring adherence per article and per checklist item. The protocol was prospectively registered at the Open Science Framework ( https://osf.io/kg9yw ). A total of 53 studies describing 71 prognostic models were identified. Overall, adherence to both guidelines was low, with significantly poorer compliance among machine learning-based studies (TRIPOD + AI) compared to conventional model studies (TRIPOD) (28.4
Approximately one million adults in the UK are estimated to have undiagnosed type 2 diabetes mellitus (T2DM), with a further 5.1 million adults with nondiabetic hyperglycaemia (NDH) that does not meet the threshold for a diabetes diagnosis. The Leicester Risk Assessment score (LRA) and Leicester Practice Risk score (LPR) are diagnostic risk prediction models that estimate an individual’s risk of undiagnosed T2DM and NDH, developed for use in community and primary care settings respectively. The LRA is also used as a prognostic model; neither model has been updated since development. This study will systematically review all applications of these models as diagnostic and prognostic tools and any published updates to evaluate their performance in different populations. This review has been registered with PROSPERO (CRD420251005841). We will implement a citation search strategy to search Scopus, Web of Science and Google Scholar, restricted to full text, English language papers. Eligible papers will validate, update or modify either model. Data will be extracted using a form based on the Critical Appraisal and Data Extraction for Systematic Reviews of Prediction Modelling Studies (CHARMS) checklist; missing information will be sought from authors or estimated from other available information where possible. Meta-analysis of predictive performance measures will be completed if sufficient data exist. Subgroup and sensitivity analyses will be used to explore between-study heterogeneity and risk-of-bias impact. This review will identify studies that have implemented, modified or validated the LRA and LPR for the risk of undiagnosed T2DM and NDH in different populations. This will allow summary measures, including level of uncertainty, of model performance to be calculated, making this highly relevant to individuals and stakeholders who recommend and implement these models. Review conclusions will also inform the potential update and recalibration of the models. This will ultimately lead to improved outcomes through earlier diagnosis and management.
OBJECTIVES:The rise in popularity and off-the-shelf availability of machine learning (ML) and AI-based methodology to develop new prediction models provides developers with ample choices to compare and select the best performing model out of many possible models. Many studies have shown that such comparisons on any particular dataset, the difference in performance between models developed using different techniques (e.g. logistic regression, vs. random forest or neural networks) can often be small, especially when looking at crude performance measures such as the area under the ROC curve. This may lead to the conclusion that such models are essentially exchangeable, and model selection is arbitrary. However, as we will illustrate using a dataset on deep venous thrombosis, prediction models with similar discriminative performance may nonetheless generate different outcome probability estimates for individual patients and potentially lead to meaningfully different decision making. METHODS:We developed diagnostic prediction models to predict the presence of deep venous thrombosis (DVT) in a large dataset of patients with leg symptoms suspected of having DVT, using five modelling techniques: unpenalized logistic regression (ULR), ridge logistic regression (RLR), random forests (RF), support vector machine (SVM) and neural network (NN). Age, sex, d-dimer, history of DVT, diagnosis alternative to DVT, and having cancer were used as a fixed set of predictors. Model performance was evaluated in terms of discrimination, calibration, and stability of individual risk prediction for a set of patients across the models. RESULTS:Of the 6,087 suspected patients, 1,146 (19%) were diagnosed with DVT based on leg ultrasound (reference test). Three prediction models (ULR, RLR, NN) had similar discrimination with AUCs point estimates of 0.84. However, the 6087 individuals' estimated probabilities of DVT varied substantially across the five different modelling techniques, highlighting differences in prediction stability. Notably, the RF model tended to overestimate individual risks, while the SVM model tended to underestimate them compared to the other models. While the estimated probabilities were more similar for ULR, RLR and NN, classification measures (sensitivity, specificity, positive and negative predictive value) did differ because of differences in estimated probabilities of individuals near the risk threshold, illustrating that differences, even when relatively small, could potentially lead to different clinical decisions. CONCLUSIONS:Prediction models developed with different modeling techniques yielded very different individuals' outcome probabilities, even though the models had similar discriminative performance in this low-dimensional setting. Part of this variation can be explained by differences in calibration but also from modelling choices as estimated risks also differed for modelling techniques with similar calibration performance. Hence, our findings highlight the impact of the choice of modelling techniques on model performance, individual estimated probabilities and consequently the impact of that choice on risk-based clinical decision making.
Abstract Background One non-randomized approach to estimate the average effect of a newly introduced treatment is to compare observed outcomes under the new treatment with predicted outcomes under the standard treatment. These counterfactual predictions are made using a model developed before the new treatment was introduced, using patient characteristics and individualized treatment information. Although the approach has been intuitively applied (e.g., as model-based clinical evaluation in radiotherapy) and can be recognized as a specific case of standardization, a method of virtual controls or a g-method, the theory and conditions required for unbiased treatment effect estimation have not been formally described. The objective of this paper is to formalize the approach and clarify these conditions. Methods We formalize the approach within the potential outcomes framework for causal inference. We explain the methodology, its necessary conditions, and approaches for assessing their validity. These conditions are furthermore illustrated through a case study from radiotherapy, estimating the benefit of proton therapy compared to photon therapy on dysphagia in patients with head and neck cancer. Results We describe a set of five sufficient conditions, including examples of violations: transportability, ignorability of treatment assignment, consistency, positivity, and correct model specification. While these conditions are largely untestable, we describe how empirical evidence, such as comparing predicted and observed outcomes in related samples, can increase confidence in their plausibility. Conclusion When the prediction model predicts well in relevant (sub)populations, the approach can yield unbiased treatment effect estimates. However, there are many possible sources of bias. Therefore, we recommend systematic consideration of all required conditions, informed by domain expertise, and the use of empirical evidence whenever possible to support their plausibility.
Abstract Background Clinical prediction models enable healthcare professionals to estimate individual outcomes using patient characteristics. Current sample size guidelines for developing or updating models with continuous outcomes aim to minimise overfitting and ensure accurate estimation of population-level parameters, but do not explicitly address the precision of predictions. This is a critical limitation, as wide confidence intervals around predictions can undermine clinical utility and fairness, particularly if precision varies across subgroups. Methods We propose methodology for calculating the sample size required to ensure precise and fair predictions when developing models with continuous outcomes using linear regression. Building on Fisher’s unit information matrix theory, our approach calculates how sample size impacts the epistemic (model-based) uncertainty of predictions and allows researchers to either (i) evaluate whether an existing dataset is sufficiently large, or (ii) determine the sample size needed to target a particular confidence interval width around predictions. The method requires real or synthetic data representing the target population. To assess fairness, the approach can evaluate prediction precision across subgroups. Extensions to prediction intervals are included to additionally address aleatoric uncertainty. Results We demonstrate the methodology by examining the sample size required to develop or update a model for predicting Forced Expiratory Volume (FEV) using age, height, and sex. Existing guidance suggests a minimum sample size of 237 participants for this setting. We show this corresponds to an anticipated mean confidence interval width of 0.206 L across all participants in the target population, and that widths may be considerably larger for some individuals. Our new approach calculates a minimum of 694 patients would be needed to ensure all anticipated interval widths were ≤ 0.3. Subgroup analysis showed comparable anticipated precision across sex subgroups. Sensitivity analysis assuming conditional independence among predictors yielded consistent results. For prediction intervals, the magnitude of the residual variance imposes a lower bound on interval width, even with very large samples. Conclusions Our methodology provides a practical framework for examining required sample sizes when developing or updating prediction models with continuous outcomes, focusing on achieving precise and equitable predictions. It supports the development of more reliable and fair models, enhancing their clinical applicability and trustworthiness.