
Introduction Electronic medication management systems (EMMS) are central to hospital digital transformation and are intended to improve patient safety, operational efficiency and staff experience, but evidence of their impact remains fragmented, inconsistent and context-dependent. Consequently, there is an urgent need to understand how local organisational structures and complex sociotechnical factors drive successful long-term integration.Methods and analysis This systematic review will search MEDLINE, EMBASE, PsycINFO, CINAHL, Scopus, ABI/INFORM Collection and Business Source Premier for empirical hospital studies published from 2010 to 2026 that centrally apply a theoretical or conceptual framework. Study quality will be appraised using the Mixed Methods Appraisal Tool (MMAT) Version 2018, and a theory-informed thematic synthesis will identify implementation mechanisms and organise them into higher-order domains. This protocol has been registered with PROSPERO (CRD420261283974).Ethics and dissemination Ethics approval is not required because only published data will be used. Findings will be disseminated through peer-reviewed publication and doctoral outputs.PROSPERO registration number CRD420261283974
Care pathways are essential to enhance the organisation of care, optimise resource use and improve patient outcomes, thereby supporting value-based healthcare. As care pathways are increasingly required, there is a need for clear methodological guidance on the identification and visualisation of care pathways. This perspective article provides an overview of methodological approaches to support the identification and visualisation of these care pathways. Two possible methods are described, including co-creation and process mining as well as different visualisation techniques that can be applied for care pathway presentation. Moreover, practical guidance as well as advantages and limitations of each identification method are discussed to further support decision-making.Whether to implement co-creation or process mining, or use a combined approach, depends on several factors, such as the intended purpose and end users of the care pathway, time availability, level of expertise at hand, the structure and quality of the data, as well as stakeholder needs and availability. These considerations will vary for different clinical conditions and care processes involved. With anticipated workforce shortages, rising healthcare costs and increasing prevalence of hybrid and virtual care, delivering efficient, patient-centred and personalised care has become imperative. Consequently, the development of care pathways is essential to eliminate unnecessary interventions or those that do not meaningfully contribute to patient outcomes.
Objective We aimed to develop and evaluate the performance of CMC-ID, an algorithm designed to identify children and youth with medical complexity (CMC) from electronic health record (EHR) data for recruitment to a transition to adult care programme, and then tested it in younger patients.Methods CMC-ID was developed iteratively at a paediatric tertiary care hospital (SickKids) in Toronto, Canada, using a data repository derived from an Epic System EHR. The algorithm captured standard clinical criteria for CMC used throughout Ontario, Canada, including (1) Complexity, (2) Chronicity, (3) Fragility and (4) Technology dependence. We refined the CMC-ID iteratively to maximise positive predictive value (PPV). In phase I, we retrospectively identified youth aged 17 years to <18 years and conducted chart reviews to confirm CMC status, and then repeated the search in a younger cohort (1 year to <17 years) in phase II.Results Among 280 unique patients identified in phase I, the last of six iterations of the algorithm identified CMC with a PPV of 85.5% (95% CI 73.3% to 93.5%). When applied to a younger cohort in phase II, 947 CMC were identified, and the algorithm performed similarly with a PPV of 89.1% (95% CI 86.9% to 91.0%).Discussion CMC-ID identified youth for recruitment to an intervention focused on transition to adult care with high PPV, and the algorithm performed similarly when applied to younger patients.Conclusion CMC-ID is a pragmatic, high-alert index to support recruitment to clinical programmes and other interventions aimed at improving CMC outcomes.
Objectives Clinical practice guidelines are a cornerstone of evidence-based medicine, yet their implementation in routine care remains inconsistent. Large language models (LLMs), particularly with Retrieval-Augmented Generation (RAG), have shown strong performance in medical question answering, but their ability to use knowledge from German-language guidelines has not been systematically evaluated due to a lack of a dedicated benchmark. We therefore developed such a benchmark and evaluated guideline-based question answering with different LLMs and retriever configurations. Methods We developed cpgQA-DE, an expert-validated benchmark dataset of 200 multiple-choice questions derived from 10 current German clinical practice guidelines across five specialties. All questions were reviewed for correctness, relevance and complexity. The dataset includes case-based and knowledge-based questions with metadata on guideline source, specialty, relevance and difficulty. We evaluated LLM performance using a RAG-based pipeline built on a corpus of German guidelines, comparing multiple model–retriever combinations. Results RAG integration substantially improved accuracy across all tested models and for most retrievers. The best-performing configuration, GPT-5 combined with the multilingual-e5-large retriever, achieved an accuracy of 95%. Notably, the open-weight model gpt-oss-120b reached 90% accuracy when used with RAG. Conclusions cpgQA-DE enables systematic and reproducible offline evaluation of guideline-aware question answering systems in the German healthcare context. The observed performance gains with RAG support the use of LLMs augmented with quality-assured external knowledge. Such systems may help bridge the evidence-practice gap by providing guideline-based recommendations at the point of care, while strong performance of open-weight models suggests potential for on-premises deployment in privacy-sensitive clinical environments.
Objectives To assess the performance of the Glasgow admission prediction score (GAPS) and ambulatory score (AmbS) for identifying emergency department (ED) attendances suitable for medical same day emergency care (SDEC) services and to derive and validate a novel tool for this purpose, the SDEC Triage Tool (SDEC-T). Methods A retrospective diagnostic study using routine healthcare data from three hospitals in a diverse urban setting (Birmingham, UK). All unplanned ED attendances by adults requiring internal medicine assessment were included. The primary outcome was suitability for SDEC, defined as discharged alive with a length of stay under 12 hours (LOS<12). The SDEC-T was derived using multivariable analysis, informed by stakeholder workshops. Results 152 877 attendances were included (median age: 58 years; 54.3% female; 68.4% White ethnicity); LOS <12 was achieved in 45.0% (n=68 752). The GAPS and AmbS had moderate predictive accuracy, with areas under the receiver operating characteristic curve (AUROCs) of 0.741 (95% CI 0.738 to 0.744) and 0.733 (95% CI 0.730 to 0.736), respectively. The SDEC-T comprised elements of the GAPS, AmbS, National Early Warning Score 2 (NEWS2) and presenting complaint and achieved an AUROC of 0.850 (95% CI 0.845 to 0.854) on internal validation. Stakeholders considered the tool acceptable and suitable for deployment across settings. Discussion The SDEC-T demonstrated improved discrimination using routinely available variables, balancing predictive performance with clinical practicality. Its design supports implementation across hospitals with varying digital maturity. Conclusions In a diverse patient cohort, the SDEC-T outperformed existing tools for identifying patients suitable for medical SDEC services.
Objectives Multidisciplinary team (MDT) meetings are key to delivering cancer care. Increasing caseload and limited resources make them less effective and unsustainable. The aim of this quality improvement project was to assess novel artificial intelligence-based clinical decision support (CDS) technology to develop and validate standard of care (SoC) to streamline the breast MDT meetings in a tertiary cancer centre.Methods A clinical governance group of the MDT approved international guidelines used to develop SoC. Deontics CDS was assessed for its suitability to apply the SoC pathway for benign and malignant breast disease with the exclusion of metastatic and recurrent cancer.Results Patients discussed over the preceding 16 months were added to the platform in cohorts of 50 women: two consisting of 50 women each diagnosed with benign disease (benign A and B: n=100) and three consisting of 50 women each diagnosed with malignant disease (cancer A, B and C: n=150). Concordance between the blinded MDT decision outcomes and SoC recommendations was analysed. This stepwise approach identified knowledge gaps in SoC and refined the CDS. Concordance improved from 82% to 100% in benign and from 94% to 100% in malignant cases.Discussion A sequential process of validating the SoC with data derived from the development of evidence-based SoC protocols based on international guidelines resulted in a final 100% concordance rate between the platform and MDT recommendations for both benign and malignant disease.Conclusions CDS technology could be a milestone in using SoC to deliver a sustainable clinical decision pathway.
Objectives Shigella remains a major cause of diarrhoea and mortality in children under five in low- and middle-income countries, where laboratory confirmation is often inaccessible and dysentery-based management lacks sensitivity. This study aimed to develop and internally evaluate machine learning models to predict microbiologically confirmed Shigella infection and secondarily to demonstrate the feasibility of translating the best-performing model into a prototype web-based decision-support application.Methods We analysed data from 3356 children with diarrhoea enrolled in the Global Enteric Multicentre Study, excluding co-infections and non-diarrhoeal controls. Multiple machine learning algorithms were trained using 32 predictors and evaluated with 10-fold cross-validation. Performance was assessed using area under the receiver operating characteristic curve (AUC), recall and Brier score, with recall prioritised to minimise false negatives. The selected model was simplified using the 10 most informative predictors and integrated into a prototype web-based application. All analyses were restricted to internal validation.Results Support vector machine (SVM) demonstrated the highest recall (0.64) with modest but potentially useful discrimination (AUC 0.74). The parsimonious SVM model achieved recall 0.67, AUC 0.77 and Brier score 0.16. A web-based prototype was developed to illustrate real-time model outputs for research purposes.Discussion The model achieved moderate internal performance while prioritising recall, supporting its potential role as a research-stage risk stratification tool in settings with limited diagnostic capacity.Conclusion This internally validated model shows potential for supporting risk stratification for shigellosis. External validation and impact evaluation are required before clinical or operational use.
OBJECTIVES:Effective clinician-patient communication is a fundamental pillar of high-quality clinical care and can influence patient engagement, treatment adherence and clinical outcomes. However, the extent to which clinical letters meet recommended readability standards remains largely unquantified. We aimed to benchmark this at scale in an ophthalmology setting-the highest-volume outpatient specialty in the United Kingdom (UK) healthcare service. METHODS:Over 4.6 million outpatient clinic letters written for 804 986 patients across 17 subspecialty services were included. Standard readability scores (including Flesch Reading Ease (FRE), SMOG Index, Automated Readability Index) were calculated for letters written between 2013 and 2025 at Moorfields Eye Hospital Foundation Trust, UK. Complex linguistic features such as lexical density and type-token ratio were also evaluated, and rule-based measures were applied to identify letters addressed directly to patients. Using the FRE as the primary outcome, readability was compared across subspecialties, and temporal trends were evaluated. RESULTS:The median FRE was 51.1 (IQR 41.9-58.5), indicating college-level readability which was substantially poorer than recommended for patient-facing materials. Overall, 95.3% of letters were less readable than recommended, and these patterns extended across all readability metrics evaluated. Readability varied across subspecialty services (p<0.001), with limited association with sociodemographic covariates such as age and ethnicity. There was a significant decline in FRE over the study period (-0.64 points/annum, p<0.001). Rate of change varied across subspecialty services and was more pronounced in some such as neuro-ophthalmology (-1.24 points/annum, p<0.001), with few being stable (optometry, p>0.05). The use of second person terms was very low (median 0 per 100 words, IQR 0.0-1.1). There was a small annual increase in the likelihood of letters being addressed directly to patients (OR 1.123, 95% CI 1.121 to 1.25), although the proportion remained low overall (2.25%). DISCUSSION:Clinical correspondence across ophthalmology services remained well above recommended readability thresholds and has become more complex over time, with substantial variability across subspecialty services. These findings will provide an empirical basis for local quality improvement initiatives. We hope that the methodological framework will also serve as a valuable resource for other specialties. CONCLUSION:Our findings highlight a persistent gap between policy recommendations and real-world practice, underscoring the need for targeted interventions and innovative approaches to support clearer patient-centred communication to optimise patient outcomes.
OBJECTIVES:To evaluate the feasibility, usability and validity of a co-designed digital questionnaire for early complexity screening in elective surgery patients. METHODS:A mixed-methods feasibility study was conducted using operational data, patient interviews and surveys. RESULTS:The average time to preoperative screening decreased from 3-4 weeks to 3.4 days. Of 244 patients, 44% were eligible for fast-tracking to low-complexity surgical hubs. A nurse-led validation process correctly identified 100% of medium-complexity and high-complexity cases. Most patients (97%) completed the questionnaire without assistance. DISCUSSION:The tool proved feasible and user-friendly, enabling earlier identification of complexity and streamlining patient triage. It also holds promise for reducing cancellations and facilitating timely prehabilitation. CONCLUSIONS:This feasibility study supports the tool's feasibility, usability and validity. Further research will assess its clinical, operational and health economic impacts within redesigned surgical pathways.
Objectives The rapid evolution of large language models (LLMs) and their growing application in clinical text processing have created an urgent need for reliable de-identification mechanisms. While LLMs show promise in identifying sensitive health information (SHI), their capabilities require rigorous evaluation. This study aims to conduct a comprehensive benchmarking analysis of various LLM-based, traditional rule-based and hybrid de-identification methods.Methods Our benchmark analysis used five datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016 and OpenDeID v1) from different countries. We developed three baseline and eight LLM-based models. The experimental setup encompassed nine different settings using various combinations of training and testing sets to assess model robustness and cross-dataset performance.Results In the baseline models, the approach trained on the combined corpus of all five datasets (setting 3) significantly outperformed the other settings, achieving a strict F1 micro-average score of 0.8172. Regarding LLM-based models, the supervised fine-tuning approach using the same combined configuration (setting 9) achieved the highest performance with a strict F1 score of 0.9447.Discussion The harmonisation of corpora ensured standardised data formatting and SHI management across five diverse datasets, highlighting the necessity for uniform categorisation to enhance the reliability of de-identification results.Conclusions Our findings indicate that while fine-tuned LLMs offer superior accuracy, the observed performance variability across heterogeneous electronic health record sources poses significant technical challenges. Real-world implementation must address these inconsistencies to overcome the ethical and technical hurdles associated with deploying LLMs for handling sensitive health data.
OBJECTIVE:This study developed and validated monolingual and bilingual sentence-bidirectional encoder representations from transformers (SBERT) models for detecting cancer recurrence within Thai-English electronic medical records (EMRs) from Thai cancer hospitals. METHOD:A multicentre dataset of 32 436 documents from 1250 patients was used for model development. External validation involved an independent dataset of 9244 documents from 384 patients across two Thai cancer hospitals. Performance was benchmarked against a fine-tuned PubMedBERT (MetBERT). RESULTS:The development dataset included breast (43.9%), colorectal (12.1%), cervical (28.0%) and head and neck (16.0%) cancers. MetBERT achieved the highest area under the precision-recall curve (AUPRC) for locoregional versus no recurrence (11.1%) and locoregional versus distant recurrence (91.7%), while monolingual-SBERT excelled at distant versus no recurrence (32.0%). External validation demonstrated MetBERT superiority for locoregional versus no recurrence (9.30%-21.50%). For distant versus no recurrence, bilingual-SBERT performed best with AUPRC 17.55%-24.39%. While MetBERT led in distinguishing locoregional versus distant recurrence (88.30%-94.70%), bilingual-SBERT demonstrated robust external validation performance (AUPRC 85.25%-91.80%). DISCUSSION:Low AUPRC values (9%-32%) reflect the extreme class imbalance in real-world data (~1% recurrence prevalence). Despite this, fine-tuned MetBERT achieved highest performance, while bilingual-SBERT demonstrated superior robustness during external validation. This validates sentence embedding models for handling mixed Thai-English medical records in multilingual clinical environments. CONCLUSION:Sentence embedding frameworks provide a practical, generalisable solution for detecting cancer recurrence within multilingual EMRs. Despite text-length constraints, these models are suitable for clinical integration as a screening tool for cancer registry workflows.
Objective To examine UK general practitioners’ (GPs) adoption of ambient artificial intelligence (AI) scribes and to assess user-reported error rates, workflow impact and consent practices in primary care. Methods We conducted a nationwide online mixed-methods survey of GPs recruited via Doctors.net.uk . Detailed analyses of use, errors and workflow impact focused on current users of ambient AI scribes. Results In August 2025, of 1003 respondents, 14% (n=141) reported current use of ambient AI scribes, 39% (n=396) intended to adopt them soon and 46% (n=466) had no plans to use them. Among users (n=141), Heidi Health predominated (86%). Most reported efficiency gains: 80% (n=112) reported reduced time spent on documentation and 70% (n=99) reduced cognitive load. Documentation quality was judged positively, with 55% (n=78) rating outputs as better than standard notes. Errors were common but usually minor: 32% (n=45) reported errors often/always, including 14% (n=20) with significant-to-critical implications. Errors were most frequent in multiparty consultations (38%), complex histories (35%) and non-English encounters (31%). Consent practices varied: 63% (n=89) routinely sought consent, with ≤10% of patients declining. Free-text responses (21% of users) highlighted benefits for workflow, alongside concerns about accuracy, ethics and system integration. Discussion Findings suggest that ambient AI scribes deliver meaningful efficiency gains and improved perceived documentation quality, but introduce non-trivial risks related to accuracy, equity and medicolegal accountability. The uneven performance in complex and multilingual consultations raises particular concerns about potential exacerbation of healthcare disparities. Conclusion Ambient AI scribes are already in use across UK primary care. Proactive regulation, consistent consent practices and independent evaluation, including patient perspectives, are urgently needed to ensure safe, equitable and sustainable implementation.
The conventional medical approach of treating symptoms as they appear with restricted screening often limits intervention to slowing disease progression rather than fully reversing it. A new approach leveraging artificial intelligence (AI) and computational technologies across expanding multimodal biomedical datasets holds the promise to enable predicting actionable future changes in health before symptom onset. This article presents a conceptual framework for a preventive paradigm in precision medicine and healthcare, integrating recent advancements in AI and biomedical datasets. Key remaining challenges facing computational systems, real-world clinical validation and implementation, and preventive interventions together with recommendations and prioritised future directions are highlighted. Such an approach could pave the way for more proactive and preventive medical interventions to effectively address the growing burden of chronic disease.
Objectives Social determinants of health (SDOH) may improve Alzheimer’s disease (AD) risk prediction by capturing upstream contextual risk beyond routinely measured clinical variables. We aimed to develop and validate an accurate, interpretable machine-learning pipeline for AD risk prediction in UK Biobank using routinely collected data.Methods Using data from 13 076 participants in the UK Biobank, we developed an automated machine-learning pipeline for AD risk prediction with feature selection and a C5.0 boosted-tree classifier. Data were split into training, development and test sets (7:2:1); missing values were imputed in the training data only, and feature selection, tuning and threshold calibration were performed using the training/development data, with final evaluation on the independent test set. Internal validation used repeated subsampling without replacement.Results During up to 16 years of follow-up, 927 participants developed AD. Feature selection reduced 3590 variables to 26 predictors spanning age, APOE4, SDOH, medical history and routine clinical measures. The final model showed good discrimination (area under the precision–recall curve 0.89) and adequate calibration (Hosmer-Lemeshow p=0.71), with stable performance under repeated subsampling. Sex-stratified models showed similar patterns.Discussion SDOH contributed useful predictive information, but their associations should be interpreted as predictive rather than causal and may reflect socioeconomic confounding and healthcare access.Conclusions This model could support scalable AD risk screening using routinely collected data, but external validation and recalibration in non-UK populations are needed before broader application.
Objective This scoping review aims to map and describe mental health indicators used in peer-reviewed studies in the WHO European Region to inform the development of a regional mental health measurement framework.Methods reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Review (PRISMA-ScR), this scoping review analysed 75 studies identified through database searches on mental health monitoring and indicators in the WHO European Region. The scope was restricted to English-language, peer-reviewed literature published from 2019 onwards. Indicators were extracted and standardised to reflect their original reported meaning, treated as distinct and pragmatically classified into nine predefined domains.Results Across the 75 studies, 450 distinct indicators were identified. Most indicators related to mental health status and mental health risk factors and/or determinants. Geographic coverage varied, with northern and western European countries most frequently represented among the included literature. Diverse measurement tools were employed.Discussion The predominance of indicators related to mental health status and determinants reflects the emphasis on burden and underlying factors, while system-level or structural indicators remain under-represented. Differences in geographic coverage likely stem from disparities in research capacity, publication practices and data availability.Conclusions This scoping review provides a descriptive overview of mental health indicators applied in recent peer-reviewed research in the Region. Greater alignment between research indicators and policy frameworks may strengthen the comparability and policy utility of mental health data across the Region.