
Objective By carefully accounting for the semantic structure of the International Classification of Diseases (ICD), this study aims to mitigate impact arising from incomplete recording of chronic diseases in ICD-coded data. Materials and methods We redefined four chronicity statuses (Groups). We then assigned chronicity status to ICD codes (from different versions: ICD-10-CM, ICD-10-CA, and ICD-11) using the semantic structures of ICD and SNOMED CT. Based on this status, an algorithm was developed to carry forward relevant codes across a patient’s multiple encounters. Codes classified as Group 2 (incurable diseases) were systematically assigned to all subsequent encounters for that patient, whereas Group 1 (chronic diseases) codes were assigned only to encounters occurring within one year of the first identified coding. We then tested the impact of this chronicity adjustment on current and adaptations of comorbidity/frailty indices (temporal stability and predictive performance) using different versions of ICD and two data sources: critically ill, heterogeneous Intensive Care Unit patients (MIMIC database) and a more homogeneous, procedure-specific hip-related cohort (SI-CPSS). Results Across each dataset and indices, we found that chronicity adjustment improved the temporal stability and precision of scores as evidenced by intra-class correlation. Adjusted indices consistently outperformed their unadjusted counterparts for hospital length-of-stay (LOS), 30-day readmission, and all-cause mortality prediction with improvement ranged, respectively in the SI-CPSS and MIMIC datasets, from 0.108 to 0.168 and 0.006 to 0.014 for readmission AUC, from 0.115 to 0.135 and 0.059 to 0.064 for mortality C-statistic and also from 0.081 to 0.330 and 0.003 to 0.010 for R2 LOS. Discussion-conclusion Our algorithm provides a better representation of comorbidities, enabling a more reliable identification, in databases, of patients with established comorbidity burden and strengthening the interpretation of patient pathways, with downstream benefits for risk stratification, cohort selection, and care planning.
Background Digital health tools (DHTs) are increasingly used in African clinical settings, yet implementation remains uneven across countries, facility levels, and technologies. Existing reviews are fragmented by disease area, tool type, or setting, and global syntheses rarely account for the unique environments shaping digital health implementation in Africa. Limited reviews have synthesized the barriers and facilitators influencing digital health implementation in African clinical care using implementation science frameworks. This umbrella review synthesized review-level evidence on these determinants using the Consolidated Framework for Implementation Research (CFIR). Methods This umbrella review included reviews published between 2013 and 2025 in PubMed, Scopus, and CINAHL that reported implementation determinants of digital health tools in African clinical care settings. The team extracted data in duplicate, appraised study quality using the JBI checklist, quantified overlap with the GROOVE tool, and deductively mapped determinants to CFIR domains with inductive coding for non-aligned themes. Results Six reviews covering 101 unique primary studies met inclusion criteria, spanning diverse African settings and digital health tools, including EMRs/EHRs, telemedicine, mHealth, AI-enabled tools, and other ICTs.None specified the implementation phase examined or explicitly reported using an implementation theory, model, or framework. Innovation-level, barriers included complexity, poor interoperability, limited usability, and cost, whereas relative advantage, adaptability, and co-design facilitated implementation. Implementer-level barriers included limited digital literacy, inadequate training, and resistance to change; training, mentorship, and early engagement were facilitators. Inner-setting barriers included unreliable infrastructure, limited organizational resources, workflow incompatibility, and inadequate technical support, while infrastructure investment, supervision, and incentives supported implementation. Outer-setting determinants included regulatory limitations, local attitudes, and partnership arrangements; community trust, government ownership, and cross-sector collaboration facilitated implementation. Implementation-process determinants were infrequently reported. Conclusion DHT implementation in African clinical settings is shaped by determinants across all CFIR domains. Sustainable implementation requires context-responsive strategies that address these factors collectively.297/300.
BACKGROUND:Although data reuse is increasingly common in healthcare research, changing regulatory frameworks are impeding these efforts. Synthetic data constitute a promising issue but also create challenges, such as the assessment of data fidelity and the ability to output identical statistical results. Conventional validation approaches often rely on univariate or bivariate comparisons, which fail to capture the complexity of associations between multivalued categorical variables. MATERIAL:We built and studied a fictitious database of 10,000 hospital stays reproducing the structure of the French Programme de Médicalisation des Systèmes d'Information database. Each stay included single-valued variables (one value per individual: sex, age in deciles, and diagnosis-related group) and multivalued variables (zero, one or several values per individual: diagnoses coded according to the International Classification of Diseases, 10th Edition, and procedures coded according to the French Classification Commune des Actes Médicaux). METHOD:All categorical variables were binarized, and thousands of pairwise association metrics (primarily odds ratios) were calculated for the reference and evaluation datasets. The results were summarized using curves, bubble charts, heatmaps, and coefficients such as exponential mean deviation. Simulated data degradations from 0% to 100% were introduced to evaluate the method's sensitivity. RESULTS:We analyzed 500 ICD-10 diagnoses and 450 CCAM procedures, representing 225,000 possible combinations. We developed and evaluated graphical representations for assessing data fidelity at a glance. In simulations of an increasing degree of data degradation, those graphical representations and comprehensive, quantitative metrics facilitated the detection of the gradual loss of data fidelity. CONCLUSION:We developed a simple, scalable, agnostic framework for assessing the fidelity of healthcare databases by systematically analyzing associations among the modalities of coded variables. This method complements existing approaches. It is particularly suitable for the evaluation of synthetic relational databases because it offers both general and granular insights into data fidelity loss.
BACKGROUND:Speech- and language-based digital biomarkers are increasingly proposed as scalable tools for mental health assessment, but their diagnostic accuracy and readiness for clinical translation in schizophrenia remain uncertain. METHODS:MEDLINE, Embase, PsycINFO, Web of Science, Scopus, IEEE Xplore, ACM Digital Library and citation searching were used without language or geographic restriction. Eligible full empirical journal articles and full conference proceedings evaluated automated speech/language classification in schizophrenia-spectrum disorders. After peer review, a second author independently re-screened random 10 % samples of title/abstract and full-text records and independently reassessed QUADAS-2. Patient-level 2 × 2 data were synthesized with a bivariate random-effects model; mixed comparators, segment-level observations and AUC-only datasets were kept separate. Additional exploratory sensitivity analyses conducted during revision addressed index-test risk of bias, patient-level splitting, model/metric selection and modality. Deeks funnel-plot asymmetry testing was performed for the primary set. RESULTS:Forty-six reports representing 47 datasets were included. The primary patient-level healthy-control set (k = 15; median total N = 100, range 16-284) yielded sensitivity 0.804 (95 % CI 0.743-0.853), specificity 0.829 (0.800-0.855) and SROC AUC 0.865. Between-study heterogeneity was substantial for sensitivity (τ2 = 0.328) but limited for false-positive rate (τ2 = 0.021), with correlation ρ = - 0.890 and prediction-region coordinate bounds of sensitivity 0.486-0.947 and false-positive rate 0.118-0.241 (specificity 0.759-0.882). After exclusion of one study with self-reported diagnoses, the exploratory psychiatric/mixed-comparator set (k = 3) yielded sensitivity 0.637 (0.505-0.751), specificity 0.882 (0.736-0.953) and SROC AUC 0.762; the random-effects correlation reached a boundary, so individual studies remain central to interpretation. Segment-level results (three studies; 940 segments) and all 12 AUC-only datasets were reported descriptively. No primary dataset provided verified independent external validation. Deeks testing did not indicate small-study asymmetry (t = 0.328, df = 13, P = 0.748). CONCLUSION:Automated speech and language models show promising internal discrimination, particularly against healthy controls, but the evidence does not establish real-world diagnostic utility. Clinically credible evaluation now requires locked-model prospective external validation in diagnostic-uncertainty populations, calibration and decision-curve reporting, standardized multilingual/multidevice acquisition, fairness assessment, interpretable outputs and explicit human oversight.
BACKGROUND:AI-based emergency department (ED) triage must balance predictive performance, clinical safety, and interpretability. Prior models suffer from target leakage due to Emergency Severity Index (ESI) scores, or prioritize accuracy while failing to detect high-risk patients. OBJECTIVE:To expose a pervasive clinical safety gap in ESI-free ED triage, accuracy-optimized models detecting fewer than 1% of high-risk patients despite AUROC near 0.80, and to show that imbalance-aware training is a broadly effective remedy across the evaluated model families. METHODS:Using MIMIC-IV-ED (N = 407,735 adult visits; ESI excluded from inputs and labels), outcome-based labels were High Risk (expired, transferred, or 30-day mortality), Medium (admitted, survived), and Low (discharged). DA-V2, a dual-attention model trained with focal loss, mixup, and five-seed ensembling, was compared against seven baselines and an identically-trained MLP; a vitals-only variant was externally validated. Attention-SHAP concordance was assessed by Spearman correlation. RESULTS:DA-V2 achieved AUROC 0.7949 (95% CI 0.7908-0.7991), outperforming all balanced tree baselines (p<0.001). An identically-trained MLP matched it (AUROC 0.7937, p=0.928; HR Recall 74.2% vs. 74.1%), showing that imbalance-aware training, not the architecture, drives detection. Focal loss with seed ensembling (mixup omitted, TabNet API) raised TabNet's HR Recall from under 1% to 44% (comparable to class-balanced TabNet, 45%) but did not reach the 74% attained only by the two dense models. Default models detected fewer than 1% of high-risk patients despite AUROC 0.796-0.800. Externally (NHAMCS-ED 2022, 12,598 adults; an approximate, proxy-label check), the vitals-only model retained HR Recall 0.759 at AUROC 0.695. Attention showed rank concordance with SHAP (Spearman ρ=0.708, p<10-12), evidence of rank agreement rather than causal faithfulness. CONCLUSIONS:A pervasive clinical safety gap, high AUROC with near-zero high-risk detection, is closed by imbalance-aware training and operating-point selection, not by any specific architecture. DA-V2 provides an interpretable, externally checked instantiation for trustworthy, ESI-free triage.
Background Predictive analytics is increasingly applied to routine and population health data, but its translation into operational health-system decisions remains uncertain. We examined implementation maturity, data integration, governance, workflow integration, uncertainty and documented decision pathways. Methods We conducted a global scoping review of peer-reviewed studies published from 2014 to 2025 across five databases. Eligible studies applied predictive or forecasting methods to routine healthcare or population-level data for health-system decision-making. Findings were synthesised descriptively and narratively and reported in accordance with PRISMA-ScR. An additional assessment of Overton, WHO IRIS and PAHO IRIS examined implementation evidence in grey literature. Results Of 2,623 screened records, 161 articles were included; 128 (79.5%) were from high-income settings. The highest documented stage was development in 139 articles (86.3%), validation in 10 (6.2%), pilot implementation in 3 (1.9%) and operational deployment in 9 (5.6%). Although 118 articles (73.3%) were positioned as relevant to resource allocation or capacity planning, only 9 (5.6%) documented an output-to-decision pathway and 14 (8.7%) reported routine workflow integration, revealing a marked claim-to-action reporting gap between stated relevance and documented action. Implementation barriers most often concerned data quality and interoperability (112; 69.6%) and validation and transportability (104; 64.6%). Only 24 articles (14.9%) described how uncertainty informed decisions. Supplementary grey literature identified three additional implementations, two operational and one pilot, supporting stroke-service planning, neighbourhood risk targeting and claims anomaly investigation. Conclusions Peer-reviewed evidence remains dominated by model development, while integration into decision pathways and routine workflows is infrequently documented. The three supplementary implementations from grey literature illustrate a practical role for predictive informatics in identifying system-level risks and directing planning, prevention or investigation. Advancing predictive analytics to operational decision support requires evaluation of the complete decision-integration chain, comprising decision actors, predictive outputs and associated uncertainty, delivery mechanisms, decision rules, actions, governance, workflow integration and lifecycle monitoring.
Objective Pathology reports are increasingly released directly to patients, but their diagnostic language is primarily designed for clinicians. This scoping review examined how large language models (LLMs) have been used for patient-facing pathology report interpretation, how generated outputs have been evaluated, and gaps related to health literacy, language, patient subgroups, and implementation. Methods This scoping review followed JBI methodology and PRISMA-ScR guidance. Six databases were searched for studies published from 1 January 2018 to 12 August 2026. Eligible studies reported empirical use or evaluation of LLMs for patient-facing interpretation, explanation, rewriting, question answering, or summarization based on pathology reports or report-like pathology texts. We extracted and descriptively synthesized study characteristics, LLM applications, patient-facing output tasks, evaluation methods, and reporting related to health literacy, language, patient subgroups, and implementation. Results Nineteen studies were included. GPT-family models were evaluated in 17 studies, and report-level transformation was the most common patient-facing task (12/19, 63.2%). Fidelity was assessed in all 19 studies, safety in ten (52.6%), readability in nine (47.4%), and comprehension and usability in six studies each (31.6%). Five studies included patients or other non-clinician participants in the evaluation. No study evaluated outcomes by health-literacy level or cross-language performance, and prospective evaluation within clinical communication workflows was not reported. Conclusions LLM-based patient-facing pathology report interpretation shows potential to bridge specialist pathology language and patient communication. Evaluation should extend beyond readability to include fidelity to the original pathology report, patient understanding, safety, and usability. Validation with patients and in clinical communication settings is needed before routine use.
OBJECTIVE:This systematic review aims to synthesise the present evidence base concerning artificial intelligence (AI)-supported total parenteral nutrition (TPN) management in neonatal intensive care units (NICUs) with respect to clinical efficacy, patient safety, and system integration. METHODS:This systematic review follows the PRISMA 2020 guideline. We searched PubMed/MEDLINE, Scopus, and Web of Science between January and March 2026, and used the PICO framework to include quantitative studies that employed AI, machine learning (ML), computerised physician order entry (CPOE), or clinical decision support systems (CDSS) in neonatal TPN management. We appraised methodological quality using the Cochrane RoB 2 tool for randomised controlled trials, the Newcastle-Ottawa Scale for observational studies, and the GRADE framework for the overall strength of the evidence. RESULTS:Sixteen records met the broad topical and technological inclusion criteria. Of these, thirteen quantitative primary studies (published 2008-2026, n = 30-9,330) form the evidence base for data synthesis and GRADE appraisal; three additional records provided historical and architectural context only. CPOE and rule-based CDSS significantly improved macronutrient target attainment, glycemic control, and medication safety. The TPN2.0 transformer model attained a Pearson R = 0.94 correlation with expert decisions, whereas classical ML algorithms achieved R2 > 0.70 in macronutrient prediction. CPOE implementation reduced the PN medication error rate from 10.8% to 3.2%. By contrast, only one-third of U.S. NICUs employed a CDSS. GRADE evidence was moderate for clinical efficacy and patient safety, and low for system integration. CONCLUSION:AI-supported TPN management is associated with favourable outcomes in NICUs with respect to clinical efficacy and patient safety, although this conclusion rests on only thirteen quantitative primary studies of predominantly moderate-to-low certainty and should be interpreted accordingly. The field is moving from CPOE-based automation towards deep learning models. We propose three priority areas for future research: multicentre randomised controlled trials measuring long-term neurodevelopmental outcomes, standardised TPN data repositories, and explainable AI design.
INTRODUCTION:Reduced sleep and mobility in acute care contribute to post-hospital syndrome, yet they are rarely monitored in this setting. Wearable devices like Fitbits can passively track these factors, but their validity in hospitalized patients remains uncertain. OBJECTIVES:To assess the validity of Fitbit Sense 2-derived heart rate, step counts, sleep and oxygen saturation in hospitalized internal medicine patients compared with validated reference devices; the secondary aim was to examine patient perceptions of continuous wearable device monitoring. METHODS:A prospective cohort pilot study was conducted over 24 h among internal medicine patients at a large academic hospital (Toronto, Canada). Fitbit data were compared against Nox T3s (sleep quality, heart rate, oxygen saturation) and StepWatch (step count) using Bland-Altman analysis and intraclass correlation coefficient (ICC). Sensitivity and positive predictive value (PPV) were calculated to detect sedentary, light and moderate activity (<40, 40-400, >400 steps/hr). Patient perceptions were collected via questionnaires. RESULTS:Of 33 patients enrolled (mean age 52.6 years), Fitbit data completeness varied: heart rate and step count 100 %, total sleep time and sleep efficiency 91 %, REM 56 %, oxygen saturation 41 %. Heart rate tracked closely with the reference (-0.3 bpm, ICC 0.998). Step counts were overestimated (bias + 29.4 steps/hr, ICC 0.64) but showed promise for sedentary detection (sensitivity 75 %; PPV 91 %). Total sleep time (bias -0.1 hrs; ICC 0.65) and sleep efficiency (bias -2.0 %; ICC 0.69) showed moderate agreement with wide limits of agreement (±4 hrs and ± 36 %, respectively). REM was overestimated (+29.4 min; ICC 0.41). Oxygen saturation showed modest agreement (bias + 1.7 %; ICC 0.47). Device usability was rated positively. CONCLUSION:In hospitalized internal medicine patients, the validity of wearables varied considerably by measure: heart rate and sedentary detection appeared promising, whereas absolute step count, sleep and oxygen saturation showed greater variability and higher rates of missing data. Patient perceptions of wearables were highly positive. These findings inform the design of larger validation studies. TRIAL IDENTIFIER:NCT07229833.
OBJECTIVE:Diagnosing atypical pathogen-induced infectious diseases is a complex and challenging problem, given their unique characteristics. A-priori identification of features using modern AI approaches can facilitate the diagnostic process, potentially leading to improved outcomes, early diagnosis, and effective resource utilization. We aimed to develop an AI-driven clinical decision support system for identifying rare and atypical infections - a composite of six microbiologically and clinically distinct diseases (blastomycosis, cryptococcosis, histoplasmosis, mucormycosis, pneumocystosis, and tuberculosis) unified by their shared pattern of diagnostic delay - in the early stages of hospital admission, which are often missed due to nonspecific symptoms, anchoring bias, and low clinical suspicion, leading to delays of 20-30 days. MATERIALS AND METHODS:We analyzed EHR data from 11,494 patients (1,494 with rare infections, 10,000 controls) admitted to Mayo Clinic hospitals (2010-2023). Machine learning models, including HistGradientBoosting, were trained on demographics, medical history, and laboratory results within three days of admission. SMOTE addressed class imbalance, and SHAP values provided interpretability. Performance was evaluated using accuracy, sensitivity, specificity, and AUC. RESULTS:The HistGradientBoosting model achieved an accuracy of 93.2 %, AUC of 0.947 (IQR 0.942-0.952), s pecificity of 97.3 %, and sensitivity of 61.9 % at the model's default classification threshold; threshold optimization for the intended flagging use case is planned prior to implementation. Key predictors included elevated phosphate, decreased lymphocyte counts, male sex, immunodeficiency, age, and calcium levels, aligning with clinical markers. CONCLUSION:The AI-driven clinical decision support system developed in this study shows strong potential for aiding the early diagnosis of rare infections. By leveraging EHR data and advanced machine learning models, this system could improve resource utilization, enable timely clinical interventions, and ultimately enhance patient outcomes. Further validation in external datasets and prospective trials is needed to confirm these findings and facilitate real-world implementation.
BACKGROUND:Generative AI lowers the technical barrier to clinician-led software development, but functional success and usability do not establish clinical or technical assurance. OBJECTIVE:To characterise defects that survived deployment in a clinician-built, AI-assisted ePROM application and the assurance activities that detected them. METHODS:Single-case retrospective development-and-assurance report of STUIapp, a browser-accessible application integrating six validated lower urinary tract symptom instruments. Assurance comprised structured post-hoc code audit, independent clinical review of a frozen 78-case scoring matrix, WCAG 2.1 parameter measurement, a 23-canary persistent-storage study, formative usability evaluation with 14 clinicians, 26 patients and 12 older adults, preliminary MDR positioning, and targeted post-hoc requirements traceability across 926 attributable author messages with independent classification. Descriptive statistics only. RESULTS:Mean SUS was 92.3 (SD 8.8) in clinicians and 92.0 (SD 10.8) in patients. Although all 78 original automated cases passed, independent review found 17 expected results requiring clinical or methodological correction. Technical findings included instrument mislabelling, a manifest that failed installability validation, and a third-party analytics tag contradicting the local-only privacy claim. Accessibility measurement found 10-px text and 2.56:1 contrast. Persistent-storage permission was denied in all browser-tab canaries (10/10) and granted in all installed canaries (13/13); no eviction was observed through 38.3 days. Home-screen addition was unaided in 7/12 older adults and 0/4 with limited smartphone use. Within selected defect classes, requirement traceability showed heterogeneous failure patterns. CONCLUSIONS:Usability and functional success did not establish clinical correctness, privacy, standards-based accessibility or regulatory preparedness. These findings support independent, domain-specific assurance of both implementation and the clinical expectations and non-functional requirements encoded in clinician-led AI-assisted clinical software.
INTRODUCTION:Non-attendance to elective surgeries contributes to inefficient healthcare delivery and utilization. Prediction models may be limited when outcome definitions combine patient-related factors with hospital-initiated cancellations, potentially obscuring the behavioral patterns underlying true patient non-attendance. This study aims to develop and evaluate machine learning models for predicting patient-related non-attendance to elective surgeries using electronic health records. METHODS:Electronic health record data from 24,507 elective surgery appointments scheduled between 2017 and 2022 at Shamir Medical Center, Israel, were analyzed. Hospital-initiated cancellations were excluded, and non-attendance outcomes were defined based on patient-related factors. Two prediction targets were evaluated: patient-related non-attendance at any time before surgery and within 24 h of the scheduled procedure. Six machine learning models, including Logistic regression, decision tree, random forest, and gradient boosting models were trained and evaluated. Model performance was assessed using the area under the receiver operating characteristic curve (AUC), precision, recall, the F1 measure, the area under the precision-recall curve, and the Brier score. RESULTS:The best-performing models, gradient boosting and random forest, achieved AUC values of approximately 0.93 for patient-related non-attendance both at any time before surgery and within 24 h of the scheduled procedure. Previous cancellation history was the strongest single predictor and alone achieved substantial predictive accuracy discrimination (AUC = 0.88), approaching that of the full machine learning models. By comparison, previously published models for surgical cancellation and non-attendance have reported AUC values that did not exceed 0.80. This improvement may be attributable to the isolation of patient-related cancellations from hospital-initiated cancellations, reducing outcome heterogeneity. CONCLUSION:Accurate prediction of patient-related non-attendance to elective surgeries is feasible using electronic health record data. These findings highlight the importance of outcome definition in healthcare prediction models and suggest that reducing outcome heterogeneity may improve identification of meaningful behavioral patterns. Simple indicators, such as previous cancellation history, may support targeted interventions to improve healthcare resource utilization.
BACKGROUND:Differentiating Kawasaki disease (KD) from other febrile illnesses remains challenging because of overlapping clinical and laboratory features. This study systematically evaluated the methodological quality, risk of bias, applicability, and predictive performance of machine learning (ML) and logistic regression-family (LR-family) models. METHODS:PubMed, Web of Science, and Embase were searched from January 1, 2006, to December 31, 2025, for studies developing or validating ML or LR-family models for differentiating KD from other febrile illnesses. Risk of bias and applicability were assessed using PROBAST + AI. The required minimum sample size for each study was formally estimated, and exploratory meta-analyses, subgroup analyses, and leave-one-out sensitivity analyses were performed. RESULTS:Twenty-eight studies (10 ML and 18 LR-family) were included. Common methodological limitations included inadequate sample size, retrospective design, inadequate handling of missing data, limited external validation, and poor calibration reporting; no ML study reported calibration metrics. Given the high risk of bias across all studies, pooled areas under the receiver operating characteristic curve (AUCs) were interpreted as exploratory quantitative summaries, showing no significant difference in internal validation performance between ML and LR-family models (0.95 [95% CI 0.91-0.97] vs. 0.92 [95% CI 0.89-0.94]; P = 0.222). CONCLUSIONS:Current evidence is limited by high risk of bias, substantial heterogeneity, and important methodological limitations, and limited independent external validation evidence precluded a reliable comparison of the external validation performance of ML and LR-family models. Accordingly, current prediction models are not yet ready for routine clinical decision-support deployment.
PURPOSE:Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS:Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS:Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION:Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.
Background Transfusion recipients are a heterogeneous group of patients, yet identifying these groups has traditionally relied on human-driven univariate analyses and domain knowledge instead of analyzing multivariate characteristics of individuals. Electronic health records (EHR) combined with unsupervised machine learning enables robust, data-driven way for phenotyping populations, providing finer-grained view on subgroup characteristics. Methods We introduce an extension to the Variational Autoencoder (VAE) framework and apply the model to EHR data of 19,629 adult transfusion recipients. The latent representation of VAEs approximates a low-dimensional manifold of input data, where patients with similar characteristics are embedded close to one another. The model integrates clustering via a Gaussian Mixture Model (GMM) prior to identify clinically relevant patient subgroups from diagnosis codes, laboratory values and demographics, while simultaneously classifying the type of transfused products. Final clusters are derived using a modified consensus clustering approach. Results We identified six patient groups with distinct diagnosis, laboratory, demographic, and transfusion profiles. These clusters, including obstetric, anemic and trauma patients as well as groups with comorbidities, provide a refined characterization of transfusion-related phenotypes, revealing distinctions among subgroups. Our model achieved moderate classification accuracy, with AUROC of 0.879, 0.806 and 0.861, and PR-AUC of 0.448, 0.357 and 0.492 for red blood cells (RBC), plasma and platelets, respectively. Clustering accuracy remains consistent across training and testing. Conclusions Deep representation learning identified clinically meaningful subgroups of transfusion recipients and yielded a level of phenotypic resolution that extends beyond prior characterization. The model helps to understand the heterogeneous nature of patients requiring transfusion and provides insights on how different blood product profiles shift cluster assignments. These findings underscore the utility of latent variable modelling for population characterization and suggest potential applications in predicting transfusion demand, supporting targeted patient blood management strategies, and improving blood supply planning. Validation in external cohorts remains unestablished.
Managing oncology drug inventories in hospitals is a critical yet complex task due to high treatment sensitivity, fluctuating patient demand, supply chain disruptions, and strict storage constraints. Traditional hospital inventory systems mainly provide descriptive monitoring and lack mechanisms for anticipatory decision support. This study proposes ODTIC (Oncology Drug Traceability and Intelligent Control), a knowledge graph-driven framework designed to enable proactive control of oncology drug inventories through semantic traceability and intelligent reasoning. The framework models relationships among oncology drugs, batches, inventory items, storage units, and hospital departments using an ontology-based knowledge graph. By combining Semantic Web Rule Language (SWRL) reasoning and SPARQL inference, the system automatically detects operational risks such as stock depletion, safety stock violations, expiration threats, and inter-department stock imbalance. When risks are identified, the framework generates semantically structured decision actions including alerts, reorder triggers, safety stock replenishment, stock redistribution, and therapeutic substitution recommendations. A case study conducted on fourteen oncology drugs distributed across multiple hospital departments demonstrates the capability of the proposed approach to detect risks and generate actionable decisions with low reasoning latency. The results highlight the potential of knowledge graph technologies to transform pharmaceutical traceability systems into proactive decision-support infrastructures.
OBJECTIVE:Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator. MATERIALS AND METHODS:We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications. RESULTS:The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training. DISCUSSION:VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends. CONCLUSION:The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.