
Unlabelled:To align with contemporary trends and place greater focus on persistent and novel challenges in the field, JMIR Medical Informatics has updated its focus and scope. Each submitted work will be evaluated against a stronger criterion for translational impact. Additionally, submissions on AI methods and applications are expected to meet standards of rigor and reporting transparency. In this editorial, we explain further what changed; why; and what this means for prospective authors, reviewers, and readers.
Background Personalized follow-up for type 2 diabetes may improve the alignment between monitoring intensity and patient needs, but operational approaches that jointly consider follow-up timing, modality, expected health benefits, and resource use remain limited. Objective This study develops a Markov decision process (MDP) framework for personalized follow-up planning, externally validates complementary risk-prediction models, and estimates the projected 12-month cost-effectiveness of model-generated follow-up strategies. Methods We retrospectively analyzed longitudinal data from 41,398 patients with type 2 diabetes managed in 10 community health centers in Nanjing, China, over the 2015-2024 calendar period. An independent Shanghai cohort included 25,506 patients from 10 communities. We developed least absolute shrinkage and selection operator (LASSO)-Cox models to predict 1-year incident complication and mortality risk and evaluated discrimination in Shanghai. Separately, a 12-cycle finite-horizon MDP used observed action-conditional state transitions with prespecified utility, cost, access, and willingness-to-pay parameters to generate personalized follow-up policies. Results For an example patient initially without recorded complications (S0), the personalized policy increased annual effectiveness by 0.48 quality-adjusted life days (QALDs; 0.0013 quality-adjusted life years [QALYs]) and cost by 33.46 CNY (1 CNY=US $0.15), yielding an incremental cost-effectiveness ratio (ICER) of 25,499 CNY/QALY. In a heterogeneous simulated cohort of 1000 patients initialized in S0, mean annual effectiveness increased from 313.52 to 314.05 QALDs (0.85896 to 0.86041 QALYs), and mean annual cost increased by 24.67 CNY, yielding an ICER of 17,016 CNY/QALY. External C-indices were 0.778 for incident complications and 0.812 for mortality. Deterministic sensitivity analyses did not materially alter the cost-effectiveness conclusion. Conclusions The MDP framework generated individualized 12-month follow-up policies with small projected QALY gains at modest incremental program cost, while the complementary risk models demonstrated external discrimination in an independent cohort. As action-conditional transitions were estimated from observational records and several economic parameters were prespecified, the modeled differences represent projections rather than causal treatment effects and require prospective implementation and economic validation before clinical adoption.
Background Augmented reality (AR) has emerged as a promising tool to enhance surgical precision during robot-assisted partial nephrectomy (RAPN), particularly by enabling the intraoperative overlay of 3D anatomical models. However, real-time AR implementation requires robust image segmentation of anatomical structures, such as the kidney, which remains technically challenging in dynamic laparoscopic environments. Objective This study aimed to develop and validate a large, annotated image dataset to train deep learning models for automated segmentation of the renal parenchyma during RAPN, as a prerequisite for real-time AR guidance. Methods We conducted a single-center, observational image annotation study using prospectively collected surgical videos from 131 RAPN procedures performed between 2022 and 2024. Patients had localized renal tumors, with 11 presenting with multifocal disease (160 tumors in total). A total of 48,000 frames were extracted based on image sharpness, diversity, and the exclusion of artifacts. A subset of 454 images was annotated by 9 contributors (surgeons, engineers, and nonexperts) after structured training. Interannotator agreement was assessed using the Dice similarity coefficient (DSC) and sensitivity against an expert reference. A convolutional neural network (AlbuNet-34) was trained using 12,546 annotated images and evaluated on a validation set of 3137 images. Model performance was analyzed overall and across surgical phases. Results Annotators achieved high agreement, with median DSC values ranging from 0.91 to 0.95 and sensitivity consistently more than 0.89. The deep learning model achieved a mean DSC of 0.75 (SD 0.23) and a sensitivity of 0.71 (SD 0.24) on the validation set. Segmentation accuracy varied significantly across surgical phases, with lower performance observed during tumor resection and tumor bed reconstruction due to increased visual complexity. Conclusions This study demonstrates the feasibility of automated renal parenchyma segmentation using deep learning in real-world intraoperative settings. Although current performance remains below that of expert-level annotations, the creation of a large, annotated dataset and the implementation of a structured multiannotator workflow represent key milestones toward reliable AR-assisted surgery. Ongoing refinements in annotation quality, dataset diversity, and neural network optimization are expected to enhance future real-time AR applications in urology.
Background:Following 2 decades of electronic medical record (EMR) adoption, most large public health systems hold comprehensive digital clinical data but lack the complementary informatics capability to return those data to clinicians, coders, and operational teams in a usable form. In South Australia (SA), the statewide Sunrise (Altera Digital Health) EMR system has digitized documentation and ordering since 2017; however, tools for back-end data extraction and enriched clinical information displays were not prioritized, and clinical, operational, and research users have reported ongoing difficulties accessing timely data. Objective:The aim of this study is to describe the development, governance, and deployment of a cloud-native health informatics system (HIS) in the Central Adelaide Local Health Network (CALHN), which serves approximately 40% of SA public patients, and to report implementation outcomes in accordance with the iCHECK-DH (Guidelines and Checklist for the Reporting on Digital Health Implementations) guidelines. Methods:The HIS comprises 6 architectural layers deployed on Microsoft Azure with Red Hat OpenShift (IBM), extracting Sunrise EMR data in near real time under a read-only model in which all patient data remain within the SA Health network at all times. Implementation proceeded through 5 overlapping phases (2021-2026), governed across 4 domains: technical (security impact assessment, information asset classification, and independent cybersecurity review), clinical (a clinical governance committee that has met quarterly since May 2022), corporate (incorporation of HeartAI Pty Ltd in 2022, with conflicts of interest declared and managed under SA Health policy), and ethical (human research ethics committee [HREC], reference 18079 with subsequent amendments). Development was funded through approximately Aus $1.84 million (Aus $1=US $0.72 as of August 24, 2026) in competitive and institutional grants under a public-private model with CALHN and AusHealth. Unlabelled:Three applications have reached deliberately different stages of maturity. The CALHN Critical Care Informatics System (CCCIS) has been implemented and evaluated: it has been deployed across the 2 CALHN intensive care units (ICUs; 54 beds, approximately 5000 admissions per year) since 2022, is in daily clinical use, and automates the submission of 115 variables to the Australian and New Zealand Intensive Care Society Centre for Outcome and Resource Evaluation (ANZICS CORE) registry. In a formal evaluation by 8 senior intensivists, its mean System Usability Scale score was 76 of 100, and the overall workload was low (NASA Task Load Index: 21/100). CODEXA, an AI-assisted clinical coding application, is in operational validation: models trained on 500,000 episodes and tested on 50,000 held-out episodes achieved a pooled F1-score of 71% across the full case mix and 57.86% (precision 64.72%; recall 55.53%) on complex acute episodes, in line with published benchmarks of 58% to 61% obtained from curated research datasets. The Patient Flow application is in co-design with CALHN's Network Operations Centre, with interface prototype evaluations completed in July 2026. In May 2026, the platform received unconditional endorsement from Digital Health SA's Technical Design Review Committee, the highest technical governance approval in SA Health, concluding a 5-year governance pathway. Conclusions:A locally developed informatics platform can be governed, deployed, and validated in a public health system; demonstrating clinical outcomes and transitioning to operational status are the next priorities.
BACKGROUND:Call abandonment is a critical barrier to patient access in health care call centers; yet, predictive modeling efforts are limited by strict privacy regulations that restrict the use of personal or behavioral data. Whether abandonment can be accurately predicted using only anonymized operational metrics remains unclear. OBJECTIVE:This study evaluated the feasibility, performance, and operational use of machine learning models trained exclusively on nonpersonal, routinely collected call center metrics to predict call abandonment across distinct organizational phases. METHODS:We analyzed 1,037,363 call records from a large academic health care system spanning 4 operational periods marked by workflow changes and skill consolidation. Features included temporal variables, skill identifiers, and rolling operational metrics (in-queue time, occupancy, handle time, after-call work time, and active agents). Random forest and CatBoost models were trained on 3 phases defined as: (T1 [January-April 2023; original workflows], T2 [May-August 2023; post skill consolidation, cross-training, and new workflows], and T3a [September-December 2023; optimized processes]) using 5-fold cross-validation with 3 imbalance-handling strategies (none, class weighting, and synthetic minority oversampling technique). Temporal generalizability was assessed by evaluating all models on all 4 phases, with T3b serving as an unseen holdout set. Performance was evaluated using area under the curve (AUC), precision-recall area under the curve (PR-AUC), Brier score, and calibration error. Shapley additive explanations (SHAP) values quantified feature contributions. RESULTS:Across all training phases and algorithms, adding operational metrics improved the area under the receiver operating characteristic curve by 0.03-0.13 vs models using only temporal and skill features. The best configuration was a CatBoost model trained on T2 with operational metrics and no imbalance correction (T3b AUC 0.767, PR-AUC 0.068, expected calibration error 0.006, and Brier 0.027). Models trained solely on preintervention data (T1) generalized poorly to postintervention periods when restricted to temporal and skill features (AUC 0.32-0.38) but achieved AUC 0.715 on T3b when operational metrics were included. SHAP analysis consistently identified in-queue time as the dominant predictor, with the number of logged-in agents, hour-of-day, and skill identifiers comprising the remaining top features. Abandonment declined from 8.7% (20,809/238,722) in T1 to 2.8% (8293/300,060) in T3b; model-based analyses of temporal features (day of week and hour of day) showed the highest risk on Mondays and between 11:00 and 16:00. Skill-level analyses showed marked improvement in high-volume imaging teams with high abandonment. CONCLUSIONS:Models trained on nonpersonal operational data predicted call abandonment on the T3b temporal holdout, with real-time queue and staffing metrics providing the dominant signal. These findings support the feasibility of interpretable, operationally grounded abandonment prediction in health care call centers and indicate that models should be retrained after major operational changes.
BACKGROUND:Background: Traditional Chinese Medicine (TCM) constitution is a structured health-state taxonomy used in preventive care, but its relationship with routinely collected health examination data, disease-related markers, and longitudinal change remains difficult to interpret in clinical informatics settings. OBJECTIVE:Objective: This study aimed to develop and evaluate a longitudinal clinical informatics framework for characterizing TCM constitution as a computable, explainable, record-based phenotype using routine health examination records. METHODS:Methods: We conducted a retrospective longitudinal analysis of 47,417 examination-constitution records from 11,355 older adults examined between 2017 and 2026. Baseline analyses used the first available record per participant, and longitudinal analyses used 32,648 adjacent annual visit pairs. The framework included bidirectional disease-constitution mapping, nonoverlapping multimarker burden modeling, lagged next-visit association models, constitution-state transition analysis, and temporal evaluation of routine-examination-based label prediction models. Additional sensitivity analyses evaluated participant overlap across calendar-year splits, participant-disjoint temporal evaluation, BMI-only and BMI-plus-waist baselines, exclusion of anthropometric predictors, adiposity adjustment of longitudinal models, and new-onset and persistence outcomes. RESULTS:Results: Phlegm-dampness showed the clearest record-based signature and was associated with higher cardiometabolic marker burden than balanced constitution (incidence rate ratio 1.65, 95% CI 1.60-1.70). In lagged models, its associations with a subsequent abdominal ultrasound abnormality flag and cardiometabolic risk clustering persisted after BMI adjustment, whereas associations with diabetes-related markers, proteinuria, dyslipidemia, and glucose abnormalities were substantially attenuated. Of 11,355 participants, 8122 appeared in at least two original calendar-year splits. In a participant-disjoint temporal sensitivity analysis with 1054 new test participants, XGBoost identified phlegm-dampness label presence with an area under the receiver operating characteristic curve of 0.927 (95% CI 0.912-0.941). A BMI-only model achieved 0.923 (95% CI 0.907-0.938), whereas XGBoost without anthropometric predictors achieved 0.685 (95% CI 0.653-0.717). CONCLUSIONS:Conclusions: Routine health examination data captured a reproducible, predominantly adiposity-centered phlegm-dampness label phenotype in this older adult cohort. High discrimination was retained in participant-disjoint evaluation but was nearly matched by BMI alone, and several longitudinal associations were explained by adiposity. The models should therefore be interpreted as decision-support and communication aids for an existing constitution assessment process, not as stand-alone diagnostic systems or comprehensive classifiers of TCM constitution. CLINICALTRIAL:
Background:Atrial fibrillation (AF) is a common arrhythmia associated with an increased risk of stroke and heart failure. To improve prevention, recent studies have used deep learning models to identify at-risk individuals early from normal sinus rhythm (NSR). However, studies using mobile electrocardiogram (mECG) in outpatient, real-world settings remain underexplored. Objective:The study aimed to develop and evaluate deep learning models using a real-world limb-lead mECG database to predict the short-term occurrence of AF from NSR recordings. Methods:mECG data were collected from real-world users of commercially available handheld mECG devices capable of capturing 6 limb leads. AF occurrence was defined as an AF event within a predefined time window (7, 14, or 31 d) from the date of the NSR recording. Transformer-based prediction models were developed for limb-lead and lead I input configurations using a multistage training approach with self-supervised pretraining and domain adaptation, drawing on both open, large-scale clinical 12-lead ECG and proprietary real-world mECG databases. The models were evaluated in an internal real-world cohort and explored in an external cohort as a proof of concept via time-to-event analysis. Results:Between March 2023 and November 2024, 386,519 mECGs were acquired from 8206 users. There were 18,949, 25,206, and 33,524 AF incidences within the 7-, 14-, and 31-day time windows. The models were pretrained with 787,257 12-lead ECGs and 202,689 mECGs, then fine-tuned to predict AF occurrence using 97,447 labeled mECGs. The limb-lead models achieved areas under the receiver operating characteristic curves (AUROCs) of 0.793, 0.785, and 0.787 for 7-, 14-, and 31-day predictions on the internal cohort, respectively, with a user-level AUROC of 0.702 for the 31-day prediction. These models significantly outperformed the lead I models (P<.001), supporting the value of multilead configurations. The multistage pretraining was essential, as single-source pretraining yielded lower AUROCs of 0.555 with mECGs only and 0.761 with 12-lead ECGs only for the 31-day prediction. In the subgroup analysis, AUROC values were consistent across age, PR interval, and corrected QT interval, but showed disparities (P<.001) by sex (0.713 in females vs 0.794 in males) and by QRS duration (0.583 in ≥120 ms vs 0.796 in <120 ms). In the external cohort (n=144), the 31-day model stratified all 5 new-onset AF events, showing significantly different survival functions between the positively and negatively predicted groups (P=.03); Cox proportional hazards regression yielded a hazard ratio of 1.49 (95% CI 1.06-2.09) per 0.1 increase in model output. Conclusions:Our findings elucidate the feasibility of deep learning-based AF risk prediction using single-NSR recordings from mobile devices, highlighting the potential for remote AF management in real-world populations. The model output may serve as a risk indicator to support opportunistic AF screening, prompting further clinical evaluation and informing decisions about more intensive monitoring.
Background:Despite advances in understanding and treating non-ST-elevation acute coronary syndrome (NSTE-ACS), patients continue to experience high rates of adverse outcomes, particularly those with non-ST-segment elevation myocardial infarction, which remains a leading cause of cardiovascular mortality. Existing risk models may not fully reflect contemporary patient populations due to substantial changes in clinical profiles. Developing new machine learning (ML)-based risk calculators may improve the prediction of in-hospital mortality (IHM) at different stages of the diagnostic process, and ultimately improve patient outcomes. Objective:This study aimed to develop predictive models for IHM in patients with NSTE-ACS using ML methods and predictor sets obtained during the diagnostic process. Methods:This retrospective observational study included 1144 patients with NSTE-ACS admitted between 2019 and 2021. IHM occurred in 94 (8.1%) of 1144 patients. Predictive models were developed using multivariable logistic regression, Random Forest, XGBoost (Extreme Gradient Boosting), and CatBoost algorithms. Model performance was evaluated using the area under the receiver operating characteristic curve (ROC-AUC), area under the precision-recall curve, calibration metrics, and decision curve analysis. Results:A key feature of the developed models was their applicability at different stages of the diagnostic process, using predictors available at each specific stage. At admission, ML models achieved ROC-AUC values up to 0.93. After incorporating laboratory and echocardiographic data, predictive performance increased to ROC-AUC values of 0.94 to 0.95. The best-performing model demonstrated an ROC-AUC of 0.951 (95% CI 0.946-0.956), a sensitivity of 0.902 (95% CI 0.889-0.916), and a specificity of 0.891 (95% CI 0.886-0.896). The area under the precision-recall curve reached 0.685, and the Brier score was 0.0331, indicating good calibration. Random Forest models demonstrated greater clinical utility than the Global Registry of Acute Coronary Events score (P<.001). Shapley Additive Explanations analysis identified the most significant predictors of mortality risk, including Killip class of acute heart failure, age, Charlson comorbidity index, creatinine level, and hematocrit level. Conclusions:ML-based models enabled accurate prediction of IHM in patients with NSTE-ACS at different stages of the diagnostic process and may improve risk stratification and clinical decision-making in real-world practice.
Abstract Background Screening for atherosclerosis is essential for early intervention, but conventional screening methods are often invasive and resource-intensive. As a result, there is growing interest in leveraging AI with noninvasive tools such as retinal fundus imaging to enable opportunistic cardiovascular risk assessment. The deep-learning funduscopic atherosclerosis score (DL-FAS) is an AI-derived biomarker, generated by a deep learning model, that was developed in a previous study to reflect the likelihood of carotid artery atherosclerosis from retinal fundus images. Objective This study aimed to externally validate DL-FAS in a multicenter health checkup population and investigate its association with coronary artery calcification to gain deeper insights into its ability to reflect the broader systemic atherosclerotic burden. Methods We used data from 108,982 participants in a Korean health checkup population who underwent retinal fundus imaging and at least one of either carotid artery sonography or coronary artery calcium scoring across 5 health-promotion centers operated by the Korea Association of Health Promotion between 2018 and 2021. Carotid atherosclerosis was defined by increased intima-media thickness (≥ 0.9 mm), atheroma, or stenosis. A coronary artery calcium score >0 indicated the presence of coronary artery calcification. DL-FAS (range 0‐1) was generated for each fundus image, and the average score from both eyes was used as the final DL-FAS when available. The discriminative performance of DL-FAS for carotid atherosclerosis was assessed using the area under the receiver operating characteristic curve. We performed multivariable logistic regression, adjusted for the Pooled Cohort Equations (PCE) 10-year cardiovascular risk score, to quantify the associations between DL-FAS and both outcomes. Results The area under the receiver operating characteristic curve for detecting carotid atherosclerosis was 0.700 (95% CI 0.697‐0.703), confirming the generalizability of DL-FAS across multiple centers. In multivariable logistic regression adjusted for the PCE score, a 10% absolute increase in DL-FAS was associated with both carotid atherosclerosis (odds ratio 1.29, 95% CI 1.28‐1.31) and coronary artery calcification (odds ratio 1.29, 95% CI 1.24‐1.33). These associations remained significant among participants aged <60 years, as well as within the PCE-defined low- and moderate-risk subgroups, highlighting the potential utility of DL-FAS in populations that may benefit most from early detection and intervention. Conclusions DL-FAS was externally validated in a large, multicenter health checkup dataset and was significantly associated with both carotid atherosclerosis and coronary artery calcification. These findings highlight its potential as a noninvasive biomarker associated with systemic atherosclerotic burden and cardiovascular risk stratification, particularly in the context of opportunistic screening.
Background:Sepsis-associated acute kidney injury (SA-AKI) is a frequent and life-threatening complication of sepsis. While the static lactate-to-albumin ratio (LAR) has prognostic value, its dynamic temporal evolution during early resuscitation and its utility for guiding clinical risk stratification remain underexplored. Objective:The aim of the study is to identify distinct dynamic trajectories of LAR, evaluate their independent associations with adverse clinical outcomes, and construct a practical prognostic nomogram for patients with SA-AKI. Methods:Data were extracted from the Medical Information Mart for Intensive Care IV database, including adult patients with SA-AKI and ≥ 3 lactate and albumin measurements within the first 72 hours of intensive care unit admission. A multicenter external validation cohort was assembled from 5 tertiary hospitals in Beijing, China. Group-based trajectory modeling identified distinct LAR trajectories. The primary outcome was 28-day mortality; secondary outcomes included 90-day mortality and continuous renal replacement therapy initiation. Trajectory-outcome associations were assessed using multivariable Cox and logistic regression models. Robustness was examined using restricted cubic splines, inverse probability of treatment weighting, weight truncation, doubly robust estimation, and subgroup analyses. Incremental predictive value beyond single baseline LAR was quantified, and a trajectory-integrated nomogram was developed and externally validated. Results:Among the 615 patients in the primary cohort, 3 LAR trajectories were identified: trajectory 1 (low-stable type), trajectory 2 (rapid-clearance type), and trajectory 3 (delayed-clearance type). In adjusted multivariable models, trajectory 3 (vs trajectory 1) was independently associated with increased 28-day mortality risk (hazard ratio 1.63, 95% CI 1.00-2.64; P=.0496), 90-day mortality (hazard ratio 1.72, 95% CI 1.14-2.60; P=.01), and continuous renal replacement therapy initiation (odds ratio 3.40, 95% CI 1.64-7.15; P=.001). Conversely, the mortality risk for trajectory 2 did not differ significantly from trajectory 1. Sensitivity analyses supported these findings, although trajectory 3 associations attenuated in inverse probability of treatment weighting-weighted analyses. Incorporating trajectories improved predictive accuracy over single baseline LAR (continuous net reclassification improvement 0.170, integrated discrimination improvement 0.023; both P=.01). In the external validation cohort (n=508), 3 analogous trajectories were identified. Trajectory 3 consistently conferred a higher mortality risk across both cohorts, whereas trajectory 2 was significantly associated with 28-day mortality only in validation. The nomogram exhibited modest discrimination in the external cohort (concordance index=0.612), acceptable calibration, and potential utility as a supplementary bedside risk stratification tool. Conclusions:Dynamic LAR trajectories are independently associated with 28- and 90-day mortality in SA-AKI. The delayed-clearance trajectory identifies a specific high-risk phenotype, suggesting that longitudinal LAR monitoring and the constructed nomogram may support early risk stratification and inform clinical decision-making.
Background:Widespread and sustained uptake of AI-based clinical decision support systems (CDSSs) in real-world health care settings is uncommon, despite their potential to improve patient care and reduce clinician burnout. Although previous studies have examined determinants of implementing AI-based CDSSs, limited evidence has synthesized barriers and facilitators identified during actual clinical implementation and use. Objective:The objectives of this scoping review were to (1) map and synthesize barriers to and facilitators of implementing AI-based CDSSs in real-world health care settings and (2) draw on this knowledge to inform future implementation strategies. Methods:Five electronic databases (MEDLINE, Embase, CINAHL, APA PsycInfo, and the Cochrane Library) were searched from inception to May 2022. Eligible studies included primary research describing real-world implementation processes or reporting determinants (barriers and facilitators) of implemented AI-based CDSSs in any health care setting. Studies focused on non-decision support tasks, non-AI CDSSs, patient-facing tools, or development or effectiveness without implementation were excluded. No study design restrictions were applied. Full texts were reviewed to extract explicit statements describing determinants influencing implementation. These determinants were classified as barriers or facilitators and mapped to the Consolidated Framework for Implementation Research (CFIR) by 2 independent reviewers. A qualitative synthesis was conducted. Results:After removing 4234 duplicate records, 10,875 articles were screened by title and abstract, which excluded 10,355 articles. After further exclusions based on full-text availability, 494 full-text articles were assessed for eligibility, of which 13 met the inclusion criteria. Nine of these studies reported explicit implementation determinants and were included in the CFIR-based synthesis. Studies were primarily conducted in the United States and involved multicenter implementation of machine learning-based CDSSs in critical care and emergency medicine settings. A total of 28 determinants (16 barriers and 12 facilitators) were identified. Barriers were most frequently mapped to the inner setting, innovation, and individuals domains, whereas facilitators were most frequently mapped to the implementation process and innovation domains. Common barriers included limited algorithm interpretability, data quality and management challenges, misalignment with clinical workflows, and insufficient user capability and motivation. Facilitators included early and ongoing assessment of end-user needs, stakeholder engagement, peer endorsement, and robust supporting evidence. Conclusions:This review identified key determinants influencing the real-world implementation of AI-based CDSSs, highlighting the importance of system design, organizational context, and implementation strategies. However, the small number of studies reporting explicit implementation determinants underscores a critical gap in the literature, suggesting that many real-world implementations do not adequately evaluate or report factors influencing adoption and sustained use. Addressing this gap will be essential for advancing the translation of AI-based CDSSs into routine clinical practice. These findings provide a foundation for developing targeted implementation strategies and emphasize the need for more rigorous, implementation-focused research in real-world health care settings.
BackgroundThe prediction of weaning from mechanical ventilation (MV) can support clinical decision-making and help reduce the risk of weaning failure in intensive care units (ICUs). Cross-silo federated learning (FL) offers a promising approach to developing robust predictive models across multiple institutions without requiring the sharing of patient-level data. ObjectiveThis study aimed to evaluate the feasibility and efficacy of FL for predicting successful weaning from MV across 5 diverse ICU databases and to compare its performance with local learning (LL) and centralized learning (CL) approaches that differ in their data-sharing requirements. MethodsWe conducted a retrospective analysis using 5 ICU databases, namely the eICU Collaborative Research Database (eICU-CRD), Medical Information Mart for Intensive Care IV (MIMIC-IV), Universitätsklinikum Augsburg (UKA), High-Resolution ICU Dataset (HiRID), and Amsterdam University Medical Centers (AUMC), transforming clinical variables into the Observational Medical Outcomes Partnership (OMOP) Common Data Model. We defined successful weaning as a sustained reduction in positive end-expiratory pressure. We compared 3 learning approaches, FL, LL, and CL, using extreme gradient boosting (XGBoost). Performance was evaluated using the area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), precision, recall, and F1-score. All data use complied with local ethical regulations and institutional review board approvals, and all databases contained deidentified patient data accessed under institutional data use agreements and governed by applicable privacy regulations. ResultsA total of 24,521 patients were included across 5 databases. The CL model achieved an AUROC of 0.81, AUPRC of 0.57, and F1-score of 0.54 on pooled test data. The FL model achieved a macroaveraged AUROC of 0.74, AUPRC of 0.56, and F1-score of 0.52. LL model performance varied across databases (AUROC=0.68-0.84, AUPRC=0.51-0.68, F1-score=0.49-0.67), reflecting differences in data distribution and class balance. ConclusionsOur findings highlight performance differences between learning approaches for MV weaning prediction. LL models achieved the highest performance within their respective institutions (AUROC=0.68-0.84), CL achieved the highest performance on pooled test data (AUROC=0.81), and FL showed lower but reasonable performance while avoiding direct sharing of patient-level data (AUROC=0.74). The choice between approaches depends on institutional data-sharing constraints, local dataset characteristics, and acceptable performance thresholds. As privacy was not formally measured, these findings should be read as a performance comparison rather than a privacy-performance trade-off.
BACKGROUND:The Charlson Comorbidity Index (CCI) is widely used to quantify comorbidity burden in observational research. Applying it to real-world data requires accurate International Classification of Diseases-based mappings. The commonly used mapping does not reflect current German coding standards. OBJECTIVE:This study aimed to develop year-specific mappings (2004 to 2026) based on the German modification of the International Classification of Diseases, 10th Revision (ICD-10-GM), for the CCI with an open-source implementation and compare them with the commonly used mapping. METHODS:For each year, mappings were curated via independent dual review with consensus. In 515,827 inpatient cases from a German tertiary center (2010-2024), year-appropriate mappings and the commonly used mapping were applied, and the results were pooled into 2 corresponding datasets that were subsequently compared. Agreement was assessed using the intraclass correlation coefficient (2-way mixed-effects model for absolute agreement based on single measurement), and the significance of the discordances was measured using the McNemar test. RESULTS:Agreement was very high; 3.1% (16,136/515,827) of cases differed. A single code was present in 48.9% (7896/16,136) of the discordant cases and was intentionally excluded in the new mappings as it was not diagnostic for the category it was assigned to. Discordant results were found in 58.8% (10/17) of the categories ("peripheral vascular disease," "mild liver disease," and "severe liver disease") being present in, respectively, 51.7% (8339/16,136), 17.4% (2813/16,136), and 16.4% (2657/16,136) of the discordant cases. After filtering the dataset to only include observations with any diagnosis of a malignant disease, 6.9% (9377/136,607) of cases differed, and 79.3% (7434/9377) of those differences were caused by the "peripheral vascular disease" category, suggesting significant discrepancies for specific research questions. CONCLUSIONS:We provide transparent, year-specific ICD-10-GM (2004-2026) mappings and an open-source tool for CCI calculation. The mappings align with current German coding practice and enable reproducible, year-appropriate comorbidity adjustment using administrative data.
Background:Overactive bladder (OAB) is a prevalent condition, particularly among women, characterized by urinary urgency, often accompanied by frequency and nocturia. Traditional risk prediction methods for OAB are limited, as they fail to fully integrate multidimensional risk factors, including female reproductive history. Machine learning offers potential for enhanced predictive accuracy by using large-scale datasets like the National Health and Nutrition Examination Survey (NHANES). Objective:This study aimed to develop and validate a machine learning-based model to predict OAB risk in women, incorporating reproductive and sociodemographic factors, and to identify key predictors using interpretable methods. Methods:This retrospective observational study analyzed data from 7884 participants across 4 consecutive cycles (2011-2018) of the National Health and Nutrition Examination Survey. LASSO (least absolute shrinkage and selection operator) regression and univariate and multivariate logistic regression analyses were applied to identify key variables in the training set. Fourteen variables were selected via LASSO regression for model construction, among which age, BMI, ratio of family income to poverty threshold (PIR), age at menarche, and number of vaginal deliveries were identified as the most significant clinical predictors. The SHAP (Shapley Additive Explanations) method interpreted the optimal model, and restricted cubic spline (RCS) curves were used for dose-response analysis. Results:Five variables were identified as significant predictors. Among the 11 ML models, random forest (RF) demonstrated the highest predictive performance. The random forest model achieved an AUROC (area under the receiver operating characteristic curve) of 0.8536 (95% CI 0.8435-0.8638) in the training set and 0.6999 (95% CI 0.6768-0.7212) in the test set, indicating moderate predictive capability. SHAP analysis identified age, BMI, and the number of vaginal deliveries as the top 3 contributors to OAB risk. Both RCS and SHAP analyses revealed a positive association of age and BMI with OAB risk and a negative association with PIR. Additionally, RCS showed that the risk of OAB was higher with an earlier age at menarche and a greater number of vaginal deliveries. Conclusions:Integrating ML with SHAP interpretability provides a robust predictive tool for OAB, facilitating early identification and clinical management.
Unlabelled:AI is rapidly expanding in health care. However, there is a significant underrepresentation of low- and middle-income countries (LMICs) in datasets used to train AI applications. Current Food and Drug Administration (FDA)-approved AI tools predominantly use data from high-income countries, with fewer than 4% reporting geographic or racial diversity. This geographic skew leads to significant performance degradation when these tools are used in LMIC populations, perpetuating an equity crisis where health burdens are highest., Addressing this disparity, we highlight the Medical Imaging Datasets for India (MIDAS) initiative as a viable model to transition LMICs from "data poverty" to "data sovereignty." MIDAS uses a rigorous, 4-domain Dataset Quality Matrix to ensure representativeness, documentation, technical fidelity, and governance, thereby creating openly benchmarked, gold-standard datasets tailored to local contexts. Initial releases, including datasets for oral and dural lesions, demonstrate the feasibility and practical value of developing robust, generalizable AI models., Furthermore, we propose a multilateral South-South Data Commons structured around 3 foundational pillars: a harmonized dataset-grading rubric, distributed custodial governance, and outcome-linked incentives. This infrastructure supports local stewardship, encourages global collaboration, and ensures that quality benchmarks drive financial incentives for dataset expansion and diversity., This proposed framework not only positions LMICs as autonomous data stewards but also enhances global AI equity. By institutionalizing quality control, interoperability, and outcome accountability, LMICs can transform from passive data consumers into active contributors to essential, trustworthy, and globally relevant biomedical datasets.
Background:The medical burden caused by stroke is increasingly severe, and a small minority of high-cost patients consume the majority of medical expenditures. Therefore, revealing the formation mechanisms of this population and exploring a scientific cost-risk stratification system are crucial for improving the quality of care and achieving the optimal allocation of medical resources. Objective:This study aimed to construct a comorbidity network for patients with stroke using standardized front-page medical record data, extract network features that reflect complex disease interactions, and develop identification models in combination with machine learning algorithms. The study focused on building a core model integrating variables from the near-discharge stage for stratifying the risk of high hospitalization costs in patients at the near-discharge stage. In addition, an early prediction model was developed using only data available at admission. Methods:We conducted a retrospective study, collecting the hospital discharge data of inpatients with stroke from a tertiary hospital in Northeast China between 2021 and 2023. The data from 2021 to 2022 were used to construct a network and extract features to capture the potential relationship between diseases and high costs. Using the 2023 data partitioned into training and testing sets, we developed 5 models to identify inpatients with stroke who incurred high hospitalization costs and compared their performance when input with different features. In addition, the Shapley Additive Explanations interpretability method was adopted to explain the global and local contributions of the model features. Results:The inclusion of network features significantly improved the model's performance, among which Extreme Gradient Boosting performed the best. The global feature importance showed that network features occupied a major proportion. The results of the Shapley Additive Explanations interaction analysis indicated potential phased changes in patient resource consumption. However, the overall performance of the early identification model constructed solely from admission data was subject to clear limitations. Conclusions:This study developed an integrated framework combining comorbidity network analysis with machine learning, which significantly improved the accuracy of identifying inpatients with stroke at high risk of incurring excessive hospitalization costs. The core model demonstrated good performance in risk stratification during the near-discharge stage, showing potential for application in the formulation of risk management strategies and the optimization of health care resource allocation. It also laid the foundation for the subsequent development of more accurate early identification models.
Unlabelled:Coiera and Fraile-Navarro question whether AI scribes are being evaluated on metrics that truly impact care. While current evaluations focus on the quality of the initial draft, signed clinical notes are dynamic, as their content can be copied, summarized, coded, and re-ingested by downstream AI tools. We argue that safety must be measured downstream, focusing on how small errors in initial documentation can compound across the patient's electronic health record.
Abstract The clinical management of acute exacerbations of chronic obstructive pulmonary disease (AECOPD) is increasingly moving from reactive treatment to proactive early warning. However, existing environmental forecasting approaches remain constrained by heterogeneous exposure-response relationships, limited AECOPD-specific evidence for digital behavioral surveillance, and vulnerability to concept drift under nonstationary social and health care conditions. This viewpoint argues for a drift-aware multimodal early warning research agenda that integrates environmental exposure data, candidate digital behavioral signals, and routinely collected clinical burden indicators. Drawing on representative literature from environmental epidemiology, respiratory medicine, digital epidemiology, and medical informatics, we discuss current methodological challenges and potential directions for future AECOPD surveillance. Particular attention is paid to multimodal data integration, concept drift, adaptive state-space modeling, regime-aware handling of structural breaks, and equity-related challenges associated with digital behavioral data. Rather than presenting a validated forecasting system, we outline conceptual design considerations for future AECOPD-specific studies, including multimodal data fusion, state-space adaptive updating, drift subtype diagnosis, and subgroup-aware validation strategies. We also highlight important evidence gaps, particularly regarding the use of internet search queries and social media signals for AECOPD prediction. Drift-aware multimodal surveillance represents a promising direction for future AECOPD early warning, but substantial methodological, clinical, and implementation challenges remain. Future research should prioritize disease-specific validation, transparent evaluation of adaptive forecasting methods, and equitable deployment across populations with differing levels of digital access and health care resources.
Background:Noninvasive brain stimulation may alleviate social media addiction, but its efficacy requires accurate individual targeting and real-time brain monitoring. The neural mechanisms underlying the effects of social media use (SMU) remain unclear, limiting the development of interventions. Understanding how different levels of SMU modulate brain activity could guide personalized neuromodulation strategies. Objective:This study investigated a cohort of young adults using a multistage design, with concurrent electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) to examine the effects of SMU on brain activity. It aimed to characterize static and dynamic neural changes at baseline and after a standardized SMU task in individuals with different daily SMU durations. Methods:Participants were all male and were divided into a heavy social media users (HSMU) group and a light social media users (LSMU) group based on self-reported daily SMU duration. All participants underwent baseline fMRI scanning, followed by an EEG-fMRI session immediately after a 2-hour controlled SMU task. Analyses were performed on the static and dynamic amplitudes of low-frequency fluctuations (sALFF and dALFF), static and dynamic functional connectivity (sFC and dFC), and EEG microstates. Results:At baseline, compared with the LSMU group, the HSMU group showed lower dALFF variability in the middle frontal gyrus. After the immediate-effect task, the LSMU group exhibited increased sALFF in the temporal lobe and decreased sALFF in the middle and superior frontal gyri. The HSMU group showed increased sALFF in the middle temporal gyrus and decreased sALFF in the inferior temporal gyrus, superior parietal gyrus, and prefrontal cortex. Regarding dALFF variability, the LSMU group showed a decrease in the superior medial frontal gyrus, whereas the HSMU group showed a decrease in the middle frontal gyrus and an increase in the calcarine cortex. sFC analysis revealed increased connectivity across nearly all networks in the LSMU group. Conversely, the HSMU group showed reduced sFC between the visual and default mode networks. The HSMU group showed significantly shorter duration of microstate A, shorter duration and lower coverage of microstate C, and longer duration and higher coverage of microstate D. Conclusions:This study is the first to characterize the distinct neural patterns associated with different levels of daily SMU, using both EEG and fMRI to assess sALFF and dALFF alterations, widespread functional connectivity changes, and EEG microstate reorganizations. These findings demonstrate the unique value of multimodal assessment in identifying potential neural targets for personalized neuromodulation in social media addiction. Future studies should explore whether modulating these identified neural markers can effectively alleviate addictive behaviors and improve clinical outcomes across diverse populations.