
Infective endocarditis (IE) is a severe cardiac infection disease. This study aimed to explore 30-day mortality risk factors in patients with IE, evaluate machine learning (ML) models, and establish an interpretable ML model. A retrospective study including 391 IE patients (January 2017 and December 2024) was conducted. Feature selection was executed using the least absolute shrinkage and selection operator (LASSO). Eleven machine learning algorithms were used to construct prediction models. The area under the curve (AUC), sensitivity, specificity, positive predictive value, negative predictive value, F1 score and calibration curve analysis were used to evaluate performance. The Shapley Additive exPlanations (SHAP) method was used to model interpretability. Sixteen independent predictors were identified. Random Forest (RF) performed best, with an AUC of 1.00 (development) and 0.819 (validation). SHAP analysis showed that the top eight important indicators were: lactate dehydrogenase to lymphocyte percentage ratio (LLPR), surge, C-reactive protein (CRP), activated partial thromboplastin time (APTT), urea nitrogen (Urea), platelets (PLT), prothrombin time-international normalized ratio (PT-INR), and triglycerides (TG). Using these top eight features, the RF model achieved its highest performance (development AUC = 1.00, validation AUC = 0.813). An RF model based on eight key features effectively predicts 30-day mortality in IE patients with good interpretability and performance.
Cardiovascular diseases (CVD) remain a leading cause of morbidity and mortality worldwide, with arrhythmia being one of their common manifestations. Traditional diagnostic methods, such as electrocardiography (ECG), rely heavily on expert interpretation for detecting irregular heart rhythms, which can lead to inaccuracies and delays in diagnosis, particularly in resource-limited settings. In recent years, machine learning (ML) has emerged as a powerful tool for improving the accuracy and efficiency of medical diagnostics. Although deep learning models have demonstrated remarkable success in various domains, their computational complexity poses challenges for real-time applications. This study investigates various quantum-classical hybrid algorithms for improving ECG-based arrhythmia classification. In this study, we employed three hybrid quantum-classical algorithms for arrhythmia detection, estimator-based Quantum Neural Network (QNN), sampler-based QNN, and Quantum Support Vector Classifier (QSVC), to enhance the accuracy of automated arrhythmia classification. The models are evaluated using the MIT-BIH arrhythmia database. The experimental results demonstrated that these three quantum models outperform traditional Recurrent Neural Networks (RNNs) and their variants in several metrics, including accuracy, sensitivity, and specificity. The QSVC (k = 16) model achieved the highest performance, with an accuracy of 97.7
To evaluate the diagnostic performance of an artificial intelligence (AI)-assisted system for detecting dental caries and staging periodontitis using periapical radiographs, and to assess its impact on diagnostic accuracy among clinicians with different levels of clinical experience. This study comprised retrospective AI model development and internal evaluation using 2700 periapical radiographs containing 10,211 annotated teeth, followed by a prospective single-center clinical reader study involving 54 patients. Dental caries detection was performed using a ConvNeXt-Tiny model, whereas periodontitis staging was conducted using a multi-stage framework integrating You Only Look Once (YOLO) instance segmentation, YOLO semantic segmentation, residual neural network (ResNet)-50 regression, and rule-based post-processing. Model performance was assessed using accuracy, sensitivity, specificity, precision, recall, F1-score, and AUC. Diagnostic performance was further compared between experienced and less experienced dentists, with and without AI assistance. AI assistance improved the diagnostic accuracy of less-experienced clinicians from 67.9
Neonatal cholestasis (NC) is an important hepatobiliary complication in preterm infants, but its short-term identification before the first clinical diagnosis remains challenging. We aimed to develop an explainable machine-learning model for short-term identification of NC using routinely collected clinical information available before the first documented diagnosis. This single-center retrospective matched case-control study included 280 preterm infants with NC and 280 controls matched by sex, gestational age, and year of hospitalization. Matched pairs were allocated to derivation (196 pairs) and held-out test (84 pairs) datasets. Boruta feature selection and hyperparameter optimization were restricted to the derivation dataset. Eight algorithms were evaluated using grouped five-fold cross-validation. Discrimination, calibration, classification performance, paired bootstrap model comparisons, sensitivity analyses, prevalence-dependent predictive values, and SHapley Additive exPlanations (SHAP) were assessed. Boruta confirmed 13 predictors spanning enteral and parenteral nutrition, inflammatory markers, hepatobiliary measurements, and hematological status. CatBoost had the highest derivation-set cross-validated area under the receiver operating characteristic curve (ROC-AUC) (0.803) and was selected before test evaluation. In the held-out test dataset, its ROC-AUC was 0.867 (95
Large language models (LLMs) show promise for automating literature screening, although their performance varies across clinical domains and further evaluation is required for broader implementation. This study aimed to assess the accuracy, reliability and feasibility of literature screening in headache research across seven LLMs, spanning compact to frontier models (GPT-4o, DeepSeek-V3, and Meta’s LLaMA-3 at multiple scales). This study is a comparative methodological evaluation of LLM-assisted title and abstract screening. Using a dataset from a real-world systematic review including randomized controlled trials on placebo and nocebo responses in migraine, we evaluated classification performance, processing runtimes, and computational costs across seven models using zero-shot prompting and IMAPR (Iterative Multi-Agent Prompt Refinement). The IMAPR system is our in-house developed, LLM-driven multi-agent framework in which a single, fixed LLM iteratively refines prompts. Each model was run five times per method, and sensitivity was assessed against a predefined threshold of ≥ 95
Fast-track colonoscopy pathways detect colorectal cancer (CRC) but strain endoscopy capacity. We developed and externally validated multivariable risk models combining faecal immunochemical test (FIT) with clinical data to triage Swedish fast-track referrals. We analysed 2,539 fast-track colonoscopies (2016–2020) and 723 as a validation cohort (2021–2022). Predictors were age, sex, eight symptoms, binary FIT, haemoglobin, and suspicious imaging. In development, missing predictors were imputed within training/test splits; validation used complete cases. Class imbalance was handled with oversampling to balance classes. Logistic regression and CatBoost models were built. External validation assessed AUC, calibration, and calibration slope. Variable influence was examined with feature contribution analysis. Decision-curve analysis (DCA) used the entry rule as a treat-all baseline. CRC prevalence was 16
Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.
Traditional pharmacy management, which relies on paper- and computer-based methods, is often inefficient and inflexible. There is an urgent need to develop support tools to address these challenges. The WeChat mini-program has been applied in multiple fields because it is readily accessible within the WeChat ecosystem. Its use in pharmacy management may support more accessible and flexible workflows. To develop and evaluate a WeChat mini-program-based Mobile Pharmacy Management Assistance System. A Pharmacy Collaboration Team consisting of front-line pharmacists, pharmacy administrators, and one informatics pharmacist was established to collect functional requirements through group discussions. The system was developed as a WeChat mini-program and implemented in routine pharmacy-management practice. After the 3-month implementation period, participating pharmacists were invited to complete a questionnaire. The questionnaire included the System Usability Scale, self-reported experience ratings comparing paper-based, computer-based, and mobile management modes, and open-ended feedback. Experience ratings were summarized using medians and interquartile ranges. Differences among the three paired management modes were analyzed using the Friedman test with Bonferroni correction. The final version of the system consisted of six modules: shift handovers, drug knowledge base, medication-error recording and analysis, drug-expiry reminders, damaged-drug recording, and medication consultations. Twenty-eight participants were invited to participate, and 27 completed valid questionnaires. The mean System Usability Scale score was 79.54 with a standard deviation of 11.33. The mobile-based mode received higher ratings for accessibility, work data sharing, and update frequency. The most frequently selected liked feature was the convenience of access anytime and anywhere, whereas the least appreciated feature was the limited interface size on mobile devices. This single-center exploratory evaluation supports the feasibility and preliminary usability of a WeChat mini-program-based mobile pharmacy management assistance system developed through a multi-role Pharmacy Collaboration Team. The system showed potential advantages in selected self-reported user-experience dimensions among pharmacists. However, the findings should be interpreted as preliminary evidence based on perceived user experience. Larger multicenter studies using objective workflow, data-quality, and medication-management outcomes are needed to evaluate its practical impact further.
To develop and validate an interpretable machine learning (ML) model integrated with inflammatory indicators for predicting ICU readmission in patients after coronary intervention (PCI/CABG). We used retrospective data from MIMIC-IV (training, 2196 patients) and MIMIC-III (external validation, 1717 patients). Patients aged > 18 years with complete hematological data and < 40
Artificial intelligence (AI) chatbots including ChatGPT provide rapid, conversational answers to complex medical questions. Despite promising performance on standardized examinations and specialty-specific tasks, the accuracy and safety of ChatGPT’s responses to real-time patient care questions remain uncertain, particularly when compared directly to practicing physicians. Our study aimed to address this gap by evaluating the accuracy, relevance, comprehensiveness, and safety of ChatGPT-generated answers compared with physician-generated answers to clinical questions arising during routine patient care. In this prospective comparative, blinded, cross-sectional study, practicing physicians in the emergency department (ED), general medicine ward, and intensive care unit (ICU) provided real-time clinical questions generated during patient care. The same questions were then independently answered by resident, fellow, or attending physicians using only non-AI resources, while ChatGPT provided separate responses to the questions. Nine attending physicians, blinded to the source of answers, evaluated all responses using a 5-point Likert scale across six domains: accuracy, clinical relevance, understandability, comprehensiveness, potential for patient harm, and potential bias. Composite mean scores from three evaluators per question served as the primary analysis unit. Comparisons were performed using the Mann–Whitney U test. A total of 285 questions were collected (33.0
To compare response accuracy between contemporary large language models (LLMs) and pediatricians across difficulty-stratified pediatric clinical vignettes, characterize LLM response times, and quantify the effect of AI-generated outputs on pediatrician decisions. In a prospective, randomized, assessor-blinded, vignette-based comparative study, 120 pediatric questions (four difficulty levels) were administered to five LLMs (ChatGPT-5, Gemini Pro 2.5, Claude Opus 4.1, Super Grok 4, DeepSeek V3) and to 30 pediatricians and to 30 pediatricians randomized 1:1 to the Control group (unassisted) or the AI-Assisted group. Models were queried in separate sessions with tools/browsing disabled using an identical prompt template. In the AI-Assisted group, pediatricians recorded an initial answer and confidence (1–10), reviewed five anonymized AI outputs presented in random order, and then submitted a final answer. A blinded adjudicator scored all responses against a prespecified answer key; p-values were adjusted via Holm–Bonferroni. Mixed-effects logistic regression was additionally performed to account for clustering of repeated responses within pediatricians and items. Response accuracy was 78.50
In-hospital mortality risk varies substantially among patients with pulmonary fibrosis (PF). Existing prognostic models have primarily focused on long-term survival and have limited ability to capture early in-hospital mortality risk or incorporate worsening oxygenation, inflammatory burden, and treatment-escalation information. This study aimed to develop and internally validate an interpretable machine learning model for in-hospital mortality risk stratification using early admission clinical data and information on ventilatory support within 24 h after admission in hospitalized patients with PF. This retrospective cohort study included hospitalized patients with PF admitted to a tertiary hospital between January 2019 and December 2023. Candidate variables included demographic characteristics, comorbidities, disease subtype, clinical symptoms, laboratory tests, arterial blood gas analysis, modified Medical Research Council (mMRC) dyspnea score, and the use of invasive or noninvasive ventilation within 24 h after admission. Nested LASSO feature selection combined with repeated 10-fold cross-validation was used for model development and internal validation. The main candidate models included logistic regression (LR), naive Bayes (NB), support vector machine (SVM), random forest (RF), neural network (NN), multilayer perceptron (MLP), adaptive boosting (AdaBoost), and gradient boosting machine (GBM). A supplementary clinically simplified model based on LASSO prescreening, recursive feature elimination, and logistic regression (RFE_LR) was also developed. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), threshold-dependent classification metrics, calibration analysis, decision curve analysis, clinical impact curves, bootstrap internal validation, and SHAP interpretation. Additional analyses included pairwise AUC comparisons with Holm adjustment for multiple testing, alternative-threshold analyses, exploratory in-sample recalibration analyses, and sensitivity analyses related to noninvasive ventilation (NIV). A total of 547 patients were included, of whom 108 met the prespecified primary mortality outcome, corresponding to an event rate of 19.74
Flatfoot is a common musculoskeletal disorder that, if not detected in its early stages, can lead to impaired posture, reduced mobility, and diminished quality of life. Conventional diagnosis relies heavily on manual interpretation of radiographic and clinical findings, which can be time-consuming, subjective, and dependent on expert experience. This study aimed to develop a multimodal deep learning framework that integrates complementary anatomical and functional information for automatic flatfoot classification. The proposed framework jointly processed three data modalities: (i) RGB plantar foot images, (ii) plantar-pressure time-series signals collected during gait, and (iii) biomechanical angular measurements. Three modality-specific branches were employed: a Swin Transformer to extract spatial representations of plantar morphology, a bidirectional long short-term memory (BiLSTM) network to model temporal plantar-pressure dynamics, and a Graph Convolutional Network (GCN) to represent structural relationships among biomechanical measurements. The modality-specific representations were projected into a common latent space and integrated using a multi-head attention-based fusion mechanism. The framework was evaluated using participant-level stratified 10-fold cross-validation on a clinically acquired multimodal dataset comprising 978 foot-level samples from 489 independent participants. Complementary publicly available pressure-related data were used exclusively during model development and were not included as independent multimodal cases in the final evaluation. The proposed multimodal framework achieved an accuracy of 94.0
Accurate prediction of in-hospital mortality in patients with acute myocardial infarction (AMI) is clinically critical. However, many existing risk models are complex, rely on variables not routinely available early during hospitalization, or have limited applicability in resource-constrained settings. Therefore, there remains a need for a clinically interpretable, parsimonious, and data-driven model with strong discrimination and calibration. This retrospective registry-based study used data from the Yazd Cardiovascular Disease Registry and included hospitalized AMI patients between 2016 and 2019. The primary outcome was in-hospital mortality. A clinically informed parsimonious logistic regression model was developed using five routinely available predictors: age, fasting blood sugar, ejection fraction, sex, and Killip class. Model performance was assessed using AUC, Brier score, calibration plots, calibration intercept, and calibration slope. Internal validation was performed using bootstrap resampling with 200 repetitions. Missing data were handled using multiple imputation by chained equations (MICE) as the primary analysis, with complete-case analysis performed as a sensitivity analysis. A total of 1,396 AMI patients were included, of whom 74 (5.3
The early and accurate detection and diagnosis of diseases play a crucial role in impeding disease progression and guiding appropriate treatment selection. Complex diseases such as autoimmune diseases (ADs), often present nonspecific symptoms that can overlap also with other conditions, leading to misdiagnosis. Medical software, known as clinical decision support systems (CDSS), is used by clinicians to categorize patients based on specific criteria. However, existing CDSS implementations, generally developed within individual hospitals, are often limited to specific data types, and reflect questionnaires that do not take advantage of machine learning (ML) methods for detection and prediction of complex patterns. To address this gap, we developed Personalis as a proof-of-concept medical software. Personalis integrates various modalities of medical datasets and applies ML models to target hospital datasets and disease-target specific prediction tasks within a unified software environment. The platform has been run with real-world data to investigate the potential of this proof-of-concept medical software in providing early clinical support for the personalized prediction of specific autoimmune diseases in individual patients. These intermediate results have been implemented to redesign the front-end to enhance user-experience, and serve the conveyed clinical needs and expectations. Personalis is a proof-of-concept medical software that demonstrated the feasibility of integrating various data modalities and machine learning methods to support clinicians in the diagnosis, treatment or management of individual patients with autoimmune diseases. By leveraging ML methods, Personalis provided prediction with a certain accuracy of a specific autoimmune disease for an individual patient, highlighting factors that contributed to the prediction results. Explainable machine learning made the models’ decisions understandable to humans, and the resulting factors were further used to provide additional clinical insights beyond the disease prediction accuracy metric. The platform was designed considering user-experience and human factors to enable actionable clinical practical adoption. Its results may provide hints on the prognosis assessment by selecting a specific patient record. Personalis provides a proof-of-concept medical software for integrating heterogeneous clinical data with configurable ML-based prediction and interpretation workflows for autoimmune diseases. It offers a mechanistic overview of the parameters that mostly influence the prediction of the machine learning models, therefore providing interpretable results and mechanistic insights essential to model transparency. The potential of the medical software Personalis is be extended to several diseases, and support in cases of challenging differential diagnoses.
Postoperative acute kidney injury (AKI) is a common and serious complication, driven by complex perioperative physiologic derangement and hemodynamic instability. There remains a need for a mechanism for deep interaction between preoperative characteristics and intraoperative dynamics, and provide clinical interpretability. We propose an extended Temporal Fusion Transformer for postoperative AKI prediction, designed for heterogeneous clinical data and intraoperative sequences, such as arterial pressure, medications, and fluid infusions. The model is built based on Long Short-Term Memory (LSTM) with the Variable Selection Networks, which adaptively select static and temporal features using gated residual networks, conditioned by static context vectors. Static covariates are encoded as multiple context vectors that guide temporal variable selection, LSTM initialization, and the temporal attention mechanism, modelling complex interactions between baseline risk and intraoperative signal fluctuations. The final representation for outcome prediction was aggregated by using the preoperative state as the attention query. The model was evaluated on two public surgical datasets, the Informative Surgical Patient dataset for Innovative Research Environment (INSPIRE) and Medical Informatics Operating Room Vitals and Events Repository (MOVER). Across all experiments, the proposed model achieved better performance and showed stable performance in subgroup analysis. The interpretability results showed the feature importance of static variables and time series at each time step, revealing that the model consistently assigned attention to periods of sustained hemodynamic instability. These findings supported the potential for explainable perioperative decision support systems and emphasized the importance of intraoperative signals in AKI risk stratification.
The increasing prevalence and clinical interrelationship of diabetes mellitus and cardiovascular disease (CVD) present substantial challenges to healthcare systems, creating a need for efficient and robust diagnostic approaches to support early intervention and continuous patient monitoring. This study presents ECNN-MmIC, an intelligent cloud-oriented Internet of Medical Things (IoMT) framework that integrates an Enhanced Convolutional Neural Network (ECNN), Bayesian Optimization (BO), and the Salp Swarm Algorithm (SSA) for joint diabetes and CVD prediction, hyperparameter optimization, and feature selection. The experimental analysis was conducted using the publicly available Diabetes in Bangladesh (DiaBD) medical dataset comprising 5,288 patient records collected from communities across 63 unions in Bangladesh, representing urban, semi-urban, and rural populations aged 21–80 years. The dataset incorporates 14 demographic, physiological, anthropometric, and clinical predictors, including age, gender, glucose level, pulse rate, systolic and diastolic blood pressure, BMI, hypertension, family histories of diabetes and hypertension, CVD, and stroke, thereby providing clinically relevant information for investigating diabetes–CVD comorbidity. The ECNN-MmIC framework was evaluated against Logistic Regression, SVM, Random Forest, DNN, conventional CNN, LSTM, and CNN-LSTM. Experimental results showed that ECNN-MmIC achieved 98.73
Many risk prediction models have been developed for major adverse cardiovascular events (MACEs) in non-cardiac surgery, but their quality and applicability in clinical practice are still unclear. PubMed, Cochrane Library, Embase, and Web of Science were thoroughly searched for studies on perioperative MACEs, encompassing cardiac death, cardiac arrest, acute coronary syndrome, myocardial infarction, heart failure, cardiac arrhythmias, myocardial injury, and stroke in non-cardiac surgery. The Prediction Model Risk of Bias Assessment Tool (PROBAST) was utilized to appraise the risk of bias. The area under the curve (AUC) value was pooled based on the true/false positive and negative numbers, and the pooled effect size was computed. Heterogeneity across studies was judged via the I2 statistic. After the screening, 93 eligible papers were included. 50 studies performed internal validation, but only 21 studies performed external validation. The overall pooled AUC was 0.80 (0.76–0.83), and the pooled AUC in different subgroups ranged from 0.70 (0.66–0.74) to 0.85 (0.81–0.87). The AUC for new machine learning (ML) was 0.85 (0.81–0.87) and for traditional logistic regression or COX regression was 0.79 (0.75–0.82). Heterogeneity was high overall. 79 studies were graded at high risk of bias. Deeks' funnel plots suggested no publication bias. 19 studies (20.4
Data-driven predictive models are increasingly used in clinical settings, but adoption hinges on both interpretability and credibility—i.e., alignment with domain knowledge. For tabular electronic health records data, conventional deep models are hard to interpret, while sparse linear models may sacrifice accuracy or disregard expert guidance. We formulate credible supervised learning as a Distributionally Robust Optimization (DRO) problem with a tailored DRO-expert norm. The norm promotes dense use of an expert-identified group of predictors while encouraging group-wise sparsity elsewhere, recovering a Group LASSO–type penalty. We handle both non-overlapping groups and overlapping groups using latent-variable decompositions. We evaluate on synthetic data by varying predictor correlation and signal-to-noise ratio (SNR), as well as on two real-world healthcare tasks: (i) ICU admission risk for COVID-19 patients and (ii) identification of uncontrolled hypertension from socio-demographic variables. Beyond AUC and F1-score, we report model sparsity and an explicit credibility metric (fraction of expert features retained). Across synthetic studies, the DRO-expert approach yields the most stable and accurate estimation—consistently improving Mean Absolute Error (MAE), Relative Risk, Relative Test Error, and Proportion of Variance Explained (PVE)—especially under high correlation or low SNR; the overlapping variant is strongest overall. On COVID-19 ICU prediction, the overlapping DRO-expert achieves an AUC of 95.85
Health research increasingly relies on observational datasets that capture systematically different patient populations, affecting the representativeness, comparability, and transportability of study findings. Yet these differences are rarely quantified across clinical domains or explored interactively, and the workflow for pairwise dataset comparison remains fragmented. No existing OHDSI tool provides an integrated visual analytics workflow for exploratory pairwise comparison of concept prevalence patterns between datasets on the OMOP Common Data Model. We present Syrona, an open-source visual analytics tool for systematic pairwise comparison of datasets and subcohorts on the OMOP Common Data Model. Syrona extracts annual prevalence for conditions, procedures, and drugs from any OMOP CDM database, computes prevalence ratios across demographic strata, and synthesizes them via multilevel meta-analysis. A coordinated dashboard - distributional overviews, stratified heatmaps, forest plots, and absolute prevalence comparisons - enables interactive exploration driven by domain-adaptive SNOMED CT and ATC filtering. Three case studies on Estonian national health data (495,000 persons, 2012–2024) demonstrate how the tool supports representativeness assessment, institutional practice comparison, and coding artifact detection. Syrona is open-source and applicable to any OMOP CDM database. By treating dataset comparison as a visual analytics problem rather than a tabular reporting task, it addresses a gap in the OHDSI ecosystem and supports systematic assessment of datasets in distributed observational research.