
Background: Intraoperative anaphylaxis remains a rare yet fatal condition that has been a challenge due to its unpredictability. The unique characteristics of physiological parameter changes that precede or occur during the early stages of anaphylaxis may enable timely identification. We aimed to develop machine learning methods for the early detection of intraoperative anaphylaxis among patients with hypotension using real-world time-series physiological data. Methods: Physiological data, extracted from the EMRs at a 10-second sampling frequency, of patients undergoing surgeries at Peking University People’s Hospital (January 1, 2011 - January 1, 2023) were analyzed. Three datasets (Datasets +1Min, +2Min, and +3Min) were constructed, each spanning 10 min before to 1, 2, and 3 min after hypotension, consisting of positive groups (intraoperative anaphylactic patients with hypotension) and negative groups (intraoperative non-anaphylactic patients with hypotension). Random forests, extreme gradient boosting, and categorical boosting (CatBoost) were employed to construct detection models. The model with the best performance was identified with a five-fold cross-validation. Results: Datasets +1Min, +2Min, and +3Min contained 49, 48, and 44 positive samples and 980, 960, and 880 negative samples, respectively. CatBoost models performed best on both Datasets +3Min and +2Min, achieving on Dataset +3Min an area under the receiver operator characteristic curve (AUROC) of 0.851 (± 0.082), an area under the precision-recall curve (AUPRC) of 0.431 (± 0.094), a sensitivity of 0.861 (± 0.210), and a specificity of 0.783 (± 0.164); corresponding values on Dataset +2Min were 0.823 (± 0.145), 0.365 (± 0.097), 0.711 (± 0.183), and 0.940 (± 0.065). Conclusions: This study suggests the feasibility of early detection of intraoperative anaphylaxis among patients with hypotension using physiological time-series data. The models could alert clinicians to suspected anaphylaxis just 2–3 min after hypotension, enabling early detection of intraoperative anaphylaxis.
Green classification of cystocele on dynamic transperineal ultrasound (TPUS) remains operator-dependent because it requires manual frame selection and landmark-based assessment of the Valsalva maneuver. We developed a workflow-oriented AI-assisted clinical decision support system for automated urethrovesical junction localization and dynamic Green classification and prospectively evaluated its standalone and reader-support performance. This diagnostic accuracy and reader study included 881 patients from a tertiary referral hospital, comprising a retrospective development cohort (n = 688) and an independent prospective test cohort (n = 193). A nested subset of 67 prospective patients was used for a reader study involving two junior and two intermediate radiologists under unaided and AI-assisted conditions. In the complete prospective test cohort, Green-AttGRU achieved a macro-averaged AUC of 0.939 (95
This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring’s content analysis guided by the FUTURE-AI framework. At each model’s best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95
Hospitals worldwide invest heavily in digital infrastructure to improve efficiency, safety, and patient-centeredness. Hospital Information Systems (HIS) form the backbone of these efforts. However, the HIS provider market remains fragmented and heterogeneous, thus potentially influencing hospitals’ ability to achieve digital transformation. Empirical evidence on how HIS choice relates to digital maturity remains limited. Using data from more than 1,600 hospitals participating in the German-wide DigitalRadar (DR) project (i.e., covering about 90
EHR data are widely used in clinical research; however, their temporal stability remains poorly understood. This study aimed to evaluate temporal variability in EHR data arising from differences in extraction time points. We conducted a retrospective, descriptive, observational study using diagnostic data from a multicenter integrated EHR database in Japan. The target period was 2022, and diagnosis records with a start date within this period were extracted at three time points: October 2023, October 2024, and October 2025. Diagnoses were mapped to the three-character level of the International Classification of Diseases, 10th Revision (ICD-10). Changes in diagnosis code distributions were quantified using the Jensen–Shannon distance (JSD). Changes in record counts per ICD-10 code between 2023 and 2025 were also evaluated. Five facilities were included in the analysis. In the integrated dataset, variations in total records and unique patients were minimal (total records: +0.06
Human oversight is widely invoked as a safeguard in clinical AI, yet it is often treated as a single, stable property. In practice, clinicians remain responsible for AI-assisted decisions without remaining able to meaningfully review the basis of the output. This tension becomes sharper as systems become multimodal, agentic, and cognitively ambitious. We propose a distinction between cognitive and normative oversight and discuss its implications for AI evaluation and deployment.
This study aimed to develop an interpretable machine learning-based scoring system for predicting sepsis and septic shock among febrile patients at emergency department (ED) triage using longitudinal data. This retrospective, single-center study included adult patients, presented to ED of tertiary academic hospital with fever from January 2016 to December 2021. Using the AutoScore framework, we developed a novel scoring system for predicting sepsis and septic shock at the triage stage, incorporating nine variables and a maximum score of 29. The predictive performance of our score was assessed by calculating the area under the receiver operating characteristic curve (AUROC), and its performance was compared with that of two existing scoring systems: the quick Sequential Organ Failure Assessment (qSOFA) and the Modified Early Warning Score (MEWS). Our model incorporated nine variables including initial vital signs, age, baseline platelet count, total bilirubin, and creatinine levels. Among these, initial systolic blood pressure was identified as the most important predictor. AUROC of our model was 0.844 (95
Evaluating clinical reasoning in large language models (LLMs) poses two open challenges: reference-oriented semantic metrics do not directly assess whether a model’s stated diagnosis is supported by the evidence in its own justification, and the increasingly popular LLM-as-judge approach rests on a largely untested assumption—that independent verifier LLMs agree with one another. We assess three generator LLMs (HuatuoGPT-o1-8B, Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.3-70B-Instruct) on 1,000 MIMIC-IV hospital-stay cases along four complementary axes (medical concept grounding, semantic similarity, semantic uncertainty, and evidence–conclusion coherence), with coherence judged independently by three frontier verifiers (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.4 mini). Two findings emerge. First, coherence reveals a dissociation that reference-oriented metrics do not capture: a model can score well on those axes yet still produce rationales that do not support its own conclusions. Second, inter-verifier agreement on coherence is consistently low (Fleiss’ κ 0.087–0.223; disagreement 62.2
Digital transformation is a key priority for modernizing China’s public hospitals. However, a standardized and context-specific framework to evaluate their digital maturity remains absent. This study aims to develop and validate a comprehensive, multidimensional evaluation framework tailored to Chinese tertiary public hospitals to support systematic assessment and inform policy decisions. Based on systematic literature review and policy analysis, we constructed a framework comprising Digital Readiness, Technology Application, and Data Management Capability, with 11 subdimensions and 65 indicators. Indicator weights were derived using a two-round Delphi consultation, analytic hierarchy process, and criteria importance through intercriteria dependence method. The framework was applied to 1,361 tertiary public hospitals across 28 mainland provincial-level divisions in China. Digital maturity scores were analyzed using global sensitivity analysis (GSA), k-means clustering, and logistic regression with Firth’s penalized likelihood. Digital Readiness received the largest combined weight (34.9
International Classification of Diseases (ICD) codes enable correct billing, insurance reimbursement, and healthcare analytics. However, manual coding is time-consuming, expensive, and error-prone, creating bottlenecks in clinical workflow and limiting scalability. Artificial intelligence (AI) has emerged as a promising solution for automated ICD code assignment from unstructured clinical text. This systematic review explores the current state of automated ICD coding research, examining models applied to diverse clinical documents including discharge summaries, electronic health records, nursing notes, and pathology reports. Following PRISMA guidelines, we searched six databases for studies published between 2019 and 2024, selecting 54 relevant studies from 4,280 initial citations. Our analysis reveals the use of diverse datasets, preprocessing techniques, and feature extraction methods, alongside a clear evolution from traditional machine learning to deep learning approaches, with substantial architectural diversity across convolutional, recurrent, transformer, and hybrid models. Performance varies considerably across dataset configurations, with models achieving higher accuracy on frequent code subsets compared to full label spaces. However, critical gaps persist: overreliance on single-language, single-institution datasets limits generalizability; difficulties in predicting rare codes remain unresolved; lack of model interpretability undermines clinical trust; and inconsistent evaluation protocols hinder meaningful comparison. To address these challenges, we propose a 5P evidence-grounded research agenda: Population Diversity, Performance Robustness, Prediction of Rare Codes, Provenance Transparency, and Practical Integration. These findings underscore AI’s potential to transform ICD coding while highlighting the need for standardized benchmarks, rigorous external validation, multilingual datasets, and explainable architectures to enable equitable and effective deployment in real-world healthcare systems.
Federated Learning (FL) emerged as a privacy-preserving paradigm for collaborative training of deep learning models across institutions without sharing patient data. This approach has been applied to complex tasks such as medical image-to-image (I2I) translation, including MRI-to-synthetic CT (sCT) generation. However, existing federated I2I frameworks often assume privacy preservation as an inherent property of FL rather than a requirement to be explicitly validated, leaving their robustness to representative adversarial threat scenarios largely unexplored. In this study, we evaluated the vulnerability of a federated MRI-to-sCT translation framework (FedSynthCT-Brain) to three representative attack classes: Deep Leakage from Gradients (DLG), Federated Membership Inference Attack (FedMIA), and data poisoning. The efficacy of corresponding defense mechanisms, such as Secure Aggregation (SecAgg) and Byzantine-robust median aggregation (FedMedian), were assessed. DLG enabled only the recovery of coarse anatomical structures, with no clinically identifiable details (SSIM ≤ 0.16, PSNR ≤ 11 dB) across clients, suggesting limited vulnerability under the evaluated DLG setting. In contrast, FedMIA achieved high membership discrimination, with AUC scores between 0.92 and 0.99, revealing a critical privacy vulnerability. The introduction of SecAgg reduced AUC values to near-random levels (0.23–0.56) across all centers without impacting synthesis quality. Under high-noise poisoning, the standard federated averaging (FedAvg) aggregation rendered the federation inoperative, while FedMedian restored performance close to the no-poisoning baseline in most scenarios, with significant residual degradation in specific center configurations. At low noise levels, the advantage of FedMedian was less consistent, as low-level noise injection may be indistinguishable from natural heterogeneity across centers, potentially enabling stealthy degradation. These findings demonstrate that federated I2I translation frameworks are not inherently secure and require explicit, multi-layered evaluation. As FL is increasingly adopted in clinical workflows, our results underscore the necessity of integrating cryptographic, algorithmic, and infrastructural safeguards for secure deployment.
Medically specialized AI systems that have obtained regulatory clearance as medical devices can be deployed for patient-specific clinical decision support under defined compliance requirements. However, it remains unclear whether medical specialization and regulatory status translate into higher-quality breast cancer treatment recommendations than those produced by a general-purpose large language model (LLM). This blinded, multicenter study compared the performance of two medically specialized AI systems with a general-purpose model in breast cancer care. Two medically specialized (Prof. Valmed and OpenEvidence) and one general-purpose system (ChatGPT-5 Thinking) were prompted to generate treatment plans for 20 standardized breast cancer patient cases. Outputs were rated and ranked by blinded, board-certified breast cancer specialists from seven university breast cancer centers for safety, guideline adherence, medical adequacy, completeness, overall quality, and logical coherence. Statistical analyses comprised descriptive statistics, inter-rater reliability assessment, non-parametric performance comparisons of rating and ranking outcomes, and correlation analyses. Mean (± standard deviation) processing time for ChatGPT-5 Thinking (159 ± 58 s) was more than fourfold higher than that of Prof. Valmed (35 ± 4) and OpenEvidence (9 ± 1). ChatGPT-5 Thinking achieved significantly higher ratings across all evaluation categories, with no significant differences between the two medically specialized systems. Treatment plans generated by ChatGPT-5 Thinking were ranked as the top choice in 96.4
Autism spectrum disorder (ASD) affects tens of millions of families worldwide, yet parents confront abundant but unreliable online advice and limited access to timely, empathetic guidance. To address this critical gap, we developed Starmate ( http://kefeng.mpu.edu.mo/starmate ), a 1.5B-parameter, domain-tuned AI assistant for ASD caregivers, using a rigorous user-centered mixed-methods framework. Informed by in-depth interviews ( n=13 ) and a Kano survey ( n=60 ) that identified “Hands-on guidance” as a must-have caregiver requirement, we engineered a novel modular architecture that integrates sentiment analysis, expert-vetted knowledge-graph-augmented retrieval (LightRAG), and a domain-fine-tuned Qwen2.5-1.5B model. In a blinded, side-by-side comparison against leading commercial LLMs, Starmate demonstrated improved performance across key metrics within this evaluation framework (86.76 vs 78.43–83.84; p < 0.001 ) and showed specific advantages in Empathy, Hands-on guidance, and Logical clarity (all p < 0.05 ). Automated benchmarking corroborated these results, with top scores for Professional accuracy (86.18), Empathy (86.79), and Hands-on guidance (82.58). These findings demonstrate the technical feasibility of a lightweight, privacy-conscious, domain-specific LLM to generate accurate, empathetic, and actionable responses in benchmarked scenarios, laying the groundwork for future real-world usability and clinical testing.
Rare disease registries in Brazil remain fragmented across federal, state, and local initiatives, limiting the availability of reliable epidemiological information to support diagnosis, care planning, research, and public policy. This study aimed to map existing rare disease registry entities and registry-related initiatives in Brazil and to propose practical guidelines for their unification into an integrated national registry. We conducted a descriptive, exploratory mapping study combining a structured literature search with documentary analysis of public policies, health information systems, registry portals, institutional reports, and legislative documents related to rare diseases in Brazil. PRISMA-S was used to report the search component, and a PRISMA-style flow diagram documented source identification and selection. We identified a rapidly evolving legislative landscape, including federal bills proposing a national monitoring system or registry and recent state-level statutes related to identification and observatories. Using an expanded, auditability-oriented inventory definition, we mapped 28 registry entities and registry-related initiatives. Of these, 24 are implemented, three are legislative proposals, and one is under development. Among the 24 implemented initiatives, 16 are national or multicentre, Brazil-based initiatives; three are state-level; four are regional/local; and one is a transnational registry with documented participation of a Brazilian cohort. Registry creation and registry-related activity accelerated after 2018, particularly between 2020 and 2026. We conclude that Brazil exhibits substantial data fragmentation across uncoordinated systems. A unified approach should integrate epidemiological data from existing networks, state notification systems, specialised hospital registries, and technology appraisal information under coordinated governance, while embedding privacy-by-design and information security safeguards.
To develop and internally validate a vectorcardiography (VCG)–augmented model for estimating 12-month major adverse cardiovascular events (MACE) in hospitalized patients with chronic heart failure (CHF). We conducted a single-centre retrospective cohort study of adults hospitalized with CHF between 31 May 2023 and 31 May 2024, with data lock on 31 May 2025. The prespecified endpoint was any MACE within 12 months after the index hospitalization. ECG-to-VCG transformation was performed using the Kors method. The prespecified primary six-predictor model included LVEDD, NYHA class, frontal, horizontal, and sagittal QRS–T angles, and the QRS-loop reversal/U-turn sign. BNP and LVEF were evaluated in full-model, comparator-model, and incremental-value analyses. Because the prediction target was fixed 12-month risk rather than time-to-event hazard, multivariable logistic regression was used as the primary modelling approach. Internal validation was performed using 1000 bootstrap resamples, with apparent, optimism-corrected, calibration, and shrinkage-adjusted performance reported. Of 201 screened, 160 were included; MACE occurred in 68/160 (42.5
Automated electronic alerts integrated into inpatient electronic health records (EHRs) may improve recognition of moderate-to-severe acute kidney injury (AKI) and support timely medication review. An AKI e-alert system integrated with medication-focused clinical decision support (CDS) was implemented to identify adults with KDIGO stage 2–3 hospital-acquired AKI in near real time while minimizing alert burden. The alert logic used KDIGO serum creatinine (SCr) criteria, defined the operational real-time inpatient baseline as the lowest SCr value within the preceding seven days, and applied suppression rules for dialysis within seven days or imminent transfer/discharge. Alerts were generated through scheduled near-real-time processing four times daily and were accompanied by secure in-hospital medication guidance; the CDS did not automatically place orders. System performance was evaluated at a 2,768-bed tertiary medical center in Taiwan using EHR data from March 2018 to May 2023 and validated against a retrospective computerized reference algorithm designed to replicate the deployed logic. The system generated 3,946 stage 2–3 AKI alerts, achieving 90.94
Objective: To evaluate the alignment between self-reported confidence of large language models (LLM) and their accuracy in answering medical multiple-choice questions. Materials and Methods: We prompted six LLMs (GPT-5, GPT-5-mini, GPT-5-nano, GPT-4o, Claude Sonnet 4.5, and Gemini 2.5) to answer MedMCQA items and report confidence scores. Based on 12,000 LLM responses, we calculated Expected Calibration Errors (ECE) by averaging absolute differences between observed accuracy and predicted confidence. Results: Mean ECE differed by model (Claude Sonnet 4.5 best: 0.06; Gpt-4o worst: 0.127) and varied across specialties (“Skin” best: 0.041; “Social Preventive Medicine” worst: 0.141). Accuracy of examined LLMs showed analogous variation between specialties. Discussion: Our results demonstrate that high accuracy does not guarantee reliable uncertainty estimation. We identified substantial heterogeneity across medical specialties, where pooled metrics masked a threefold ECE increase between best- and worst-performing domains. Conclusion: We recommend incorporating calibration reporting into LLM evaluations, as larger models exhibit improved “self-knowledge”, but uneven overconfidence persists.