
Reliable preoperative risk stratification remains challenging for meningioma patients receiving adjuvant radiotherapy, particularly in real-world settings where complete multimodal datasets are difficult to assemble. This retrospective multicenter study developed and externally validated a multimodal artificial intelligence model integrating peritumoral multiparametric MRI and clinicopathological data to predict post-radiotherapy recurrence. A total of 250 meningioma patients treated with adjuvant radiotherapy across three neurosurgical centers were included. Preoperative T1-weighted, T2-weighted, and contrast-enhanced T1-weighted MRI were analyzed using a 2.5D ResNet-50 framework across three spatial contexts: tumor only, tumor plus 1-cm peritumoral margin, and tumor plus 2-cm peritumoral margin. The 1-cm peritumoral region provided the most transferable imaging representation and was selected for downstream modeling. Deep learning, radiomics, and clinicopathological features were evaluated alone and in late-fusion Cox models. Radiomics showed high apparent training performance but limited external generalizability, whereas deep learning features demonstrated more stable cross-center performance. The final clinical–deep learning fusion model achieved the best overall discrimination, with a mean C-index of 0.856 across cohorts, showed favorable calibration and clinical net benefit, and stratified patients into clinically distinct recurrence-risk groups. These findings support peritumoral MRI-based multimodal AI as a practical tool for recurrence risk stratification after adjuvant radiotherapy in meningioma.
While artificial intelligence (AI) and large language models (LLMs) have shown promise in identifying and classifying suicidal ideation, their generalizability and equity in the presence of heterogeneous clinical data remain largely unexplored. This study hypothesized a subgroup disparity in a crude AI classifier of clinician-rated suicidal ideation because of the linguistic heterogeneity and proposed a factorization approach to decompose complex data into simpler components by reducing topic dimensions of clinical transcripts. Results showed that topic-specific classifiers reduced subgroup disparity compared to the topic-general classifier, with ΔAUC decreasing from 0.11 to 0.01 and 0.05—a noticeable reduction of 0.10 and 0.06, respectively. More specifically, with the topic-general classifier, the odds of missing a suicidal case increased by 2.39 times for alexithymia individuals, compared to non-alexithymia individuals (OR = 2.39, p = 0.002). These findings underscore the significance of data heterogeneity on AI classifiers of suicidal ideation and demonstrate the potential of the proposed factorization approach.
Predictive biomarkers for immune-checkpoint inhibitor (ICI) therapy response often fail to show stable predictive performance because tumor microenvironment (TME) heterogeneity can create biologically confounding tumor states, where similar molecular patterns can lead to divergent therapeutic outcomes. When analyzed together, these biologically distinct states can act as out-of-distribution (OOD) samples, complicating model training and limiting predictive consistency. Here, we present a stepwise framework that improves target-based ICI-response prediction by excluding senescence-associated tumor states prior to model training, thereby reducing biological heterogeneity that can obscure the relationship between checkpoint activity and therapeutic response. The framework leverages network-based representations of immune-checkpoint and senescence pathways to identify senescence-associated non-responders (SNRs). Across multiple melanoma, gastric, bladder, and lung cancer cohorts, this approach improved accuracy, AUROC, precision, and specificity in within-cohort evaluations, and showed consistent performance in external validation using independent melanoma cohorts. The excluded tumors were enriched for senescence-associated markers, suggesting that they exhibit senescence-related transcriptional features that may act as biological confounders limiting the predictive capacity of standard target-based models. These results suggest that confounder-aware filtering of tumor microenvironment states can provide a practical approach to improving ICI therapy response prediction.
Artificial intelligence (AI) screening of placental slides and patient data uncovers decisive biomarkers, most notably decidual vasculopathy, that predict recurrent preeclampsia and adverse maternal–child outcomes1–4. Clinical adoption of such screening methods remains limited because current models often misclassify cases and lack biologically grounded interpretability5 for their predictions. In our study, interpretability refers to the ability to detect differences in the spatial organization of extravillous trophoblast cells (EVT) relative to red blood cells (RBC) within placental vessels rather than relying on data‑driven heuristics that yield black‑box diagnoses. This study introduces the optical density morphology mapping technique (ODMMT), an AI framework that delivers biologically interpretable unsupervised classification, automates the correction of misclassifications, and makes AI pipelines more deployable for real‑time diagnosis. The framework comprises three steps: extracting EVT–RBC organizations via optical density masking, converting them into z‑scores with a normalizing flow model, and calculating a morphology separation score for clustering and biological interpretation. ODMMT surpasses other AI frameworks, achieving two distinct clusters in distinguishing healthy from diseased vessels on the experimental dataset. This advancement automates AI misclassification correction and quantifies RBC–EVT spatial organization, providing biomarker insights from unlabeled images to improve diagnosis and better understand the maternal-fetal relationship.
Artificial intelligence (AI)-enabled infectious disease surveillance platforms are expanding rapidly, but their roles within health systems remain poorly characterized. This review examined 20 platforms according to primary data inputs, surveillance function, AI methods, data privacy, geographic scope and pathogen focus. Three archetypes emerged: Early Warning Networks, Situational Awareness Platforms and Integrated Surveillance Platforms. Their comparative analysis highlighted how data type, integration and localization shape platform contributions to public-health decision making.
Speech production integrates respiratory, laryngeal, articulatory, prosodic, linguistic and executive control, so neurodegeneration can leave measurable acoustic traces before conventional clinical scales change. Two decades of work have identified credible candidate speech and voice biomarkers for amyotrophic lateral sclerosis (ALS) and Parkinson’s disease (PD), yet the field remains fragmented into small, single-condition, single-language models, many of which do not reproduce when evaluated on speakers entirely unseen during training (all recordings from each participant are confined to one data partition, preventing identity leakage between training and testing). In this Perspective we argue that a testable next step is a clinically grounded voice-biomarker foundation model (a single self-supervised backbone, pretrained on large, diverse, ethically sourced, multi-condition speech and adapted to explicit contexts of use), rather than further bespoke classifiers. We propose ALS bulbar-progression monitoring as a suitable lead context of use because some speech-derived measures appear more responsive than the coarse ALSFRS-R speech item; one ALS speech-analytics platform has received Breakthrough Device designation, an expedited-review status that is neither marketing authorization nor endpoint qualification. We use PD screening as a test case carrying the field’s central cautionary lessons about data leakage, modest real-world operating points, and the potential value of articulation-rich over phonation-only tasks. We separate established evidence from inference and proposed research, define what would justify the “foundation model” label, specify minimum methodological and governance standards, state the limitations candidly, and outline a prospective-validation and regulatory roadmap. To our knowledge, no speech- or voice-derived endpoint for ALS or PD was qualified by FDA or EMA as of the time of this publication; the foundation-model case is therefore presented as a research agenda, not an achieved capability.
Cardiac magnetic resonance (CMR) is the clinical gold standard for assessing cardiovascular conditions, yet translating image semantic representations into specific disease diagnoses remains challenging. Current large language models (LLMs) often lack the interpretability required for clinical trust. Here, we present CMR-R, a large reasoning model that utilizes CMR quantitative parameters and semantic descriptions from CMR reports as inputs to provide explicit, interpretable diagnostic chains for CMR semantic diagnosis. We curated a multi-center dataset of 16,104 cases and developed a multi-stage adaptive-learning training framework for enhancing reasoning ability, mitigating long-tail distribution bias, and enabling prospective continuous learning with unlabeled data. CMR-R achieved diagnostic accuracy (ACC: 0.858, AUC: 0.944) across eight categories, surpassing that of radiologists (with >10 years of experience) and other LLMs, including in rare diseases such as left ventricular non-compaction (LVNC) and cardiac amyloidosis. Furthermore, CMR-R reveals underlying associations between CMR imaging features and cardiac diseases. This work has the potential to enhance CMR diagnostic accuracy and improve clinical efficiency.
Artificial intelligence (AI) has advanced rapidly across diagnostic, prognostic, and clinical decision-support applications, yet the pathway from laboratory performance to demonstrable clinical benefit remains fragmented and inconsistently defined. Existing evaluations rely heavily on retrospective testing and algorithm-centric metrics, while current guidelines emphasize reporting standards rather than specifying validation across stages of model maturity. This study proposes a five-phase evaluation framework for medical AI, supported by a dynamic evaluation architecture reflecting the nonlinear, iterative nature of AI systems. The framework integrates technical validation, operational robustness validation, controlled interaction validation, clinical evidence validation, and real-world integration validation, while incorporating phase-gating criteria and local and systemic fall-back triggers. These mechanisms enable re-entry into earlier phases based on drift, version updates, or safety signals, and accommodate parallel activities such as implementation research informing clinical trials. By systematically mapping multicenter external validation, shadow-mode testing, human-AI comparison and cooperation studies, randomized controlled trials, real-world evaluations, and adaptive designs into a coherent lifecycle pathway, the framework addresses persistent gaps between laboratory performance and clinical benefit. It provides researchers, clinical institutions, and regulators with an operational, scalable approach aligned with evolving regulatory expectations, supporting trustworthy, ethically aligned, and lifecycle-based evidence generation for medical AI systems.
The rapid integration of foundation models into clinical practice and their use for public health inquiries necessitates a rigorous evaluation of their true clinical reasoning capabilities, which extends beyond success on narrow examinations. Current benchmarks, often based on medical licensing exams or curated vignettes, fail to capture the integrated, multimodal reasoning required in real-world patient care. To address this gap, we developed the Bones and Joints (B&J) Benchmark, a comprehensive evaluation framework comprising 1245 questions derived from real-world patient cases in orthopedics and sports medicine. This benchmark assesses models across seven core tasks that mirror the clinical reasoning pathway, including knowledge recall, text interpretation, image interpretation, diagnosis generation, treatment planning, and the underlying rationale. We evaluated 14 vision-language models (VLMs) and six large language models (LLMs), comparing their performance against expert-derived ground truth. Our findings reveal a pronounced performance gap. While state-of-the-art models achieved high accuracy, exceeding 90% on structured multiple-choice questions, their performance markedly declined on open-ended tasks requiring multimodal integration, with accuracy scarcely reaching 60%. VLMs demonstrated substantial limitations in interpreting medical images and frequently exhibited text-driven hallucinations. Notably, medical-specific models showed no consistent advantage over general-purpose counterparts. These results indicate that current foundation models face significant challenges in achieving independent clinical competence within highly specialized musculoskeletal fields. Their safe deployment should be limited to supportive, text-based roles, while advancement in core clinical tasks awaits fundamental breakthroughs in multimodal integration and visual understanding.
The use of artificial intelligence (AI) to draft responses to patient portal messages has been proposed to reduce provider in-basket burden. However, little is known about its effects on patient–provider communication. In this retrospective observational study, we evaluated demographic differences in the tone of AI-generated draft replies (AI-GDRs) and care team responses to patient messages. Our study included 12,202 message triads comprising patient messages, AI-GDRs, and care team responses from three internal and family medicine practices in New York City. We found differences in tone across patient demographics. AI-GDRs had lower odds of including polite language in responses to Hispanic patients, compared to White patients. AI-GDRs also had lower odds of conveying positive affect in responses to patients who were Hispanic, preferred a non-English language, assigned female sex at birth, or lived in an area with a lower average income. Some of these tone differences were also found in care team responses. Physicians and advanced practice providers had lower odds of conveying positive affect when responding to patients who were Hispanic or preferred a non-English language. Our findings highlight that careful implementation of AI drafting is needed to ensure that this technology does not introduce or amplify inequities in patient–provider communication.
Artificial intelligence (AI) could help identify patients at risk of medication non-adherence, but the clinical readiness of published prediction models is uncertain. We systematically reviewed 41 adherence prediction modelling studies and assessed model development and evaluation using PROBAST + AI across participants/data sources, predictors, outcomes, and analyses. Most models showed great concern for development quality (71%) and high risk of bias in evaluation (80%), commonly due to poorly defined adherence outcomes, inadequate handling of missing data, and limited validation. Reported discrimination did not consistently improve with more complex algorithms, indicating that methodological rigour, rather than model type, is the key barrier to translation. We provide framework-guided recommendations to improve robustness, interpretability, and clinical actionability of future adherence prediction models.
The rapid growth of biomedical literature has rendered traditional systematic reviews unsustainable. Although large language models (LLMs) offer automation potential, citation fabrication and unreliable evidence discrimination remain critical barriers. Here, using PubChat as a PubMed E-utilities-grounded retrieval framework, we tested whether source-verifiable AI-assisted retrieval could preserve recall, criterion-driven relevance stratification, and multilingual accessibility in systematic-review benchmarking. The system comprises Phase I, hierarchical decomposition of the research question into five relevance levels, and Phase II, multi-round retrieval with embedding-based pre-filtering and three-round LLM verification. Validated against 20 Cochrane systematic reviews (585 ground-truth articles) across eight languages, PubChat was benchmarked against four general LLMs (GPT-5.2-Thinking, Gemini 3.0 Pro, Grok-4.1-Thinking, Qwen3-Max), one search-augmented retrieval tool (Perplexity-Sonar), and three specialized retrieval tools (Elicit, ASTA, and Consensus). PubChat produced no fabricated citations in this benchmark and showed distinct recall–precision profiles across its three operating modes. Among the evaluated configurations, PubChat-Broad achieved the highest observed recall and nDCG, whereas PubChat-Core achieved the highest observed F1- and F2-scores. In a user evaluation of 279 biomedical researchers across 18 countries, PubChat scored above 80/100 for reliability, innovation, efficiency, user experience, and self-reported preference over the evaluated alternatives (78% vs. specialized tools and 72% vs. manual search). An exploratory MIDE case study further illustrates how PubChat-derived corpora can be organized into evidence-traceable research-gap candidates for hypothesis prioritization. PubChat provides a benchmark-validated PubMed-grounded framework for source-faithful biomedical literature retrieval, relevance stratification, and structured evidence organization.
Digital health technologies (DHTs) enable continuous and remote measurement in clinical research, but DHT-derived measures require evidence of accuracy, usability, generalizability, and regulatory suitability. Without structured guidance, studies risk generating data that are technically sophisticated but clinically ambiguous and difficult to reproduce. This Perspective presents a six-component framework for selecting, validating, and optimizing DHT-derived measures to the appropriate research question, endpoint role, context of use, and evidentiary purpose.
Intracapsular hip fractures are a leading cause of disability in young and middle-aged adults and are associated with high failure and reoperation rates. Accurate preoperative prediction of conversion to arthroplasty is crucial for quantifying individual prognosis and potentially informing individualized, clinician-led treatment decisions. Yet models built solely on structured electronic health records (EHRs) are insufficient. Here, we introduce MMHIP, trained on a four-center development cohort (n = 951) and externally tested in three independent centers (n = 240). To our knowledge, this is the largest multi-center cohort integrating 3D CT and long-term follow-up scans for this task. Advancing beyond conventional approaches, MMHIP integrates preoperative pelvic CT with EHR data, achieving AUROC 0.920 in development and 0.873 on external validation, significantly outperforming representative image-only and EHR-only baselines (p < 0.05). Incorporating CT markedly improved discrimination over EHR-based models, with absolute AUROC gains of 0.150–0.252 in development and 0.151–0.221 externally. Using a development-derived PPV-targeted operating threshold, MMHIP stratified patients into lower- and higher-predicted conversion-risk groups. At this threshold, MMHIP achieved a PPV of 0.901, sensitivity of 0.654, specificity of 0.980, and NPV of 0.910 in the development set; when applied unchanged to the external testing set, it maintained a PPV of 0.900, sensitivity of 0.529, specificity of 0.984, and NPV of 0.886, outperforming the single-modal comparators in threshold-based risk stratification. These findings support MMHIP as an interpretable preoperative risk-stratification model that may provide clinically relevant prognostic information for individualized, clinician-led treatment decisions.
Closed-loop deep brain stimulation (DBS) relies on continuous neural biomarker sensing, yet clinical utility is often limited by signal dropout, stimulation artifacts, and hardware constraints in subcortical recordings. Here, we develop a deep learning framework combining spectral processing with generative diffusion models to digitally reconstruct deep brain signals from cortical electrocorticography (ECoG), enabling continuous subcortical biomarker inference without direct deep brain sensing. We validate this approach across 723 h of simultaneous cortico-subcortical recordings from 49 patients with movement disorders (Parkinson’s disease, dystonia, Tourette syndrome) across three international centers. The framework decodes subcortical activity across multiple deep brain targets (subthalamic nucleus, globus pallidus internus, thalamus), behavioral states (rest, movement, sleep), and therapeutic conditions (medication and stimulation ON and OFF), with performance remaining above chance in every condition tested. Using generative diffusion models, we achieve raw signal reconstruction that preserves clinically relevant neural features, including beta burst dynamics that correlate with motor symptom severity (UPDRS-III R² = 0.70). We demonstrate clinical utility by showing that cortically-derived signals can rescue state detection during DBS recording failures and augment limited sensing configurations. This digital approach to deep brain inference could expand the applicability of adaptive neuromodulation therapies and enable closed-loop control for emerging non-invasive stimulation techniques.
In this perspective paper, we introduce multimodal multi-task federated foundation models (M3T FedFMs) as a paradigm for privacy-preserving and distributed learning over biomedical sensing and imaging data. We outline their architecture, applications across medical sectors, key challenges, and future research directions, as well as the metrics and datasets that can facilitate their benchmarking.
Spatial omics technologies map molecular information within intact tissue architecture, revealing how cellular organization and interactions shape tumor biology, therapeutic responses, resistance, and relapse. Advances in artificial intelligence and computational pathology are bridging research discovery and clinical practice, enabling prognostic extraction from routine histology and cost-effective molecular inference. Nevertheless, barriers including high costs, protocol complexity, and lack of standardized workflows remain. We outline how these converging fields can deliver translational value in oncology and identify key enablers for clinical adoption.
Generating novel and functional molecules is an essential task in drug discovery, particularly in addressing the critical challenges of antibiotic resistance and the scarcity of effective treatments for major diseases such as cancer. Three-dimensional (3D) structure design can directly reflect a molecule’s biological function, while its complexity and topology irregularity make de novo 3D molecule generation highly difficult under valid geometric constraints. Conditional or controllable design is a promising solution for this challenging task in a more accurate and quick expectation manner. In this study, we propose TDmol-a Text-guided De novo 3D molecule generation approach based on a multimodal diffusion model. TDmol designs a new two-stage modality alignment contrastive deep learning pipeline to extract the cross-modality shared knowledge and unique features from different modalities, including molecular textual descriptions and molecular 2D/3D structures, enabling conditional and controllable molecule generation. To the best of our knowledge, TDmol is the first to realize text-3D modality alignment for conditional molecule generation. Experimental results demonstrate that TDmol achieves a significant performance enhancement for generating valid and reliable 3D molecular structures. This work highlights the potential of multimodal foundation models in digital medicine, offering a scalable and generalizable framework for 3D molecule generation that can be adapted across diverse clinical settings. A user-friendly webserver of TDmol has been deployed for academic use at http://www.csbio.sjtu.edu.cn/bioinf/TDMol.
Individuals with similar body mass index (BMI) often exhibit different cardiometabolic risk profiles, yet current approaches require multiple circulating biomarkers, limiting scalability. Here, we develop ROSA (Retinal-based Obesity Subtyping Algorithm), a deep learning framework that uses knowledge distillation to transfer multimodal cardiometabolic knowledge, derived from latent profile analysis of circulating biomarkers and retinal foundation model features, into a model requiring only a single fundus image. Applied to 63,693 participants across the UK Biobank (UKB; n = 37,137) and the Beijing Health Management Cohort external validation cohort (BHMC-EV; n = 26,556), ROSA classifies participants into three obesity subtypes identified by the multimodal latent profile analysis: baseline concordant, discordant inflammatory, and discordant hyperglycaemic. In the independent external BHMC-EV cohort, ROSA achieved a macro-averaged AUROC of 0.66 (95% CI: 0.64–0.67) and macro-averaged sensitivity of 0.45; class-specific external sensitivity and PPV were 0.84 and 0.87 for BC, 0.02 and 0.73 for DIS, and 0.51 and 0.34 for DHG, respectively. Performance in UKB-CHV was higher, with a macro-averaged AUROC of 0.80 (0.79–0.82) and sensitivity of 0.59. Over 13 years of follow-up in UKB-CHV, the hyperglycaemic subtype was associated with increased type 2 diabetes risk (HR 4.72, 95% CI: 3.70–6.02), consistent with its defining glycaemic profile, and accelerated progression to first cardiometabolic disease (HR 3.21, 2.53–4.06). Multi-omics analyses spanning the phenome, metabolome, proteome, and genome identified distinct biological associations with each subtype. ROSA demonstrates that retinal imaging can capture part of the cardiometabolic heterogeneity within obesity, while its limited sensitivity indicates that further calibration and validation are needed before clinical deployment.
The rapid growth of digital therapeutics (DTx) development is expected to enhance global health access and expand treatment options for broader patient populations. Nevertheless, because high-income countries remain the primary drivers of digital healthcare innovation, global analyses on DTx often overlook patterns in middle- and low-income countries. Without careful consideration to foster balanced growth, innovation gaps and disparities in accessibility may deepen. Accordingly, this study explores DTx development trends across therapeutic areas and technologies in countries with varying economic levels. Specifically, we examined 1672 patents filed with the United States Patent and Trademark Office and 7439 clinical trials registered in ClinicalTrials.gov from 2011 to 2025, analyzing patterns using large language models, network analysis, and the Bass diffusion model. Within these records, developed economies demonstrate broad coverage across therapeutic areas and technologies, positioning themselves as key innovators, whereas less developed economies tend to concentrate on underserved conditions with engagement limited to established technologies. We outline practical strategies tailored to different economic contexts and highlight the importance of international collaboration for the sustainable advancement of DTx to enhance balanced digital health access.