Surgical decision-making is complex and requires understanding causal relationships between patient characteristics, interventions, and outcomes. In high-stakes settings like spinal fusion or scoliosis correction, accurate estimation of individualized treatment effects (ITEs) remains limited due to the reliance on traditional statistical methods that struggle with complex and heterogeneous data. In this study, we develop a multi-task meta-learning framework, X-MultiTask, for ITE estimation that models each surgical decision (e.g., anterior vs. posterior approach, surgery vs. no surgery) as a distinct task while learning shared representations across tasks. To strengthen causal validity, we incorporate the inverse probability weighting (IPW) into the training objective. We evaluate our approach on two datasets: (1) a public spinal fusion dataset (1017 patients) to assess the effect of anterior vs. posterior approaches on complication severity; and (2) a private AIS dataset (368 patients) to analyze the impact of posterior spinal fusion (PSF) vs. non-surgical management on patient-reported outcomes (PROs). Our model achieves the highest average AUC (0.84) in the anterior group and maintains competitive performance in the posterior group (0.77). It outperforms baselines in treatment effect estimation. Similarly, when predicting PROs in AIS, X-MultiTask consistently shows superior performance across all domains. By providing robust, patient-specific causal estimates, X-MultiTask offers a powerful tool to advance personalized surgical care and improve patient outcomes. The code is available at https://github.com/Wizaaard/X-MultiTask.
Artificial intelligence (AI) in health care is shifting from incremental experimentation to systemic integration, reshaping clinical workflows, care delivery, and operational strategy. This article identifies eight converging technology trends across a unified data-to-decision pipeline defining the future of biomedical AI, driven by two megatrends: AI and automation.
To validate a clinically accessible approach for quantifying the Upper Extremity Reachable Workspace (UERW) using a single (monocular) camera and Artificial Intelligence (AI)-driven Markerless Motion Capture (MMC) for biomechanical analysis. Objective assessment and validation of these techniques for specific clinically oriented tasks are crucial for their adoption in clinical motion analysis. AI-driven monocular MMC reduces the barriers to adoption in the clinic and has the potential to reduce the overhead for analysis of this common clinical assessment. Nine adult participants with no impairments performed the standardized UERW task, which entails reaching targets distributed across a virtual sphere centered on the torso, with targets displayed in a VR headset. Movements were simultaneously captured using a marker-based motion capture system and a set of eight FLIR cameras. We performed monocular video analysis on two of these video camera views to compare a frontal and offset camera configurations. The frontal camera orientation demonstrated strong agreement with the marker-based reference, exhibiting a minimal mean bias of $0.61 \pm 0.12$ \% reachspace reached per octanct (mean $\pm$ standard deviation). In contrast, the offset camera view underestimated the percent workspace reached ($-5.66 \pm 0.45$ \% reachspace reached). Conclusion: The findings support the feasibility of a frontal monocular camera configuration for UERW assessment, particularly for anterior workspace evaluation where agreement with marker-based motion capture was highest. The overall performance demonstrates clinical potential for practical, single-camera assessments. This study provides the first validation of monocular MMC system for the assessment of the UERW task. By reducing technical complexity, this approach enables broader implementation of quantitative upper extremity mobility assessment.
Accurate prediction of breast cancer recurrence remains a critical challenge in oncology, with profound implications for treatment planning and long term outcomes. We present PharmacoKinetics-inspired cross-Attention(PKAttn), a multimodal deep learning framework that integrates pre-operative dynamic contrast enhanced MRI (DCE-MRI) with 84 routinely collected clinicopathologic features. PKAttn employs curve tokenization of DCE dynamics for low cost temporal modeling, phase-aware cross-attention that improves classification without a heavy 4D transformer, and simple squeeze-and-excitation (SE) style clinical gating. Trained on a cohort of 920 patients from the Duke Breast-MRI dataset, PKAttn achieved an F1-score of 0.9259, a recall of 0.9257, and a precision of 0.9338 on an independent test set outperforming single modality models (MRI-only $F 1=0.7995$; clinical only $F 1=0.8960$). Ablation studies confirm the essential role of temporal modeling, crossattention mechanism, multi-modality fusion, and backbone choice in predictive performance. These findings highlight the promise of synergistic imaging-clinical models for recurrence prediction of breast cancer.
Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics presents unique challenges with its complex biochemical pathways, heterogeneous identifier systems, and fragmented databases. To systematically evaluate LLM capabilities in this domain, we introduce MetaBench, the first benchmark for metabolomics assessment. Curated from authoritative public resources, MetaBench evaluates five capabilities essential for metabolomics research: knowledge, understanding, grounding, reasoning, and research. Our evaluation of 25 open- and closed-source LLMs reveals distinct performance patterns across metabolomics tasks: while models perform well on text generation tasks, cross-database identifier grounding remains challenging even with retrieval augmentation. Model performance also decreases on long-tail metabolites with sparse annotations. With MetaBench, we provide essential infrastructure for developing and evaluating metabolomics AI systems, enabling systematic progress toward reliable computational tools for metabolomics research.
We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. MedAgentGym comprises 72,413 task instances across 129 categories derived from 12 authentic real-world biomedical scenarios. Tasks are encapsulated within executable sandbox environments, each featuring detailed task specifications, interactive feedback mechanisms, verifiable ground truth annotations, and scalable training trajectory generation. Extensive benchmarking of 29 LLMs reveals substantial performance disparities in biomedical data science between commercial and open-source LLMs. Leveraging efficient multi-threaded and multi-turn trajectory sampling in MedAgentGym, Med-Copilot achieves performance gains of +43.02% and +45.28% from offline and online reinforcement learning, respectively, demonstrating MedAgentGym as an effective training ground while establishing itself as a cost-effective, privacy-preserving alternative competitive with proprietary LLMs (gpt-4o). By offering a unified execution environment with a comprehensive benchmark and accessible, extensible training resources, MedAgentGym delivers an integrated platform to develop LLM-based coding assistants for advanced biomedical data science.
Contribution: This article presents a modularized, implementation-ready problem-based learning (PBL) framework tailored to biomedical artificial intelligence (AI), with safe use of generative AI (GenAI) for knowledge summarization and coding support. A replication package (syllabi, milestones, rubrics, team procedures, and AI-usage templates) accompanies the framework. Background: PBL has significantly impacted biomedical engineering (BME) education since its introduction in the early 2000s, effectively enhancing student learning through critical thinking and real-world problem-solving. However, traditional PBL approaches demand considerable faculty resources, domain-specific expertise, and continuous curricular updates to align with rapidly evolving biomedical technologies. Recent advancements, including the 2024 Nobel Prizes awarded for AI-enabled discoveries, highlight the importance of comprehensive student training in AI. As BME rapidly converges with AI, integrating effective AI education into established curricula is necessary but faces challenges. These include diverse student backgrounds, limited personalized mentoring, constrained computational resources, and difficulties in safely scaling hands-on experiments due to privacy and ethical concerns associated with biomedical data. Intended outcomes: This article aims to improve students' grasp of biomedical AI concepts, increase proactive learning and teamwork, and address equity and scalability concerns in AI education through authentic real-world problem-solving experiences. Application design: A three-year case study from 2021 to 2023 across the Georgia Institute of Technology and Emory University engaged 248 students in interdisciplinary teams to address real biomedical AI problems. The curriculum positioned GenAI as both a class topic and a learning tool, governed by strict disclosure, source-anchoring, verification, and version-logging policies. Findings: This implementation coincided with measurable improvements in learning outcomes, evidenced by high problem-solving productivity (16 student-authored peer-reviewed publications), positive peer evaluations, and the successful development of innovative computational methods for real biomedical problems. This strategy not only prepares students for future healthcare innovation but also systematically addresses educational disparities and resource limitations inherent to conventional learning approaches. The study presents a practical and scalable roadmap for BME departments aiming to integrate robust AI education into their curricula.
The fragmentation of Electronic Health Records (EHR) and clinical information systems remains a significant challenge for an integrated healthcare system. In addition, the limitations associated with unidirectional data interactions of Extended Reality (XR) platforms, further constrain interoperability. Together, such factors hinder the establishment of a complete closed-loop pipeline integrating data acquisition, clinical decision making, and documentation within immersive clinical environments. To address this gap, we propose the Bidirectional Reality and Informatics Data GatEway (BRIDGE) platform, designed to achieve a fully integrated closed-loop pipeline. Centered on workflow continuity and a human-in-the-loop paradigm, BRIDGE adopts a distributed architecture that integrates Fast Healthcare Interoperability Resources (FHIR) Release 4 server implementation and Python-based medical image segmentation services. The system seamlessly integrates implements a complete workflow within a Virtual Reality (VR) environment on Meta Quest 3, encompassing patient retrieval, EHR browsing, advanced two-dimensional and three-dimensional medical image visualization, Artificial Intelligence (AI) assisted segmentation verification, and structured diagnostic feedback recording. End-to-end functional validation was conducted using synthetic EHR data and publicly available Neuroimaging Informatics Technology Initiative (NIfTI) image datasets. The results demonstrate the feasibility of maintaining a closed-loop workflow between the VR client and the FHIR server, including patient retrieval, medical image visualization, segmentation result loading, and structured write-back. Clinical deployment, clinician-facing usability evaluation, quantitative segmentation benchmarking, and privacy-preserving deployment remain subjects for future work.
Medical artificial intelligence systems require comprehensive knowledge integration to provide safe and reliable clinical decision support; however, many existing single-source RAG approaches offer limited coverage and may introduce bias. Here, we present MultiMed-RAG, a multi-agent framework that synthesizes heterogeneous medical knowledge from structured graphs, curated textual databases, web resources, and LLM internal knowledge to generate evidence-grounded clinical responses. The framework orchestrates specialized agents that decompose complex queries, dynamically select optimal knowledge sources, validate retrieved evidence against safety and accuracy criteria, and generate responses with transparent source attribution. Across nine medical benchmarks spanning diverse tasks, MultiMed-RAG achieves state-of-the-art performance, demonstrating consistent improvements of 4–7% over direct generation and 2–4% over existing RAG methods. Qualitative evaluation further indicates superior adherence to clinical safety, trustworthiness, actionability, and responsibility (STAR) principles. By addressing key challenges of knowledge heterogeneity, query complexity, and evidence transparency, MultiMed-RAG establishes a principled pathway toward trustworthy AI-assisted clinical decision support systems.
Contrastive learning, which is a powerful technique for learning image-level representations from unlabeled data, leads a promising direction to dealing with the dilemma between large-scale pre-training and limited labeled data. However, most existing contrastive learning strategies are designed mainly for downstream tasks of natural images, therefore they are sub-optimal and even worse than learning from scratch when directly applied to medical images whose downstream tasks are usually segmentation. In this work, we propose a novel asymmetric contrastive learning framework named JCL for medical image segmentation with self-supervised pre-training. Specifically, (1) A novel asymmetric contrastive learning strategy is proposed to pre-train both encoder and decoder simultaneously in one-stage to provide better initialization for segmentation models. (2) A multi-level contrastive loss is designed to take the correspondence among feature-level, image-level and pixel-level projections, respectively into account to make sure multi-level representations can be learned by the encoder and decoder during pre-training. (3) Experiments on multiple medical image datasets indicate our JCL framework outperforms existing SOTA contrastive learning strategies.
Foundation models for electroencephalography (EEG) signals have recently demonstrated success in learning generalized representations of EEGs, outperforming specialized models in various downstream tasks. However, many of these models lack transparency in their pretraining dynamics and offer limited insight into how well EEG information is preserved within their embeddings. For successful clinical integration, EEG foundation models must ensure transparency in pretraining, downstream fine-tuning, and the interpretability of learned representations. Current approaches primarily operate in the temporal domain, overlooking advancements in digital signal processing that enable the extraction of deterministic and traceable features, such as wavelet-based representations. We propose MENDR (Manifold Explainable Neural Data Representations), a filter bank-based EEG foundation model built on a novel Riemannian Manifold Transformer architecture to resolve these issues. MENDR learns symmetric positive definite matrix embeddings of EEG signals and is pretrained on a large corpus comprising over 4,000 hours of EEG data, decomposed via discrete wavelet packet transforms into multi-resolution coefficients. MENDR significantly enhances interpretability by visualizing symmetric positive definite embeddings as geometric ellipsoids and supports accurate reconstruction of EEG signals from learned embeddings. Evaluations across multiple clinical EEG tasks demonstrate that MENDR achieves near state-of-the-art performance with substantially fewer parameters, underscoring its potential for efficient, interpretable, and clinically applicable EEG analysis.
Accurate quantification of spinal curvature from radiographs is essential for diagnosing scoliosis and guiding treatment, yet manual Cobb angle measurement remains time-consuming and prone to observer variability. Existing automated methods typically emphasize either global spinal morphology or local vertebral landmarks but rarely integrate both sources of information. To address this limitation, we propose a hybrid deep learning framework that combines global spine context with local landmark detection to enhance the accuracy of Cobb angle estimation. By leveraging the complementary strengths of both approaches, our model improves measurement precision compared to using each component in isolation. Applied to full-length anterior-posterior radiographs, the framework delivers fully automated Cobb angle measurements with a Symmetric Mean Absolute Percentage Error (SMAPE) of 8.58%, outperforming prior methods that report SMAPE values between roughly 9.5% and 51%, highlighting its potential to streamline clinical workflows and support data-driven spine care.
We introduce MedAgentGYM, the first publicly available training environment designed to enhance coding-based medical reasoning capabilities in large language model (LLM) agents. MedAgentGYM comprises 72,413 task instances across 129 categories derived from authentic real-world biomedical scenarios. Tasks are encapsulated within executable coding environments, each featuring detailed task descriptions, interactive feedback mechanisms, verifiable ground-truth annotations, and scalable training trajectory generation. Extensive benchmarking of over 30 LLMs reveals a notable performance disparity between commercial API-based models and open-source counterparts. Leveraging MedAgentGYM, Med-Copilot-7B achieves substantial performance gains through supervised fine-tuning (+36.44 (+42.47 competitive with gpt-4o. By offering both a comprehensive benchmark and accessible, expandable training resources within unified execution environments, MedAgentGYM delivers an integrated platform to develop LLM-based coding assistants for advanced biomedical research and practice.
BackgroundAdolescent idiopathic scoliosis (AIS) is the most common type of scoliosis, affecting 1-4% of adolescents. The Scoliosis Research Society-22R (SRS-22R), a health-related quality-of-life instrument for AIS, has allowed orthopedists to measure subjective patient outcomes before and after corrective surgery beyond objective radiographic measurements. However, research has revealed that there is no significant correlation between the correction rate in major radiographic parameters and improvements in patient-reported outcomes (PROs), making it difficult to incorporate PROs into personalized surgical planning.MethodsThe objective of this study is to develop an artificial intelligence (AI)-enabled surgical planning and counseling support system for post-operative patient rehabilitation outcomes prediction in order to facilitate personalized AIS patient care. A unique multi-site cohort of 455 pediatric patients undergoing spinal fusion surgery at two Shriners Children's hospitals from 2010 is investigated in our analysis. In total, 171 pre-operative clinical features are used to train six machine-learning models for post-operative outcomes prediction. We further employ explainability analysis to quantify the contribution of pre-operative radiographic and questionnaire parameters in predicting patient surgical outcomes. Moreover, we enable responsible AI by calibrating model confidence for human intervention and mitigating health disparities for algorithm fairness.ResultsThe best prediction model achieves an area under receiver operating curve (AUROC) performance of 0.86, 0.85, and 0.83 for individual SRS-22R question response prediction over three-time horizons from pre-operation to 6-month, 1-year, and 2-year post-operation, respectively. Additionally, we demonstrate the efficacy of our proposed prediction method to predict other patient rehabilitation outcomes based on minimal clinically important differences (MCID) and correction rates across all three-time horizons.ConclusionsBased on the relationship analysis, we suggest additional attention to sagittal parameters (e.g., lordosis, sagittal vertical axis) and patient self-image beyond major Cobb angles to improve surgical decision-making for AIS patients. In the age of personalized medicine, the proposed responsible AI-enabled clinical decision-support system may facilitate pre-operative counseling and shared decision-making within real-world clinical settings.
Polygenic risk scores (PRSs) are increasingly being used to predict disease risk from genetic data. While promising in research, their clinical utility—especially when combined with non-genetic (NG) data such as lab results, physical measurements, and diagnostic history—remains uncertain. Myocardial infarction (MI), a leading cause of morbidity and mortality, is a key use case for assessing the incremental value of PRSs in risk models. Using UK Biobank data, we evaluated the added value of PRSs for 10-year MI risk prediction. We trained models with NG data alone and in combination with PRSs, varying model complexity and the NG feature space. Two modeling frameworks were used: logistic regression and a neural network. NG data was defined using two feature sets: NG1, which included established MI risk factors from structured fields; and NG2, a high-dimensional dataset derived from millions of diagnostic codes across five linked UK Biobank electronic health records (EHR) datasets combined with NG1 features. NG2 was generated using a deep representation learning approach that produced low-dimensional embeddings capturing latent medical concepts and disease co-occurrence patterns. Each model was trained with and without PRSs and evaluated using metrics such as the area under the ROC curve (AUC). PRSs add minimal predictive value when used alone. In contrast, diagnostic data from EHRs significantly improve performance. The best results are achieved using a multimodal neural network combining NG1, NG2, and PRSs. PRSs provide limited standalone utility for MI prediction compared to detailed diagnostic data. Their clinical value likely lies in integration with EHR-based models. Future work should focus on multi-modal approaches that contextualize PRS information within broader clinical data. Polygenic risk scores are a measure of how likely it is a person will get a particular disease based on the genes they inherited from their parents. We studied whether polygenic risk scores (PRSs) help predict heart attack occurrence when combined with clinical data such as information from blood tests and medical history. We tested two different computational models using two types of clinical data: one with common risk factors for heart attacks and another with broader data. We found that PRSs added little benefit by themselves, instead, detailed clinical information was much more useful. The best predictions came from combining all data types. This information and our models could be helpful to better identify people who will have heart attacks. Isgut et al. compare the effectiveness of polygenic risk score (PRS) versus electronic health records (EHR) in predicting 10-year myocardial infarction risk using a machine learning approach. While PRS adds minimal value compared to EHR data, integrating both could optimise risk stratification.
Sleep disorders, particularly Obstructive Sleep Apnea (OSA), have a considerable effect on an individual's health and quality of life. Accurate sleep stage classification and prediction of OSA are crucial for timely diagnosis and effective management of sleep disorders. In this study, we develop a sequential network that enhances sleep stage classification by incorporating self-attention mechanisms and Conditional Random Fields (CRF) into a deep learning model comprising multi-kernel Convolutional Neural Networks (CNNs) and Transformer-based encoders. The self-attention mechanism enables the model to focus on the most discriminative features extracted from single-channel electroencephalography (EEG) recordings, while the CRF module captures the temporal dependencies between sleep stages, improving the model's ability to learn more plausible sleep stage sequences. Moreover, we explore the relationship between sleep stages and OSA severity by utilizing the predicted sleep stage features to train various regression models for Apnea-Hypopnea Index (AHI) prediction. Our experiments demonstrate an improved sleep stage classification performance of 78.7%, particularly on datasets with diverse AHI values, and highlight the potential of leveraging sleep stage information for monitoring OSA. By employing advanced deep learning techniques, we thoroughly explore the intricate relationship between sleep stages and sleep apnea, laying the foundation for more precise and automated diagnostics of sleep disorders.
Colon cancer is one of the deadliest types of cancer in the United States, with close to 50,000 projected deaths in 2024. The disease requires early diagnosis to optimize chances of survival by enabling timely administration of treatment. To investigate the key non-genetic (NG) factors influencing the onset of colon cancer and evaluate how genetic factors enhance the performance of machine learning (ML) models in predicting incidence, we incorporated polygenic risk scores (PRSs) alongside NG data in ML models to predict 10-year incident risk prediction of colon cancer using data from the UK Biobank. This approach enabled us to assess the added predictive value of PRSs in multi-modal models in estimating the 10-year risk of developing colon cancer over NG data alone. Moreover, our research focused on identifying the most relevant and predictive PRS and validating them using a robust ML framework. To ensure the robustness, we restricted the cohort to White British individuals to minimize ancestry-related heterogeneity. PRSs have proven effective in enhancing disease prediction for conditions such as breast cancer, myocardial infarction, and schizophrenia, reinforcing their relevance in clinical research. Exploring six PRSs, our goal was to minimize false negatives while simultaneously maximizing area under the receiver-operating characteristic curve (AUC), in order to improve early detection rates by identifying those who are at risk for colon cancer. This research shows that PRSs can be used to enhance overall predictive ability of ML models in colon cancer research over NG factors alone, bolstering the argument for incorporating PRSs into routine clinical practice. PRSs can also help minimize false negatives, a key feature for disease prediction models, as missed potential diagnoses are life-threatening.
Retinal vessel segmentation is a vital early detection method for several severe ocular diseases. Despite significant progress in retinal vessel segmentation with the advancement of Neural Networks, there are still challenges to overcome. Specifically, retinal vessel segmentation aims to predict the class label for every pixel within a fundus image, with a primary focus on intra-image discrimination, making it vital for models to extract more discriminative features. Nevertheless, existing methods primarily focus on minimizing the difference between the output from the decoder and the label, but ignore fully using feature-level fine-grained representations from the encoder. To address these issues, we propose a novel Attention U-shaped Kolmogorov-Arnold Network named AttUKAN along with a novel Label-guided Pixel-wise Contrastive Loss for retinal vessel segmentation. Specifically, we implement Attention Gates into Kolmogorov-Arnold Networks to enhance model sensitivity by suppressing irrelevant feature activations and model interpretability by non-linear modeling of KAN blocks. Additionally, we also design a novel Label-guided Pixel-wise Contrastive Loss to supervise our proposed AttUKAN to extract more discriminative features by distinguishing between foreground vessel-pixel pairs and background pairs. Experiments are conducted across four public datasets including DRIVE, STARE, CHASE_DB1, HRF and our private dataset. AttUKAN achieves F1 scores of 82.50%, 81.14%, 81.34%, 80.21% and 80.09%, along with MIoU scores of 70.24%, 68.64%, 68.59%, 67.21% and 66.94% in the above datasets, which are the highest compared to 11 networks for retinal vessel segmentation. Quantitative and qualitative results show that our AttUKAN achieves state-of-the-art performance and outperforms existing retinal vessel segmentation methods. Our code will be available at https://github.com/stevezs315/AttUKAN.