BACKGROUND:Macrophages are key players in the pathogenesis of atherosclerosis. They trigger immune responses through their cell-surface receptors. However, how macrophages regulate those receptors in response to proinflammatory stimuli is not completely understood. Endocytic membrane trafficking involving receptor internalization, followed by endosomal transport and recycling of the internalized receptors, plays essential roles in balancing cell-surface receptors to meet cellular needs. Here, we explored the role of the endocytic regulator EHD (c-terminal Eps15 [Epidermal growth factor receptor substrate 15] homology domain) 1 in immune responses in macrophages and determined its contribution to atherosclerosis progression. METHODS:EHD1 expression profiles in mouse and human plaques were determined by single-cell RNA sequencing and immunofluorescence staining. Bone marrow transplantation by transplanting bone marrow cells from Ehd1-/- or littermate wild-type mice to irradiated Ldlr-/- mice was performed to determine the effect of EHD1 deletion on atherosclerosis progression. In vitro mechanistic studies, including inflammation signaling and endocytosis assays, were performed in bone marrow-derived macrophages. RESULTS:EHD1 expression in macrophages is enhanced as atherosclerosis progresses in both mice and humans. Histological analysis of aortic root sections from bone marrow transplantation mice showed that EHD1 deletion reduces lesion size. Single-cell RNA sequencing of aortic CD (cluster of differentiation) 45+ cells demonstrated that EHD1 deletion attenuates proinflammatory responses and cell-cell interactions. Mechanistic studies revealed that EHD1 accelerates the endocytic recycling of TNFR2 (tumor necrosis factor receptor 2) and activates NF-κB (nuclear factor kappa B), leading to increased expression of inflammatory cytokines. Moreover, EHD1 interacts with retromer and stabilizes sortilin, a retrograde cargo of retromer and a risk factor for atherosclerosis. CONCLUSIONS:EHD1 promotes inflammation by enhancing TNFR2-NF-κB signaling and stabilizing sortilin, leading to accelerated atherosclerosis. Our study reveals novel roles for EHD1-mediated membrane trafficking in macrophage function and paves the way to innovative therapeutic strategies that aim to address dysregulated membrane trafficking in atherosclerosis.
Ambulatory electrocardiograms (ECG) provides continuous monitoring of the heart's electrical activity. However, many existing machine learning and artificial intelligence models for analyzing ambulatory ECG traces are often unimodal and do not incorporate patient clinical context. In this study, we propose a multimodal framework integrating ambulatory ECG-derived representations with clinical text embeddings to predict two cardiac outcomes: sudden cardiac death and pump failure death. Ambulatory ECG traces are preprocessed, segmented, and encoded via a multiple instance learning and temporal convolutional neural network framework. In parallel, patient clinical features are parsed into structured prompts, which are passed through a large language model to generate clinical reasoning; this reasoning passes through a biomedical language encoder to generate a text embedding. With the ECG and text embeddings, we systematically evaluate multiple fusion strategies, including concatenation- and gating-based approaches, to integrate these two data modalities. Our results demonstrate that multimodal models consistently outperform unimodal baselines, with adaptive fusion mechanisms providing the greatest improvements in predictive performance. Decision curve analysis highlights the potential clinical utility of the proposed framework for risk stratification. Finally, we visualize model attention across modalities, including ECG attention patterns, segment-level saliency, heart rate variability features, and clinical reasoning, to contextualize patient-specific predictions.
The ACM KDD 2025 Health Day theme, ''Harnessing AI Opportunities in Biomedicine and Healthcare'' highlights the transformative potential of AI-driven applications in healthcare, translational biomedical research, and basic biological research. This extended abstract discusses recent advancements, challenges, and future directions, focusing on integrating AI-ready data sets, interdisciplinary collaborations, and ethical AI practices. It aims to catalyze discussions on the potential of AI ecosystems in revolutionizing healthcare and related fields.
In this work, we study the problem pertaining to personalized classification of subclinical atherosclerosis by developing a hierarchical graph neural network framework to leverage two characteristic modalities of a patient: clinical features within the context of the cohort, and molecular data unique to individual patients. Current graph-based methods for disease classification detect patient-specific molecular fingerprints, but lack consistency and comprehension regarding cohort-wide features, which are an essential requirement for understanding pathogenic phenotypes across diverse atherosclerotic trajectories. Furthermore, understanding patient subtypes often considers clinical feature similarity in isolation, without integration of shared pathogenic interdependencies among patients. To address these challenges, we introduce ATHENA: Atherosclerosis Through Hierarchical Explainable Neural Network Analysis, which constructs a novel hierarchical network representation through integrated modality learning; subsequently, it optimizes learned patient-specific molecular fingerprints that reflect individual omics data, enforcing consistency with cohort-wide patterns. With a primary clinical dataset of 391 patients (PESA study), their respective transcriptomics signatures, as well as their STRING PPI profiles, we demonstrate that this heterogeneous alignment of clinical features with molecular interaction patterns has significantly boosted subclinical atherosclerosis classification performance across various baselines by up to 13% in area under the receiver operating curve (AUC) and 20% in F1 score. We further validated ATHENA on a secondary clinical dataset on atherosclerosis (by Steenman and Espitia et al.). Taken together, ATHENA enables mechanistically-informed patient subtype discovery through explainable AI (XAI)-driven subnetwork clustering; this novel integration framework strengthens personalized intervention strategies, thereby improving the prediction of atherosclerotic disease progression and management of their clinical actionable outcomes.
The ACM KDD 2025 Health Day theme, "Harnessing AI Opportunities in Biomedicine and Healthcare" highlights the transformative potential of AI-driven applications in healthcare, translational biomedical research, and basic biological research. This extended abstract discusses recent advancements, challenges, and future directions, focusing on integrating AI-ready data sets, interdisciplinary collaborations, and ethical AI practices. It aims to catalyze discussions on the potential of AI ecosystems in revolutionizing healthcare and related fields.
Existing machine learning methods for molecular (e.g., gene) embeddings are restricted to specific tasks or data modalities, limiting their effectiveness within narrow domains. As a result, they fail to capture the full breadth of gene functions and interactions across diverse biological contexts. In this study, we have systematically evaluated knowledge representations of biomolecules across multiple dimensions representing a task-agnostic manner spanning three major data sources, including omics experimental data, literature-derived text data, and knowledge graph-based representations. To distinguish between meaningful biological signals from chance correlations, we devised an adjusted variant of Singular Vector Canonical Correlation Analysis (SVCCA) that quantifies signal redundancy and complementarity across different data modalities and sources. These analyses reveal that existing embeddings capture largely non-overlapping molecular signals, highlighting the value of embedding integration. Building on this insight, we propose Platform for Representation and Integration of multimodal Molecular Embeddings (PRISME), a machine learning based workflow using an autoencoder to integrate these heterogeneous embeddings into a unified multimodal representation. We validated this approach across various benchmark tasks, where PRISME demonstrated consistent performance, and outperformed individual embedding methods in missing value imputations. This new framework supports comprehensive modeling of biomolecules, advancing the development of robust, broadly applicable multimodal embeddings optimized for downstream biomedical machine learning applications.
BackgroundStatin-associated muscle symptoms (SAMS) are a significant clinical issue, and their exact cause is not well understood. Immunological mechanisms have been suggested but have not been confirmed. This study is a rare, longitudinal case-based analysis that uses transcriptomics to explore immune-related gene expression changes in peripheral blood mononuclear cells (PBMCs) in response to exercise before, during, and after the onset and resolution of SAMS.MethodsA healthy volunteer (HV1) enrolled in an exercise immuno-fitness study underwent cardiopulmonary exercise testing (CPX) with blood collected at three timepoints: pre-exercise (TP1), peak exercise (TP2), and 1 hour post-exercise (TP3). After baseline testing (Visit 1), the participant began statin therapy on their own, developed SAMS, and had repeat CPX testing during the symptomatic phase (Visit 2) and partial recovery phase (Visit 3). RNA was extracted from PBMCs and analyzed using next-generation RNA sequencing. The data were evaluated using differential gene expression analysis and Weighted Gene Co-expression Network Analysis (WGCNA). Pathway and gene ontology enrichment were used to identify immunologic signatures associated with SAMS.ResultsThe PBMC gene expression profiles showed distinct changes during SAMS compared to the baseline and recovery phases. WGCNA identified 39 co-expression modules. Several modules had high expression at peak exercise in the healthy state (V1), which was attenuated in SAMS (V2) and partially restored in recovery (V3). Gene ontology and Reactome analyses of key modules identified 16 genes that were differentially expressed at peak exercise and may be involved in specific immune pathways in SAMS pathogenesis.ConclusionThis case study suggests that profiling the exercise-induced immune transcriptome can reveal dynamic immunological changes related to statin-induced myopathy. These findings support the hypothesis of an immune-mediated component in SAMS and provide a basis for future studies to validate transcriptomic biomarkers for the early detection and management of SAMS.
Objective: As AI becomes increasingly central to healthcare, there is a pressing need for bioinformatics and biomedical training systems that are personalized and adaptable. Materials and Methods: The NIH Bridge2AI Training, Recruitment, and Mentoring (TRM) Working Group developed a cross-disciplinary curriculum grounded in collaborative innovation, ethical data stewardship, and professional development within an adapted Learning Health System (LHS) framework. Results: The curriculum integrates foundational AI modules, real-world projects, and a structured mentee-mentor network spanning Bridge2AI Grand Challenges and the Bridge Center. Guided by six learner personas, the program tailors educational pathways to individual needs while supporting scalability. Discussion: Iterative refinement driven by continuous feedback ensures that content remains responsive to learner progress and emerging trends. Conclusion: With over 30 scholars and 100 mentors engaged across North America, the TRM model demonstrates how adaptive, persona-informed training can build interdisciplinary competencies and foster an integrative, ethically grounded AI education in biomedical contexts.
The vast amount of biomedical information available today presents a significant challenge for investigators seeking to digest, process, and understand these findings effectively. Large Language Models (LLMs) have emerged as powerful tools to navigate this complex and challenging data landscape. However, LLMs may lead to hallucinatory responses, making Retrieval Augmented Generation (RAG) crucial for achieving accurate information. In this protocol, we present RUGGED (Retrieval Under Graph-Guided Explainable disease Distinction), a comprehensive workflow designed to support investigators with knowledge integration and hypothesis generation, identifying validated paths forward. Relevant biomedical information from publications and knowledge bases are reviewed, integrated, and extracted via text-mining association analysis and explainable graph prediction models on disease nodes, forecasting potential links among drugs and diseases. These analyses, along with biomedical texts, are integrated into a framework that facilitates user-directed mechanism elucidation as well as hypothesis exploration through RAG-enabled LLMs. A clinical use-case demonstrates RUGGED's ability to evaluate and recommend therapeutics for Arrhythmogenic Cardiomyopathy (ACM) and Dilated Cardiomyopathy (DCM), analyzing prescribed drugs for molecular interactions and unexplored uses. The platform minimizes LLM hallucinations, offers actionable insights, and improves the investigation of novel therapeutics.
The scale of biomedical knowledge, spanning scientific literature and curated knowledge bases, poses a significant challenge for investigators in processing, evaluating, and interpreting findings effectively. Large Language Models (LLMs) have emerged as powerful tools for navigating this complex knowledge landscape but may produce hallucinatory responses. Retrieval-Augmented Generation (RAG) is essential for identifying relevant information to enhance accuracy and reliability. This protocol introduces RUGGED (Retrieval Under Graph-Guided Explainable disease Distinction), a comprehensive workflow designed to support knowledge integration, to mitigate bias, and to explore and validate new research directions. Biomedical information from publications and knowledge bases are synthesized and analyzed through text-mining association analysis and explainable graph prediction models to uncover potential drug-disease relationships. These findings, along with the source text corpus and knowledge bases, are incorporated into a framework that employs RAG-enhanced LLMs to enables users to explore hypotheses and investigate underlying mechanisms. A clinical use case demonstrates RUGGED's capability in evaluating and recommending therapeutics for Arrhythmogenic Cardiomyopathy (ACM) and Dilated Cardiomyopathy (DCM), analyzing prescribed drugs for molecular interactions and potential new applications. The platform reduces LLM hallucinations, highlights actionable insights, and streamlines the investigation of novel therapeutics.
Knowledge graphs (KGs) have emerged as a powerful framework for representing and integrating complex biomedical information. However, assembling KGs from diverse sources remains a significant challenge in several aspects, including entity alignment, scalability, and the need for continuous updates to keep pace with scientific advancements. Moreover, the representative power of KGs is often limited by the scarcity of multi-modal data integration. To overcome these challenges, we propose Know2BIO, a general-purpose heterogeneous KG benchmark for the biomedical domain. Know2BIO integrates data from 30 diverse sources, capturing intricate relationships across 11 biomedical categories. It currently consists of ~219,000 nodes and ~6,200,000 edges. Know2BIO is capable of user-directed automated updating to reflect the latest knowledge in biomedical science. Furthermore, Know2BIO is accompanied by multi-modal data: node features including text descriptions, protein and compound sequences and structures, enabling the utilization of emerging natural language processing methods and multi-modal data integration strategies. We evaluate KG representation models on Know2BIO, demonstrating its effectiveness as a benchmark for KG representation learning in the biomedical field. Data and source code of Know2BIO are available at https://github.com/Yijia-Xiao/Know2BIO/.
The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency and accessibility of medical care. While LLMs have shown some promise in the medical domain, their application in clinical diagnosis remains underexplored, especially in real-world clinical practice, where highly sophisticated, patient-specific decisions need to be made. Current evaluations of LLMs in this field are often narrow in scope, focusing on specific diseases or specialties and employing simplified diagnostic tasks. To bridge this gap, we introduce CliBench, a novel benchmark developed from the MIMIC IV dataset, offering a comprehensive and realistic assessment of LLMs' capabilities in clinical diagnosis. This benchmark not only covers diagnoses from a diverse range of medical cases across various specialties but also incorporates tasks of clinical significance: treatment procedure identification, lab test ordering and medication prescriptions. Supported by structured output ontologies, CliBench enables a precise and multi-granular evaluation, offering an in-depth understanding of LLM's capability on diverse clinical tasks of desired granularity. We conduct a zero-shot evaluation of leading LLMs to assess their proficiency in clinical decision-making. Our preliminary results shed light on the potential and limitations of current LLMs in clinical settings, providing valuable insights for future advancements in LLM-powered healthcare.
Background: Cardiorespiratory fitness positively correlates with longevity and immune health. Regular exercise may provide health benefits by reducing systemic inflammation. In chronic disease conditions, such as chronic heart failure and chronic fatigue syndrome, mechanistic links have been postulated between inflammation, muscle weakness, frailty, catabolic/anabolic imbalance, and aberrant chronic activation of immunity with monocyte upregulation. We hypothesize that (1) temporal changes in transcriptome profiles of peripheral blood mononuclear cells during strenuous acute bouts of exercise using cardiopulmonary exercise testing are present in adult subjects, (2) these temporal dynamic changes are different between healthy persons and heart failure patients and correlate with clinical exercise-parameters and (3) they portend prognostic information. Methods: In total, 16 Heart Failure (HF) patients and 4 healthy volunteers (HV) were included in our proof-of-concept study. All participants underwent upright bicycle cardiopulmonary exercise testing. Blood samples were collected at three time points (TP) (TP1: 30 min before, TP2: peak exercise, TP3: 1 h after peak exercise). We divided 20 participants into 3 clinically relevant groups of cardiorespiratory fitness, defined by peak VO2: HV (n = 4, VO2 ≥ 22 mL/kg/min), mild HF (HF1) (n = 7, 14 < VO2 < 22 mL/kg/min), and severe HF (HF2) (n = 9, VO2 ≤ 14 mL/kg/min). Results: Based on the statistical analysis with 20–100% restriction, FDR correction (p-value 0.05) and 2.0-fold change across the three time points (TP1, TP2, TP3) criteria, we obtained 11 differentially expressed genes (DEG). Out of these 11 genes, the median Gene Expression Profile value decreased from TP1 to TP2 in 10 genes. The only gene that did not follow this pattern was CCDC181. By performing 1-way ANOVA, we identified 8/11 genes in each of the two groups (HV versus HF) while 5 of the genes (TTC34, TMEM119, C19orf33, ID1, TKTL2) overlapped between the two groups. We found 265 genes which are differentially expressed between those who survived and those who died. Conclusions: From our proof-of-concept heart failure study, we conclude that gene expression correlates with VO2 peak in both healthy individuals and HF patients, potentially by regulating various physiological processes involved in oxygen uptake and utilization during exercise. Multi-omics profiling may help identify novel biomarkers for assessing exercise capacity and prognosis in HF patients, as well as potential targets for therapeutic intervention to improve VO2 peak and quality of life. We anticipate that our results will provide a novel metric for classifying immune health.
The ACM KDD 2024 Health Day theme, "Building Health AI Ecosystem: From Data Harmonization to Knowledge Discovery," highlights the transformative potential of AI-driven ecosystems in healthcare, translational biomedical research, and basic biological research. This extended abstract discusses recent advancements, challenges, and future directions, focusing on integrating AI-ready data sets, interdisciplinary collaborations, and ethical AI practices. It aims to catalyze discussions on the potential of AI ecosystems in revolutionizing healthcare and related fields.
Wei Wang (王薇)合作论文数Department of Computer Science, University of California at Los Angeles;Department of Computational Medicine, University of California at Los Angeles;Scalable Analytics Institute, University of California at Los Angeles40