Artificial intelligence agents are emerging as powerful applications of large language models (LLMs), automating complex tasks and enabling scientific data exploration. However, their use in biomedical data analysis remains limited by the difficulty of handling specialized tools and multistep reasoning. Here we introduce BioMedAgent, a self-evolving LLM multi-agent framework, which learns to use diverse bioinformatics tools and chain them into executable workflows through interactive exploration and memory retrieval algorithms. It allows biomedical users to initiate tasks using natural language, without requiring computational expertise. Evaluated on our newly released BioMed-AQA benchmark comprising 327 biomedical data tasks, BioMedAgent achieved a 77% success rate, surpassing other LLM agents, and generalized robustly to the external BixBench dataset. Beyond benchmarks, it autonomously performs cross-omics analysis, machine-learning modelling and pathology image segmentation, highlighting its potential to advance biomedical research and extend to other scientific domains requiring complex tool integration and multistep reasoning.
BackgroundColorectal cancer (CRC) remains a leading cause of global cancer mortality, highlighting the need for precise survival prediction to guide clinical decisions. Although tissue-level multi-omics is widely utilized for survival prediction, its limited resolution cannot capture tumor heterogeneity. Single-cell RNA sequencing (scRNA-seq) enables dissection of the tumor microenvironment (TME) at cellular resolution, supporting personalized prognostic assessment.MethodsWe collected 213 CRC scRNA-seq samples and established a CRC-specific TME atlas comprising 339,060 cells. Using this atlas as a reference, we deconvolved bulk RNA-seq data from TCGA-CRC cohort with the EcoTyper algorithm to reconstruct TME features. Clinical, genomic, and transcriptomic data were obtained from the Xena platform; microbial data were sourced from the BIC database. We integrated TME and multi-omics features through a self-normalizing neural network to construct a deep learning model (single-cell resolution TME ecosystem with multi-omics data [SCMO]) for survival prediction. To enhance interpretability, we utilized the Integrated Gradients algorithm and spatial transcriptomic data to analyze multi-omics and TME features. We performed anticancer drug screening with tumor necrosis factor receptor-associated protein 1 (TRAP1), a critical feature according to the Integrated Gradients algorithm, as a potential target.ResultsWe identified 13 survival-related TME features from the CRC-specific atlas: 12 cell states and one multi-cellular ecosystem. SCMO, which combined TME and multi-omics features, improved survival prediction and outperformed existing methods, achieving a concordance index of 0.762. The SCMO demonstrated robust performance for long-term predictions, achieving areas under the curve (AUCs) of 0.752, 0.772, and 0.869 for 1-, 3-, and 5-year predictions in the training set, with corresponding test set AUCs of 0.639, 0.756, and 0.772. TME features from the SCMO model revealed that ecosystem density increased with CRC malignancy. Multi-omics features included TRAP1 as a potential drug target. Drug screening identified saikosaponin A as a novel TRAP1 inhibitor, and its anticancer activity was validated in vitro. We developed SCMO-Lite, a simplified model incorporating 12 high-attribution-weight multi-omics features, which demonstrated robust risk stratification.ConclusionsSCMO combines analytical precision with biological interpretability, offering novel insights for oncology survival prediction.
Breast cancer remains a major global health burden, underscoring the urgent need for reliable early detection strategies. Exosomes, as mediators of intercellular communication, have shown promise in early tumor screening through Raman spectroscopy and gene expression profiling in pancreatic and colorectal cancers. However, the application of exosomal gene expression profiles for breast cancer prediction remains largely unexplored. Exosomal mRNA profiles were obtained from exoRBase 3.0 (242 breast cancer, 244 healthy controls). Sample sex was inferred using XIST and UTY expression, yielding 337 female samples for analysis. A nested cross-validation framework (20 repetitions, fivefold) was implemented, with differential expression analysis and feature selection performed exclusively within each training fold to prevent information leakage. Ten machine learning classifiers were evaluated on an independent held-out test set. Model performance was assessed using accuracy, precision, recall, and F1-score. Feature selection demonstrated high stability (average Jaccard score 0.5912), with 9 genes consistently selected across all 100 iterations and a set of robust feature genes was identified. Among classifiers, xgbTree achieved the best performance (AUC 0.992, accuracy 0.970, F1 0.979) on the independent test set, supporting exosomal mRNA profiles as a promising non-invasive approach for the early breast cancer detection.
Neoadjuvant chemoradiotherapy (nCRT) is the standard treatment for locally advanced rectal cancer (LARC), yet clinically validated biomarkers for predicting response remain lacking. This study aimed to identify candidate molecular events associated with nCRT response and to develop a pretreatment prediction framework integrating genomic and pathological information. Whole-exome sequencing (WES) was performed on pretreatment tumors from 67 patients with LARC, and an additional 22 published WES cases were integrated to compare genomic differences between responders (R) and nonresponders (NR). Using histopathological whole-slide images (WSIs; n = 106) and genome-derived features, a weakly supervised, multimodal deep learning fusion model was developed to predict nCRT response. Multiomics profiling was used for exploratory pathway characterization, and functional assays were conducted in colorectal cancer cell lines and mouse models harboring KRASG12V or KRASA146T. WES identified 41 response-associated hotspot codon events. KRASG12V and KRASA146T were detected in the NR group in this cohort and were directionally aligned with poor response, indicating an association with nCRT resistance. Because these events are low-frequency alterations, and the study is a single-center retrospective cohort with limited numbers of carriers, multivariable adjustment for key covariates (including stage and T/N status) was not feasible; thus, these findings should be interpreted as exploratory candidate signals. The multimodal fusion model showed good discrimination within the cohort (AUC = 0.882). Mechanistically, exploratory multiomics analyses and orthogonal functional assays were consistent with KRAS variants being associated with altered DNA damage repair signaling and increased repair capacity, with the functional assays providing the main support for this interpretation. The proposed genome–pathology fusion model provides a research-oriented framework for pretreatment prediction and risk stratification of nCRT response in LARC. KRASG12V and KRASA146T are presented as candidate molecular events aligned with poor response, but their independent predictive value and the clinical usability of the model require validation in larger, multicenter prospective cohorts that include external WSI data, together with systematic evaluation of thresholding and calibration before clinical translation.
Artificial intelligence (AI) is reshaping medical informatics from a discipline of data management into a science of integration, inference, and translation. As biomedical data proliferate across physiological, clinical, and molecular domains, AI functions as the integrative engine that transforms complexity into actionable understanding. In this survey, we synthesize recent advances spanning data representation, algorithmic innovation, and clinical deployment, emphasizing the transition from isolated tasks to cohesive systems that link discovery and care. We highlight how advances in medical AI algorithms across clinical data, medical imaging, and multi-omics are beginning to converge with applications in clinical diagnosis, drug discovery, precision medicine, and surgery. Looking ahead, medical AI is moving toward a self-reflective and collaborative paradigm, where progress in multi-modality, trustworthiness, human-machine synergy, and ethical reasoning may allow intelligence to be woven into the pipeline of clinical practice and fulfill its translational promise.
Network pharmacology (NP) explores pharmacological mechanisms through biological networks. Multi-omics data enable multi-layer network construction under diverse conditions, requiring integration into NP analyses. We developed POINT, a novel NP platform enhanced by multi-omics biological networks, advanced algorithms, and knowledge graphs (KGs) featuring network-based and KG-based analytical functions. In the network-based analysis, users can perform NP studies flexibly using 1,158 multi-omics biological networks encompassing proteins, transcription factors, and non-coding RNAs across diverse cell line-, tissue- and disease-specific conditions. Network-based analysis-including random walk with restart (RWR), GSEA, and diffusion profile (DP) similarity algorithms-supports tasks such as target prediction, functional enrichment, and drug screening. We merged networks from experimental sources to generate a pre-integrated multi-layer human network for evaluation. RWR demonstrated superior performance with a 33.1 second-best algorithm, PageRank, in identifying known targets across 2,002 drugs. Additionally, multi-layer networks significantly improve the ability to identify FDA-approved drug-disease pairs compared to the single-layer network. For KG-based analysis, we compiled three high-quality KGs to construct POINT KG, which cross-references over 90 illustrated the platform's capabilities through two case studies. POINT bridges the gap between multi-omics networks and drug discovery; it is freely accessible at http://point.gene.ac/.
The intratumoral microbiome is an emerging hallmark of cancer, yet its multi-kingdom host-microbiome ecosystem in colorectal cancer (CRC) remains poorly characterized. Here, we conducted an integrated analysis using deep shotgun metagenomics and proteomics on 185 tissue samples, including adenoma (A), paired tumor (T), and para-tumor (P). We identified 4057 bacterial, 61 fungal, 108 archaeal, and 374 viral species in tissues and revealed distinct intratumor microbiota dysbiosis, indicating a CRC-specific multi-kingdom microbial ecosystem. Proteomic profiling uncovered four CRC subtypes (C1-C4), each with unique clinical prognoses and molecular signatures. We further discovered that host-microbiome interactions are dynamically reorganized during carcinogenesis, where different microbial taxa converge on common host pathways through distinct proteins. Leveraging this interplay, we identified 14 multi-kingdom microbial and 8 protein markers that strongly distinguished A from T samples (area under the receiver operating characteristic curve (AUROC) = 0.962), with external validation in two independent datasets (AUROC = 0.920 and 0.735). Moreover, we constructed an early- versus advanced-stage classifier using 8 microbial and 4 protein markers, which demonstrated high diagnostic accuracy (AUROC = 0.926) and was validated externally (AUROC = 0.659-0.744). Functional validation in patient-derived organoids and murine allograft models confirmed that enterotoxigenic Bacteroides fragilis and Fusobacterium nucleatum promoted tumor growth by activating Wnt/β-catenin and NF-κB signaling pathways, corroborating the functional potential of these biomarkers. Together, these findings reveal dynamic host-microbiome interactions at the protein level, tracing the transition from adenoma to carcinoma and offering potential diagnostic and therapeutic targets for CRC.
Clinical trials and meta-analyses are considered high-level medical evidence with solid credibility. However, such clinical evidence for traditional Chinese medicine (TCM) is scattered, requiring a unified entrance to navigate all available evaluations on TCM therapies under modern standards. Besides, novel experimental evidence has continuously accumulated for TCM since the publication of HERB 1.0. Therefore, we updated the HERB database to integrate four types of evidence for TCM: (i) we curated 8558 clinical trials and 8032 meta-analyses information for TCM and extracted clear clinical conclusions for 1941 clinical trials and 593 meta-analyses with companion supporting papers. (ii) we updated experimental evidence for TCM, increased the number of high-throughput experiments to 2231, and curated references to 6 644. We newly added high-throughput experiments for 376 diseases and evaluated all pairwise similarities among TCM herbs/ingredients/formulae, modern drugs and diseases. (iii) we provide an automatic analyzing interface for users to upload their gene expression profiles and map them to our curated datasets. (iv) we built knowledge graph representations of HERB entities and relationships to retrieve TCM knowledge better. In summary, HERB 2.0 represents rich data type, content, utilization, and visualization improvements to support TCM research and guide modern drug discovery. It is accessible through http://herb.ac.cn/v2 or http://47.92.70.12.
Glioma, a malignant intracranial tumor with high invasiveness and heterogeneity, significantly impacts patient survival. This study integrates multi-omics data to improve prognostic prediction and identify therapeutic targets. Using single-cell data from glioblastoma (GBM) and low-grade glioma (LGG) samples, we identified 55 distinct cell states via the EcoTyper framework, validated for stability and prognostic impact in an independent cohort. We constructed multi-omics datasets of 620 samples, integrating transcriptomic, copy number variation (CNV), somatic mutation (MUT), Microbe (MIC), EcoTyper result data. A scRNA-seq enhanced Self-Normalizing Network-based glioma prognosis model achieved a C-index of 0.822 (training) and 0.817 (test), with AUC values of 0.867, 0.876, and 0.844 at 1, 3, and 5 years in the training set, and 0.820, 0.947, and 0.936 in the test set. Gradient attribution analysis enhanced the interpretability of the model and identified key molecular markers. The classification into high- and low-risk groups was validated as an independent prognostic factor. HDAC inhibitors are proposed as potential treatments. This study demonstrates the potential of integrating scRNA-seq and multi-omics data for robust glioma prognosis and clinical decision-making support.
Single-cell transcriptome sequencing technology has been applied to decode the cell types and functional states of immune cells, revealing their tissue-specific gene expression patterns and functions in cancer immunity. Comprehensive assessments of immune cells within and across tissues will provide us with a deeper understanding of the tumor immune system in general. Here, we present Cross-tissue Immune cell type or state Enrichment analysis of gene lists for Cancer (CIEC), the first web-based application that integrates database and enrichment analysis to estimate the cross-tissue immune cell types or states. CIEC version 1.0 consists of 480 samples covering primary tumor, adjacent normal tissue, lymph node, metastasis tissue, and peripheral blood from 323 cancer patients. By applying integrative analysis, we constructed an immune cell type/state map for each context, and adopted our previously developed Kyoto Encyclopedia of Genes and Genomes (KEGG) Orthology Based Annotation System (KOBAS) algorithm to estimate the enrichment for context-specific immune cell types/states. In addition, CIEC also provides an easy-to-use online interface for users to comprehensively analyze the immune cell characteristics mapped across multiple tissues, including expression map, correlation, similar gene detection, signature score, and expression comparison. We believe that CIEC will be a valuable resource for exploring the intrinsic characteristics of immune cells in cancer patients and for potentially guiding novel cancer-immune biomarker development and immunotherapy strategies. CIEC is freely accessible at http://ciec.gene.ac/.
Large language models can lighten the workload of clinicians and patients, yet their responses often include fabricated evidence, outdated knowledge, and insufficient medical specificity. We introduce a general retrieval-augmented question-answering framework that continuously gathers up-to-date, high-quality medical knowledge and generates evidence-traceable responses. Here we show that this approach significantly improves the evidence validity, medical expertise, and timeliness of large language model outputs, thereby enhancing their overall quality and credibility. Evaluation against 15,530 objective questions, together with two physician-curated clinical test sets covering evidence-based medical practice and medical order explanation, confirms the improvements. In blinded trials, resident physicians indicate meaningful assistance in 87.00% of evidence-based medical scenarios, and lay users find it helpful in 90.09% of medical order explanations. These findings demonstrate a practical route to trustworthy, general-purpose language assistants for clinical applications.
Electronic Medical. Records (EMRs), extensively recognized as a significant repository of clinical experience and medical knowledge, often jeopardize their utility by being typically penned in free text, resulting in under-structured information. This lack of structural organization poses a major hurdle in fully capitalizing on the medical data embedded in these EMRs. Medical information extraction (IE) can fill this gap by converting the clinic text to structured data. In this paper, we propose Marker LAttice Transformer (MAT), a strong framework for medical IE. This framework is composed of three separate models, each designed for a specific task: medical entity recognition (MER), medical relation extraction (MRE), and medical attribute extraction (MAE). All the models are deeply based on markers embedded in the input text, with which the models compute representations from bottom to top layers. This allows the representations to encode deep semantic information, leading to better outputs. In addition, we enhance the models by lattice-style incorporation of medical dictionary information, further pre-training on large-scale EMRs, and auxiliary inputs of medical departments and EMR sections. We evaluated MAT using the HwaMei-500 dataset, the most comprehensive and current collection of Chinese electronic medical records. The framework demonstrated exceptional performance, achieving F-1 scores of 93.0%, 71.5%, and 88.9% for MER, MRE, and MAE tasks respectively. These scores outperform the baseline by large margins and also surpass the results of recent state-of-the-art relation extraction models. As a medical IE framework, we provide a practical example for other information extraction efforts, particularly those focusing on Chinese electronic medical record data.
The"one drug-multiple targets"paradigm has revolutionized therapeutic development for complex diseases by addressing the limitations of single-target approaches[1].However,elucidating multi-target synergism remains a major challenge.Network phar-macology(NP)enables polypharmacological investigations through biological networks[2].Recent advances have highlighted the utility of NP across diverse diseases through protein-protein interaction(PPI)networks.Although PPI networks are widely used in NP,the increasing recognition of intracellular regulatory ele-ments has expanded potential drug targets beyond proteins to include biomolecules such as transcription factors(TFs)and non-coding RNAs(ncRNAs)[3].These findings demonstrate the limita-tions of relying solely on PPI networks for target discovery and highlight the need for multi-layer networks that integrate biomo-lecules across diverse regulatory levels.
BACKGROUND:Structural variations (SVs) are common genetic alterations in the human genome. However, the profile and clinical relevance of SVs in patients with hereditary breast and ovarian cancer (HBOC) syndrome (germline BRCA1/2 mutations) remains to be fully elucidated. METHODS:Twenty HBOC-related cancer samples (5 breast and 15 ovarian cancers) were studied by optical genome mapping (OGM) and next-generation sequencing (NGS) assays. RESULTS:The SV landscape in the 5 HBOC-related breast cancer samples was comprehensively investigated to determine the impact of intratumor SV heterogeneity on clinicopathological features and on the pattern of genetic alteration. SVs and copy number variations (CNVs) were common genetic events in HBOC-related breast cancer, with a median of 212 SVs and 107 CNVs per sample. The most frequently detected type of SV was insertion, followed by deletion. The 5 HBOC-related breast cancer samples were divided into SVhigh and SVlow groups according to the intratumor heterogeneity of SVs. SVhigh tumors were associated with higher Ki-67 expression, higher homologous recombination deficiency (HRD) scores, more mutated genes, and altered signaling pathways. Moreover, 60% of the HBOC-related breast cancer samples displayed chromothripsis, and 8 novel gene fusion events were identified by OGM and validated by transcriptome data. CONCLUSIONS:These findings suggest that OGM is a promising tool for the detection of SVs and CNVs in HBOC-related breast cancer. Furthermore, OGM can efficiently characterize chromothripsis events and novel gene fusions. SVhigh HBOC-related breast cancers were associated with unfavorable clinicopathological features. SVs may therefore have predictive and therapeutic significance for HBOC-related breast cancers in the clinic.
Deciphering universal gene regulatory mechanisms in diverse organisms holds great potential for advancing our knowledge of fundamental life processes and facilitating clinical applications. However, the traditional research paradigm primarily focuses on individual model organisms and does not integrate various cell types across species. Recent breakthroughs in single-cell sequencing and deep learning techniques present an unprecedented opportunity to address this challenge. In this study, we built an extensive dataset of over 120 million human and mouse single-cell transcriptomes. After data preprocessing, we obtained 101,768,420 single-cell transcriptomes and developed a knowledge-informed cross-species foundation model, named GeneCompass. During pre-training, GeneCompass effectively integrated four types of prior biological knowledge to enhance our understanding of gene regulatory mechanisms in a self-supervised manner. By fine-tuning for multiple downstream tasks, GeneCompass outperformed state-of-the-art models in diverse applications for a single species and unlocked new realms of cross-species biological investigations. We also employed GeneCompass to search for key factors associated with cell fate transition and showed that the predicted candidate genes could successfully induce the differentiation of human embryonic stem cells into the gonadal fate. Overall, GeneCompass demonstrates the advantages of using artificial intelligence technology to decipher universal gene regulatory mechanisms and shows tremendous potential for accelerating the discovery of critical cell fate regulators and candidate drug targets.
Digital pathology slides can serve medical practitioners or aid in computer-assisted diagnosis and treatment. Collection personnel typically employ hyperspectral microscopes to scan pathology slides into Whole Slide Images (WSI) with pixel counts reaching the million level. However, this process incurs significant acquisition time and data storage costs. Utilizing super-resolution imaging techniques to enhance low-resolution pathological images enables downstream analysis of pathological tissue slice data under low-resource and cost-effective medical conditions. Nevertheless, existing super-resolution methods cannot integrate attention information containing variable receptive fields and effective means to handle distortions and artifacts in the output data. This leads to differences between super-resolution images and authentic images depicting cell contours and tissue morphology. We propose a method named MiHATP: A Multi(Mi)-Hybrid(H) Attention(A) Network Based on Transformation(T) Pool(P) Contrastive Learning to address these challenges. By constructing contrastive losses through reversible image transformation and irreversible low-quality image transformation, MiHATP effectively reduces distortion in super-resolution pathological images. Within MiHATP, we also design a Multi-Hybrid Attention structure to ensure strong modeling capability for long-distance and short-distance information. This ensures that the super-resolution network can obtain richer image information. The experimental results show that MiHATP achieves the best performance in both the super-image reconstruction and downstream cell segmentation and phenotypes tasks. The implementation code will be available at https://github.com/rabberk/MiHATP.git.
A complex and vast biological network regulates all biological functions in the human body in a sophisticated manner, and abnormalities in this network can lead to disease and even cancer. The construction of a high-quality human molecular interaction network is possible with the development of experimental techniques that facilitate the interpretation of the mechanisms of drug treatment for cancer. We collected 11 molecular interaction databases based on experimental sources and constructed a human protein-protein interaction (PPI) network and a human transcriptional regulatory network (HTRN). A random walk-based graph embedding method was used to calculate the diffusion profiles of drugs and cancers, and a pipeline was constructed by using five similarity comparison metrics combined with a rank aggregation algorithm, which can be implemented for drug screening and biomarker gene prediction. Taking NSCLC as an example, curcumin was identified as a potentially promising anticancer drug from 5450 natural small molecules, and combined with differentially expressed genes, survival analysis, and topological ranking, we obtained BIRC5 (survivin), which is both a biomarker for NSCLC and a key target for curcumin. Finally, the binding mode of curcumin and survivin was explored using molecular docking. This work has a guiding significance for antitumor drug screening and the identification of tumor markers.
[This corrects the article DOI: 10.1016/j.scr.2023.103115.].
RNA-binding proteins (RBPs) are key post-transcriptional regulators, and the malfunctions of RBP-RNA binding lead to diverse human diseases. However, prediction of RBP binding sites is largely based on RNA sequence features, whereas in vivo RNA structural features based on high-throughput sequencing are rarely incorporated. Here, we designed a deep bimodal information fusion network called DeepFusion for unraveling protein-RNA interactions by incorporating structural features derived from DMS-seq data. DeepFusion integrates two sub-models to extract local motif-like information and long-term context information. We show that DeepFusion performs best compared with other cutting-edge methods with only sequence inputs on two datasets. DeepFusion’s performance is further improved with bimodal input after adding in vivo DMS-seq structural features. Furthermore, DeepFusion can be used for analyzing RNA degradation, demonstrating significantly different RBP-binding scores in genes with slow degradation rates versus those with rapid degradation rates. DeepFusion thus provides enhanced abilities for further analysis of functional RNAs. DeepFusion’s code and data are available at http://bioinfo.org/deepfusion/.