Motivation Multiomics data analysis is essential for scientific discovery in precision medicine. However, translating analysis results of omics data analysis into novel scientific hypotheses remains a significant challenge. Human experts must manually review analysis results and generate new hypotheses based on extensive and interconnected biomedical prior knowledge, which is subjective and not scalable. While large language models can accelerate the discovery, their reasoning improves when grounded in structured, auditable, and comprehensive biomedical prior knowledge. However, biomedical knowledge is scattered across heterogeneous databases that use diverse and inconsistent nomenclature systems, making it difficult to integrate resources into a unified format for scalable analysis. This fragmentation limits the ability of artificial intelligence systems to fully leverage biomedical data for scientific discovery.Results We developed BioMedGraphica, a novel all-in-one platform that harmonizes fragmented biomedical resources by integrating 11 entity types and 30 relation types from 43 databases into a unified textual prior knowledge graph containing 2 306 921 entities and 27 232 091 relations. In addition, we present a novel textual-numeric graph (TNG) data structure concept, where textual information captures prior biological knowledge (e.g. transcription start sites, functions, mechanisms), numeric values represent quantitative biomedical features, and the integrated relations can help uncover mechanisms. By bridging prior knowledge with user-specific data, TNG is a novel and ideal data structure for developing novel graph analysis models.Availability and implementation The code is available at: https://github.com/FuhaiLiAiLab/BioMedGraphica and BioMedGraphica knowledge graph database can be downloaded from huggingface dataset: https://huggingface.co/datasets/FuhaiLiAiLab/BioMedGraphica
Abstract In single-cell (sc)-based scientific discovery, text-formatted biomedical prior knowledge and signaling graphs are essential for annotating and interpreting numeric sc-omics data and for generating novel testable hypotheses. A major limitation of existing single-cell large language models (scLLMs) is that they rely on numeric expression data with gene names as the only textual signal, while comprehensive biomedical priors — cellular localization, gene function, disease associations, and signaling interaction patterns — remain absent from the model input. We introduce CellTosg2Sequence , a textual-prior- and signaling-graph-augmented cell-omics-sentence language model. A lightweight heterogeneous graph encoder maps a curated 62,507-node biomedical knowledge graph (KG) into compact virtual tokens that are prepended to each cell sentence, allowing the language model to condition on biological structure with minimal sequence-length overhead. We train CellTosg2Sequence with a three-stage objective: Stage I anchors the KG channel under autoregressive language-model pretraining, leveraging Qwen2.5-32B’s own language reasoning for rapid KG alignment; Stage II aligns labels via supervised fine-tuning with KG-anchored InfoNCE; Stage III applies Group Relative Policy Optimization (GRPO) with an ontology-hierarchy reward, enabling free-generation cell-type prediction that generalizes beyond the closed training vocabulary. Across multiple benchmarks and ablation experiments, CellTosg2Sequence outperforms strong baselines. All results are achieved with lightweight LoRA training and a single unified checkpoint. Data ethics This work uses publicly available single-cell datasets from the Human Cell Atlas ( https://www.humancellatlas.org ) and the Tahoe-100M consortium. All HCA constituent studies were collected under appropriate donor consent and institutional oversight as described in their original publications; we perform computational re-analysis only and introduce no new human subjects data. HCA data access follows the HCA Data Portal terms of use. No new patient data are collected in this study; no additional IRB approval is required for this secondary computational analysis.
In precision medicine, quantitative multi-omic features, topological context, and textual biological knowledge play vital roles in identifying disease-critical signaling pathways and targets, guiding the discovery of novel therapeutics and effective treatment strategies. Existing pipelines capture only one or two of these—numerical omics ignore topological context, text-centric LLMs lack quantitative grounded reasoning, and graph-only models underuse rich node semantics and the generalization power of LLMs—thereby limiting mechanistic interpretability. Although Process Reward Models (PRMs) aim to guide reasoning in LLMs, they remain limited by coarse step definitions, unreliable intermediate evaluation, and vulnerability to reward hacking with added computational cost. These gaps motivate jointly integrating quantitative multi-omic signals, topological structure with node annotations, and literature-scale text via LLMs, using subgraph reasoning as the principle bridge linking numeric evidence, topological knowledge and language context. To resolve this challenge, we propose GALAX (Graph Augmented LAnguage model with eXplainability), an innovative framework that integrates pretrained Graph Neural Networks (GNNs) into Large Language Models (LLMs) via reinforcement learning guided by a Graph Process Reward Model (GPRM), which generates disease-relevant subgraphs in a step-wise manner initiated by an LLM and iteratively evaluated by a pretrained GNN and schema-based rule check, enabling process-level supervision without explicit labels. As an application, we also introduced Target-QA, a benchmark combining CRISPR-identified targets, multi-omic profiles, and biomedical graph knowledge across diverse cancer cell lines, which enables GNN pretraining for supervising step-wise graph construction and supports long-context reasoning over text-numeric graphs (TNGs), providing a scalable and biologically grounded framework for explainable, reinforcement-guided subgraph reasoning toward reliable and interpretable target and pathway discovery in precision medicine.
In biomedical scientific discovery, synthesizing prior knowledge from the literature is an essential component of interpreting numerical omics data analyses for disease target identification and drug discovery. Large language models (LLMs) alone can rapidly retrieve disease mechanisms from biomedical text, but text-only outputs are general and unreliable for target and drug prioritization without cohort-specific quantitative evidence. Herein, we propose a provenance-aware Text-to-Target framework that couples schema-constrained multi-model LLM retrieval with numeric omics data analysis. The key design is a modality-aware fusion step: candidates are partitioned into overlap-supported anchors, retrieval-only hidden hubs, and network-emergent novelty nodes, then propagated into staged hypothesis and strategy generation under topology constraints. We evaluate the model in Alzheimer's disease (AD) and pancreatic ductal adenocarcinoma (PDAC). In PDAC, the workflow produced a balanced 75-gene candidate universe and a 23-strategy portfolio, with significant DepMap support at both target level and strategy level. In AD, stricter candidate controls yielded a compact 34-gene universe and 14 strategies; under an expanded CRISPRbrain registry, both target-level axes were significant , with strong strategy-level enrichment. Across both diseases, final strategies preserved full provenance closure to the candidate pool, enabling end-to-end auditability from retrieval artifacts to validation outputs. These results support a transferable discovery architecture in which omics evidence constrains biological activity, LLM retrieval expands mechanistic search space, and network-aware fusion preserves interpretability. The framework provides a reproducible basis for dual-disease target prioritization and motivates continuous literature-mechanism concordance with agentic evidence-refresh loops.
With the rapid growth of large-scale single-cell omic datasets, omic foundation models (FMs) have emerged as powerful tools for advancing research in life sciences and precision medicine. However, most existing omic FMs rely primarily on numerical transcriptomic data by sorting genes as sequences, while lacking explicit integration of biomedical prior knowledge and signaling interactions that are critical for scientific discovery. Here, we introduce the Text-Omic Signaling Graph (TOSG), a novel data structure that unifies human-interpretable biomedical textual knowledge, quantitative omic data, and signaling network information. Using this framework, we construct OmniCellTOSG, a large-scale resource comprising approximately half million meta-cell TOSGs derived from around 80 million single-cell and single-nucleus RNA-seq profiles across organs and diseases. We further develop CellTOSG-FM, a multimodal graph language FM, to jointly analyze textual, omic and signaling network context. Across diverse downstream tasks, CellTOSG-FM outperforms existing omic FMs, and provides interpretable insights into disease-associated targets and signaling pathways.
Background: Alternative promoter usage contributes to isoform diversity and gene regulation in mammals but remains difficult to study at scale. Cap Analysis of Gene Expression precisely maps transcription start sites, but its cost limits large-scale application. Alternatively, ProActiv, Salmon, and DEXSeq can be utilized with widely available RNA sequencing (RNA-seq) data to infer promoter activity. However, there is currently no framework available to automate the generation of reproducible results for these methods. Results: SnakeAltPromoter, a scalable end-to-end Snakemake workflow, has been developed to automate alternative promoter analysis from raw RNA-seq data. The workflow performs quality control, alignment, and promoter quantification using 3 complementary RNA-seq analysis methods (junction-based, transcript-based, and first-exon-based), followed by promoter classification and differential activity or usage analysis. SnakeAltPromoter supports both command-line and graphical user interface usage and utilizes standardized modules to enhance reproducibility. A built-in benchmarking module compares promoter activities inferred from RNA-seq data to matched Cap Analysis of Gene Expression data to evaluate quantification performance. Our analyses revealed robust and complementary performance profiles among the evaluated methods across tissues and cell types, highlighting the value of a unified framework for promoter analysis. Conclusions: To our knowledge, SnakeAltPromoter is the first unified and reproducible framework that combines scalable execution and guided method selection for RNA-seq-based promoter analysis. By standardizing and integrating existing RNA-seq-based promoter analysis tools, it provides researchers with a robust and accessible platform to investigate promoter-level regulation and alternative promoter activity and will enhance the research value of large, public RNA-seq repositories. Code is freely available at https://github.com/YidanSunResearchLab/SnakeAltPromoter.git.
In recent years, the rapid advancement of high-throughput technologies has led to the generation of vast and complex multi-omics datasets that are valuable for characterizing and understanding complex cell signaling network systems. On the other hand, large language models (LLMs), domain-specific foundation models (FMs) and AI agents, have achieved significant breakthroughs and have been revolutionizing scientific research. The convergence of these two trends is catalyzing a new era for biomedical research to augment and speed up scientific discovery and the development of precision medicine. In this study, we examine the large-scale omics datasets, emerging applications, and challenges at the intersection of massive omic datasets and related AI models and agents, highlighting how their integration is reshaping the landscape of biomedical research and precision medicine.
Medical records and omics data are rapidly becoming standard in healthcare settings, which characterize the whole-person from dysfunctional molecules to phenotypes, and thus offer potential for precise disease diagnosis and target discovery. Whereas, it remains an open problem to systematically integrate and interprete medical record and omics data of individual patients. In this study, for the first time, we propose a novel graph AI model framework, Graph in Graph (GiG), to integrate and interpret the whole-person medical and omic datasets. Specifically, the medical record data is modeled using a person-phenotype graph, followed by omics signaling graphs of invidival patients, which enables the integration of information learned from omic-signaling graph and medical phenotype features to characterize individual patients and to prioritize important omic biomarkers and phenotypes. As an exploratory study, we applied and evaluated the GiG model to study the type 2 diabetes (T2D) and pre-T2D vs healthy using the Long Life Family Study (LLFS) cohort, which enrolls families with exceptional longevity to uncover biological mechanisms of healthy aging with medical and omics data. The evaluation results showed that GiG not only achieve a high prediction but also can interpret the prediction by ranking the essential clinical and omic biomarkers. The GiG framework can be applied to other studies by effectively integrating and interpreting medical and omic datasets for disease diagnosis and pathogenesis discovery.
The convergence of large language models (LLMs), AIagents, and large-scale omic datasets-such as single-cell omics, marks the arrival of a critical inflection point in biomedical research, via autonomous data mining and novel hypothesis generation. However, there is no specifically designed agentic AI model that can systematically integrate large-scale single-cell (sc) RNAseq (covering diverse diseases and cell types), omic data analytic tools, accumulated biomedical knowledge, and literature search to facilitate autonomous scientific discovery in precision medicine. In this study, we develop a novel agentic AI, OmniCellAgent, to empower non-computational-expert users-such as patients and family members, clinicians, and wet-lab researchers-to conduct scRNA-seq data-driven biomedical research like experts, uncovering molecular disease mechanisms and identifying effective precision therapies. The code of omniCellAgent is publicly accessible at: https://fuhailiailab.github.io/.
The integration of multi-omic data is pivotal for understanding complex diseases, but its high dimensionality and noise present significant challenges. Graph Neural Networks (GNNs) offer a robust framework for analyzing large-scale signaling pathways and protein-protein interaction networks, yet they face limitations in expressivity when capturing intricate biological relationships. To address this, we propose Graph Sequence Language Model (GraphSeqLM), a framework that enhances GNNs with biological sequence embeddings generated by Large Language Models (LLMs). These embeddings encode structural and biological properties of DNA, RNA, and proteins, augmenting GNNs with enriched features for analyzing sample-specific multi-omic data. By integrating topological, sequence-derived, and biological information, GraphSeqLM demonstrates superior predictive accuracy and outperforms existing methods, paving the way for more effective multi-omic data integration in precision medicine.
In this editorial, we summarize the 2023 International Conference on Intelligent Biology and Medicine (ICIBM 2023) conference which was held on July 16-19, 2023 in Tampa, Florida, USA. We then briefly describe the nine research articles included in this special issue. ICIBM 2023 scientific program included four tutorials and workshops, four keynote lectures, four eminent scholars’ presentations, 11 concurrent scientific sessions, and a poster session. We had total of 88 scientific oral presentations, including 62 regular oral presentations and 26 flash talks, as well as 46 poster presentations. The topics of these presentations covered artificial intelligence (AI), data sciences, bioinformatics, computational biology, genomics, biomedical informatics, among others. This special issue included nine peer reviewed manuscripts selected from those submitted to our conference. These articles cover a range of topics on the methods and applications of ML and AI models in the biomedical domain.
Multi-omic dataset can better characterize complex cellular signaling pathways from multiple views compared to individual omic data. However, integrative multi-omic data analysis to rank key disease biomarkers and infer core signaling pathways remains an open problem. In this study, we developed a novel graph AI model, mosGraphFlow, for analyzing multi-omic signaling graphs (mosGraphs), 2) analyzed multi-omic mosGraph datasets of Alzheimers’ Disease (AD), and 3) developed a visualization tool to facilitate the visualization of identified disease associated signaling biomarkers and network. The comparison results show that the proposed model not only achieves the best classification accuracy but also identifies important AD disease biomarkers and signaling interactions. In the visualization, the signaling sources are highlighted at specific omic levels to facilitate the understanding of disease pathogenesis. The proposed model can also be applied and expanded for other multi-omic data-driven studies. The code of the model is publicly accessible via GitHub: https://github.com/FuhaiLiAiLab/mosGraphFlow .
In real-world scientific discovery, human beings always make use of the accumulated prior knowledge with imagination pick select one or a few most promising hypotheses from large and noisy data analysis results. In this study, we introduce a new type of graph structure, the text-numeric graph (TNG), which is defined as graph entities and associations have both text-attributed information and numeric information. The TNG is an ideal data structure model for novel scientific discovery via graph reasoning because it integrates human-understandable textual annotations or prior knowledge, with numeric values that represent the observed or activation levels of graph entities or associations in different samples. Together both the textual information and numeric values determine the importance of graph entities and associations in graph reasoning for novel scientific knowledge discovery. We further propose integrating large language models (LLMs) and graph neural networks (GNNs) to analyze the TNGs for graph understanding and reasoning. To demonstrate the utility, we generated the text-omic(numeric) signaling graphs (TOSG), as one type of TNGs, in which all graphs have the same entities, associations and annotations, but have sample-specific entity numeric (omic) values using single cell RNAseq (scRNAseq) datasets of different diseases. We proposed joint LLM-GNN models for key entity mining and signaling pathway mining on the TOSGs. The evaluation results showed the LLM-GNN and TNGs models significantly improve classification accuracy and network inference. In conclusion, the TNGs and joint LLM-GNN models are important approaches for scientific discovery.
Multi-omic data-driven studies are at the forefront of precision medicine by characterizing complex disease signaling systems across multiple views and levels. The integration and interpretation of multi-omic data are critical for identifying disease targets and deciphering disease signaling pathways. However, it remains an open problem due to the complex signaling interactions among many proteins. Herein, we propose a multi-scale multi-hop multi-omic network flow model, M3NetFlow, to facilitate both hypothesis-guided and generic multi-omic data analysis tasks. We evaluated M3NetFlow using two independent case studies: (1) uncovering mechanisms of synergy of drug combinations (hypothesis/anchor-target guided multi-omic analysis) and (2) identifying biomarkers of Alzheimer's disease (generic multi-omic analysis). The evaluation and comparison results showed that M3NetFlow achieved the best prediction accuracy and identified a set of drug combination synergy- and disease-associated targets. The model can be directly applied to other multi-omic data-driven studies.
Background:Early alterations in cerebrospinal fluid (CSF) may play a critical role in the progression of acute ischemic stroke (AIS). This study aimed to develop a CSF-based clinical-radiomics model for predicting the functional outcomes in AIS patients following intravenous thrombolysis (IVT). Methods:This study included 308 AIS patients within 6 h of onset, divided into a training set (n=246) and a hold-out test set (n=62). Functional outcome was assessed at discharge using the modified Rankin Scale (mRS), dichotomized into good (mRS ≤2) and poor (mRS >2) outcomes. CSF regions were automatically segmented on non-contrast computed tomography images, and radiomics features were extracted. After feature selection was sequentially performed using minimum redundancy maximum relevance algorithm followed by least absolute shrinkage and selection operator, the radiomics signature model was constructed. Clinical features were selected through univariate and multivariate logistic regressions and subsequently integrated with radiomics features to develop combined models. Three machine learning classifiers, including Nu Support Vector Classification (NuSVC), logistic regression, and random forest, were trained and tested for prognostic prediction. Model performance was evaluated using receiver operating characteristic curve analysis, decision curve analysis (DCA), and calibration curves, with internal validation via five-fold cross-validation. Results:Among the 308 patients included, 155 patients had a good outcome and 153 had a poor outcome. A total of 1,874 radiomics features were extracted, among which 21 were selected for the radiomics model. Three clinical features were also identified for the clinical model. The combined clinical-radiomics model significantly outperformed single-modality models across all classifiers. The NuSVC-based combined model achieved the best performance, with an area under the curve (AUC) of 0.870 [95% confidence interval (CI): 0.850-0.889] in cross-validation and 0.893 (95% CI: 0.817-0.968) in the test cohort. DCA further confirmed the superior clinical utility of the combined model. Conclusions:This study demonstrates the value of CSF radiomics signatures in predicting functional outcomes in AIS patients after IVT. The CSF-based combined model exhibited strong prognostic performance and clinical utility, highlighting its potential for supporting treatment decision-making.
Named Entity Recognition (NER) is a fundamental task in Natural Language Processing (NLP) that plays a crucial role in information extraction, question answering, and knowledge-based systems. Traditional deep learning-based NER models often struggle with domain-specific generalization and suffer from data sparsity issues. In this work, we introduce Knowledge Graph distilled for Named Entity Recognition (KoGNER), a novel approach that integrates Knowledge Graph (KG) distillation into NER models to enhance entity recognition performance. Our framework leverages structured knowledge representations from KGs to enrich contextual embeddings, thereby improving entity classification and reducing ambiguity in entity detection. KoGNER employs a two-step process: (1) Knowledge Distillation, where external knowledge sources are distilled into a lightweight representation for seamless integration with NER models, and (2) Entity-Aware Augmentation, which integrates contextual embeddings that have been enriched with knowledge graph information directly into GNN, thereby improving the model's ability to understand and represent entity relationships. Experimental results on benchmark datasets demonstrate that KoGNER achieves state-of-the-art performance, outperforming finetuned NER models and LLMs by a significant margin. These findings suggest that leveraging knowledge graphs as auxiliary information can significantly improve NER accuracy, making KoGNER a promising direction for future research in knowledge-aware NLP.
Complex signaling pathways are believed to be responsible for drug resistance. Drug combinations perturbing multiple signaling targets have the potential to reduce drug resistance. The large-scale multi-omic datasets and experimental drug combination synergistic score data are valuable resources to study mechanisms of synergy (MoS) to guide the development of precision drug combinations. However, signaling patterns of MoS are complex and remain unclear, and thus it is challenging to identify synergistic drug combinations in clinical. Herein, we proposed a novel integrative and interpretable graph AI model, DeepSignalingFlow, to uncover the MoS by integrating and mining multi-omic data. The major innovation is that we uncover MoS by modeling the signaling flow from multi-omic features of essential disease proteins to the drug targets, which has not been introduced by the existing models. The model performance was assessed utilizing four distinct drug combination synergy evaluation datasets, i.e., NCI ALMANAC, O'Neil, DrugComb, and DrugCombDB. The comparison results showed that the proposed model outperformed existing graph AI models in terms of synergy score prediction, and can interpret MoS using the core signaling flows. The code is publicly accessible via Github: https://github.com/FuhaiLiAiLab/DeepSignalingFlow.
The expressivity of Graph Neural Networks (GNNs) has been studied broadly in recent years to reveal the design principles for more powerful GNNs. Graph canonization is known as a typical approach to distinguish non-isomorphic graphs, yet rarely adopted when developing expressive GNNs. This paper proposes to maximize the expressivity of GNNs by graph canonization, then the power of such GNNs is studies from the perspective of model stability. A stable GNN will map similar graphs to close graph representations in the vectorial space, and the stability of GNNs is critical to generalize their performance to unseen graphs. We theoretically reveal the trade-off of expressivity and stability in graph-canonization-enhanced GNNs. Then we introduce a notion of universal graph canonization as the general solution to address the trade-off and characterize a widely applicable sufficient condition to solve the universal graph canonization. A comprehensive set of experiments demonstrates the effectiveness of the proposed method. In many popular graph benchmark datasets, graph canonization successfully enhances GNNs and provides highly competitive performance, indicating the capability and great potential of proposed method in general graph representation learning. In graph datasets where the sufficient condition holds, GNNs enhanced by universal graph canonization consistently outperform GNN baselines and successfully improve the SOTA performance up to $31$%, providing the optimal solution to numerous challenging real-world graph analytical tasks like gene network representation learning in bioinformatics.
Adopting hybrids with high drought tolerance is an effective strategy to sustain maize (Zea mays L.) production under water shortage. Optimizing plant density is one of the important management practices for improving maize yield. The objective of this study was to clarify the grain yield and water use responses to plant density in drought-tolerant (DT) maize hybrid under irrigated and rainfed conditions, and to compare soil water depletion between maize hybrids differing in drought tolerance. A two-year field trial was conducted to investigate grain yield (GY), evapotranspiration (ET), and water-use efficiency (WUE) in one DT hybrid (ZD958) and one drought-sensitive (DS; ZY309) hybrid under two water treatments (irrigation and rainfed) and three planting densities (6, 7.5, 9 plants m−2). The differences in soil water depletion for DT and DS hybrids under irrigated and rainfed conditions were also assessed. On average, yield reduction due to water stress was greater for ZY309 (51%) than ZD958 (41%). Under rainfed, GY, ET and WUE were 21.4–29.7%, 7.3–10.5% and 9.9–20.3% greater in ZD958 than ZY309 across three planting densities, respectively. When planting density increased from 6 to 9 plants m−2, ZD958 and ZY309’s ET did not change, but their GY and WUE first increased and then decreased and ultimately reached their maximum values at 7.5 plants m−2. Under irrigation, increases in densities from 6 to 9 plants m−2 led to a significant increase in GY and WUE for ZD958, but for ZY309, GY and WUE were not significantly impacted by increasing densities. Yield averaged across seasons did not differ between the two hybrids at 6 and 7.5 plants m−2, and at 9 plants m−2, ZD958 had yield advantage (10.2% across seasons) over ZY309. Soil water extraction under medium density (7.5 plants m−2) and rainfed was greater for ZD958 than ZY309, which might be associated with its greater tolerance to drought stress. The findings imply that the DT hybrid showed greater tolerance to high plant density than the DS hybrid and had a higher optimal density of 9 plants m−2 under irrigation. Under rainfed, the tolerance to high plant density and also optimum plant density were similar in DT and DS hybrids, while DT hybrid had a stronger tolerance to drought in comparison to DS hybrid, which could explain DT hybrid’s greater GY and WUE. The results of this study indicate that, compared to DS hybrid, DT hybrid showed higher and more stable grain yields across rainfed and irrigated environments when grown at optimal planting densities. Thus, high and sustainable maize yields could be achieved by adopting DT hybrids and optimizing planting density for different growing environments, which would help to achieve food security in the study region.
Recently, large-scale scRNA-seq datasets have been generated to understand the complex signaling mechanisms within the microenvironment of Alzheimer’s Disease (AD), which are critical for identifying novel therapeutic targets and precision medicine. However, the background signaling networks are highly complex and interactive. It remains challenging to infer the core intra- and inter-multi-cell signaling communication networks using scRNA-seq data. In this study, we introduced a novel graph transformer model, PathFinder, to infer multi-cell intra- and inter-cellular signaling pathways and communications among multi-cell types. Compared with existing models, the novel and unique design of PathFinder is based on the divide-and-conquer strategy. This model divides complex signaling networks into signaling paths, which are then scored and ranked using a novel graph transformer architecture to infer intra- and inter-cell signaling communications. We evaluated the performance of PathFinder using two scRNA-seq data cohorts. The first cohort is an APOE4 genotype-specific AD, and the second is a human cirrhosis cohort. The evaluation confirms the promising potential of using PathFinder as a general signaling network inference model.