Immunotherapy has seen success in treating patients with cancer, but variable responses underscore the need for effective patient stratification and therapy planning. Computational tools integrating multi-omics, imaging and machine learning have advanced, yet reliable personalized predictions remain challenging. This review analyzes the field through four converging paradigms: classical machine learning, deep learning, graph and network modeling, and mechanistic systems biology. We examine the evolution from correlational features to representation learning, relational inference, and causal simulation of tumor-immune dynamics, highlighting the shift towards multi-modal fusion and interpretable, clinically deployable models. By providing an integrated review of these computational tools, we hope to bring the community closer to achieving precision immuno-oncology for personalized cancer treatments.
Artificial intelligence (AI) is increasingly applied to biomedical research, but most current systems remain limited to specific tasks, data types, or biological scales. This makes it difficult to connect molecular alterations, organelle dysfunction, cellular behavior, tissue remodeling, organ physiology, systemic regulation, and whole-body phenotypes into coherent biological reasoning. In this Perspective, we propose the Full-Body AI Agent as a hypothetical multi-agent framework and conceptual blueprint for future systemic biology and precision medicine, rather than a fully implemented software platform. This framework envisions a supervisory Full-Body AI Agent coordinating seven biological-level agents, namely Molecule, Organelle, Cell, Tissue, Organ, Organ System, and Body System AI Agents, to standardize biomedical data, decompose cross-scale questions, assign level-specific tasks, and integrate outputs through iterative feedback. We further outline the data commons, harmonization mechanisms, uncertainty handling, arbitration strategies, and traceability safeguards required for biologically grounded cross-scale reasoning. Two hypothetical scenarios, metastasis analysis and drug development, illustrate how this framework could organize multilevel evidence from molecular changes to systemic phenotypes and therapeutic responses. This Perspective aims to clarify the conceptual basis of full-body AI and provide a foundation for transparent, physiology-constrained, cross-scale AI systems in disease analysis, therapeutic evaluation, and personalized medicine.
We envision the Full-Body AI Agent as a comprehensive AI system designed to simulate, analyze, and optimize the dynamic processes of the human body across multiple biological levels. By integrating computational models, machine learning tools, and experimental platforms, this system aims to replicate and predict both physiological and pathological processes, ranging from molecules and cells to tissues, organs, and entire body systems. Central to the Full-Body AI Agent is its emphasis on integration and coordination across these biological levels, enabling analysis of how molecular changes influence cellular behaviors, tissue responses, organ function, and systemic outcomes. With a focus on biological functionality, the system is designed to advance the understanding of disease mechanisms, support the development of therapeutic interventions, and enhance personalized medicine. We propose two specialized implementations to demonstrate the utility of this framework: (1) the metastasis AI Agent, a multi-scale metastasis scoring system that characterizes tumor progression across the initiation, dissemination, and colonization phases by integrating molecular, cellular, and systemic signals; and (2) the drug AI Agent, a system-level drug development paradigm in which a drug AI-Agent dynamically guides preclinical evaluations, including organoids and chip-based models, by providing full-body physiological constraints. This approach enables the predictive modeling of long-term efficacy and toxicity beyond what localized models alone can achieve. These two agents illustrate the potential of Full-Body AI Agent to address complex biomedical challenges through multi-level integration and cross-scale reasoning.
Understanding disease progression and sophisticated tumor ecosystems is imperative for investigating tumorigenesis mechanisms and developing novel prevention strategies. Here, we dissected heterogeneous microenvironments during malignant transitions by leveraging data from 1396 samples spanning 13 major tissues. Within transitional stem-like subpopulations highly enriched in precancers and cancers, we identified 30 recurring cellular states strongly linked to malignancy, including hypoxia and epithelial senescence, revealing a high degree of plasticity in epithelial stem cells. By characterizing dynamics in stem-cell crosstalk with the microenvironment along the pseudotime axis, we found differential roles of ANXA1 at different stages of tumor development. In precancerous stages, reduced ANXA1 levels promoted monocyte differentiation toward M1 macrophages and inflammatory responses, whereas during malignant progression, upregulated ANXA1 fostered M2 macrophage polarization and cancer-associated fibroblast transformation by increasing TGF-β production. Our spatiotemporal analysis further provided insights into mechanisms responsible for immunosuppression and a potential target to control evolution of precancer and mitigate the risk for cancer development.
The categories of diabetic retinopathy (DR) are interrelated, and different ophthalmologists often give different results for the same fundus image. Automatic cross image retrieval of DR can provide an effective diagnostic solution for ophthalmologists and is of great significance in clinical practice. Cross-image (i.e. left and right fundus images for a patient) information is highly correlated and complementary and can be harnessed to improve various computer vision tasks such as image classification, object detection, image segmentation, and image retrieval. Previous studies did not explore the correlation between lesion areas in left and right fundus images of patients, limiting the effective diagnosis of DR. In this study, we proposed a cross-image siamese graph convolutional network(CIS-GCN) to retrieve fine-grained diabetic retinopathy fundus images. First, we constructed a global-specific structure to obtain the specific features of the left and right eyes. Then, we passed the specific features through the pathological localization network to obtain the location features of the lesion. Finally, a graph convolutional neural network was introduced to construct node sets for the left and right eyes, respectively, to represent relatively consistent regions in the fundus images of patients and learn their correlations. We tested our method using Diabetic Retinopathy Detection datasets and the results showed that our algorithm outperforms other state-of-the-art methods by 2.2 % similar to 3.7 % in image data retrieval.
The COVID-19 pandemic, caused by the coronavirus SARS-CoV-2, has resulted in the loss of millions of lives and severe global economic consequences. Every time SARS-CoV-2 replicates, the viruses acquire new mutations in their genomes. Mutations in SARS-CoV-2 genomes led to increased transmissibility, severe disease outcomes, evasion of the immune response, changes in clinical manifestations and reducing the efficacy of vaccines or treatments. To date, the multiple resources provide lists of detected mutations without key functional annotations. There is a lack of research examining the relationship between mutations and various factors such as disease severity, pathogenicity, patient age, patient gender, cross-species transmission, viral immune escape, immune response level, viral transmission capability, viral evolution, host adaptability, viral protein structure, viral protein function, viral protein stability and concurrent mutations. Deep understanding the relationship between mutation sites and these factors is crucial for advancing our knowledge of SARS-CoV-2 and for developing effective responses. To fill this gap, we built COV2Var, a function annotation database of SARS-CoV-2 genetic variation, available at http://biomedbdc.wchscu.cn/COV2Var/. COV2Var aims to identify common mutations in SARS-CoV-2 variants and assess their effects, providing a valuable resource for intensive functional annotations of common mutations among SARS-CoV-2 variants.
Glioblastoma multiforme (GBM)is the most common and aggressive primary brain tumor. Although temozolomide (TMZ)-based radiochemotherapy improves overall GBM patients’ survival, it also increases the frequency of false positive post-treatment magnetic resonance imaging (MRI) assessments for tumor progression. Pseudo-progression (PsP) is a treatment-related reaction with an increased contrast-enhancing lesion size at the tumor site or resection margins miming tumor recurrence on MRI. The accurate and reliable prognostication of GBM progression is urgently needed in the clinical management of GBM patients. Clinical data analysis indicates that the patients with PsP had superior overall and progression-free survival rates. In this study, we aimed to develop a prognostic model to evaluate the tumor progression potential of GBM patients following standard therapies. We applied a dictionary learning scheme to obtain imaging features of GBM patients with PsP or true tumor progression (TTP) from the Wake dataset. Based on these radiographic features, we conducted a radiogenomics analysis to identify the significantly associated genes. These significantly associated genes were used as features to construct a 2YS (2-year survival rate) logistic regression model. GBM patients were classified into low- and high-survival risk groups based on the individual 2YS scores derived from this model. We tested our model using an independent The Cancer Genome Atlas Program (TCGA) dataset and found that 2YS scores were significantly associated with the patient’s overall survival. We used two cohorts of the TCGA data to train and test our model. Our results show that the 2YS scores-based classification results from the training and testing TCGA datasets were significantly associated with the overall survival of patients. We also analyzed the survival prediction ability of other clinical factors (gender, age, KPS (Karnofsky performance status), normal cell ratio) and found that these factors were unrelated or weakly correlated with patients’ survival. Overall, our studies have demonstrated the effectiveness and robustness of the 2YS model in predicting the clinical outcomes of GBM patients after standard therapies.
StemDriver is a comprehensive knowledgebase dedicated to the functional annotation of genes participating in the determination of hematopoietic stem cell fate, available at http://biomedbdc.wchscu.cn/StemDriver/. By utilizing single-cell RNA sequencing data, StemDriver has successfully assembled a comprehensive lineage map of hematopoiesis, capturing the entire continuum from the initial formation of hematopoietic stem cells to the fully developed mature cells. Extensive exploration and characterization were conducted on gene expression features corresponding to each lineage commitment. At the current version, StemDriver integrates data from 42 studies, encompassing a diverse range of 14 tissue types spanning from the embryonic phase to adulthood. In order to ensure uniformity and reliability, all data undergo a standardized pipeline, which includes quality data pre-processing, cell type annotation, differential gene expression analysis, identification of gene categories correlated with differentiation, analysis of highly variable genes along pseudo-time, and exploration of gene expression regulatory networks. In total, StemDriver assessed the function of 23 839 genes for human samples and 29 533 genes for mouse samples. Simultaneously, StemDriver also provided users with reference datasets and models for cell annotation. We believe that StemDriver will offer valuable assistance to research focused on cellular development and hematopoiesis.
Objective The early stages of chronic disease typically progress slowly, so symptoms are usually only noticed until the disease is advanced. Slow progression and heterogeneous manifestations make it challenging to model the transition from normal to disease status. As patient conditions are only observed at discrete timestamps with varying intervals, an incomplete understanding of disease progression and heterogeneity affects clinical practice and drug development. Materials and Methods We developed the Gaussian Process for Stage Inference (GPSI) approach to uncover chronic disease progression patterns and assess the dynamic contribution of clinical features. We tested the ability of the GPSI to reliably stratify synthetic and real-world data for osteoarthritis (OA) in the Osteoarthritis Initiative (OAI), bipolar disorder (BP) in the Adolescent Brain Cognitive Development Study (ABCD), and hepatocellular carcinoma (HCC) in the UTHealth and The Cancer Genome Atlas (TCGA). Results First, GPSI identified two subgroups of OA based on image features, where these subgroups corresponded to different genotypes, indicating the bone-remodeling and overweight-related pathways. Second, GPSI differentiated BP into two distinct developmental patterns and defined the contribution of specific brain region atrophy from early to advanced disease stages, demonstrating the ability of the GPSI to identify diagnostic subgroups. Third, HCC progression patterns were well reproduced in the two independent UTHealth and TCGA datasets. Conclusion Our study demonstrated that an unsupervised approach can disentangle temporal and phenotypic heterogeneity and identify population subgroups with common patterns of disease progression. Based on the differences in these features across stages, physicians can better tailor treatment plans and medications to individual patients.
Background Sexual differences across molecular levels profoundly impact cancer biology and outcomes. Patient gender significantly influences drug responses, with divergent reactions between men and women to the same drugs. Despite databases on sex differences in human tissues, understanding regulations of sex disparities in cancer is limited. These resources lack detailed mechanistic studies on sex-biased molecules. Methods In this study, we conducted a comprehensive examination of molecular distinctions and regulatory networks across 27 cancer types, delving into sex-biased effects. Our analyses encompassed sex-biased competitive endogenous RNA networks, regulatory networks involving sex-biased RNA binding protein-exon skipping events, sex-biased transcription factor-gene regulatory networks, as well as sex-biased expression quantitative trait loci, sex-biased expression quantitative trait methylation, sex-biased splicing quantitative trait loci, and the identification of sex-biased cancer therapeutic drug target genes. All findings from these analyses are accessible on SexAnnoDB (https://ccsm.uth.edu/SexAnnoDB/). Results From these analyses, we defined 126 cancer therapeutic target sex-associated genes. Among them, 9 genes showed sex-biased at both the mRNA and protein levels. Specifically, S100A9 was the target of five drugs, of which calcium has been approved by the FDA for the treatment of colon and rectal cancers. Transcription factor (TF)-gene regulatory network analysis suggested that four TFs in the SARC male group targeted S100A9 and upregulated the expression of S100A9 in these patients. Promoter region methylation status was only associated with S100A9 expression in KIRP female patients. Hypermethylation inhibited S100A9 expression and was responsible for the downregulation of S100A9 in these female patients. Conclusions Comprehensive network and association analyses indicated that the sex differences at the transcriptome level were partially the result of corresponding sex-biased epigenetic and genetic molecules. Overall, SexAnnoDB offers a discipline-specific search platform that could potentially assist basic experimental researchers or physicians in developing personalized treatment plans.
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the etiologic agent of coronavirus disease 19 (COVID-19), has caused a global health crisis. Despite ongoing efforts to treat patients, there is no universal prevention or cure available. One of the feasible approaches will be identifying the key genes from SARS-CoV-2-infected cells. SARS-CoV-2-infected in vitro model, allows easy control of the experimental conditions, obtaining reproducible results, and monitoring of infection progression. Currently, accumulating RNA-seq data from SARS-CoV-2 in vitro models urgently needs systematic translation and interpretation. To fill this gap, we built COVIDanno, COVID-19 annotation in humans, available at http://biomedbdc.wchscu.cn/COVIDanno/. The aim of this resource is to provide a reference resource of intensive functional annotations of differentially expressed genes (DEGs) among different time points of COVID-19 infection in human in vitro models. To do this, we performed differential expression analysis for 136 individual datasets across 13 tissue types. In total, we identified 4,935 DEGs. We performed multiple bioinformatics/computational biology studies for these DEGs. Furthermore, we developed a novel tool to help users predict the status of SARS-CoV-2 infection for a given sample. COVIDanno will be a valuable resource for identifying SARS-CoV-2-related genes and understanding their potential functional roles in different time points and multiple tissue types.
Lung adenocarcinoma (LUAD) is a deadly tumor with dynamic evolutionary process. Although much endeavors have been made in identifying the temporal patterns of cancer progression, it remains challenging to infer and interpret the molecular alterations associated with cancer development and progression. To this end, we developed a computational approach to infer the progression trajectory based on cross-sectional transcriptomic data. Analysis of the LUAD data using our approach revealed a linear trajectory with three different branches for malignant progression, and the results showed consistency in three independent cohorts. We used the progression model to elucidate the potential molecular events in LUAD progression. Further analysis showed that overexpression of BUB1B, BUB1 and BUB3 promoted tumor cell proliferation and metastases by disturbing the spindle assembly checkpoint (SAC) in the mitosis. Aberrant mitotic spindle checkpoint signaling appeared to be one of the key factors promoting LUAD progression. We found the inferred cancer trajectory allows to identify LUAD susceptibility genetic variations using genome-wide association analysis. This result shows the opportunity for combining analysis of candidate genetic factors with disease progression. Furthermore, the trajectory showed clear evident mutation accumulation and clonal expansion along with the LUAD progression. Understanding how tumors evolve and identifying mutated genes will help guide cancer management. We investigated the clonal architectures and identified distinct clones and subclones in different LUAD branches. Validation of the model in multiple independent data sets and correlation analysis with clinical results demonstrate that our method is effective and unbiased.
Objectives: The association between sleep pattern and chronic kidney disease (CKD) incidence, and whether the association is dependent on the genetic backgrounds has not been addressed. We sought to investigate the association of multidimensional sleep pattern with CKD in consideration of genetic polymorphisms. Methods: In this prospective cohort study of 157,175 participants from the UK Biobank, sleep patterns were derived by multiple correspondence analysis (MCA) and k-means clustering of individual sleep traits (sleep duration, insomnia, chronotype, daytime sleepiness, snoring, and night shift status). Cox proportional hazard regression was used to estimate the association between sleep patterns and CKD incidence. Gene-environmentwide interaction study (GEWIS) was performed to detect whether gene polymorphisms were modifiers on this association. Results: Compared with "healthy sleep" pattern, increased CKD incidence was observed in the clusters with "long sleep duration" (hazard ratios (HR) 1.42, 95% confidence intervals (CI), 1.18-1.72) and "night shift" (HR 1.23, 95% CI, 1.05-1.45) patterns, but not with the "short sleep duration" pattern. By GEWIS, we identified 167 SNPs as suggestive effect modifiers that interacted with unhealthy sleep patterns and affected the risk of CKD. Conclusions: Unhealthy sleep patterns, with features of long sleep duration and night shift, may increase the risk of CKD. The study highlights the interaction of sleep and individual genetic risk to affect health outcomes.
BACKGROUND:Medical images have already become an essential tool for the diagnosis of many diseases. Thus a large number of medical images are being generated due to the daily routine inspection. An efficient image-based disease retrieval system will not only make full use of existing data, but also help physicians to prognosis the diseases. Medical image retrieval is represented by the classification and localization of common thorax diseases in x-ray images. Although extensive efforts have been put into this field, there are still many challenges. PURPOSE:Most of the existing fine-grained image research methods just apply existing deep learning frameworks in extracting the image features. However, these high-level features mainly focus on the global representations of the object, rather than simultaneously considering the local ones. It requires fine-grained details to classify the images with similar lesion areas. Thus, it is necessary to combine the global features and local ones to make the features more discriminative. On the other hand, training CNN models based on current existing strategies have a high time complexity, and is hard to get the discriminative features mentioned above. In addition, the visual retrieval method of fine-grained medical images still has the problem of insufficient sample data with accurate annotation information. METHODS:To address above challenges, we introduced a novel fine-grained medical images retrieval method. First, a centralized contrastive loss (CCLoss) is proposed as our metric learning loss function. Parameters are updated by using the center point, which not only improves the distinguishing performance of features, but also effectively reduces the time complexity of the algorithm. In addition, a weakly supervised progressive feature extraction method is proposed to gradually extract the combined features. And the attention mechanism module is applied to screen the target information after the initial positioning for fine refinement, so as to separate the features with a high degree of discrimination. The retrieval of 14 different chest diseases is evaluated on the chest x-ray datasets. RESULTS:Compared with the existing research methods, the proposed method shows a better retrieval result for Recall@8 by 2.26 % ∼ 4.6 % $\%{\sim }4.6\%$ and achieves a very efficient training speed which is 100 times faster than the pair-wise loss-based training strategy. We also assessed the effects of Recall@k (k = 2, 4, 6, 8) for progressive features extracted from different steps to obtain a model with the best retrieval performance. CONCLUSIONS:The proposed model is capable of learning discriminative representations from chest x-ray datasets, and it achieves better performance compared with other state-of-the-art methods. Therefore, the developed model would be useful in the diagnosis of common thorax disease or unknown chest disease.
Immunotherapies have revolutionized cancer treatment modalities; however, predicting clinical response accurately and reliably remains challenging. Neoantigen load is considered as a fundamental genetic determinant of therapeutic response. However, only a few predicted neoantigens are highly immunogenic, with little focus on intratumor heterogeneity (ITH) in the neoantigen landscape and its link with different features in the tumor microenvironment. To address this issue, we comprehensively characterized neoantigens arising from nonsynonymous mutations and gene fusions in lung cancer and melanoma. We developed a composite NEO2IS to characterize interplays between cancer and CD8+ T-cell populations. NEO2IS improved prediction accuracy of patient responses to immune-checkpoint blockades (ICBs). We found that TCR repertoire diversity was consistent with the neoantigen heterogeneity under evolutionary selections. Our defined neoantigen ITH score (NEOITHS) reflected infiltration degree of CD8+ T lymphocytes with different differentiation states and manifested the impact of negative selection pressure on CD8+ T-cell lineage heterogeneity or tumor ecosystem plasticity. We classified tumors into distinct immune subtypes and examined how neoantigen-T cells interactions affected disease progression and treatment response. Overall, our integrated framework helps profile neoantigen patterns that elicit T-cell immunoreactivity, enhance the understanding of evolving tumor-immune interplays and improve prediction of ICBs efficacy.
It is extremely important to identify patients with acute pancreatitis who are at high risk for developing persistent organ failures early in the course of the disease. Due to the irregularity of longitudinal data and the poor interpretability of complex models, many models used to identify acute pancreatitis patients with a high risk of organ failure tended to rely on simple statistical models and limited their application to the early stages of patient admission. With the success of recurrent neural networks in modeling longitudinal medical data and the development of interpretable algorithms, these problems can be well addressed. In this study, we developed a novel model named Multi-task and Time-aware Gated Recurrent Unit RNN (MT-GRU) to directly predict organ failure in patients with acute pancreatitis based on irregular medical EMR data. Our proposed end-to-end multi-task model achieved significantly better performance compared to two-stage models. In addition, our model not only provided an accurate early warning of organ failure for patients throughout their hospital stay, but also demonstrated individual and population-level important variables, allowing physicians to understand the sci-entific basis of the model for decision-making. By providing early warning of the risk of organ failure, our proposed model is expected to assist physicians in improving outcomes for patients with acute pancreatitis.
Aging is a complex process that accompanied by molecular and cellular alterations. The identification of tissue-/cell type-specific biomarkers of aging and elucidation of the detailed biological mechanisms of aging-related genes at the single-cell level can help to understand the heterogeneous aging process and design targeted anti-aging therapeutics. Here, we built AgeAnno (https://relab.xidian.edu.cn/AgeAnno/#/), a knowledgebase of single cell annotation of aging in human, aiming to provide comprehensive characterizations for aging-related genes across diverse tissue-cell types in human by using single-cell RNA and ATAC sequencing data (scRNA and scATAC). The current version of AgeAnno houses 1 678 610 cells from 28 healthy tissue samples with ages ranging from 0 to 110 years. We collected 5580 aging-related genes from previous resources and performed dynamic functional annotations of the cellular context. For the scRNA data, we performed analyses include differential gene expression, gene variation coefficient, cell communication network, transcription factor (TF) regulatory network, and immune cell proportionc. AgeAnno also provides differential chromatin accessibility analysis, motif/TF enrichment and footprint analysis, and co-accessibility peak analysis for scATAC data. AgeAnno will be a unique resource to systematically characterize aging-related genes across diverse tissue-cell types in human, and it could facilitate antiaging and aging-related disease research.
Tumors are often polyclonal due to copy number alteration (CNA) events. Through the CNA profile, we can understand the tumor heterogeneity and consistency. CNA information is usually obtained through DNA sequencing. However, many existing studies have shown a positive correlation between the gene expression and gene copy number identified from DNA sequencing. With the development of spatial transcriptome technologies, it is urgent to develop new tools to identify genomic variation from the spatial transcriptome. Therefore, in this study, we developed CVAM, a tool to infer the CNA profile from spatial transcriptome data. Compared with existing tools, CVAM integrates the spatial information with the spot's gene expression information together and the spatial information is indirectly introduced into the CNA inference. By applying CVAM to simulated and real spatial transcriptome data, we found that CVAM performed better in identifying CNA events. In addition, we analyzed the potential co-occurrence and mutual exclusion between CNA events in tumor clusters, which is helpful to analyze the potential interaction between genes in mutation. Last but not least, Ripley's K-function is also applied to CNA multi-distance spatial pattern analysis so that we can figure out the differences of different gene CNA events in spatial distribution, which is helpful for tumor analysis and implementing more effective treatment measures based on spatial characteristics of genes.
The emergence of RNA velocity has enriched our understanding of the dynamic transcriptional landscape within individual cells. In light of this breakthrough, we embarked on integrating RNA velocity with cellular pseudotime inference, aiming to improve the prediction of cell orders along biological trajectories beyond existing methods. Here, we developed LVPT, a novel method for pseudotime and trajectory inference. LVPT introduces a lazy probability to indicate the probability that the cell stays in the original state and calculates the transition matrix based on RNA velocity to provide the probability and direction of cell differentiation. LVPT shows better and comparable performance of pseudotime inference compared with other existing methods on both simulated datasets with different structures and real datasets. The validation results were consistent with prior knowledge, indicating that LVPT is an accurate and efficient method for pseudotime inference.
The explosive growth of biomedical Big Data presents both significant opportunities and challenges in the realm of knowledge discovery and translational applications within precision medicine. Efficient management, analysis, and interpretation of big data can pave the way for groundbreaking advancements in precision medicine. However, the unprecedented strides in the automated collection of large-scale molecular and clinical data have also introduced formidable challenges in terms of data analysis and interpretation, necessitating the development of novel computational approaches. Some potential challenges include the curse of dimensionality, data heterogeneity, missing data, class imbalance, and scalability issues. This overview article focuses on the recent progress and breakthroughs in the application of big data within precision medicine. Key aspects are summarized, including content, data sources, technologies, tools, challenges, and existing gaps. Nine fields-Datawarehouse and data management, electronic medical record, biomedical imaging informatics, Artificial intelligence-aided surgical design and surgery optimization, omics data, health monitoring data, knowledge graph, public health informatics, and security and privacy-are discussed.