Knowledge graphs have emerged as a powerful paradigm for structuring, organizing and reasoning over complex scientific knowledge, and are increasingly recognized as catalysts for accelerating AI for science. This study provides a comprehensive survey of scientific knowledge graphs (SciKGs), covering their construction methodologies and diverse applications across biology, chemistry and materials science. We examine how SciKGs support tasks such as drug development, omics analysis, reaction prediction and materials design, and highlight how the synergistic integration of SciKGs and large language models (LLMs) forms a knowledge- and language-driven framework for scientific discovery, in which SciKGs serve as the foundational knowledge infrastructure and LLMs act as dynamic semantic engines. We further identify key challenges and outline emerging opportunities for building auditable, interoperable and self-evolving SciKGs. Looking forward, we envision a new generation of SciKG-centered ecosystems where self-updating graphs, co-evolving with LLMs and embodied within AI scientists, become core infrastructures that autonomously drive, verify and accelerate scientific discovery.
Virtual cell (VC) models aim to predict cellular responses to any perturbations in silico and have emerged as a promising approach for drug discovery and precision medicine. Yet, a clear gap still remains: while models routinely reported impressive results on standard benchmarks, it is unclear whether their predictions are truly meaningful in practice. This is mainly due to limitations in current evaluation setups, which are often overly simplified or inconsistent, and do not reflect the complexity and variability of real biological systems. Here, we introduce a standardized and modular benchmarking framework for virtual cell prediction. Our framework evaluates diverse models under in-the-wild challenging scenarios, including unseen cell contexts, unseen perturbations, and cross-dataset generalization, which better reflect practical applications. Our analysis shows that model performance is highly context-dependent and shaped by task design and evaluation criteria. In commonly used setups, performance is often overestimated, and naive dataset aggregation can even reduce performance. When evaluated under more strict conditions, model performance drops markedly, indicating limited robustness to shifts across cellular contexts. In unseen perturbation settings, models including simple linear approaches capture global transcriptional trends but fail to recover fine-grained perturbation-specific effects. In addition, different evaluation metrics focus on different biological properties, leading to substantially different model rankings. Together, our framework provides a more reliable and biologically grounded evaluation, offering clearer guidance for applying virtual cell models in real scenarios.
Background The tumor immune microenvironment, containing a variety of immune cells with both tumor-promoting and anti-tumoral functions, plays a significant role in tumor immune surveillance and immunological evasion. Characterizing the landscape of the tumor immune microenvironment at the single-cell level is crucial for both cancer diagnosis and treatment strategy design. While current efforts to develop single-cell tumor immune atlases have laid a foundation for understanding the complexity and heterogeneity of the tumor immune microenvironment, existing atlases vary in their data sources, integration and annotation strategies, and the number and definition of cell types, posing challenges in selection. Results We systematically benchmarked five single-cell tumor immune atlases, comprising two pan-cancer and three cancer-specific ones. We first assessed their similarities and distinct characteristics of major immune cell subpopulations, including T, NK, B, macrophage, and dendritic cells. Next, we utilized each atlas as a reference to perform supervised annotation of six single-cell immuno-transcriptomics datasets, two with expert manual labels and four without. We evaluated annotation performance based on agreement with manual labels, mapping success, accuracy, clusterability, annotatability, and stability. Notably, supervised annotations consistently outperformed unsupervised clustering in identifying cell states related to immunotherapy response. Conclusions Our study provides insights into the characteristics and quality of existing atlases, demonstrating their utility in delivering harmonized annotations across datasets and uncovering crucial immune components associated with immunotherapy response. It also highlights key requirements and directions for the development of future atlases.
e15734 Background: Microsatellite stable colorectal cancer (CRC) is historically poorly responsive to immunotherapy, due in part to low neoantigen burden, an immunosuppressive tumor microenvironment, and poor T-cell infiltration. Dipeptidase-1 (DPEP1) is a GPI-anchored protein involved in glutathione and leukotriene metabolism and was identified as a part of a 4-gene immune cell exclusion signature associated with worse overall and progression free survival. Here we investigated the incidence and distribution of DPEP1 isoforms in data from The Cancer Genome Atlas (TCGA) and the in vitro properties of both isoforms. Methods: To determine the prevalence of DPEP1 Isoform B in human CRCs, we aligned TCGA Colon Adenocarcinoma (COAD) and Rectal Adenocarcinoma (READ) sample transcripts to the RefSeq (NCBI) genome database which has annotation of both DPEP1 isoforms. Murine CRC MC38 cells overexpressing human DPEP1 isoform A or isoform B were generated via lentiviral infection. The resulting cells underwent RNA sequencing followed by genome ontology and KEGG pathway over-representation analysis using the WebGestalt R package with FDR q-values ≤ 0.05. These DPEP1-expressing cells were injected into the tail veins of host mice to determine the effect of each isoform on metastasis, differences in means determined via Welch’s T-test. Results: Bulk RNA sequencing of 640 samples from 615 patients revealed DPEP1 A was found in 61% of CRC samples, while DPEP1 B was found in 91% of CRC samples, making DPEP1 B the predominant isoform. Furthermore, samples expressing high levels of DPEP1 B as compared to low or no expression, were associated with left-sided primary tumor location, MSS status, younger age at diagnosis, and no prior history of colon polyps. Transcriptional profiling revealed three distinct gene expression clusters that were correlated with DPEP1 isoform expression. Cluster one was associated with low to no DPEP1 A but high DPEP1 B. This cluster was enriched for non-coding RNAs and alternative spicing genes including RUNX1T1 and the small nuclear spliceosome RNAs. Cluster two was associated with high levels of both DPEP1 A and B expression. This was enriched for multiple G2/M cell cycle and Myc-associated genes. RNA sequencing revealed an upregulation in Myc-targets, EMT, and TNFα pathways in MC38 cells expressing isoform B. This expression of isoform B led to an increased metastatic burden in a mouse model compared to isoform A (p < 0.05). Conclusions: This marks the discovery of a novel isoform of DPEP1 that is upregulated in colorectal cancer patients. These data support the continued exploration of DPEP1 as a predictive biomarker for response to immune checkpoint inhibitors in MSS CRC as well as an emerging therapeutic target given its association with Myc and increased metastatic burden in preclinical models.
CRISPR/Cas9 specificity is critically affected by off-target effects. However, the complex patterns of mismatches and their combinations at off-target sites remain difficult to capture, and existing approaches show limited capacity to identify informative features. Here, we present CrisprPr, a hybrid-driven off-target prediction framework that integrates both prior information and data-driven modeling to improve the characterization of off-target activity. CrisprPr employs a synchronous updating strategy that jointly optimizes prior-knowledge and deep-learning modules, together with multi-source integration, to deliver accurate and stable off-target predictions. Evaluations on independent test sets indicate that CrisprPr achieves competitive predictive performance and generalization compared with existing deep learning methods, with statistically significant improvements observed on several datasets. Beyond predictive performance, its analysis module examines the patterns of prior embedding-space updates to reveal distinctive target-site features supported by literature evidence. Overall, CrisprPr proposes a novel framework that demonstrates competitive predictive performance while offering new insights into the characteristics of off-target effects.
Head and neck squamous cell carcinoma (HNSC) is a common and clinically diverse malignancy associated with poor outcomes. Recent evidence suggests that cytotoxic T lymphocyte evasion-related genes (CEGs) are pivotal regulators of tumor immune escape, yet their overall prognostic significance and influence on the tumor microenvironment (TME) in HNSC remain poorly defined. Here, a systematic analysis of 31 CEGs was conducted across TCGA and GEO HNSC datasets to identify molecular subtypes via consensus clustering. A risk score based on this signature was developed, validated in external cohorts, and integrated with clinical data into a nomogram. Associations between the risk score and TME features, immune checkpoints, somatic mutations, cancer stemness indices, and drug sensitivity were comprehensively assessed. The identified molecular subtypes exhibited markedly different immune infiltration patterns and survival outcomes. The developed model enabled effective patient stratification into high- and low-risk groups that differed significantly in overall survival. High-risk patients displayed upregulated immune checkpoint expression and heightened sensitivity to several chemotherapeutic agents. Knockdown of SERPINE1 in HNSC cell lines led to marked inhibition of both proliferation and colony formation. The findings demonstrate the critical involvement of CEGs in HNSC progression and immune modulation, and proposes a novel three-gene signature as a robust prognostic biomarker and potential guide for individualized therapy.
Background and objective The F 1 score and its generalized F β score are widely used to evaluate machine learning and artificial intelligence (AI) models in healthcare, particularly for imbalanced clinical datasets. In practice, competing prediction models are commonly evaluated on the same patient cohort, resulting in correlated classifier decisions. However, existing approaches for statistical inference of F 1 -related metrics typically assume independent classifier decisions or lack integrated procedures for comparative evaluation, power analysis, and sample size determination in paired validation studies. Methods We propose psF1pair, a unified framework for confidence interval estimation, hypothesis testing, and power and sample size calculation for comparative F 1 and F β scores under paired evaluation designs. Dependence between classifiers is modeled using a four-component multinomial representation of the joint decision process, allowing explicit estimation and incorporation of classifier correlations commonly encountered when AI models are evaluated on the same patient cohort. Exact distributions are used for small sample settings, while asymptotic approximations are employed for computational efficiency in large studies. Results Simulation studies demonstrated that the proposed confidence intervals achieved nominal coverage probabilities across a wide range of sample sizes and correlation settings. Estimated power closely agreed with empirical power, with discrepancies generally below 3%. Compared with existing methods, psF1pair showed competitive or superior statistical power while maintaining appropriate type I error rates across a broad range of scenarios. Applications to skin cancer classification and breast cancer screening demonstrated that accounting for classifier correlation produced narrower confidence intervals and improved statistical efficiency. Conclusions psF1pair provides a practical and rigorous framework for evaluation and study planning of medical AI systems using F 1 and F β metrics. The method supports comparative benchmarking, uncertainty quantification, and sample size determination for future validation studies. An open-source R package is freely available.
Identifying descriptors associated with gaseous arsenic adsorption by metal oxides remains challenging because literature data are heterogeneous and incomplete. A database of 280 experimental records and 20 descriptors from 17 studies was compiled to predict adsorption capacity and interpret descriptor-performance relationships. Mean, k-nearest neighbor (KNN), and inference-based imputation strategies were combined with gradient boosting decision tree (GBDT) and particle swarm optimization-tuned GBDT (PSO-GBDT) models. Among six configurations, the PSO-GBDT model trained on the inference-imputed dataset achieved the lowest five-fold cross-validation RMSE of 2.10 mg/g and test-set R2, RMSE, and MAE values of 0.98, 1.54 mg/g, and 0.71 mg/g, respectively. Permutation feature importance (PFI) and Shapley additive explanations (SHAP) showed that operating and gas-phase descriptors dominated predictions, with H2O concentration and adsorption time ranked highest, followed by adsorption temperature, As2O3 concentration, Fe content, and average pore diameter. Partial dependence plots (PDPs) associated higher predicted capacities with longer adsorption times, higher As2O3 concentrations, larger pore diameters, and lower adsorption temperatures within the compiled data range. As an exploratory application, the model prioritized Fe-Mn adsorbents containing 63%-81.5% Fe and 18.5%-37% Mn with pore diameters of 20-24 nm, corresponding to predicted capacities of 20-21.2 mg/g. Subsequent screening over broader operating ranges predicted a high-capacity region of 53-54.1 mg/g at As2O3 concentrations of 110-200 ppm, adsorption times of 48-240 min, and temperatures of 300-600 °C. Overall, the framework supports interpretable prediction and hypothesis generation for gaseous arsenic adsorption by metal oxides.
Accurate identification and classification of nucleus subtypes is crucial for cell tracking and uncovering patterns across cell types, such as local cell neighborhoods. Multiplexed immunofluorescence (MxIF) imaging is a process that involves staining, imaging, and then bleaching the same tissue multiple times. Repeating MxIF staining with different marker combinations enables subclassification of cells. However, repeated cycles of staining and bleaching can cause deformation, movement, and tissue loss, resulting in misalignment of markers at the nucleus level. This misalignment can lead to the exclusion of a significant number of cells during downstream analysis. We propose that applying a post hoc deep learning-based deformable registration technique (VoxelMorph) on the respective 4′ ,6-diamidino-2- phenylindole (DAPI) image for each round of staining can reduce the number of nuclei that are excluded due to spatial misalignment across successive staining rounds. By applying the registration transformations from different DAPI rounds to their corresponding stains, we achieve stain registration at pixel-level. To tackle the challenge of large image sizes, we propose a patch-based training and inference strategy. By analyzing residual displacement from bidirectional registrations, we are able to mask out areas in the tissue with high residual displacement to indicate image regions that should not be included for downstream analyses. For validation, we used a deterministic decision tree, based on biological domain knowledge, to classify MxIF nuclei into either one of 13 different classes or an undefined class. Our proposed registration approach effectively reduced the number of undefined nuclei, and we observed a 17.6% increase in the number of successfully classified nuclei compared to a baseline rigid registration. Our code is available at https://github.com/MASILab/MxIF_Registration
Epidermal growth factor receptor (EGFR) is an oncogenic driver in multiple cancers and a therapeutic target of tyrosine kinase inhibitors and neutralizing monoclonal antibodies. However, resistance to EGFR-targeted therapies, particularly the anti-EGFR antibody cetuximab, remains a clinical challenge in colorectal (CRC) and head and neck (HNSCC) cancers. Cetuximab exerts its antitumor activity by blocking ligand-dependent EGFR signaling and by engaging immune effector mechanisms. To investigate cetuximab resistance mechanisms, we cultured the cetuximab-sensitive CRC cell line DiFi in 3D with cetuximab, generating the cetuximab-resistant derivative (DiFi-CR). Genomic and transcriptomic profiling revealed that DiFi-CR cells harbor a mutation of the S442 residue within the EGFR ectodomain. Patient samples revealed recurrent EGFR S442 mutations following anti-EGFR therapy, suggesting S442 as a potential resistance hotspot. For mechanistic analyses, we reconstituted the EGFR S442I mutation, using a doxycycline-inducible system, and showed that it was necessary and sufficient to induce cetuximab resistance in CRC and HNSCC cells using in vitro cultures and in vivo mouse experiments. In silico studies, live-cell binding assays, and antibody enrichment in nude mice xenografts revealed that the S442I mutation leads to weaker EGFR-cetuximab binding. Weaker cetuximab binding was also predicted in silico for other S442 patient mutations. We found that mutant EGFR-driven resistance could be overcome by targeting the EGFR family member, ERBB2, with trastuzumab-deruxtecan. This combinatorial response required a physical interaction between EGFR and ERBB2, determined by co-immunoprecipitation. Our study supports EGFR S442 mutations as cetuximab resistance drivers and highlights co-targeting ERBB2 as a therapeutic strategy to restore anti-EGFR efficacy.
We propose Open Immune Oncology (OpenIO), a framework integrating generative AI and omics to advance precision oncology. By leveraging biological scaling laws and foundation models, we aim to transition immunotherapy from empirical screening to rational, AI-native engineering of therapeutic interventions.
Immune checkpoint blockade (ICB) is an effective treatment for microsatellite instability-high (MSI-H) colorectal cancers (CRCs) that are highly infiltrated by CD8+ T cells. Microsatellite stable (MSS) CRCs are unresponsive to ICB, at least in part, due to the paucity of intratumoral CD8+ T cells. We recently identified Discoidin Domain Receptor 1 (DDR1) as one of four genes associated with CD8+ T-cell exclusion in MSS CRC. There are conflicting reports about the presence and role of the cleaved ectodomain (cECD) of DDR1 in mouse models of breast and pancreatic cancer. To explore the role of the DDR1 cECD in human CRC, we developed Collagen Alignment and Spatial Transcriptomics Analysis (CASTA), which revealed that genes involved in fibroblast contractility were associated with both DDR1 tumor expression and collagen alignment as determined by label-free second harmonic generation (2HG) imaging of the tumor collagen. Using 3D collagen co-cultures of CRC spheroids and fibroblasts, we show that DDR1 promotes collagen alignment and CD8+ T-cell exclusion. We found large amounts of DDR1 cECD in MSS CRC supermeres, 25-35 nm secreted amembranous nanoparticles. Supermeres containing DDR1 cECD were sufficient to induce contraction of human colonic fibroblasts. We propose a model in which supermeres containing DDR1 cECD promote the contraction of stromal fibroblasts in MSS CRC, leading to collagen alignment and CD8+ T-cell exclusion. These results support DDR1 cECD as an attractive therapeutic target in MSS CRC.
Spatial transcriptomics (ST) has enabled direct interrogation of cell-cell communication (CCC) within intact tissues, providing critical spatial context that is lost in single-cell RNA-sequencing-based inference and allowing more accurate identification of physically plausible and spatially organized interactions. A rapidly expanding community of computational tools has emerged to decode CCC from ST data. Here, we provide a comprehensive review of the conceptual evolution and methodological landscape of spatial CCC inference, classifying existing approaches into two major trajectories. One trajectory, spatial pattern-based methods, assumes CCC events manifest as identifiable spatial patterns, such as colocalization, coordinated spatial signals, or higher-order spatial organization captured by deep learning models. The other trajectory, expression modulation-based approaches, assumes that CCC events influence the transcriptomic state of receiver cells. We systematically dissect their biological assumptions, statistical and deep learning frameworks, strengths, and limitations, and highlight emerging challenges in validation, benchmarking, multimodal integration, and tissue-specific modeling. Finally, we outline future directions toward achieving dynamic, multilayered reconstruction of inter- and intracellular communication, de novo signaling, and integrative multi-omics modeling.
The ability of T-cell receptors (TCRs) to recognize neoantigens is fundamental to the initiation and maintenance of adaptive immune responses. In TCR-based immunotherapies, elucidating the recognition patterns of TCRs for peptides and accurately identifying therapeutically relevant TCR-peptide pairs remain critical challenges. Here, we present a novel dual-pathway network model, ProTCR, which integrates the protein language model ProtT5 with deep learning methods. By incorporating both global and local feature extraction mechanisms, ProTCR enables efficient representation of amino acid sequences, thereby enhancing the model's generalizability across diverse data distributions and improving its biological interpretability. ProTCR demonstrates robust performance and broad applicability across various datasets, including neoantigens, previously unseen peptides, and MHC class II-restricted epitopes, overcoming the reliance on known peptide-TCR pairs observed in previous studies. It also offers new insights for predicting diverse classes of antigenic peptides. We applied ProTCR to several clinically relevant scenarios, including immunotherapeutic target identification in acute myeloid leukemia, neoantigen-targeted immunotherapy in solid tumours, and antigen-specific T cell recognition against pathogens such as influenza and severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Across these complex settings, ProTCR consistently maintained high accuracy and stability, demonstrating strong cross-task adaptability and broad potential for clinical application. This work not only provides a powerful tool for elucidating immune response mechanisms but also offers a solid computational foundation for the design of neoantigen or TCR based precision immunotherapy strategies.
Tumour cells evolve dynamically during cancer progression and treatment. Characterizing such complex cellular dynamics and accordingly developing sequential therapy that targets the evolving landscape of the tumour present the fundamental requirements for tumour therapy. Here we introduce SequenTx, a proof-of-concept artificial-intelligence-virtual-cell-inspired computational framework that integrates a tumour cell model with reinforcement learning to design sequential drug treatments across diverse drugs and tumour types considering dynamic-therapy-induced transitions in tumour cellular states with transcriptome-based therapeutic perturbation data. Large-scale in vitro experiments across various solid tumour types confirmed the effectiveness of SequenTx, achieving a 33% success rate (34 out of 102). Extending these findings in vivo, bromodomain and extra-terminal motif inhibitor pretreatment enhanced oxaliplatin sensitivity in a melanoma xenograft model. Mechanistic analysis via transcriptome data indicated that the initially administered drugs induced continuous alterations in the cancer cell transcriptome, leading to enhanced responses to subsequent treatments to achieve synergy. Additionally, SequenTx revealed a rationale for sequential therapy involving epigenetic inhibitors followed by other drugs, which unlocks the full therapeutic potential of these epigenetic drugs, thereby enhancing their clinical feasibility in cancer treatment. Overall, SequenTx provides a proof-of-concept framework, inspired by the artificial intelligence virtual cell paradigm, to rationally design effective sequential drug treatments for tumours, offering computational insights into tumour therapy to overcome transcription-dependent tumour evolution, resistance and heterogeneity.
Gene therapies that selectively eliminate tumor cells within a heterogenous population remain an elusive goal in oncology, largely due to the difficulty to achieve cancer specificity by strict definition. Here, we engineer a therapeutic gene circuit capable of identifying cells that harbor oncogenic gain-of-function aberrations in E26 transformation-specific (ETS) transcription factors. Using a machine learning-guided random forest framework, we develop a cancer-selective promoter PETS∗ that only activates during ETS overexpression and/or gene fusion events, while remaining inactive under physiological RAF-MEK-ERK signaling or in rapidly proliferating healthy tissues. When delivered using adenoviral vectors, PETS∗ enables tumor-restricted viral replication in vivo as well as long-lasting tumor suppression and complete survival of treated mice. Specifically, intratracheal delivery of PETS∗-driven adenoviruses achieves sustained control of metastatic lung tumors for over 140 days. This work overcomes key barriers of synthetic biology and oncolytic virotherapy and could open up important avenues for future cancer treatment.
Drug combination therapy is highly regarded in cancer treatment. Computational methods offer a time- and cost-effective opportunity to explore the vast combination space. Although deep learning-based prediction methods lead the field, their generalization ability remains unsatisfactory. Few previous studies have the ability to finely characterize drugs and cell lines at both the micro-scale and macro-scale. Furthermore, the interaction of cross-scale information is often overlooked. These two points limit models' ability of predicting the synergism of drug combinations in cell lines. To address the issues, we propose a novel anticancer synergistic drug combination prediction method termed MMFSynergy in this article. The construction of MMFSynergy involves three phases. First, MMFSynergy pretrains two micro encoders and a macro graph encoder, which can capture micro- or macro-scale information from large volumes of unlabeled data and generate generic features for drugs and proteins. Second, it represents drugs and proteins by fusing cross-scale information through a self-supervised task. Finally, it employs a Transformer Encoder-based model to predict synergy scores, taking representations of drugs in the combinations and the associated proteins of cell lines as input. We compared our method with eight advanced methods across three typical scenarios based on two public datasets. The results consistently demonstrated that the proposed method's generalization ability outperforms six advanced methods'. We also conducted experiments including but not limited to ablation study and case study to further exhibit the effectiveness of MMFSynergy.