AI single-cell foundation models (scFMs) are believed to be able to learn essential relations in cell transcriptomics with the attention modules in Transformer, but there is no method to reveal what they actually learned. We observed that different models may grasp different aspects of relations. To unravel the mystery, we propose scGeneLens, a framework for dissecting how scFMs perceive cells. We employed a sparse block attention to replace the original attention mechanism to concentrate attentions into a few dominant gene-gene relations, used attention propagation to trace how the relations propagate across Transformer layers, and used integrated gradients to disentangle the relative contributions of gene identity and expression in cell representations. We applied it to scFoundation and scGPT and show that they exhibit pluralistic perceptions of cells: scFoundation emphasizes relations among cell-type marker genes, resulting in stronger cell-type separability, whereas scGPT focuses more on genes involved in shared cellular pathways and core biological activities, leading to representations that generalize across conditions. The framework provides a unified lens for probing what scFMs learn about cells and offers actionable insights for the design of future cellular foundation models. Our code can be seen in \url{https://anonymous.4open.science/r/scGeneLens-B771/} . ### Competing Interest Statement The authors have declared no competing interest. The work is sponsored by National Natural Science Foundation of China (92470105, 62373210), the National Key R&D Program of China (2025YFC3409300).
MOTIVATION:Single-cell transcriptomic datasets annotate cell types with diverse schemes and varying resolution. This poses challenges in building unified hierarchical cell-type structures and hinders integration of large-scale datasets. To address this, several computational methods have been developed to harmonize cell type annotations across datasets and build data-driven hierarchies of cell types. RESULTS:Here, we benchmarked three state-of-the-art methods: scHPL, treeArches, and CellHint. We evaluated these methods across five simulated scenarios and five real-world scenarios across cell types and organs. To assess harmonization results, we designed three metrics, Annotation Harmonization F1-score (AH-F1), Tree Edit Distance Similarity and Parent-Children Branches Similarity, comparing the constructed cell-type hierarchies and the knowledge-based ones. Based on the benchmarking results, we found that methods performed well in simulated scenarios but still have room for improvement in complex real-world data. Thus, we developed OTHarmonizer, a tool based on partial optimal transport (OT) for cell-type harmonization and hierarchy construction. OTHarmonizer excels in accurately capturing equivalent and hierarchical relationships between cell types, offering a more effective approach for the cell-type hierarchy construction across datasets. AVAILABILITY AND IMPLEMENTATION:The simulated and real-world datasets in the benchmark are available on https://figshare.com/articles/dataset/OTHarmonizer/28243205. The source codes for the benchmark and OTHarmonizer are available online on GitHub at https://github.com/Duck-Boss/OTHarmonizer.
Although CRISPR-Cas9 holds therapeutic promise, broader application demands an understanding of complications in vast non-coding regions. We found that CRISPR-Cas9 can cause premature differentiation of neural stem cells in vivo and mouse embryonic stem cells in vitro, even when cleavage occurred at distant sites tens of kilobases away from the nearest regulatory elements. To investigate this, we employed an integrated assay for transposase-accessible chromatin (ATAC)/RNA sequencing (AR-seq) approach and identified editing-induced chromatin accessibility changes, with their scale varying by cell type. Cells with stemness are most affected, experiencing perturbations that extend over a hundred kilobases. Furthermore, even local DNA perturbations can disrupt CTCF- and condensate-associated chromatin architecture, causing distal transcriptional rewiring and, ultimately, loss of stemness identity. To minimize chromatin perturbations and preserve cell identity, we refined gene-editing strategies, including distance-aware sgRNA design, pharmacological attenuation of DNA resection, and alternative editing systems. This work paves the way for the safer and broader application of genome-editing technologies.
Motivation Recent advances in large language models have enabled the emergence of AI scientists that aim to autonomously analyze biological data and assist scientific discovery. Despite rapid progress, it remains unclear to what extent these systems can extract meaningful biological insights from real experimental data. Existing benchmarks either evaluate reasoning in the absence of data or focus on predefined analytical outputs, failing to reflect realistic, data-driven biological research.Results Here, we introduce BAISBench (Biological AI Scientist Benchmark), a benchmark for evaluating AI scientists on real single-cell transcriptomic datasets. BAISBench comprises two tasks: cell type annotation across 15 expert-labeled datasets, and scientific discovery through 193 multiple-choice questions derived from biological conclusions reported in 41 published single-cell studies. We evaluated several representative AI scientists using BAISBench and, to provide a human performance baseline, invited five graduate-level bioinformaticians to collectively complete the same tasks. The results show that while current AI scientists fall short of fully autonomous biological discovery, they already demonstrate substantial potential in supporting data-driven biological research. These results position BAISBench as a practical benchmark for characterizing the current capabilities and limitations of AI scientists in biological research. We expect BAISBench to serve as a practical evaluation framework for guiding the development of more capable AI scientists and for helping biologists identify AI systems that can effectively support real-world research workflows.Availability and implementation https://github.com/EperLuo/BAISBench, https://huggingface.co/datasets/EperLuo/BaisBench.
Single-cell multi-omics technologies offer unprecedented opportunities to decipher complex cellular mechanisms. To overcome experimental limitations in scale, cost, and coverage, powerful computational methods are essential for integrating diverse data modalities and generating high-fidelity in-silico data. In this paper, we present scDiffusion-X, a latent diffusion model for multi-omics data integration, generation, and translation. The core innovation is a Dual-Cross-Attention (DCA) module that adaptively captures intricate, hidden relationships between molecular modalities, offering a more flexible and interpretable approach than existing integration strategies. Extensive benchmarking experiments demonstrate that scDiffusion-X excels at generating realistic multi-omics data, preserving cellular heterogeneity and global data structures with excellent scalability. Beyond simulation, scDiffusion-X uniquely enables accurate modality translation, predicting one molecular modality from another with robust uncertainty quantification. Furthermore, we designed a gradient-based interpretation framework to transform the DCA module into a discovery tool, enabling inference of comprehensive cell-type-specific heterogeneous gene regulatory networks (GRNs). By integrating state-of-the-art generative modeling with biological interpretability, scDiffusion-X serves as a powerful tool for dissecting regulatory relationships, predicting perturbation responses, and accelerating discovery in single-cell multi-omics research.
Background: Vulnerable plaques are the immediate cause of acute coronary syndromes such as acute myocardial infarction (AMI). However, current preventive strategies including lipid-lowering and imaging-based approaches focus on overall cardiovascular risk and lack precision in identifying high-risk plaques before clinical events. One major limitation is the insufficient understanding of cell-type heterogeneity and dynamic changes across the full course of atherosclerosis. Methods and Results: We established a cross-species, multi-stage framework integrating human single-cell data from advanced stable and unstable plaques with multi-timepoint mouse models. This allowed reconstruction of a continuous trajectory of atherosclerosis and led to identification of a monocyte-derived population, Plaque Instability Monocytes (PIMs), enriched in unstable lesions and characterized by inflammatory and matrix-remodeling programs. To validate their relevance, we profiled additional lesions across multiple arteries and disease states, including in-stent restenosis. PIM signatures were consistently recovered, and spatial transcriptomics revealed preferential localization to fibrous caps, confirmed by confocal imaging. Circulating PIM-like cells were also elevated in patients with unstable plaques. From the PIM transcriptome, we derived a 362-gene blood signature, and used interpretable deep learning to identify a stable 20-gene panel distinguishing unstable from stable cases (AUC = 0.907). A composite risk score built using L1-logistic regression was validated in an independent cohort (n = 338, AUC = 0.92). Cytokine genes within the panel showed elevated plasma levels in AMI patients (n = 26), and UK Biobank proteomics (n = 3,798) confirmed upregulation of key panel proteins in myocardial infarction. Conclusions: By bridging single-cell discoveries with peripheral biomarker modeling and clinical validation, this study identifies a conserved monocyte subset associated with plaque vulnerability and provides a practical tool for early detection and risk stratification in coronary artery disease.
Rare diseases, affecting ~350 million people worldwide, pose significant challenges in clinical diagnosis due to the lack of experienced physicians and the complexity of differentiating between numerous rare diseases. To address these challenges, we introduce PhenoBrain, a fully automated artificial intelligence pipeline. PhenoBrain utilizes a BERT-based natural language processing model to extract phenotypes from clinical texts in EHRs and employs five new diagnostic models for differential diagnoses of rare diseases. The AI system was developed and evaluated on diverse, multi-country rare disease datasets, comprising 2271 cases with 431 rare diseases. In 1936 test cases, PhenoBrain achieved an average predicted top-3 recall of 0.513 and a top-10 recall of 0.654, surpassing 13 leading prediction methods. In a human-computer study with 75 cases, PhenoBrain exhibited exceptional performance with a top-3 recall of 0.613 and a top-10 recall of 0.813, surpassing the performance of 50 specialist physicians and large language models like ChatGPT and GPT-4. Combining PhenoBrain's predictions with specialists increased the top-3 recall to 0.768, demonstrating its potential to enhance diagnostic accuracy in clinical workflows.
Determining tumor progression status is critical for early-stage lung adenocarcinoma (esLUAD) diagnosis and treatment, yet histopathology-based grading often overlooks heterogeneity within grades. We propose RadioTrace, a deep contrastive learning framework integrating radiomic and pathological information to learn a radiomic trajectory for quantifying esLUAD progression. Across four multi-institutional cohorts, RadioTrace well predicted tumor phenotypes including spread through air spaces (STAS) and lymph node metastasis (LNM). Survival analyses demonstrated it as an independent prognostic factor (log-rank test p < 0.004 across all cohorts). Within the same pathological grade, it revealed significant survival heterogeneity (p < 0.02 across all cohorts), underscoring the limitations of current grading criteria. Genomic and transcriptomic analyses confirmed associations with progression-related molecular features. Longitudinal analysis of patients with multiple CT follow-ups further showed consistency with continuous progression. These findings demonstrate that RadioTrace enables quantitative, interpretable assessment of esLUAD progression, providing insights beyond histopathology and assisting clinical decision-making.
Reprogramming cell state transitions provides the potential for cell engineering and regenerative therapy. Finding the reprogramming transcription factors (TFs) and their combinations that can direct the desired state transition is crucial for the task. Computational methods have been developed to identify such reprogramming TFs. However, most of them can only generate a ranked list of individual TFs and ignore the identification of TF combinations. Even for individual reprogramming TF identification, current methods often fail to put the real effective reprogramming TFs at the top. To address these challenges, we developed TFcomb, a computational method that leverages single-cell multiomics data to identify reprogramming TFs and TF combinations. We modeled the task of finding reprogramming TFs and their combinations as an inverse problem, and used Tikhonov regularization to guarantee the generalization ability of solutions. For the coefficient matrix of the model, we designed a graph attention network to augment gene regulatory networks built with single-cell RNA-seq and ATAC-seq data. Benchmarking experiments on data of human embryonic stem cells demonstrate superior performance of TFcomb against existing methods for identifying individual TFs. We curate data sets of multiple cell reprogramming cases and demonstrate that TFcomb can efficiently identify reprogramming TF combinations from vast potential combinations. We apply TFcomb on a data set of mouse hair follicle development and find key TFs in cell differentiation. All experiments show that TFcomb is powerful in identifying reprogramming TFs and TF combinations from single-cell data sets to empower future cell engineering.
SUMMARY:In single-cell transcriptomics, inconsistent cell type annotations due to varied naming conventions and hierarchical granularity impede data integration, machine learning applications, and meaningful evaluations. To address this challenge, we developed the unified Hierarchical Annotation Framework (uHAF), which includes organ-specific hierarchical cell type trees (uHAF-T) and a mapping tool (uHAF-Agent) based on large language models. uHAF-T provides standardized hierarchical references for 38 organs, allowing for consistent label unification and analysis at different levels of granularity. uHAF-Agent leverages GPT-4 to accurately map diverse and informal cell type labels onto uHAF-T nodes, streamlining the harmonization process. By simplifying label unification, uHAF enhances data integration, supports machine learning applications, and enables biologically meaningful evaluations of annotation methods. Our framework serves as an essential resource for standardizing cell type annotations and fostering collaborative refinement in the single-cell research community. AVAILABILITY AND IMPLEMENTATION:uHAF is publicly available at: https://uhaf.unifiedcellatlas.org and https://github.com/SuperBianC/uhaf.
The design of neoepitope-based cancer immunotherapy requires predicting the immunogenicity of each patient-specific neoepitope candidate. We designed NeoGuider, a machine-learning model to address nonlinearity and class imbalance in such prediction. Central to NeoGuider is a supervised feature-transformation approach, which estimates odds using our custom kernel density estimation followed by centered isotonic regression. We implemented NeoGuider as a bioinformatics pipeline, which detects neoepitope candidates from sequencing data and prioritizes such candidates by predicted immunogenicity. Benchmarking on 7 cohorts, 113 patients, and 635 immunogenic candidates showed that NeoGuider outperformed existing methods in neoepitope prediction. NeoGuider is open-source at https://github.com/XuegongLab/neoguider.
Learning spatial context of cells through pretraining on spatial transcriptomics (ST) data may empower us to decipher tissue organization and cellular interactions. Yet, transformer-based generative models often focus on modeling individual cells, overlooking the intricate spatial relationships within them. To address this limitation, we develop GeST, a deep transformer model pretrained by a novel spatially informed generation task: Predicting cellular expression profile of a given location based on the information from its neighboring cells. GeST integrates a specialized spatial attention mechanism for efficient pretraining, a flexible serialization strategy for sequentializing ST data, and a cell tokenization method for quantizing gene expression profiles. We pretrained GeST on large-scale ST datasets across multiple ST technologies, achieving superior performance in generating previously unseen spatial cell profiles, extracting spatial niche embeddings in a zero-shot manner, and annotating spatial regions. Furthermore, GeST can simulate gene expression changes in response to perturbations of cells within spatial context, closely matching existing experimental results. GeST offers a powerful generative pre-training framework for learning spatial contexts.
Predicting cellular responses to perturbations is essential for understanding gene regulation and advancing drug development. Most existing in silico perturbation models treat perturbation–cell interactions with simplistic fusion strategies that overlook the hierarchical nature of regulatory processes, leading to inaccurate predictions and poor generalization across perturbation types. We present X-Pert, a universal in silico perturbation model that explicitly captures both gene–perturbation interactions and gene–gene dependencies through attention mechanisms. X-Pert flexibly handles diverse perturbation inputs across different types and combinations, and quantitatively models their dosage- and efficacy-dependent effects, all within a unified representation space. This enables accurate prediction of unseen, combinatorial, and dose- or efficacy-dependent responses of various perturbation types. Across benchmarks spanning genetic and chemical perturbations, X-Pert consistently demonstrates superior performance at both gene and pathway levels. Moreover, its unified representation space enables downstream analyses such as perturbation retrieval and drug–gene association discovery. By integrating data across perturbation types, experimental platforms, and cell contexts, X-Pert establishes a versatile and generalizable foundation for in silico perturbation, enabling broad applications in biological and therapeutic discovery. ### Competing Interest Statement The authors have declared no competing interest. The National Key R&D Program of China, 2025YFC3409300 National Natural Science Foundation of China, 62373210, 62433001, 92470105 Tsinghua-Toyota Joint Research Institute Inter-Disciplinary Program, 20243930093
Single-cell resolution spatial transcriptomics (ST) provides a great opportunity to explore the complex cellular contents in different tissues. In this study, we introduce SpatialZoomer, a spectral graph-based method that applies a set of low-pass filters to efficiently extract spatial molecular features from ST data at multiple resolutions or scales, including the single cell scale, the niche scale with dozens of closely interacting cells, and the domain scale with spatially-organized cell contents. The corresponding "critical" scales can be automatically identified by partitioning a cross-scale similarity map. Results show that SpatialZoomer can identify disease-progression signals at specific scales in Alzheimer's Disease, and spatial context dependent cell subtypes in tumor microenvironment. The extracted multi-scale features also uncovered spatially heterogeneous niches in cancer. SpatialZoomer also takes advantage of high computational efficiency and low hardware requirements.
BACKGROUND:The human lung is a highly complex organ characterised by extensive cellular heterogeneity, making it susceptible to a broad range of diseases. Single-cell transcriptomics has shed light on disease-specific cellular features, but previous studies have been fragmented, limiting a unified understanding of cellular mechanisms across various lung diseases. Furthermore, high-resolution reference atlases for lung cells are lacking, impeding the effective integration of spatial omics data and the exploration of shared pathogenic mechanisms. METHODS:We constructed uniLUNG, the most comprehensive single-cell RNA sequencing atlas of the human lung, by integrating 62 published datasets comprising 9.2 million cells from 1807 donors across health and 17 disease conditions. We leveraged this high-resolution cell atlas to seek distinct cell types across diverse lung pathologies. By integrating spatial transcriptomics, we also identified transitional cell populations in lung cancer, and provided new insights into tumour evolution and the associated microenvironment. FINDINGS:We present a comprehensive lung cell atlas, encompassing cellular data across major lung diseases and health states. Using this resource, we identified distinct cell populations, such as Lym-monocytes and T-like B cells, which are specifically enriched in certain lung diseases and linked to immune dysregulation. Furthermore, our spatially resolved multi-omics analysis revealed a transitional malignant subpopulation, NSCLC-like SCLC, which plays a key role in the transformation from non-small cell lung cancer (NSCLC) to small cell lung cancer (SCLC), driving tumour microenvironment remodelling. INTERPRETATION:We offered a high-resolution, cross-disease lung cell reference that uncovers distinct cell types and cellular transitions critical to disease progression and therapeutic resistance. This resource provides essential insights into lung disease mechanisms and has important implications for the development of targeted therapeutic strategies, particularly in the context of lung cancer. FUNDING:This work was supported by the grants of National Key R&D Program of China (2021YFF1200900 and 2021YFF1200903), National Natural Science Foundation of China (92474107), Guangdong Basic and Applied Basic Research Foundation of China (2022B1515120077), Major Project of Guangzhou National Laboratory of China (GZNL2024A01003), and Support Scheme of Guangzhou for Leading Talents in Innovation and Entrepreneurship of China (2020007).
Background: Atrial fibrillation Better Care (ABC) pathway is recommended by guidelines on atrial fibrillation (AF) and exerts a protective role against adverse outcomes of AF patients. We hypothesize that cluster analysis, an unsupervised machine learning technique, could comprehensively evaluate multiple clinical characteristics of patients, and define clusters of ABC criteria efficacy in patients with AF. Methods: We used data from an observational cohort that included 2,016 patients with AF. We utilized 46 baseline variables for cluster analysis and got the optimal clusters through the K-prototypes algorithm. We evaluated the management patterns and adverse outcomes of identified phenotypes. We assessed the effectiveness of the ABC criteria in reducing adverse outcomes of these phenotypes. Results: Cluster analysis identified AF patients into three distinct groups with markedly different clinical characteristics and outcomes. The clusters were as followed: Cluster 1, old patients with atherosclerotic-comorbidities (n = 964); Cluster 2, young females with valve-comorbidities (n = 407), and Cluster 3, paroxysmal AF patients with low comorbidities (n = 644). The clusters showed significant differences in MACNE, all-cause death, stroke, cardiovascular death, and hospital readmission for heart failure. All clusters showed that full adherence to the ABC pathway was associated with a significant reduction in the risk of MACNE (all P< 0.05). Adherence to the different ‘A’/’B’/’C’ criteria alone showed differential clinic impact for three clusters. Conclusion: Cluster analysis of the Chinese AF cohort further elucidated the heterogeneity of AF. We proposed suggestions for optimizing risk stratification and integrated management of AF patients.
The liver performs several vital functions such as metabolism, toxin removal, and glucose storage through the coordination of various cell types. With the recent breakthrough of the single-cell/single-nucleus RNA-seq (sc/snRNA-seq) techniques, there is a great opportunity to establish a reference cell map of the liver at single-cell resolution with transcriptome-wise features. In this study, we build a unified liver cell atlas uniLIVER (http://lifeome.net/database/uniliver) by integrative analysis of a large-scale sc/snRNA-seq data collection of normal human liver with 331,125 cells and 79 samples from 6 datasets. Moreover, we introduce LiverCT, a machine learning based method for mapping any query dataset to the liver reference map by introducing the definition of "variant" cellular states analogous to the sequence variants in genomic analysis. Applying LiverCT on liver cancer datasets, we find that the "deviated" states of T cells are highly correlated with the stress pathway activities in hepatocellular carcinoma, and the enrichments of tumor cells with the hepatocyte-cholangiocyte "intermediate" states significantly indicate poor prognosis. Besides, we find that the tumor cells of different patients have different zonation tendencies and this zonation tendency is also significantly associated with the prognosis. This reference atlas mapping framework can also be extended to any other tissues.
Vision-language models (VLMs) exhibit strong zero-shot generalization on natural images and show early promise in interpretable medical image analysis. However, existing benchmarks do not systematically evaluate whether these models truly reason like human clinicians or merely imitate superficial patterns. To address this gap, we propose DrVD-Bench, the first multimodal benchmark for clinical visual reasoning. DrVD-Bench consists of three modules: Visual Evidence Comprehension, Reasoning Trajectory Assessment, and Report Generation Evaluation, comprising a total of 7,789 image-question pairs. Our benchmark covers 20 task types, 17 diagnostic categories, and five imaging modalities-CT, MRI, ultrasound, radiography, and pathology. DrVD-Bench is explicitly structured to reflect the clinical reasoning workflow from modality recognition to lesion identification and diagnosis. We benchmark 19 VLMs, including general-purpose and medical-specific, open-source and proprietary models, and observe that performance drops sharply as reasoning complexity increases. While some models begin to exhibit traces of human-like reasoning, they often still rely on shortcut correlations rather than grounded visual understanding. DrVD-Bench offers a rigorous and structured evaluation framework to guide the development of clinically trustworthy VLMs.
Node importance estimation (NIE) is the task of inferring the importance scores of the nodes in a graph. Due to the availability of richer data and knowledge, recent research interests of NIE have been dedicated to knowledge graphs (KGs) for predicting future or missing node importance scores. Existing state-of-the-art NIE methods train the model by available labels, and they consider every interested node equally before training. However, the nodes with higher importance often require or receive more attention in real-world scenarios, e.g., people may care more about the movies or webpages with higher importance. To this end, we introduce Label Informed ContrAstive Pretraining (LICAP) to the NIE problem for being better aware of the nodes with high importance scores. Specifically, LICAP is a novel type of contrastive learning (CL) framework that aims to fully utilize continuous labels to generate contrastive samples for pretraining embeddings. Considering the NIE problem, LICAP adopts a novel sampling strategy called top nodes preferred hierarchical sampling to first group all interested nodes into a top bin and a nontop bin based on node importance scores, and then divide the nodes within the top bin into several finer bins also based on the scores. The contrastive samples are generated from those bins and are then used to pretrain node embeddings of KGs via a newly proposed predicate-aware graph attention networks (PreGATs), so as to better separate the top nodes from nontop nodes, and distinguish the top nodes within the top bin by keeping the relative order among finer bins. Extensive experiments demonstrate that the LICAP pretrained embeddings can further boost the performance of existing NIE methods and achieve new state-of-the-art performance regarding both regression and ranking metrics. The source code for reproducibility is available at https://github.com/zhangtia16/LICAP.