To jointly capture antigen-binding affinity and paired heavy-light-chain sequences at scale has remained a bottleneck for monoclonal antibody discovery. Here, we present Antigen Affinity BCR-seq (AAB-seq), a high-throughput single-cell sequencing platform that can obtain the relative antibody-antigen affinity of thousands of paired native BCR sequences. AAB-seq employs dual-labeled antigens and DNA-barcoded anti-light-chain antibodies to compute an AAB score that is proportional to antibody-antigen binding strength. Integrated with a rapid, low-cost direct cloning workflow, it enables affinity-guided antibody retrieval without de novo antibody gene synthesis. Validated against ovalbumin and SARS-COV-2 RBD, AAB-seq discovered potent antibodies whose AAB score correlates strongly with ELISA, including novel SARS-COV-2 neutralizing antibodies with potent effector functions. Together, AAB-seq accelerates antibody screening and potentially provides large-scale sequence-affinity datasets for machine learning-driven therapeutic antibody design and development.
In the realm of high-dimensional single-cell sequencing data analysis, the accurate measurement of similarity between cells is pivotal. However, conventional metrics like Euclidean distance after L-1 -normalization may fail by losing distinguishable information when handling high-dimensional data, where the distance between different observations gradually converges to a shrinking interval. In this article, we use distance entropy to quantify the amount of information contained in the distances, and discuss the influence of normalization by different $p$ -norrns and the defect of Euclidean distance. We discover that observation differences are better preserved when normalizing data by a higher p-norm and using geodesic distance rather than Euclidean distance as the similarity measurement. We further identify that L-2-normalization onto the hypersphere is often sufficient in preserving delicate differences even in relatively high dimensional data while maintaining computational efficiency. Subsequently, we present hypersphere t-distributed stochastic neighbor embedding (HS-SNE), a hypersphere-representation-system-based augmentation to t-distributed stochastic neighbor embedding (t-SNE), which effectively addresses the intricacy of high-dimensional data visualization and similarity measurement. Our results on multiple single-cell sequencing datasets show that this hypersphere representation system has improved resolution to identify more subtle differences between high-dimensional data points, while balancing distance entropy preservation and computational efficiency.
Protein sequence and structure similarity-based search is an important task, which underpins protein annotation, evolutionary analysis, large-scale functional inference, and the exploration of the protein “dark space”. The rapid growth of sequence and predicted structure databases has spurred diverse search methods, yet their evaluation remains limited to fold-level similarity and inconsistent benchmarking protocols. We present a comprehensive benchmark for protein sequence and structure search. Using this framework, we evaluate 14 representative methods spanning sequence alignment, structure alignment, and representation-based approaches across multiple biologically relevant scenarios. Our results show pronounced and context-dependent differences among methods. Structure alignment methods excel at detecting fold-level and geometric similarity, while representation-based searching approaches show advantages in capturing functional similarity under low sequence identity and robustness to predicted structures. Notably, all evaluated methods show limited effectiveness on intrinsically disordered proteins. This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.
Geometry-preserving dimension reduction is critical for single-cell transcriptomics, where low-dimensional distances should reflect biological divergence between cell types along the transcriptomic manifold. Due to inadequate metrics, the global structure is not sufficiently preserved in the low-dimensional manifold in standard dimension reduction regimes. We model RNA counts as Multinomial samples, leveraging their hierarchical closure property: gene-level counts refine functional gene-group counts via nested Multinomial distributions. Extending Chentsov's Theorem, we show that the Fisher-Rao metric on coarse (gene-group) and fine (gene) statistical manifolds is isometric. Following this isometry property, we propose InfoGlobe, an information-preserving statistical manifold learning framework that projects cells from high-dimensional hyperspheres (full transcriptome) to low-dimensional hyperspheres (functional groups) while preserving information geometry. Embeddings on the low-dimensional sphere explicitly represent Multinomial distributions by functional gene groups. Benchmarks demonstrate superior preservation of local-and-global cell-type geodesic distances, automatic and robust gene-group discovery, nuanced cell subtype resolution without manual feature engineering and natural batch effect mitigation without explicit alignments.
Abstract Cellular identity and fate transitions are governed by continuous molecular processes that form dynamic trajectories within a high-dimensional transcriptomic landscape. Existing methods attempt to model these dynamics from two complementary perspectives: trajectory inference and velocity modeling. Ideally, velocity and trajectory are dual aspects of transcriptomic dynamics where velocity is tangent to trajectory everywhere. This inherent connection between velocity and trajectory is currently absent in transcriptomic analysis. Splicing velocity are precision-limited to inadequately-sequenced genes, while trajectory inference prioritizes the modeling of global trends while omitting local dynamics. This divergence breaks the geometric continuity between local velocities and global trajectories, hindering the reliable interpretation of developmental dynamics. To reconcile trajectory inference and RNA velocity, we introduce VeloTrace, a framework that unifies them through Neural Ordinary Differential Equations (NeuralODEs). VeloTrace learns a continuous-time velocity field whose integral curves constitute the trajectory itself, while ensuring that velocities are tangent to integral paths everywhere. Leveraging a splicing quality score, VeloTrace incorporates high-quality splicing velocity as partial supervision for velocity orientation and grounding. During optimization, VeloTrace incorporates a Monte Carlo multi–time-frame supervision strategy to ensure coherence between local and global trajectorys and suppress sequencing-induced stochastic diffusion. Through refining the velocity field and cell-specific parameters for pseudo-time, expression, and velocity, VeloTrace reconstructs a smooth, local-and-global-coherent velocity-vector-guided flow in the transcriptomic latent space. This strategy ensures a complementary integration of velocity and trajectory, imputing the transcriptional kinetics for genes of insufficient strength, whose kinetics cannot be accurately portrayed by splicing velocity. In simulation benchmarks, VeloTrace captured the transcriptional dynamics of all expressed genes, even those with inadequate sequencing coverage, producing velocity directions that were most consistent with the true direction and every-where tangential across the entire process, outperforming state-of-the-art methods, including scVelo, UniTVelo, VeloVI and scTour. VeloTrace uniquely reconciles RNA velocity and trajectory inference, creating a velocity field where each cell can infer past and future transitions from its current state. Moreover, VeloTrace extends reliable velocity estimation to a broader set of genes. When applied to mouse neural stem cell differentiation data, it successfully recovers dynamics of driver genes for two developmental lineages, including those with low expression, shedding light on their regulatory roles during differentiation. This unified framework lays the foundation for more accurate modeling of gene regulation and cell fate decisions in complex biological systems.
Background Metabolic rewiring influences macrophage functions in tumours. However, metabolic heterogeneity of macrophages in early‐stage lung adenocarcinoma (LUAD) has not been fully understood. Methods Based on single‐cell transcriptomic analysis on lung tissues from human subsolid pulmonary nodules and from Kras G12D and Kras G12D Tgfbr2 −/− mice, we unveiled alterations of macrophage states in LUAD. We applied cell migration and invasion assays to examine effects of ATP binding cassette subfamily A member 1 (ABCA1) in macrophages on tumour progression. In addition, using coculture system, we explored its role in macrophage polarisation. Results Major macrophage subsets underwent shifts during tumourigenesis, as alveolar macrophages reduced sharply and their interstitial counterparts increased. Of note, tumour‐associated macrophages (TAMs) are metabolically rewired, with a subset enhancing lipid efflux predominated in both human and murine tumour. It was characterised by high expression of the lipid transporter, ABCA1. Targeting ABCA1 by its inhibitor probucol suppressed TAM capacity to induce tumour cell migration and invasion. In addition, inhibiting ABCA1 was associated with dampened immunosuppression of TAMs by shifting them to a M1‐like phenotype. Conclusions We uncovered TAM heterogeneity in early‐stage LUAD and proposed ABCA1 as a potential target for metabolic rewiring, which manipulated macrophage polarisation states and its effects on tumour aggressiveness.
The precise delineation of cell types is fundamental to single-cell transcriptomics, yet current clustering pipelines often violate an axiomatic principle: hierarchical consistency. Existing methods measure cell-to-cell distances within a fixed global feature space, disregarding the fact that biological distinctions are inherently context-dependent lineage separation requires different gene programs than subtype resolution. Mathematically, this implies that the similarity metric itself should not be a static functional, but a pair-dependent energy functional evaluated within a specific Hilbert subspace determined by the biological comparison at hand. The challenge lies in the fact that allowing pair-dependent metrics typically destroys the global geometric consistency required for downstream analysis, unless the family of Hilbert subspaces is given strong biological structure. To resolve this geometric dilemma, we introduce GeCCo (Gene Co-expression Constructed identity), which constructs identities by projecting cells onto a rigorously derived hierarchy of gene programs. To construct this hierarchy, GeCCo first quantifies Boolean regulatory logic via the $\phi$ coefficient, and subsequently employs a greedy topological inference to organize genes based on their synergistic and antagonistic relationships. Benchmarking on human immune atlases demonstrates that GeCCo achieves superior hierarchical consistency, ensuring that globally inferred cell identities rigorously match locally refined subtypes. Furthermore, in pancreatic endocrine progenitors, GeCCo resolves a hidden mitotic bridge state, suggesting a concentrated division phase prior to differentiation. Ultimately, GeCCo shifts the paradigm from ad hoc clustering to programmatic cell typing, offering a mathematically grounded framework for scalable atlases of cellular discovery.
BACKGROUND:Early tumour vascular invasion contributes to cancer progression. Tip cells, a subset of tumour endothelial cells, significantly decline after anti-angiogenic therapy. However, their behaviour and the roles of their signature genes during early invasion are incompletely understood. METHODS:This study employed single-cell transcriptomic analysis and 10x Genomics Visium spatial transcriptomics on fresh lung tissues from patients with pulmonary nodules and from KrasG12D (K) and KrasG12DTgfbr2-/- (KT) mice. The role of plasma vesicle-associated protein (PLVAP), a tip cell marker, was further examined using survival databases, immunofluorescence, in vitro co-culture, cell migration, invasion assays and endothelial tube formation. RESULTS:Tip cell proportions were elevated in early-stage lung adenocarcinoma (LUAD) tissues and KT mice, with evidence suggesting they arise from capillaries type I. PLVAP expression was enriched in tumour endothelial cells, induced by TGFβ1, and negatively correlated with patient prognosis. Functionally, PLVAP promoted endothelial cell invasion, migration and angiogenesis, and regulated tumour cell invasiveness. Intercellular analysis revealed that some tip cells also expressed TGFβ1, which may act on adjacent tumour cells to enhance invasion during early tumour development. CONCLUSION:Tip cells increased during early LUAD progression and likely evolved from capillaries type I. Their marker PLVAP was associated with poor prognosis and pro-invasive endothelial behaviour. Tumour-secreted TGFβ1 upregulated PLVAP in endothelial cells, promoting angiogenesis and tumour invasion. Additionally, tip-cell-derived TGFβ1 may further stimulate tumour aggressiveness, highlighting a reciprocal interaction that contributes to early tumour progression. KEY POINTS:Tip cells expand during early LUAD progression and likely originate from capillary type I endothelial cells. Tumour-derived TGFβ1 induces PLVAP expression in endothelial cells, linking tumour signals to vascular activation. PLVAP enhances endothelial cell migration, invasion and angiogenic capacity. Endothelial PLVAP promotes tumour cell invasiveness, revealing a reciprocal endothelial-tumour interaction that drives early tumour progression.
As DNA data storage gains popularity, efficient trace reconstruction algorithms are crucial for fast decoding of data from noisy sequenced reads (or "traces"). Existing approaches, often adaptations of multiple sequence alignment or read correction methods, rely on strict assumptions of fixed error rates, showing limited generalizability to more complex datasets and with slower running times. We introduce a probabilistic formulation of the trace reconstruction problem by modeling traces as observations from a k-th order Markov chain. Instead of doing alignment, we identify the sequence most likely generated by the Markov chain as the consensus. This inspires bidirectional beam search (BBS), an algorithm that reconstructs the consensus in linear time with respect to its length. Experiments on multiple public Nanopore sequencing datasets demonstrate that BBS achieves top-tier accuracy while being approximately 20× faster than existing methods, showing its potential to enhance the efficiency and reliability of DNA data storage systems.
We present a multi-stage pipeline for BioLaySumm 2025 Subtask 1.1 that improves readability, relevance, and factuality. First, we select the top-5 relevant sections and generate summaries with BioBART. Next, we retrieve a Kshot demonstration using BGE embeddings to prompt Llama 3 8B and fine-tune it with LoRA. We then merge section summaries via a second BioBART pass. Finally, we apply reinforcement learning (PPO and GRPO) with a composite reward combining factuality (AlignScore, SummaC), relevance (ROUGE-L, BERTScore), and readability (LENS, FKGL, DCRS, CLI). On PLOS and eLife validation sets, our pipeline reduces DCRS from 9.23 to 8.56 and CLI from 12.98 to 12.65, and boosts AlignScore from 0.722 to 0.862, demonstrating balanced gains in lay-summary quality.
Prostate cancer (PCa) is a common and serious health issue among older men globally. Metabolic reprogramming, particularly involving lactate and mitochondria, plays a key role in PCa progression, but studies linking these factors to prognosis are limited. To identify novel prognostic markers of PCa based on lactate-mitochondria-related genes (LMRGs), RNA sequencing data and clinical information of PCa from The Cancer Genome Atlas (TCGA) and the cBioPortal database were used to construct a lactate-mitochondria-related risk signature. Here, we established a novel nine-LMRG risk signature for PCa, and Kaplan-Meier curves confirmed a worse prognosis for high-risk subgroups in the TCGA dataset. Meanwhile, a nomogram that effectively predicts the prognosis of PCa patients was also constructed. Next, close associations between the lactate-mitochondria-related signature and the immune microenvironment were examined to clarify the role of LMRGs in shaping the immune landscape. Furthermore, as the only lactate-related gene among the nine key prognostic risk genes, myeloperoxidase (MPO) was identified as a key factor that mediates lactate production in vitro and in vivo through attenuation of the glycolytic pathway. More importantly, MPO significantly inhibited PCa cell migration, invasion, and epithelial-mesenchymal transition (EMT), indicating its potential as an anticancer gene. Additionally, PCa with high MPO expression is highly sensitive to chemotherapeutic agents and mitochondrial inhibitors, highlighting its potential as an improved therapeutic strategy for PCa management.
ST-elevation myocardial infarction (STEMI) remains a leading cause of cardiovascular morbidity and mortality worldwide, and accurate early risk stratification is critical for implementing precision therapies in clinical practice. However, existing clinical risk scores and manually derived imaging biomarkers have limited accuracy in predicting post-STEMI outcomes. To address this gap, we developed DeepSTEMI, an end-to-end deep learning system that integrates multi-sequence cardiac magnetic resonance (CMR) images with clinical parameters for predicting 2-year major adverse cardiovascular events (MACE). The system comprised two key algorithmic modules: a U-Net module that automatically segments heart regions from raw CMR images and a Transformer-based module that predicted future cardiovascular events. DeepSTEMI was developed using a multicenter dataset (n = 610; 20,618 images) from STEMI patients enrolled in the EARLY-MYO-CMR registry (NCT03768453), with external validation performed in 334 patients (9944 images) from three independent cardiac centers. In external validation, DeepSTEMI demonstrated superior predictive performance compared to conventional clinical risk scores and manual CMR parameters (AUC 0.894, 95% CI: 0.823-0.965; overall accuracy 94.3%). The model identified high-risk patients who exhibited a 20-fold MACE risk compared to low-risk counterparts (HR 20.43, log-rank P < 0.001). SHapley Additive exPlanations (SHAP) analysis revealed that DeepSTEMI's predictive power stems from clinical-imaging synergy, enabling it to capture complex pathological patterns. DeepSTEMI achieved consistently superior performance over the Eitel score across all subgroups, with the greatest benefit observed in women (NRI 1.597) and in patients imaged 4-7 d post-STEMI (NRI 1.442). Overall, DeepSTEMI serves as an automated, scalable, and interpretable clinical copilot, which advances post-STEMI risk stratification beyond the limitations of current paradigms.
Single-cell and spatial transcriptomics enable high-resolution characterization of cellular states, but standard analyses often rely on Euclidean or log-transformed distances that distort cell-to-cell relationships. Euclidean distances on normalized counts overemphasize highly expressed genes, while log transformations amplify qualitative on/off differences and are sensitive to sequencing depth. To overcome these limitations, we introduce GAIA (Geometric Analysis from an Information Aspect), an information-geometric framework that models each cell as a multinomial distribution over genes. Distances between cells are measured using the Fisher-Rao metric, which reduces to angular distances on a unit hypersphere after a square-root transformation. This geodesic approach provides a principled, interpretable, and computationally efficient similarity measure. GAIA naturally reconciles qualitative and quantitative gene expression variation: subtle quantitative changes correspond to smooth displacements along the manifold, while qualitative transitions induce larger geodesic separations. It preserves robust and consistent cell-to-cell relationships, mitigates sequencing-depth effects and reduces the need for labor-intensive gene selection. In spatial transcriptomics, GAIA amplifies nuanced transcriptomic differences between spots, improving domain segmentation. Overall, GAIA offers a knowledge-lean, variance-stabilizing framework for analyzing single-cell and spatial transcriptomic data, enhancing the resolution of cell type and state identification.
In single-cell perturbation prediction, a central task is to forecast the effects of perturbing a gene unseen in the training data. The efficacy of such predictions depends on two factors: (1) the similarity of the target gene to those covered in the training data, which informs model (epistemic) uncertainty, and (2) the quality of the corresponding training data, which reflects data (aleatoric) uncertainty. Both factors are critical for determining the reliability of a prediction, particularly as gene perturbation is an inherently stochastic biochemical process. In this paper, we propose PRESCRIBE (PREdicting Single-Cell Response wIth Bayesian Estimation), a multivariate deep evidential regression framework designed to measure both sources of uncertainty jointly. Our analysis demonstrates that PRESCRIBE effectively estimates a confidence score for each prediction, which strongly correlates with its empirical accuracy. This capability enables the filtering of untrustworthy results, and in our experiments, it achieves steady accuracy improvements of over 3% compared to comparable baselines.
Modeling cellular dynamics from single-cell RNA sequencing (scRNA-seq) data is critical for understanding cell development and underlying gene regulatory relationships. Many current methods rely on single-cell velocity to obtain pseudotime, which can lead to inconsistencies between pseudotime and velocity. It is challenging to simultaneously infer cell pseudotime and gene interaction networks, especially in multi-branch differentiation scenarios. We present single-cell Piecewise Network (scPN), a novel high-dimensional dynamical modeling approach that iteratively extracts temporal patterns and inter-gene relationships from scRNA-seq data. To tackle multi-branch differentiation challenges, scPN models gene regulatory dynamics using piecewise gene-gene interaction networks, offering an interpretable framework for deciphering complex gene regulation patterns over time. Results on synthetic data and multiple scRNA-seq datasets demonstrate the superior performance of scPN in reconstructing cellular dynamics and identifying key transcription factors involved in development compared to existing methods. To the best of our knowledge, scPN is the first attempt at modeling that can recover pseudotime, velocity fields, and gene interactions all at once on multi-branch datasets.
This study aims to evaluate the feasibility of large language model (LLM) in answering pathology questions based on pathology reports (PRs) of colorectal cancer (CRC). Four common questions (CQs) and corresponding answers about pathology were retrieved from public webpages. These questions were input as prompts for Chat Generative Pretrained Transformer (ChatGPT) (gpt-3.5-turbo). The quality indicators (understanding, scientificity, satisfaction) of all answers were evaluated by gastroenterologists. Standard PRs from 5 CRC patients who received radical surgeries in Shanghai Changzheng Hospital were selected. Six report questions (RQs) and corresponding answers were generated by a gastroenterologist and a pathologist. We developed an interactive PRs interpretation system which allows users to upload standard PRs as JPG images. Then the ChatGPT's responses to the RQs were generated. The quality indicators of all answers were evaluated by gastroenterologists and out-patients. As for CQs, gastroenterologists rated AI answers similarly to non-AI answers in understanding, scientificity, and satisfaction. As for RQ1-3, gastroenterologists and patients rated the AI mean scores higher than non-AI scores among the quality indicators. However, as for RQ4-6, gastroenterologists rated the AI mean scores lower than non-AI scores in understanding and satisfaction. In RQ4, gastroenterologists rated the AI scores lower than non-AI scores in scientificity (P = 0.011); patients rated the AI scores lower than non-AI scores in understanding (P = 0.004) and satisfaction (P = 0.011). In conclusion, LLM could generate credible answers to common pathology questions and conceptual questions on the PRs. It holds great potential in improving doctor-patient communication.
With the development of social economy, the incidence of gout is increasing, which is closely related to people’s increasingly rich diet. Eating a diet high in purine, fat, sugar and low-fibre for a long time further aggravates gout by affecting uric acid metabolism. The renal metabolism mechanism of uric acid has been thoroughly studied. To find a new treatment method for gout, increasing studies have recently been conducted on the mechanism of intestinal excretion, metabolism and absorption of uric acid. The most important research is the relationship between intestinal microbiota and the risk of gout. Gut microbiota represent bacteria that reside in a host’s gastrointestinal tract. The composition of the gut microbiota is associated with protection against pathogen colonization and disease occurrence. This review focuses on how gut microbiota affects gout through uric acid and discusses the types of bacteria that may be involved in the occurrence and progression of gout. We also describe potential therapy for gout by restoring gut microbiota homeostasis and reducing uric acid levels. We hold the perspective that changing intestinal microbiota may become a vital method for effectively preventing or treating gout.
Distinguishable metric of similarity plays a fundamental role in unsupervised learning, particularly in manifold learning and high-dimensional data visualization tasks, by which differentiate between observations without labels. However, conventional metrics like Euclidean distance after L1-normalization may fail by losing distinguishable information when handling high-dimensional data, where the distance between different observations gradually converges to a shrinking interval. In this article, we discuss the influence of normalization by different p-norms and the defect of Euclidean distance. We discover that observation differences are better preserved when normalizing data by a higher p-norm and using geodesic distance rather than Euclidean distance as the similarity measurement. We further identify that L2-normalization onto the hypersphere is often sufficient in preserving delicate differences even in relatively high dimensional data while maintaining computational efficiency. Subsequently, we present HS-SNE (HyperSphere-SNE), a hypersphere-representation-system-based augmentation to t-SNE, which effectively addresses the intricacy of high-dimensional data visualization and similarity measurement. Our results show that this hypersphere representation system has improved resolution to identify more subtle differences in high-dimensional data, while balancing information preservation and computational efficiency.