The discrepancy between computational modeling and experimental performance remains a major challenge in protein engineering. We present ORI (Ontology Reinforcement Iteration), a scalable framework integrating ontology-conditioned decoding with reinforcement learning from experimental feedback (RLWF). ORI leverages structured ontologies as semantic prompts to impose multi-level constraints, enabling controllable and interpretable protein generation. A closed-loop iterative workflow-comprising generation, experimental measurement, and model updating-enables continuous optimization under real-world objectives. We demonstrate ORI's practical applicability through diverse tasks, including enzymatic activity optimization, thermal stability enhancement, and multifunctional protein engineering. Using this framework, we engineer variants with substantial improvements over natural baselines, such as a lysozyme with 100-fold higher activity, a chitinase stable at 85 °C, and dual-function enzymes exhibiting both lysozyme and chitinase activities. These results establish ORI as a robust technical platform for efficient, multi-objective protein engineering in real-world experimental settings.
Artificial intelligence (AI)-driven histopathological image analysis has shown significant advantages for disease diagnosis, prognosis, and treatment planning, and it is receiving growing attention in modern healthcare. Due to the gigapixel size of whole slide images (WSIs), multiple instance learning (MIL) methods are widely employed in their analysis. Existing MIL approaches primarily rely on either instance-level or bag-level supervision, each facing challenges related to noisy pseudo-labels and suboptimal feature aggregation, respectively. In this paper, we present a novel MIL method for WSI analysis, termed CIB-MIL, which integrates collaborative instance-level and bag-level supervision. We introduce a label disambiguation module within the instance-level supervision channel that employs a noisy-label learning strategy to refine instance pseudo-labels and mitigate the impact of noisy labels. Additionally, we propose a collaborative supervision framework that promotes communication and interaction between the attention mechanism in the bag-level supervision channel and the pseudo-label mechanism in the instance-level supervision channel, enabling cooperative optimization of supervision in both channels. Extensive experiments conducted on five datasets, including three public datasets and two in-house datasets, demonstrate the state-of-the-art performance of CIB-MIL. The code is available at https://github.com/TencentAILabHealthcare/CIB-MIL.
Prioritizing T-cell receptor (TCR) candidates for defined peptide-HLA targets is an important step in TCR-based immunotherapy development, but it still relies heavily on laborious and expensive experimental screening. Recent advancements in generative artificial intelligence have demonstrated promising power in protein design and engineering. In this regard, we propose a pre-trained transformer model, termed Epitope-Receptor-Transformer (ERTransformer), for the epitope-conditioned generation of candidate TCR β-chain CDR3 sequences. ERTransformer is built on EpitopeBERT and ReceptorBERT, which are trained using 1.9 million epitope sequences and 33.1 million TCR sequences, respectively. To demonstrate the model capability, we generate 1,000 candidate TCR β-chain CDR3 sequences for each of the five epitopes with known natural TCRs. The generated candidates show low sequence similarity to natural TCR β-chains while retaining plausible CDR3 length, amino-acid composition, and conservative substitution patterns. We further conduct wet-lab experiments using flow cytometry in defined TCR/pMHC contexts and find that the level of T cell activation induced by selected artificial TCRs is either comparable to or even surpasses that of natural ones. Our work suggests that ERTransformer can expand and prioritize candidate TCR β-chain CDR3 sequences for downstream experimental screening in defined peptide-HLA and TCR-chain contexts.
Homology search plays a fundamental role in computational biology, enabling the identification of evolutionary relationships and functional similarities among biological sequences. However, current homology search methods, including BLAST, Foldseek and MMseqs2, often struggle to efficiently and accurately process the vast scale of biological databases. Here we introduce ERAST (efficient retrieval-augmented search tool), a solution designed to handle approximately 1 billion biological sequences within the largest vector database to date. ERAST combines large language models and vector database technology to provide both efficient and precise searches for homologous biological sequences. It enhances search quality by integrating preretrieval, retrieval and postretrieval optimization stages, and supports both nucleotide and protein sequences. Through advanced indexing techniques, fine-grained segmentation and metadata integration, ERAST achieves better precision while operating approximately 50 times faster than Foldseek and 50,000 times faster than TM-align. This performance allows ERAST to conduct accurate searches against billions of biological sequences in mere milliseconds. The vector database integrated with ERAST can be accessed at https://ai4s.tencent.com/erast .
Rheumatoid arthritis (RA) progression is driven by the pathogenic transformation of stromal fibroblast-like synoviocytes (FLSs). Using artificial intelligence-guided analysis of a synovial single-cell transcriptomic dataset, we identify O-GlcNAc transferase (OGT) as a key regulator enriched in RA-FLSs. Gain- and loss-of-function studies demonstrate that OGT is necessary to drive aggressive FLS phenotypes and exacerbate experimental arthritis. Mechanistically, OGT-mediated O-GlcNAcylation stabilizes SAP130, promoting histone deacetylation at the BTG2 promoter. This facilitates the deposition of repressive histone methylation marks, leading to BTG2 epigenetic silencing and ultimately fueling FLS aggression. To translate this mechanistic insight, we develop an FLS-targeted proteolysis-targeting chimera (PROTAC) by conjugating the OGT inhibitor OSMI-1 to the AS1411 aptamer, which selectively binds surface nucleolin (NCL) overexpressed on pathogenic FLSs and, upon internalization, employs NCL as a molecular bridge to recruit the E3 ligase MDM2 for cell-selective OGT degradation. This PROTAC suppresses FLS pathogenicity, attenuates arthritis in mice, and demonstrates additive therapeutic benefit when combined with the TNF-α inhibitor etanercept. Collectively, our work establishes a translational paradigm that progresses from AI-driven target discovery and mechanistic elucidation to the rational design of a cell-type-specific degradation therapy, offering a strategy to overcome the stromal-driven therapeutic barrier in RA.
Experimentally validated prospective, blinded benchmarks are needed to separate durable advances from hype in computational antibody design. Here AIntibody, a challenge inspired by the Critical Assessment of Structure Prediction, tests 511 artificial intelligence (AI)-designed or predicted antibodies from 29 organizations on three tasks: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain complementarity-determining region 3 (HCDR3) clusters of a selection output and CDR design of proteins not included in a selection output. Validated with diverse experimental assays, several groups produced developable antibodies with affinities <100 pM. However, these successes were exceptions that did not transfer across tasks. Affinity-matured antibodies were modeled effectively. Except for one model, predicting high-affinity clones from clustered HCDR3 datasets was worse than random clone picking. Out-of-library design was highly variable for most method submissions, with many failing to outperform standard selections. The AIntibody challenge shows that AI can optimize antibodies in defined, biologically grounded regimes, in addition to highlighting critical gaps including affinity prediction and library-inspired antibody design and cross-task generalization.
Antibody optimization is a fundamental challenge, and the identification of antibody-antigen interactions is crucial in the optimization process. However, current methods cannot accurately predict antibody-antigen interactions due to the lack of all atom modeling, thus being unable to improve the time-consuming and costly traditional optimization techniques. We present InterAb, a novel model developed for predicting antibody-antigen interactions and optimizing antibodies through all-atom modeling and antibody language models. Leveraging the proposed all-atom modeling approach, AtomInter, and pretrained antibody language models, InterAb outperforms existing methods in predicting antibody specificity and antibody-antigen binding affinity. In the antibody library we constructed, InterAb successfully identified antibodies capable of binding to influenza A virus. An antibody optimization framework, InterAb-Opt, was further developed for the optimization of broadly neutralizing antibodies. For R1-32 antibody, biolayer interferometry results reveal that 85%, 80%, 90%, and 67.5% of the 40 optimized antibodies exhibit enhanced binding affinities to wild-type SARS-CoV-2, Lambda, BQ.1.1, and EG.5.1, respectively, with a maximum improvement of up to 96-fold. For the newly emerging BA.2.86 and KP.3, 55% and 52.5% of the optimized antibodies notably transition from non-binding to binding. Neutralization assays demonstrated that the optimized antibody exhibited enhanced neutralization activity across multiple targets, highlighting the capability of InterAb-Opt in engineering broadly neutralizing antibodies. This technology enables precise analysis of antibody-antigen interactions and optimization of broadly neutralizing antibodies, holding promise for addressing challenges in immune evasion and vaccine design. ### Competing Interest Statement The authors have declared no competing interest.
Mutually exclusive gene expression, where gene pairs are expressed in strict alternation within individual cells, reflects fundamental inter-gene regulatory mechanisms and can reveal shifts in transcriptional programs during development or disease. Detecting such patterns is critical for resolving rare cellular subpopulations, temporally discrete states along pseudotime, and spatially segregated neighborhoods in single-cell and spatial multi-omics data. However, the sparsity and dropout inherent to single-cell data make mutually exclusive expression difficult to detect, leading conventional feature selection methods to overlook subtle yet functionally important genes. We present MULE, an unbiased framework that systematically organizes collective mutual exclusivity into a hierarchical taxonomy. Applying MULE to cardiac datasets, we uncovered robust upregulation of SPOCK1 and SLC6A6 in dilated cardiomyopathy, previously obscured by the inability to resolve pathological cardiomyocytes. In vivo and in vitro experiments demonstrated that stress-induced SLC6A6 upregulation serves a cardiomyocyte self-protective mechanism. Taurine supplementation reduced oxidative stress, restored calcium homeostasis, prevented cell death, and improved cardiac function post-injury. These findings elucidate a novel cardiomyocyte stress response and highlight the therapeutic promise of taurine supplementation for the treatment of dilated cardiomyopathy.
Accurate modeling of antibody-antigen complex structures holds significant potential for advancing biomedical research and the design of therapeutic antibodies. Compared to general proteins, progress in antibody structure prediction and design has been slow, and antibody discovery is still based on time-consuming animal immunization or library screening methods. Here, we present tFold System, a high-throughput computational workflow that integrates antibody structure prediction (tFold-Ab), antibody-antigen complex modeling (tFold-Ag), structure-guided virtual screening, and de novo epitope-specific antibody design. Using this system, we de novo design monoclonal antibodies (mAbs) against four therapeutically relevant antigens: influenza hemagglutinin (Flu A), PD-1, PD-L1, and SARS-CoV-2 RBD (SC2RBD). Experimental validation by surface plasmon resonance (SPR) following high-throughput screening via phage display shows the designed antibodies achieve nanomolar binding affinities and precise epitope targeting, demonstrating the efficiency of the integrated computational-experimental pipeline. Our results demonstrate that tFold System overcomes key limitations of existing methods by enabling rapid, high-throughput antibody discovery against user-defined epitopes.
T cell receptors(TCRs)serve key roles in the adaptive immune system by enabling recognition and response to pathogens and irregular cells.Various methods have been developed for TCR construction from single-cell RNA sequencing(scRNA-seq)datasets,each with its unique characteristics.Yet,a comprehensive evaluation of their relative performance under different conditions remains elusive.In this study,we conducted a benchmark analysis utilizing experimental single-cell immune profiling datasets.Additionally,we introduced a novel simulator,YASIM-scTCR(Yet Another SIMulator for single-cell TCR),capable of generating scTCR-seq reads containing diverse TCR-derived sequences with different sequencing depths and read lengths.Our results consistently showed that TRUST4 and MiXCR outperformed others across multiple datasets,while DeRR demonstrated considerable accu-racy.We also discovered that the sequencing depth inherently imposes a critical constraint on successful TCR construction from scRNA-seq data.In summary,we present a benchmark study to aid researchers in choosing the appropriate method for reconstructing TCRs from scRNA-seq data.
Antibodies are indispensable components of the immune system, yet the design of high-affinity antibodies remains a time-consuming and experimentally intensive process. To address this challenge, we present IgGM, a novel generative foundation model designed to accelerate high-affinity antibody engineering. IgGM learns the complex relationships underlying the binding interactions between antigens and antibodies, as well as the mapping between antibody sequences and structures. By conditioning on different inputs, IgGM supports a wide range of antibody design tasks, including complex structure prediction, inverse design, affinity maturation, framework optimization, humanization, and de novo antibody design. It is compatible with both conventional antibodies and nanobodies, and allows user-defined CDR loop lengths for flexible design. To prioritize candidates, we introduce a frequency-based computational screening strategy that enhances design efficiency. Extensive evaluation through both in silico benchmarks and in vitro experiments across diverse antigens such as PD-L1, Protein A, TNF- α , IL-33, SARS-CoV-2 RBD and its variants demonstrates that IgGM consistently generates antibodies or nanobodies with high measured affinity. These results underscore IgGM’s versatility and effectiveness as a powerful tool for next-generation antibody discovery and optimization. ### Competing Interest Statement The authors have declared no competing interest.
Alpha-beta T cell receptor (alpha-beta TCR) recognition of peptide-major histocompatibility complexes (pMHCs) is a cornerstone of the adaptive immune system. Fast and accurate modeling of TCR-pMHC structures is crucial for understanding TCR recognition of pMHCs at the molecular level, which is essential for the development of TCR-based therapeutics and vaccines. Despite significant interest, this challenge remains unresolved due to the diversity of TCR-pMHC interactions and limited structural data. Here, we present tFold-TCR, a high-throughput, end-to-end universal model for predicting three-dimensional (3D) atomic-level structures of TCR-pMHC complexes, capable of predicting TCRs of different classes and MHC structures from diverse systems. tFold-TCR leverages a specially trained, protein-protein interaction-sensitive large protein language model to extract intra- and inter-chain residue contact information and evolutionary relationships, bypassing the need for multiple sequence alignment (MSA) searches. It also features innovative structure prediction and flexible docking modules to enhance accuracy, particularly for interacting contacts. Compared to existing methods, including AlphaFold-3, tFold-TCR demonstrates a 30.7% increase in prediction success rate evaluated by DockQ and is over 25 times faster. These advancements enable large-scale structural characterization of TCRs and their interactions with pMHCs. Utilizing this capability, we constructed TCRStructDB, the largest database of TCR-pMHC structures to date, encompassing 2.2 million TCRs, 0.8 million pMHCs, and 45,000 TCR-pMHC complexes. TCRStructDB provides unprecedented insights into one of the most diverse receptor-ligand interactions in biology. ### Competing Interest Statement The authors have declared no competing interest.
Protein Language Models (PLMs), pre-trained on extensive evolutionary data from natural proteins, have emerged as indispensable tools for protein design. While powerful, PLMs often struggle to produce proteins with precisely specified functionalities or properties due to inherent challenges in controlling their outputs. In this work, we investigate the potential of Activation Steering, a technique originally developed for controlling text generation in Large Language Models (LLMs), to direct PLMs toward generating protein sequences with targeted properties. We propose a simple yet effective method that employs activation editing to steer PLM outputs, and extend this approach to protein optimization through a novel editing site identification module. Through comprehensive experiments on lysozyme-like sequence generation and optimization, we demonstrate that our methods can be seamlessly integrated into both auto-encoding and autoregressive PLMs without requiring additional training. These results highlight a promising direction for precise protein engineering using foundation models.
Recent advances in single-cell technology enable the simultaneous capture of T cell receptor (TCR) sequences and gene expression (GEX), providing an integrated view of T cell function. However, linking TCRαβ information and T cell phenotypes at the population level to elucidate their disease association remains an unaddressed gap. Here, by constructing a large-scale reference of paired single-cell RNA/TCR sequencing (scRNA/TCR-seq) comprising more than 2 million T cells from 70 studies, 1017 biological samples, 583 individuals, and 46 disease conditions, along with their single-cell transcriptome, full-length paired TCR, and human leukocyte antigen (HLA) genotypes, we revealed the intrinsic features of germline-encoded TCR-major histocompatibility complex (MHC) restriction in CD4+/CD8+ lineages. We also observed widely existing public TCRαβs across the population, associated with higher clonal expansion levels and shared HLA alleles. The most publicly shared TCRs are likely to target epitopes from common viruses, such as Epstein-Barr virus (EBV), cytomegalovirus (CMV), and influenza A virus (IAV). Furthermore, we introduced TCR-DeepInsight, a computational framework to identify HLA-shared and disease-associated TCRαβ clusters that exhibit similar TCR sequence and GEX profiles, extensible for researchers to incorporate their data with our reference and characterize potentially functional TCRs. In summary, our work presents a panoramic scTCRαβ reference and computational methods for TCR study.
One individual human’s immune repertoire consists of a huge set of adaptive immune receptors at a certain time point, representing the individual's adaptive immune state. Immune repertoire classification and associated receptor identification have the potential to make a transformative contribution to the development of novel vaccines and therapies. The vast number of instances and exceedingly low witness rate pose a great challenge to the immune repertoire classification, which can be formulated as a Massive Multiple Instance Learning (MMIL) problem. Traditional MIL methods, at both bag-level and instance-level, confront the issues of substantial computational burden or supervision ambiguity when handling massive instances. To address these issues, we propose a novel label disambiguation-based multimodal massive multiple instance learning approach (LaDM³IL) for immune repertoire classification. LaDM³IL adapts the instance-level MIL paradigm to deal with the issue of high computational cost and employs a specially-designed label disambiguation module for label correction, mitigating the impact of misleading supervision. To achieve a more comprehensive representation of each receptor, LaDM³IL leverages a multimodal fusion module with gating-based attention and tensor-fusion to integrate the information from gene segments and amino acid (AA) sequences of each immune receptor. Extensive experiments on the Cytomegalovirus (CMV) and Cancer datasets demonstrate the superior performance of the proposed LaDM³IL for both immune repertoire classification and associated receptor identification tasks. The code is publicly available at https://github.com/Josie-xufan/LaDM3IL.
Current cancer vaccines using T cell epitopes activate antitumor T cell immunity through dendritic cell/macrophage-mediated antigen presentation, but they lack the ability to promote B/CD4 T cell crosstalk, limiting their anticancer efficacy. We developed antigen-clustered nanovaccine (ACNVax) to achieve long-term tumor remission by promoting B/CD4 T cell crosstalk. The topographic features of ACNVax were achieved using an iron nanoparticle core attached with an optimal number of gold nanoparticles, where the clusters of HER2 B/CD4 T cell epitopes were conjugated on the gold surface with an optimal intercluster distance of 5-10 nm. ACNVax effectively trafficked to lymph nodes and cross-linked with BCR, which are essential for stimulating B cell antigen presentation-mediated B/CD4 T cell crosstalk in vitro and in vivo. ACNVax, combined with anti-PD-1, achieved long-term tumor remission (>200 days) with 80% complete response in mice with HER2+ breast cancer. ACNVax not only remodeled the tumor immune microenvironment but also induced a long-term immune memory, as evidenced by complete rejection of tumor rechallenge and a high level of antigen-specific memory B, CD4, and CD8 cells in mice (>200 days). This study provides a cancer vaccine design strategy, using B/CD4 T cell epitopes in an antigen clustered topography, to achieve long-term durable anticancer efficacy through promoting B/CD4 T cell crosstalk.
Accurate prediction of antibody-antigen complex structures holds significant potential for advancing biomedical research and the design of therapeutic antibodies. Currently, structure prediction for protein monomers has achieved considerable success, and promising progress has been made in extending this achievement to the prediction of protein complexes. However, despite these advancements, fast and accurate prediction of antibody-antigen complex structures remains a challenging and unresolved issue. Existing end-to-end prediction methods, which rely on homology and templates, exhibit sub-optimal accuracy due to the absence of co-evolutionary constraints. Meanwhile, conventional docking-based methods face difficulties in identifying the contact interface between the antigen and antibody and require known structures of individual components as inputs. In this study, we present a fully end-to-end approach for three-dimensional (3D) atomic-level structure predictions of antibodies and antibody-antigen complexes, referred to as tFold-Ab and tFold-Ag, respectively. tFold leverages a large protein language model to extract both intra-chain and inter-chain residue-residue contact information, as well as evolutionary relationships, avoiding the time-consuming multiple sequence alignment (MSA) search. Combined with specially designed modules such as the AI-driven flexible docking module, it achieves superior performance and significantly enhanced speed in predicting both antibody (1.6% RMSD reduction in the CDR-H3 region, thousand times faster) and antibody-antigen complex structures (37% increase in DockQ score, over 10 times faster), compared to AlphaFold-Multimer. Given the performance and speed advantages, we further extend the capability of tFold for structure-based virtual screening of binding antibodies, as well as de novo co-design of both structure and sequence for therapeutic antibodies. The experiment results demonstrate the potential of tFold as a high-throughput tool to enhance processes involved in these tasks. To facilitate public access, we release code and offer a web service for antibody and antigen-antibody complex structure prediction, which is available at . ### Competing Interest Statement The authors have declared no competing interest.
CD8+ T cells exhibit remarkable phenotypic diversity in inflammation and cancer. However, a comprehensive understanding of their clonal landscape and dynamics remains elusive. Here we introduce scAtlasVAE, a deep-learning-based model for the integration of large-scale single-cell RNA sequencing data and cross-atlas comparisons. scAtlasVAE has enabled us to construct an extensive human CD8+ T cell atlas, comprising 1,151,678 cells from 961 samples across 68 studies and 42 disease conditions, with paired T cell receptor information. Through incorporating information in T cell receptor clonal expansion and sharing, we have successfully established connections between distinct cell subtypes and shed light on their phenotypic and functional transitions. Notably, our approach characterizes three distinct exhausted T cell subtypes and reveals diverse transcriptome and clonal sharing patterns in autoimmune and immune-related adverse event inflammation. Furthermore, scAtlasVAE facilitates the automatic annotation of CD8+ T cell subtypes in query single-cell RNA sequencing datasets, enabling unbiased and scalable analyses. In conclusion, our work presents a comprehensive single-cell reference and computational framework for CD8+ T cell research. scAtlasVAE is a deep learning-based model for cross-atlas integration. Here it enables the development of a large-scale human CD8+ T cell atlas with integrated T cell receptor data.