Type 1 diabetes (T1D) is a T cell-mediated disease with a strong immunogenetic human leukocyte antigen (HLA) dependence. HLA allelic influence on the T cell receptor (TCR) repertoire shapes thymic selection and controls activation of diabetogenic clones yet remains largely unresolved in T1D. We sequenced the circulating TCRβ chain repertoire from 2250 HLA-typed participants across three cross-sectional cohorts, including individuals with T1D and healthy related and unrelated controls. We found that HLA risk alleles show higher restriction of TCR repertoires in individuals with T1D. We leveraged deep learning to identify T1D-associated TCR subsequence motifs that were also observed in independent TCR cohorts residing in pancreas-draining lymph nodes of individuals with T1D. Collectively, our data demonstrate T1D-related TCR motif enrichment based on genetic risk, offering a potential metric for autoreactivity and groundwork for TCR-based diagnostics and therapeutics.
ABSTRACT Protein aggregation is central to amyloid-related disorders and remains a major developability challenge for protein therapeutics. Over the past two decades, significant advances have been made to predict aggregation-prone regions (APRs) and estimate aggregation propensity in proteins and peptides. In contrast, the prediction of aggregation kinetics has received relatively less attention due to the limited availability and heterogeneity of experimental data. Consequently, aggregation propensities from APR prediction algorithms were widely accepted as a means to predict relative changes in the aggregation kinetics of proteins and mutants. Previous studies have demonstrated, using large-scale datasets, that aggregation propensity shows a weak or inconsistent correlation with aggregation kinetics. In the present study, we have integrated complementary state-of-the-art mechanistic and kinetic prediction tools for protein aggregation into a unified, user-friendly web framework entitled “Amylo-Pipe”. Amylo-Pipe also implements practical features that are especially useful for protein engineering, such as gatekeeper-residue mutational scanning to support the design of aggregation-resistant variants. By consolidating multiple prediction tasks in a single interface, Amylo-Pipe enables a more comprehensive assessment of aggregation behavior than APR-only workflows. The web server is freely accessible at: https://web.iitm.ac.in/bioinfo2/amylopipe/ .
Abstract Human epidermal growth factor receptor 2 (HER2) is an oncogenic receptor tyrosine kinase in breast cancer and other malignancies. A subset of HER2-positive tumours expresses 611-CTF-p95HER2, a tumour-specific, hyperactive truncated isoform associated with metastasis and treatment resistance that lacks most of the extracellular domain targeted by conventional HER2-directed antibodies. We previously developed NAZ-mAb (formerly known as Oslo-2), a monoclonal antibody against 611-CTF-p95HER2. Here, we describe a computational antibody-engineering workflow for designing variants of NAZ-mAb. Starting from the sequence alone, we modeled the NAZ-mAb–611-CTF-p95HER2 complex, generated a combinatorial mutational landscape using FoldX 5.0, and prioritized candidate variants using predicted interaction energy and developability criteria. Two variants representing distinct design strategies were selected for validation: an aromatic double mutant, NAZ-mAb v1 (L:S31W/L:H107W), and a conservative single mutant, NAZ-mAb v2 (L:S31M). Both variants were successfully expressed as recombinant IgGs; NAZ-mAb v2 achieved a five-fold higher recombinant expression yield than parental NAZ-mAb, while both variants retained antigen binding with a higher apparent signal than the parental antibody in indirect ELISA. However, Biacore two-state kinetic analysis revealed weaker affinities than the parental antibody (K D NAZ-mAb v1: 32.6 nM, NAZ-mAb v2: 9.45 nM vs. parental NAZ-mAb: 5.33 nM). These findings show that the computational workflow can generate experimentally tractable, antigen-engaging NAZ-mAb variants, while also highlighting the limitations of fixed-backbone interaction-energy ranking as a predictor of binding affinity and yield. This study provides a practical framework for computationally driven, developability-aware antibody optimization in the absence of experimental structural data.
Supervised machine learning models depend on training datasets containing positive and negative examples: dataset composition directly impacts model performance and bias. Given the importance of machine learning for immunotherapeutic design, we examined how different negative class definitions affect model generalization and rule discovery for antibody–antigen binding. Using synthetic-structure-based binding data, we evaluated models trained with various definitions of negative sets. Our findings reveal that high out-of-distribution performance can be achieved when the negative dataset contains more similar samples to the positive dataset, despite lower in-distribution performance. Furthermore, by leveraging ground-truth information, we show that binding rules associated with positive data change based on the negative data used. Validation on experimental data supported simulation-based observations. This work underscores the role of dataset composition in creating robust, generalizable and biology-aware sequence-based ML models. Negative data composition critically shapes machine learning robustness in sequence-based biological tasks. Training data composition and its implications are investigated on biological rule discoveries.
Light chain amyloidosis is a medical condition characterized by the aggregation of misfolded antibody light chains into insoluble amyloid fibrils in the target organs, causing organ dysfunction, organ failure, and death. Despite extensive research to understand the factors contributing to amyloidogenesis, accurately predicting whether a given protein will form amyloids under specific conditions remains a formidable challenge. In this study, we have conducted a comprehensive analysis to understand the amyloidogenic tendencies within a dataset containing 1828 (348 amyloidogenic and 1480 non-amyloidogenic) antibody light chain variable region (VL) sequences obtained from the AL-Base database. Physicochemical and structural features often associated with protein aggregation, such as net charge, isoelectric point (pI), and solvent-exposed hydrophobic regions did not reveal a consistent association with the aggregation capability of the antibody light chains. However, the solvent-exposed aggregation-prone regions (APRs) occur with higher frequencies among the amyloidogenic light chains when compared with the non-amyloidogenic ones, with the difference ranging from 2% to 15% at various relative solvent-accessible surface area (rASA) cutoffs. We have, for the first time, identified structural gatekeeping residues around the APRs and assessed their impact on the amyloidogenicity of the antibody light chains. The non-amyloidogenic light chains contain these structural gatekeeper residues vicinal to their APRs more often than the amyloidogenic ones. We observed that the rASA cutoff of 35% is optimal for identifying the surface-exposed APRs, and a 4 Å distance cutoff from the APR motif(s) is optimal for identifying the structural gatekeeper residues. Moreover, lambda light chains were found to contain solvent-exposed APRs more often and surrounded by fewer gatekeepers, rendering them more susceptible to aggregation. The insights gained from this report have significant implications for understanding the molecular origins of light-chain amyloidosis in humans and the design of aggregation-resistant therapeutic antibodies.
The adaptive immune receptor repertoire (AIRR) encompasses an immense diversity of antibody and T-cell receptor sequences, whose collective organization – how receptors are distributed, clustered, and interrelated across sequence and functional (e.g., antigen-binding) dimensions – remains poorly characterized. Representing AIRRs in continuous representation spaces that capture sequence, biochemical, and structural similarity between receptors may enable comparisons beyond discrete sequence features. Using both one-hot encodings and protein language model (PLM) embeddings, we developed a quantitative framework to map immune receptor organization at global (sequence-set-level) and local (single-sequence-level) scales. Applying the geometry-aware Wasserstein-2 distance, we show that the global structure of the AIRR space can be recovered from as few as ∼10 5 sequence embeddings, at least 10 orders of magnitude smaller than the theoretical immune receptor diversity. We found that immune receptor sequences annotated with different antigen specificities occupy distinct regions of representation space. To resolve local relationships, we introduce a spatial homogeneity metric that quantifies the extent of functional clustering. We found higher spatial homogeneity in embedding spaces than in sequence space for diverse antigen-specific datasets. Our framework establishes a foundation for quantitative mapping of adaptive immune repertoire organization.
Therapeutic antibodies have gained prominence in recent years due to their precision in targeting specific diseases. As these molecules become increasingly essential in modern medicine, comprehensive data tracking and analysis are critical for advancing research and ensuring successful clinical outcomes. YAbS, The Antibody Society’s Antibody Therapeutics Database, serves as a vital resource for monitoring the development and clinical progress of therapeutic antibodies. The database catalogs detailed information on over 2,900 commercially sponsored investigational antibody candidates that have entered clinical study since 2000, as well as all approved antibody therapeutics. Data for the late-stage clinical pipeline and antibody therapeutics in regulatory review or approved (over 450 molecules) are openly accessible (https://db.antibodysociety.org). Antibody-related information includes molecular format, targeted antigen, current development status, indications studied, and the clinical development timeline of the antibodies, as well as the geographical region of company sponsors. Furthermore, the database supports in-depth industry trends analysis, facilitating the identification of innovative developments and the assessment of success rates within the field. This resource is continually updated and refined, providing invaluable insights to researchers, clinicians, and industry professionals engaged in antibody therapeutics development.
B cell and T cell receptor repertoires compose the adaptive immune receptor repertoire (AIRR) of an individual. The AIRR is a unique collection of antigen-specific receptors that drives adaptive immune responses, which in turn is imprinted in each individual AIRR. This supports the concept that the AIRR could determine disease outcomes, for example in autoimmunity, infectious disease and cancer. AIRR analysis could therefore assist the diagnosis, prognosis and treatment of human diseases towards personalized medicine. High-throughput sequencing, high-dimensional statistical analysis, computational structural biology and machine learning are currently employed to study the shaping and dynamics of the AIRR as a function of time and antigenic challenges. This Primer provides an overview of concepts and state-of-the-art methods that underlie experimental and computational AIRR analysis and illustrates the diversity of relevant applications. The Primer also addresses some of the outstanding challenges in AIRR analysis, such as sampling, sequencing depth, experimental variations and computational biases, while discussing prospects of future AIRR analysis applications for understanding and predicting adaptive immune responses. The adaptive immune receptor repertoire (AIRR) drives adaptive immune responses, which could determine disease outcomes, infectious disease and cancer. Mhanna, Bashour et al. outline the approaches and challenges in AIRR analysis, as well as future developments towards predicting adaptive immune responses.
There is currently considerable interest in the field of de novo antibody design, and deep learning techniques are now regularly applied to optimise antibody properties such as binding affinity. However, robust baselines within this field have not kept up with recent developments. In this study, we generate a dataset of over 524,000 Trastuzumab variants and use this to show that standard computational methods such as BLOSUM, AbLang, ESM, and Protein-MPNN can be used to design diverse antibody libraries from just a single starting sequence. These novel libraries are predicted to be enriched in binding variants and experimental validation of 700 of these designs is ongoing. We also demonstrate that, even with only a very small number of experimental data points, simple machine learning classifiers can be trained in seconds to accurately pre-screen future designs. This pre-screening maintains library diversity and saves experimental time and money. ### Competing Interest Statement V.G. declares advisory board positions in aiNET GmbH, Enpicom B.V, Absci, Omniscope, and Diagonal Therapeutics. V.G. is a consultant for Adaptive Biosystems, Specifica Inc, Roche/Genentech, immunai, LabGenius, and FairJourney Biologics. J.R.J. is employed by GlaxoSmithKline plc. The remaining authors declare no competing interests.
Designing effective monoclonal antibody (mAb) therapeutics faces a multi-parameter optimization challenge known as “developability”, which reflects an antibody’s ability to progress through development stages based on its physicochemical properties. While natural antibodies may provide valuable guidance for mAb selection, we lack a comprehensive understanding of natural developability parameter (DP) plasticity (redundancy, predictability, sensitivity) and how the DP landscapes of human-engineered and natural antibodies relate to one another. These gaps hinder fundamental developability profile cartography. To chart natural and engineered DP landscapes, we computed 40 sequence- and 46 structure-based DPs of over two million native and human-engineered single-chain antibody sequences. We find lower redundancy among structure-based compared to sequence-based DPs. Sequence DP sensitivity to single amino acid substitutions varied by antibody region and DP, and structure DP values varied across the conformational ensemble of antibody structures. We show that sequence DPs are more predictable than structure-based ones across different machine-learning tasks and embeddings, indicating a constrained sequence-based design space. Human-engineered antibodies localize within the developability and sequence landscapes of natural antibodies, suggesting that human-engineered antibodies explore mere subspaces of the natural one. Our work quantifies the plasticity of antibody developability, providing a fundamental resource for multi-parameter therapeutic mAb design. Analysis of 2 million native antibodies reveals that human-engineered antibodies form subspaces of the natural developability space. This large-scale analysis allows the quantification of developability plasticity, accelerating antibody drug design.
Summary:High-throughput sequencing (HTS) offers a modern, fast, and explorative solution to unveil the full potential of display techniques, like antibody phage display, in molecular biology. However, a significant challenge lies in the processing and analysis of such data. Furthermore, there is a notable absence of open-access user-friendly software tools that can be utilized by scientists lacking programming expertise. Here, we present ExpoSeq as an easy-to-use tool to explore, process, and visualize HTS data from antibody discovery campaigns like an expert while only requiring a beginner's knowledge. Availability and implementation:The pipeline is distributed via GitHub and PyPI, and it can either be installed as a package with pip or the user can choose to clone the repository.
In the version of the article initially published, there were errors in Figs.1234.In Fig. 1, in the upper box, "CDR3" initially appeared as "CDR2" and "IGLJ(10)" read "IGLV( 10)".In the lower box, "TRBD" was missing from above "CDR3".In Fig. 2a, "class-switch recombination" originally read "class switching".In Fig. 3d, the uppermost arrow, " + /-UMI" and the lower "MTPX PCR" were all missing, and the second instance of " + /-UMI" appeared as "UMI".In Fig. 4c, the AIRR 1 image was missing the TCR from the surface of one of the cells.
COVID-19 has resulted in millions of deaths and severe impact on economies worldwide. Moreover, the emergence of SARS-CoV-2 variants presented significant challenges in controlling the pandemic, particularly their potential to avoid the immune system and evade vaccine immunity. This has led to a growing need for research to predict how mutations in SARS-CoV-2 reduces the ability of antibodies to neutralize the virus. In this study, we assembled a set of 1813 mutations from the interface of SARS-CoV-2 spike protein's receptor binding domain (RBD) and neutralizing antibody complexes and developed a machine learning model to classify high or low escape mutations using interaction energy, inter-residue contacts and predicted binding free energy change. Our approach achieved an Area under the Receiver Operating Characteristics (ROC) Curve (AUC) of 0.91 using the Random Forest classifier on the test dataset with 217 mutations. The model was further utilized to predict the escape mutations on a dataset of 29,165 mutations located at the interface of 83 RBD-neutralizing antibody complexes. A small subset of this dataset was also validated based on available experimental data. We found that top 10 % high escape mutations were dominated by charged to nonpolar mutations whereas low escape mutations were dominated by polar to nonpolar mutations. We believe that the present method will allow prioritization of high/low escape mutations in the context of neutralizing antibodies targeting SARS-CoV-2 RBD region and assist antibody design for current and emerging variants.
Antibodies are canonically Y-shaped multimeric proteins capable of highly specific molecular recognition. The CDRH3 region located at the tip of variable chains of an antibody dominates antigen-binding specificity. Therefore, it is a priority to design optimal antigen-specific CDRH3 regions to develop therapeutic antibodies. However, the combinatorial nature of CDRH3 sequence space makes it impossible to search for an optimal binding sequence exhaustively and efficiently using computational approaches. Here, we present \texttt{AntBO}: a combinatorial Bayesian optimisation framework enabling efficient \textit{in silico} design of the CDRH3 region. Ideally, antibodies are expected to have high target specificity and developability. We introduce a CDRH3 trust region that restricts the search to sequences with favourable developability scores to achieve this goal. For benchmarking, \texttt{AntBO} uses the \texttt{Absolut!} software suite as a black-box oracle to score the target specificity and affinity of designed antibodies \textit{in silico} in an unconstrained fashion~\citep{robert2021one}. The experiments performed for $159$ discretised antigens used in \texttt{Absolut!} demonstrate the benefit of \texttt{AntBO} in designing CDRH3 regions with diverse biophysical properties. In under $200$ calls to black-box oracle, \texttt{AntBO} can suggest antibody sequences that outperform the best binding sequence drawn from 6.9 million experimentally obtained CDRH3s and a commonly used genetic algorithm baseline. Additionally, \texttt{AntBO} finds very-high affinity CDRH3 sequences in only 38 protein designs whilst requiring no domain knowledge. We conclude \texttt{AntBO} brings automated antibody design methods closer to what is practically viable for in vitro experimentation.
We have developed a database, Ab-CoV, which contains manually curated experimental interaction profiles of 1780 coronavirus-related neutralizing antibodies. It contains more than 3200 datapoints on half maximal inhibitory concentration (IC50), half maximal effective concentration (EC50) and binding affinity (K-D). Each data with experimentally known three-dimensional structures are complemented with predicted change in stability and affinity of all possible point mutations of interface residues. Ab-CoV also includes information on epitopes and paratopes, structural features of viral proteins, sequentially similar therapeutic antibodies and Collier de Perles plots. It has the feasibility for structure visualization and options to search, display and download the data.
The expression of human epidermal growth factor receptor 2 (HER2) is a key classification factor in breast cancer. Many breast cancers express isoforms of HER2 with truncated carboxy-terminal fragments (CTF), collectively known as p95HER2. A common p95HER2 isoform, 611-CTF, is a biomarker for aggressive disease and confers resistance to therapy. Contrary to full-length HER2, 611-p95HER2 has negligible normal tissue expression. There is currently no approved diagnostic assay to identify this subgroup and no therapy targeting this mechanism of tumor escape. The purpose of this study was to develop a monoclonal antibody (mAb) against 611-CTF-p95HER2. Hybridomas were generated from rats immunized with cells expressing 611-CTF. A hybridoma producing a highly specific Ab was identified and cloned further as a mAb. This mAb, called Oslo-2, gave strong staining for 611-CTF and no binding to full-length HER2, as assessed in cell lines and tissues by flow cytometry, immunohistochemistry and immunofluorescence. No cross-reactivity against HER2 negative controls was detected. Surface plasmon resonance analysis demonstrated a high binding affinity (equilibrium dissociation constant 2 nM). The target epitope was identified at the N-terminal end, using experimental alanine scanning. Further, the mAb paratope was identified and characterized with hydrogen-deuterium-exchange, and a molecular model for the (Oslo-2 mAb:611-CTF-p95HER2) complex was generated by an experimental-information-driven docking approach. We conclude that the Oslo-2 mAb has a high affinity and is highly specific for 611-CTF-p95HER2. The Ab may be used to develop potent and safe therapies, overcoming p95HER2-mediated tumor escape, as well as for developing diagnostic assays.
Machine learning (ML) is a key technology for accurate prediction of antibody-antigen binding. Two orthogonal problems hinder the application of ML to antibody-specificity prediction and the benchmarking thereof: The lack of a unified ML formalization of immunological antibody specificity prediction problems and the unavailability of large-scale synthetic benchmarking datasets of real-world relevance. Here, we developed the Absolut! software suite that enables parameter-based unconstrained generation of synthetic lattice-based 3D-antibody-antigen binding structures with ground-truth access to conformational paratope, epitope, and affinity. We formalized common immunological antibody specificity prediction problems as ML tasks and confirmed that for both sequence and structure-based tasks, accuracy-based rankings of ML methods trained on experimental data hold for ML methods trained on Absolut!-generated data. The Absolut! framework thus enables real-world relevant development and benchmarking of ML strategies for biotherapeutics design. Graphical abstract The software framework Absolut! enables (A,B) the generation of virtually arbitrarily large numbers of synthetic 3D-antibody-antigen structures, (C,D) the formalization of antibody specificity as machine learning (ML) tasks as well as the exploration of ML strategies for real-world antibody-antigen binding or paratope-epitope prediction. Highlights Software framework Absolut! to generate an arbitrarily large number of synthetic 3D-antibody-antigen structures that contain biological layers of antibody-antigen binding complexity that render ML predictions challenging Immunological antibody specificity prediction problems formalized as machine learning tasks for which the in silico complexes are immediately usable as benchmark datasets Exploration of machine learning prediction accuracy as a function of architecture, dataset size, choice of negatives, and sequence-structure encoding Relative ML performance learnt on Absolut! datasets transfers to experimental datasets