Abstract Preclinical antibody discovery relies on progressive screening and down-selection of candidate antibodies from large immune repertoires, yet this critical process is poorly represented in existing public databases. Here we introduce KyDab (Kymouse Antibody Database), a well-curated database of antibody discovery selection data generated using standardized workflows on the Kymouse humanized mouse platform. The current release includes 11 Kymouse platform mice immunisation studies covering 51 immunogens, more than 120,000 paired heavy–light chain sequences, and binding measurements for a selected subset of experimentally characterized clones. By capturing full-funnel selection data with consistent metadata and both positive and negative experimental outcomes, KyDab provides a valuable data resource for the development and evaluation of artificial intelligence models for antibody discovery. KyDab is accessible at https://kydab.naturalantibody.com/ , and the database will be continuously updated as new datasets become available.
Computational antibody design has seen many recent advances pioneered via the use of language models and advanced structure prediction tools. Developing a de novo antibody against a specific antigen requires structural awareness that most language models lack. A prominent class of machine learning methods combining the best of language model and structural worlds is inverse folding. This approach aims to predict a sequence that would fit a given structure. Such methods are now increasingly used to predict alternate sequences given a structure of a binder. It is known that, just like language models, such methods have certain predictive power in identifying binders. Here we performed a set of tests to reveal where, if at all, such methods provide value in the realistic setting of antibody discovery.
The development of computational models addressing therapeutic antibodies faces significant challenges. Particularly, the prediction of binding affinity across a diverse set of measurements, due to the scarcity of data. A critical data element is the set of antibody-antigen interaction pairs associated with sequences. To address this issue, we developed the Antigen Specific Antibody Database (ASD, https://naturalantibody.com/agab/), a database aggregating antibody-antigen interaction data from multiple studies with standardized formatting and annotations. Our dataset compilation strategy resulted in data from 15 distinct sources, resulting in 1,097,946 unique antibody-antigen interactions (with 9575 unique antigens). The ASD captures diverse affinity measures and qualitative binding assessment, along with metadata including UniProt and PDB identifiers, target protein names, confidence levels, and experimental conditions such as type of measured affinity, source organism, and germline genes. Through this integration drive, we make available an ample resource of interaction data gathered from the public domain to act as a foundation for model development and further data generation.
Studying the interactions between antibodies and antigens is fundamental to the development of novel therapeutic biologics. Predictions of such interactions start with data collection. Though there exist reliable resources to identify antibody structures in the Protein Data Bank (PDB), such data still requires substantial processing to be usable in predictive tasks. Redundancy in sequences needs to be removed to avoid data leakages between train, test, and validation sets. Descriptors such as surface accessibility, secondary structure, and antibody region information need to be additionally annotated. Information on inter- and intra-molecular contacts, which is crucial to studying paratope/epitope information, needs to be collected. The specialized immunoglobulin format of Nanobodies® requires a separate dataset mirroring that of antibodies, given that their structure contains only a single VHH chain. Because antibody-antigen structures account for a small amount of all protein-protein contacts, having a molecular contact reference from other proteins is also desired. To address these issues, we introduce NAStructuralDB (https://naturalantibody.com/na-structural/), a dataset of processed structures of antibodies, Nanobodies®, proteins, and their complexes with molecular contact information and associated annotations. We use the opportunity of having collected the contact data to provide a reference of binding propensities of different residues across distinct contact types.
Antibodies are naturally evolved molecular recognition scaffolds that can bind a variety of surfaces. Their designability is crucial to the development of biologics, with computational methods holding promise in accelerating the delivery of medicines to the clinic. Modeling antibody-antigen recognition is prohibitively difficult, with data paucity being one of the biggest hurdles. Current affinity datasets comprise a small number of experimental measurements, which are often not standardized between molecules. Here, we address these issues by creating a dataset of seven antigens with two antibodies each, for which we introduce a heterogeneous set of mutations to the CDR-H3 measured by ELISA. Each of the parental complexes has a known crystal structure. We perform benchmarking of state-of-the-art affinity prediction algorithms to gauge their effectiveness. Current computational methods exhibit substantial limitations in accurately predicting the effects of single-point mutations. In contrast, the older empirical, physics-based method FoldX performs well in identifying mutants that retain binding. These findings highlight the need for more resources like the one presented here, i.e. large, molecularly diverse, and experimentally consistent datasets.
Antibodies devoid of light chains are a promising class of biotherapeutics. Computational methods that address these molecules are crucially needed to accelerate the traditional, long and expensive experimental process of their discovery. Inverse folding, wherein one is tasked to predict a sequence given molecular coordinates, is an established method in scaffold-based protein design. Here we develop an inverse folding method specific to nanobodies. We demonstrate its application in nanobody-engineering scenarios of enriching binders from next-generation sequencing experiments and novel binder design. ### Competing Interest Statement The authors have declared no competing interest.
Antibody discovery has been successful in designing and progressing molecules to the clinic and market based on largely empirical methods and human experience. The field is now transitioning from classical monospecific antibodies to innovative smart biologics that employ diverse mechanisms of action, such as targeting, antagonism, agonism, and target-independent function. This evolution is being assisted, augmented, and potentially disrupted by artificial intelligence and machine learning (AI/ML) technologies. This perspective is focused on bringing clarity to the strategy and thinking that is required when designing antibody drug candidates and how emerging AI/ML strategies can address the real-world challenges of drug discovery and continue to improve performance.
Understanding the pairing preferences and structural interactions between antibody heavy and light chains can enhance our ability to design more effective and specific therapeutic antibodies. Insights from natural antibody repertoires and conserved contact sites help reduce autoreactivity and improve drug safety and efficacy. Current databases represent only a limited portion of the estimated diversity of unique paired antibody molecules. To address this, we introduce PairedAbNGS, a novel database with paired heavy/light antibody chains. To our knowledge, this is the largest resource for paired natural antibody sequences with 58 bioprojects and over 14 million assembled productive sequences. Using this dataset, we investigated heavy and light chain variable (V) gene pairing preferences and found significant biases beyond gene usage frequencies, possibly due to receptor editing favoring less autoreactive combinations. Analyzing the available antibody structures from the Protein Data Bank, we studied conserved contact residues between heavy and light chains, particularly interactions between the CDR3 region of one chain and the FWR2 region of the opposite chain. Examination of amino acid pairs at key contact sites revealed significant deviations of amino acids distributions compared to random pairings, in the heavy chain's CDR3 region contacting the opposite chain, indicating specific interactions might be crucial for proper chain pairing. This observation is further reinforced by preferential IGHV-IGLJ and IGLV-IGHJ pairing preferences. We hope that both our resources and the findings would contribute to improving the engineering of biological drugs. We make the database accessible at https://naturalantibody.com/paired-ab-ngs as a valuable tool for biological and machine-learning applications.
Machine learning applications in protein sciences have ushered in a new era for designing molecules in silico. Antibodies, which currently form the largest group of biologics in clinical use, stand to benefit greatly from this shift. Despite the proliferation of these protein design tools, their direct application to antibodies is often limited by the unique structural biology of these molecules. We note that multiple methods attempting antibody design focus on the discovery of an antigen-specific antibody. Here, we review the current computational methods for antibody design, focusing on binder discovery, contextualizing their role in the drug discovery process.
Antibodies represent the largest and fastest growing class of biologic therapeutics, yet forecasting their clinical performance, particularly immunogenicity, remains a major hurdle in drug development. Despite hundreds of antibody-based drugs progressing through clinical pipelines, systematic integration of their clinical outcomes has been limited by fragmented and heterogeneous data. Here, we present the Therapeutic Antibody Database, a comprehensive and curated resource that links therapeutic antibodies to clinical trial outcomes, with a dedicated focus on immunogenicity. Our dataset is sourced from approximately 11,500 anti-drug antibody (ADA) measurements across diverse molecules and indications, offering an unprecedented view into the clinical manifestation of immune responses to biologics. In order to evaluate the main drivers of ADA, we evaluate gathered immunogenicity incidence and prevalence data against various therapeutic descriptors which includes sequence, structure and contextual features related to therapeutics. We find that most tools have very poor performance, and we pinpoint the causes of it, demonstrating the need for systems immunology approaches incorporating clinical metadata beyond biochemical properties of the molecules alone. ### Competing Interest Statement The authors have declared no competing interest.
Motivation:Nanobodies are a subclass of immunoglobulins, whose binding site consists of only one peptide chain, bestowing favorable biophysical properties. Recently, the first nanobody therapy was approved, paving the way for further clinical applications of this antibody format. Further development of nanobody-based therapeutics could be streamlined by computational methods. One of such methods is infilling-positional prediction of biologically feasible mutations in nanobodies. Being able to identify possible positional substitutions based on sequence context, facilitates functional design of such molecules.Results:Here we present nanoBERT, a nanobody-specific transformer to predict amino acids in a given position in a query sequence. We demonstrate the need to develop such machine-learning based protocol as opposed to gene-specific positional statistics since appropriate genetic reference is not available. We benchmark nanoBERT with respect to human-based language models and ESM-2, demonstrating the benefit for domain-specific language models. We also demonstrate the benefit of employing nanobody-specific predictions for fine-tuning on experimentally measured thermostability dataset. We hope that nanoBERT will help engineers in a range of predictive tasks for designing therapeutic nanobodies.Availability and implementation:https://huggingface.co/NaturalAntibody/.
Antibodies are a cornerstone of the immune system, playing a pivotal role in identifying and neutralizing infections caused by bacteria, viruses, and other pathogens. Understanding their structure, and function, can provide insights into both the body's natural defenses and the principles behind many therapeutic interventions, including vaccines and antibody-based drugs. The analysis and annotation of antibody sequences, including the identification of variable, diversity, joining, and constant genes, as well as the delineation of framework regions and complementarity-determining regions, is essential for understanding their structure and function. Currently analyzing large volumes of antibody sequences is routine in antibody discovery, requiring fast and accurate tools. While there are existing tools designed for the annotation and numbering of antibody sequences, they often have limitations such as being restricted to either nucleotide or amino acid sequences; slow execution times; or reliance on germline databases that are closed, frequently changed, or have sparse coverage for some species. Here, we present the Rapid Immunoglobulin Overview Tool (RIOT), a novel open-source solution for antibody numbering that addresses these shortcomings. RIOT handles nucleotide and amino acid sequence processing, comes integrated with an Open Germline Receptor Database, and is computationally efficient. We hope that the tool will facilitate rapid annotation of antibody sequencing outputs for the benefit of understanding antibody biology and discovering novel therapeutics.
Antibody-based therapeutics must not undergo chemical modifications that would impair their efficacy or hinder their developability. A commonly used technique to de-risk lead biotherapeutic candidates annotates chemical liability motifs on their sequence. By analyzing sequences from all major sources of data (therapeutics, patents, GenBank, literature, and next-generation sequencing outputs), we find that almost all antibodies contain an average of 3-4 such liability motifs in their paratopes, irrespective of the source dataset. This is in line with the common wisdom that liability motif annotation is over-predictive. Therefore, we have compiled three computational flags to prioritize liability motifs for removal from lead drug candidates: 1. germline, to reflect naturally occurring motifs, 2. therapeutic, reflecting chemical liability motifs found in therapeutic antibodies, and 3. surface, indicative of structural accessibility for chemical modification. We show that these flags annotate approximately 60% of liability motifs as benign, that is, the flagged liabilities have a smaller probability of undergoing degradation as benchmarked on two experimental datasets covering deamidation, isomerization, and oxidation. We combined the liability detection and flags into a tool called Liability Antibody Profiler (LAP), publicly available at lap.naturalantibody.com. We anticipate that LAP will save time and effort in de-risking therapeutic molecules.
The naïve human antibody repertoire has theoretical access to an estimated > 1015 antibodies. Identifying subsets of this prohibitively large space where therapeutically relevant antibodies may be found is useful for development of these agents. It was previously demonstrated that, despite the immense sequence space, different individuals can produce the same antibodies. It was also shown that therapeutic antibodies, which typically follow seemingly unnatural development processes, can arise independently naturally. To check for biases in how the sequence space is explored, we data mined public repositories to identify 220 bioprojects with a combined seven billion reads. Of these, we created a subset of human bioprojects that we make available as the AbNGS database (https://naturalantibody.com/ngs/). AbNGS contains 135 bioprojects with four billion productive human heavy variable region sequences and 385 million unique complementarity-determining region (CDR)-H3s. We find that 270,000 (0.07% of 385 million) unique CDR-H3s are highly public in that they occur in at least five of 135 bioprojects. Of 700 unique therapeutic CDR-H3, a total of 6% has direct matches in the small set of 270,000. This observation extends to a match between CDR-H3 and V-gene call as well. Thus, the subspace of shared ('public') CDR-H3s shows utility for serving as a starting point for therapeutic antibody design.
Antibodies are proteins produced by our immune system that have been harnessed as biotherapeutics. The discovery of antibody-based therapeutics relies on analyzing large volumes of diverse sequences coming from phage display or animal immunizations. Identification of suitable therapeutic candidates is achieved by grouping the sequences by their similarity and subsequent selection of a diverse set of antibodies for further tests. Such groupings are typically created using sequence-similarity measures alone. Maximizing diversity in selected candidates is crucial to reducing the number of tests of molecules with near-identical properties. With the advances in structural modeling and machine learning, antibodies can now be grouped across other diversity dimensions, such as predicted paratopes or three-dimensional structures. Here we benchmarked antibody grouping methods using clonotype, sequence, paratope prediction, structure prediction, and embedding information. The results were benchmarked on two tasks: binder detection and epitope mapping. We demonstrate that on binder detection no method appears to outperform the others, while on epitope mapping, clonotype, paratope, and embedding clusterings are top performers. Most importantly, all the methods propose orthogonal groupings, offering more diverse pools of candidates when using multiple methods than any single method alone. To facilitate exploring the diversity of antibodies using different methods, we have created an online tool-CLAP-available at (clap.naturalantibody.com) that allows users to group, contrast, and visualize antibodies using the different grouping methods.
Designing effective monoclonal antibody (mAb) therapeutics faces a multi-parameter optimization challenge known as “developability”, which reflects an antibody’s ability to progress through development stages based on its physicochemical properties. While natural antibodies may provide valuable guidance for mAb selection, we lack a comprehensive understanding of natural developability parameter (DP) plasticity (redundancy, predictability, sensitivity) and how the DP landscapes of human-engineered and natural antibodies relate to one another. These gaps hinder fundamental developability profile cartography. To chart natural and engineered DP landscapes, we computed 40 sequence- and 46 structure-based DPs of over two million native and human-engineered single-chain antibody sequences. We find lower redundancy among structure-based compared to sequence-based DPs. Sequence DP sensitivity to single amino acid substitutions varied by antibody region and DP, and structure DP values varied across the conformational ensemble of antibody structures. We show that sequence DPs are more predictable than structure-based ones across different machine-learning tasks and embeddings, indicating a constrained sequence-based design space. Human-engineered antibodies localize within the developability and sequence landscapes of natural antibodies, suggesting that human-engineered antibodies explore mere subspaces of the natural one. Our work quantifies the plasticity of antibody developability, providing a fundamental resource for multi-parameter therapeutic mAb design. Analysis of 2 million native antibodies reveals that human-engineered antibodies form subspaces of the natural developability space. This large-scale analysis allows the quantification of developability plasticity, accelerating antibody drug design.
Mass spectrometry-based proteomics facilitates the identification and quantification of thousands of proteins but encounters challenges in measuring human antibodies due to their vast diversity. Bottom-up proteomics methods primarily rely on database searches, comparing experimental peptide values to theoretical database sequences. While the human body can produce millions of distinct antibodies, current databases, such as UniProtKB/Swiss-Prot, contain only 1095 sequences (as of January 2024), potentially hindering antibody identification via mass spectrometry. Therefore, expanding the database is crucial for discovering new antibodies. Recent genomic studies have amassed millions of human antibody sequences in the Observed Antibody Space (OAS) database, yet this data remains underutilized. Leveraging this vast collection, we conduct efficient database searches in publicly available proteomics data, focusing on SARS-CoV-2. In our study, thirty million heavy antibody sequences from 146 SARS-CoV-2 patients in the OAS database were digested in silico to obtain 18 million unique peptides. These peptides form the basis for new bottom-up proteomics databases. We used those databases for searching new antibody peptides in publicly available SARS-CoV-2 human plasma samples in the Proteomics Identification Database (PRIDE). This approach avoids false positives in antibody peptide identification as confirmed by searching against negative controls (brain samples) and employing different database sizes. We show that new antibody peptides were found in previous plasma samples and expect that the newly discovered antibody peptides can be further employed to develop therapeutic antibodies. The method will be broadly applicable to find characteristic antibodies for other diseases.
AlphaFold2 has hallmarked a generational improvement in protein structure prediction. In particular, advances in antibody structure prediction have provided a highly translatable impact on drug discovery. Though AlphaFold2 laid the groundwork for all proteins, antibody-specific applications require adjustments tailored to these molecules, which has resulted in a handful of deep learning antibody structure predictors. Herein, we review the recent advances in antibody structure prediction and relate them to their role in advancing biologics discovery.
Background Machine learning (ML) technologies, especially deep learning (DL), have gained increasing attention in predictive mass spectrometry (MS) for enhancing the data-processing pipeline from raw data analysis to end-user predictions and rescoring. ML models need large-scale datasets for training and repurposing, which can be obtained from a range of public data repositories. However, applying ML to public MS datasets on larger scales is challenging, as they vary widely in terms of data acquisition methods, biological systems, and experimental designs. Results We aim to facilitate ML efforts in MS data by conducting a systematic analysis of the potential sources of variability in public MS repositories. We also examine how these factors affect ML performance and perform a comprehensive transfer learning to evaluate the benefits of current best practice methods in the field for transfer learning. Conclusions Our findings show significantly higher levels of homogeneity within a project than between projects, which indicates that it is important to construct datasets most closely resembling future test cases, as transferability is severely limited for unseen datasets. We also found that transfer learning, although it did increase model performance, did not increase model performance compared to a non-pretrained model.