Machine learning applications in protein sciences have ushered in a new era for designing molecules in silico. Antibodies, which currently form the largest group of biologics in clinical use, stand to benefit greatly from this shift. Despite the proliferation of these protein design tools, their direct application to antibodies is often limited by the unique structural biology of these molecules. We note that multiple methods attempting antibody design focus on the discovery of an antigen-specific antibody. Here, we review the current computational methods for antibody design, focusing on binder discovery, contextualizing their role in the drug discovery process.
Computational prediction of molecule-protein interactions has been key for developing new molecules to interact with a target protein for therapeutics development. Previous work includes two independent streams of approaches: (1) predicting protein-protein interactions (PPIs) between naturally occurring proteins and (2) predicting binding affinities between proteins and small-molecule ligands [also known as drug-target interaction (DTI)]. Studying the two problems in isolation has limited the ability of these computational models to generalize across the PPI and DTI tasks, both of which ultimately involve noncovalent interactions with a protein target. In this work, we developed Equivariant Graph of Graphs neural Network (EGGNet), a geometric deep learning (GDL) framework, for molecule-protein binding predictions that can handle three types of molecules for interacting with a target protein: (1) small molecules, (2) synthetic peptides, and (3) natural proteins. EGGNet leverages a graph of graphs (GoG) representation constructed from the molecular structures at atomic resolution and utilizes a multiresolution equivariant graph neural network to learn from such representations. In addition, EGGNet leverages the underlying biophysics and makes use of both atom- and residue-level interactions, which improve EGGNet's ability to rank candidate poses from blind docking. EGGNet achieves competitive performance on both a public protein-small-molecule binding affinity prediction task (80.2% top 1 success rate on CASF-2016) and a synthetic protein interface prediction task (88.4% area under the precision-recall curve). We envision that the proposed GDL framework can generalize to many other protein interaction prediction problems, such as binding site prediction and molecular docking, helping accelerate protein engineering and structure-based drug development.
Carbohydrates and glycoproteins modulate key biological functions. However, experimental structure determination of sugar polymers is notoriously difficult. Computational approaches can aid in carbohydrate structure prediction, structure determination, and design. In this work, we developed a glycan-modeling algorithm, GlycanTreeModeler, that computationally builds glycans layer-by-layer, using adaptive kernel density estimates (KDE) of common glycan conformations derived from data in the Protein Data Bank (PDB) and from quantum mechanics (QM) calculations. GlycanTreeModeler was benchmarked on a test set of glycan structures of varying lengths, or “trees”. Structures predicted by GlycanTreeModeler agreed with native structures at high accuracy for both de novo modeling and experimental density-guided building. We employed these tools to design de novo glycan trees into a protein nanoparticle vaccine to shield regions of the scaffold from antibody recognition, and experimentally verified shielding. This work will inform glycoprotein model prediction, glycan masking, and further aid computational methods in experimental structure determination and refinement.
Understanding how proteins evolve under selective pressure is a longstanding challenge. The immensity of the search space has limited efforts to systematically evaluate the impact of multiple simultaneous mutations, so mutations have typically been assessed individually. However, epistasis, or the way in which mutations interact, prevents accurate prediction of combinatorial mutations based on measurements of individual mutations. Here, we use artificial intelligence to define the entire functional sequence landscape of a protein binding site in silico, and we call this approach Complete Combinatorial Mutational Enumeration (CCME). By leveraging CCME, we are able to construct a comprehensive map of the evolutionary connectivity within this functional sequence landscape. As a proof of concept, we applied CCME to the ACE2 binding site of the SARS-CoV-2 spike protein receptor binding domain. We selected representative variants from across the functional sequence landscape for testing in the laboratory. We identified variants that retained functionality to bind ACE2 despite changing over 40% of evaluated residue positions, and the variants now escape binding and neutralization by monoclonal antibodies. This work represents a crucial initial stride towards achieving precise predictions of pathogen evolution, opening avenues for proactive mitigation.
The human infectious disease COVID-19 caused by the SARS-CoV-2 virus has become a major threat to global public health. Developing a vaccine is the preferred prophylactic response to epidemics and pandemics. However, for individuals who have contracted the disease, the rapid design of antibodies that can target the SARS-CoV-2 virus fulfils a critical need. Further, discovering antibodies that bind multiple variants of SARS-CoV-2 can aid in the development of rapid antigen tests (RATs) which are critical for the identification and isolation of individuals currently carrying COVID-19. Here we provide a proof-of-concept study for the computational design of high-affinity antibodies that bind to multiple variants of the SARS-CoV-2 spike protein using RosettaAntibodyDesign (RAbD). Well characterized antibodies that bind with high affinity to the SARS-CoV-1 (but not SARS-CoV-2) spike protein were used as templates and re-designed to bind the SARS-CoV-2 spike protein with high affinity, resulting in a specificity switch. A panel of designed antibodies were experimentally validated. One design bound to a broad range of vari-ants of concern including the Omicron, Delta, Wuhan, and South African spike protein variants.
[This corrects the article DOI: 10.1016/j.heliyon.2023.e15032.].
Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) vaccines may target epitopes that reduce durability or increase the potential for escape from vaccine-induced immunity. Using synthetic vaccinology, we have developed rationally immune-focused SARS-CoV-2 Spike-based vaccines. Glycans can be employed to alter antibody responses to infection and vaccines. Utilizing computational modeling and in vitro screening, we have incorporated glycans into the receptor-binding domain (RBD) and assessed antigenic profiles. We demonstrate that glycan-coated RBD immunogens elicit stronger neutralizing antibodies and have engineered seven multivalent configurations. Advanced DNA delivery of engineered nanoparticle vaccines rapidly elicits potent neutralizing antibodies in guinea pigs, hamsters, and multiple mouse models, including human ACE2 and human antibody repertoire transgenics. RBD nanoparticles induce high levels of cross-neutralizing antibodies against variants of concern with durable titers beyond 6 months. Single, low-dose immunization protects against a lethal SARS-CoV-2 challenge. Single-dose coronavirus vaccines via DNA-launched nanoparticles provide a platform for rapid clinical translation of potent and durable coronavirus vaccines.
Host-pathogen interactions drive an evolutionary game of cat-and-mouse between a pathogen’s protein virulence factors, the host’s adaptive immune system, and therapeutics targeting the pathogen. There is an urgent need for treatments and prophylactics that remain effective as a pathogen evolves, and the ability to predict pathogen evolution is a longstanding challenge. Therefore, a common strategy has been to target conserved epitopes, but strong selective pressures can drive pathogens to evolve resistance nonetheless. Here, we report a novel, generally-applicable approach called Deep Evolutionary Forecasting that predicts protein evolution using artificial intelligence and molecular modeling. The first step is to perform a complete enumeration of the functional sequence landscape in silico for a target protein. Then, we construct a graph where the edges between sequence variants are weighted by evolutionary probability. Protein evolution is forecasted by traversing this graph. We chose the SARS-CoV-2 receptor binding domain (RBD) as a model system because highly-mutated viral variants have continued to emerge that escape available therapeutics and vaccines. The RBD variants that we forecasted carry up to 11 concurrent amino acid substitutions at the host receptor binding site. Pseudoviruses harboring forecasted RBDs are active and escape binding and neutralization by FDA-approved monoclonal antibody therapeutics. We identified bottlenecks in the evolutionary landscape of SARS-CoV-2 that are promising targets for therapeutics that preempt evolution.
Antibody complementarity determining regions (CDRs) are loops within antibodies responsible for engaging antigens during the immune response and in antibody therapeutics and laboratory reagents. Since the 1980s, the conformations of the hypervariable CDRs have been structurally classified into a number of “canonical conformations” by Chothia, Lesk, Thornton, and others. In 2011 (North et al, J Mol Biol. 2011), we produced a quantitative clustering of approximately 300 structures of each CDR based on their length, a dihedral angle metric, and an affinity propagation algorithm. The data have been made available on our PyIgClassify website since 2015 and have been widely used in assigning conformational labels to antibodies in new structures and in molecular dynamics simulations. In the years since, it is has become apparent that many of the clusters are not “canonical” since they have not grown in size and still contain few sequences. Some clusters represent multiple conformations, given the assignment method we have used since 2015. Electron density calculations indicate that some clusters are due to misfitting of coordinates to electron density. In this work, we have performed a new statistical clustering of antibody CDR conformations. We used Electron Density in Atoms (EDIA, Meyder et al., 2017) to produce data sets with different levels of electron density validation. Clusters were chosen by their presence in high electron density cutoff data sets and with sufficient sequences (≥10) across the entire PDB (no EDIA cutoff). About half of the North et al. clusters have been “retired” and 13 new clusters have been identified. We also include clustering of the H4 and L4 CDRs, otherwise known as the “DE loop” which connects strands D and E of the variable domain. The DE loop sometimes contacts antigens and affects the structure of neighboring CDR1 and CDR2 loops. The current database contains 6,486 PDB antibody entries. The new clustering will be useful in the analysis and development of new antibody structure prediction and design algorithms based on rapidly emerging techniques in deep learning. The new clustering data are available at http://dunbrack2.fccc.edu/PyIgClassify2 .
Coronavirus disease 2019, caused by SARS-CoV-2, remains an on-going pandemic, partly due to the emergence of variant viruses that can "break-through" the protection of the current vaccines and neutralizing antibodies (nAbs), highlighting the needs for broadly nAbs and next-generation vaccines. We report an antibody that exhibits breadth and potency in binding the receptor-binding domain (RBD) of the virus spike glycoprotein across SARS coronaviruses. Initially, a lead antibody was computationally discovered and crystallographically validated that binds to a highly conserved surface of the RBD of wild-type SARS-CoV-2. Subsequently, through experimental affinity enhancement and computational affinity maturation, it was further developed to bind the RBD of all concerning SARS-CoV-2 variants, SARS-CoV-1 and pangolin coronavirus with pico-molar binding affinities, consistently exhibited strong neutralization activity against wild-type SARS-CoV-2 and the Alpha and Delta variants. These results identify a vulnerable target site on coronaviruses for development of pan-sarbecovirus nAbs and vaccines.
Biomolecular structure drives function, and computational capabilities have progressed such that the prediction and computational design of biomolecular structures is increasingly feasible. Because computational biophysics attracts students from many different backgrounds and with different levels of resources, teaching the subject can be challenging. One strategy to teach diverse learners is with interactive multimedia material that promotes self-paced, active learning. We have created a hands-on education strategy with a set of sixteen modules that teach topics in biomolecular structure and design, from fundamentals of conformational sampling and energy evaluation to applications like protein docking, antibody design, and RNA structure prediction. Our modules are based on PyRosetta, a Python library that encapsulates all computational modules and methods in the Rosetta software package. The workshop-style modules are implemented as Jupyter Notebooks that can be executed in the Google Colaboratory, allowing learners access with just a web browser. The digital format of Jupyter Notebooks allows us to embed images, molecular visualization movies, and interactive coding exercises. This multimodal approach may better reach students from different disciplines and experience levels as well as attract more researchers from smaller labs and cognate backgrounds to leverage PyRosetta in their science and engineering research. All materials are freely available at https://github.com/RosettaCommons/PyRosetta.notebooks.
Carbohydrate chains are ubiquitous in the complex molecular processes of life. These highly diverse chains are recognized by a variety of protein receptors, enabling glycans to regulate many biological functions. High-resolution structures of protein-glycoligand complexes reveal the atomic details necessary to understand this level of molecular recognition and inform application-focused scientific and engineering pursuits. When experimental challenges hinder high-throughput determination of quality structures, computational tools can, in principle, fill the gap. In this work, we introduce GlycanDock, a residue-centric protein-glycoligand docking refinement algorithm developed within the Rosetta macromolecular modeling and design software suite. We performed a benchmark docking assessment using a set of 109 experimentally determined protein-glycoligand complexes as well as 62 unbound protein structures. The GlycanDock algorithm can sample and discriminate among protein-glycoligand models of native-like structural accuracy with statistical reliability from starting structures of up to 7 Å root-mean-square deviation in the glycoligand ring atoms. We show that GlycanDock-refined models qualitatively replicated the known binding specificity of a bacterial carbohydrate-binding module. Finally, we present a protein-glycoligand docking pipeline for generating putative protein-glycoligand complexes when only the glycoligand sequence and unbound protein structure are known. In combination with other carbohydrate modeling tools, the GlycanDock docking refinement algorithm will accelerate research in the glycosciences.
Structure-based antibody and antigen design has advanced greatly in recent years, due not only to the increasing availability of experimentally determined structures but also to improved computational methods for both prediction and design. Constant improvements in performance within the Rosetta software suite for biomolecular modeling have given rise to a greater breadth of structure prediction, including docking and design application cases for antibody and antigen modeling. Here, we present an overview of current protocols for antibody and antigen modeling using Rosetta and exemplify those by detailed tutorials originally developed for a Rosetta workshop at Vanderbilt University. These tutorials cover antibody structure prediction, docking, and design and antigen design strategies, including the addition of glycans in Rosetta. We expect that these materials will allow novice users to apply Rosetta in their own projects for modeling antibodies and antigens.
Epoxide hydrolases catalyze the conversion of epoxides to vicinal diols in a range of cellular processes such as signaling, detoxification, and virulence. These enzymes typically utilize a pair of tyrosine residues to orient the substrate epoxide ring in the active site and stabilize the hydrolysis intermediate. A new subclass of epoxide hydrolases that utilize a histidine in place of one of the tyrosines was established with the discovery of the CFTR Inhibitory Factor (Cif) from Pseudomonas aeruginosa. Although the presence of such Cif-like epoxide hydrolases was predicted in other opportunistic pathogens based on sequence analyses, only Cif and its homolog aCif from Acinetobacter nosocomialis have been characterized. Here we report the biochemical and structural characteristics of Cfl1 and Cfl2, two Cif-like epoxide hydrolases from Burkholderia cenocepacia. Cfl1 is able to hydrolyze xenobiotic as well as biological epoxides that might be encountered in the environment or during infection. In contrast, Cfl2 shows very low activity against a diverse set of epoxides. The crystal structures of the two proteins reveal quaternary structures that build on the well-known dimeric assembly of the α/β hydrolase domain, but broaden our understanding of the structural diversity encoded in novel oligomer interfaces. Analysis of the interfaces reveals both similarities and key differences in sequence conservation between the two assemblies, and between the canonical dimer and the novel oligomer interfaces of each assembly. Finally, we discuss the effects of these higher-order assemblies on the intra-monomer flexibility of Cfl1 and Cfl2 and their possible roles in regulating enzymatic activity.
Each year vast international resources are wasted on irreproducible research. The scientific community has been slow to adopt standard software engineering practices, despite the increases in high-dimensional data, complexities of workflows, and computational environments. Here we show how scientific software applications can be created in a reproducible manner when simple design goals for reproducibility are met. We describe the implementation of a test server framework and 40 scientific benchmarks, covering numerous applications in Rosetta bio-macromolecular modeling. High performance computing cluster integration allows these benchmarks to run continuously and automatically. Detailed protocol captures are useful for developers and users of Rosetta and other macromolecular modeling tools. The framework and design concepts presented here are valuable for developers and users of any type of scientific software and for the scientific community to create reproducible methods. Specific examples highlight the utility of this framework and the comprehensive documentation illustrates the ease of adding new tests in a matter of hours.
AbstractCarbohydrates and glycoproteins modulate key biological functions. Computational approaches inform function to aid in carbohydrate structure prediction, structure determination, and design. However, experimental structure determination of sugar polymers is notoriously difficult as glycans can sample a wide range of low energy conformations, thus limiting the study of glycan-mediated molecular interactions. In this work, we expanded theRosettaCarbohydrateframework, developed and benchmarked effective tools for glycan modeling and design, and extended the Rosetta software suite to better aid in structural analysis and benchmarking tasks through the SimpleMetrics framework. We developed a glycan-modeling algorithm,GlycanTreeModeler, that computationally builds glycans layer-by-layer, using adaptive kernel density estimates (KDE) of common glycan conformations derived from data in the Protein Data Bank (PDB) and from quantum mechanics (QM) calculations. After a rigorous optimization of kinematic and energetic considerations to improve near-native sampling enrichment and decoy discrimination,GlycanTreeModelerwas benchmarked on a test set of diverse glycan structures, or “trees”. Structures predicted byGlycanTreeModeleragreed with native structures at high accuracy for bothde novomodeling and experimental density-guided building.GlycanTreeModeleralgorithms and associated tools were employed to designde novoglycan trees into a protein nanoparticle vaccine that are able to direct the immune response by shielding regions of the scaffold from antibody recognition. This work will inform glycoprotein model prediction, aid in both X-ray and electron microscopy density solutions and refinement, and help lead the way towards a new era of computational glycobiology.
Many scientific disciplines rely on computational methods for data analysis, model generation, and prediction. Implementing these methods is often accomplished by researchers with domain expertise but without formal training in software engineering or computer science. This arrangement has led to underappreciation of sustainability and maintainability of scientific software tools developed in academic environments. Some software tools have avoided this fate, including the scientific library Rosetta. We use this software and its community as a case study to show how modern software development can be accomplished successfully, irrespective of subject area. Rosetta is one of the largest software suites for macromolecular modeling, with 3.1 million lines of code and many state-of-the-art applications. Since the mid 1990s, the software has been developed collaboratively by the RosettaCommons, a community of academics from over 60 institutions worldwide with diverse backgrounds including chemistry, biology, physiology, physics, engineering, mathematics, and computer science. Developing this software suite has provided us with more than two decades of experience in how to effectively develop advanced scientific software in a global community with hundreds of contributors. Here we illustrate the functioning of this development community by addressing technical aspects (like version control, testing, and maintenance), community-building strategies, diversity efforts, software dissemination, and user support. We demonstrate how modern computational research can thrive in a distributed collaborative community. The practices described here are independent of subject area and can be readily adopted by other software development communities.