In CASP16, we assessed the ability of computational methods to predict the distribution of relative orientations of two domains tethered by a flexible linker. The range of interdomain distances and orientations (poses) of such domain-linker-domain (D-L-D) proteins can play an important role in protein function, allostery, aggregation, and the thermodynamics of binding. The CASP16 Conformational Ensembles Experiment included two challenges to predict the interdomain pose distribution of a Staphylococcal protein A (SpA) D-L-D construct, called ZLBT-C, in which two of SpA's five nearly identical domains are connected by either (1) a six-residue wild-type (WT) linker (KADNKF), or (2) an all-glycine (Gly6) linker. The wild-type linker has a highly conserved sequence and is thought to contribute to the energetic barrier for binding with host antibodies. Ground truth was provided by nuclear magnetic resonance (NMR) residual dipolar coupling (RDC) data on WT protein and small angle X-ray scattering (SAXS) data on both proteins in solution. Twenty-five predictor groups submitted 35 sets of predicted conformational distributions, in the form of population-weighted finite ensembles of discrete structures. Unlike traditional CASP assessments that compare predicted atomic models to experimental atomic models, the accuracy of these predictions was assessed by back-calculating NMR RDCs and SAXS curves from each ensemble of atomic models and comparing these results to respective experimental data. Accuracy was also assessed by using kernelization to compare ensembles to the continuous orientational distributions optimally fit to experimental data. In our assessment, predictions spanned a wide range of accuracy, but none were close fits to the combined NMR and SAXS data. In addition, none were able to recapitulate the observed difference between WT and Gly6 proteins, as observed in the SAXS data. These results, and our analysis, highlighted strengths and weaknesses, plus complementarity of NMR RDC and SAXS analysis.
Due to the favorable chemical properties of mirrored chiral centers (such as improved stability, bioavailability, and membrane permeability) the computational design of D-peptides targeting biological L-proteins is a valuable area of research. To design these structures in silico , a computational workflow should correctly dock and fold a peptide while maintaining chiral centers. The latest AlphaFold 3 (AF3) from Abramson et al. (2024) enforces a strict chiral violation penalty to maintain chiral centers from model inputs and is reported to have a low chiral violation rate of only 4.4% on a PoseBusters benchmark containing diverse chiral molecules. Herein, we report the results of 3,255 experiments with AF3 to evaluate its ability to predict the fold, chirality, and binding pose of D-peptides in heterochiral complexes. Despite our inputs specifying explicit D-stereocenters, we report that the AF3 chiral violation for D-peptide binders is much higher at 51% across all evaluated predictions; on average the model is as accurate as chance (random chirality choice, L or D, for each peptide residue). Increasing the number of seeds failed to improve this violation rate. The AF3 predictions exhibit incorrect folds and binding poses, with D-peptides commonly oriented incorrectly in the L-protein binding pocket. Confidence metrics returned by AF3 also fail to distinguish predictions with low chirality violation and correct docking vs. predictions with high chirality violation and incorrect docking. We conclude that AF3 is a poor predictor of D-peptide chirality, fold, and binding pose. Summary:A crucial task in computational protein design is predicting fold, as this property determines the structure and function of a protein. Abramson et al. 1 published in Nature on AlphaFold 3 (AF3), a powerful deep learning framework for predicting chemical structures in both bound and unbound states. This architecture is tuned to respect chiral centers, which are atoms (in proteins, backbone α -carbons) covalently bound to four different chemical species 2 . These centers adopt two non-superposable forms, often called "handedness," termed L (all biological proteins adopt this form) and D (the mirror image of L). L and D chiral centers exert significant influence on chemical function; changing the chirality of even a single residue can dramatically alter chemical properties such as enantioselective binding (e.g., antifolate resistance 3 ) and stability 4 . Additionally, D-peptides (small proteins containing exclusively D chiral centers) exhibit many advantages compared to their L-peptide counterparts, such as protease evasion 5 , and are therefore therapeutically relevant modalities. Due to vastly differing chemical properties, an algorithm should respect chiral center inputs and exhibit an error rate of 0%. Although Abramson et al. 1 reports a low 4.4% chirality violation across diverse chiral centers, we have found that the chiral violation rate for D-peptides with D chiral center inputs explicitly specified is much higher at 51%. Increasing the number of seeds fails to improve this rate. Our data highlights a crucial structural prediction error in AF3 and demonstrates the model is as accurate on average as chance (random chirality choice, L or D, for each peptide residue). Compared to empirical structures, AF3 is also highly inaccurate when folding and docking D-peptide:L-protein complexes. The failure of AF3 to accurately predict these chemical interactions indicates more work is need for high-quality prediction of D-peptides.
Accurate binding affinity prediction is crucial to structure-based drug design. Recent work used computational topology to obtain an effective representation of protein-ligand interactions. Although persistent homology encodes geometric features, previous works on binding affinity prediction using persistent homology employed uninterpretable machine learning models and failed to explain the underlying geometric and topological features that drive accurate binding affinity prediction. In this work, we propose a novel, interpretable algorithm for protein-ligand binding affinity prediction. Our algorithm achieves interpretability through an effective embedding of distances across bipartite matchings of the protein and ligand atoms into real-valued functions by summing Gaussians centered at features constructed by persistent homology. We name these functions internuclear persistent contours (IPCs) . Next, we introduce persistence fingerprints , a vector with 10 components that sketches the distances of different bipartite matching between protein and ligand atoms, refined from IPCs. Let the number of protein atoms in the protein-ligand complex be n , number of ligand atoms be m , and ω ≈ 2.4 be the matrix multiplication exponent. We show that for any 0 < ε < 1, after an 𝒪 ( mn log( mn )) preprocessing procedure, we can compute an ε -accurate approximation to the persistence fingerprint in 𝒪 ( m log 6 ω ( m/" )) time, independent of protein size. This is an improvement in time complexity by a factor of 𝒪 (( m + n ) 3 ) over any previous binding affinity prediction that uses persistent homology. We show that the representational power of persistence fingerprint generalizes to protein-ligand binding datasets beyond the training dataset. Then, we introduce PATH , Predicting Affinity Through Homology, an interpretable, small ensemble of shallow regression trees for binding affinity prediction from persistence fingerprints. We show that despite using 1,400-fold fewer features, PATH has comparable performance to a previous state-of-the-art binding affinity prediction algorithm that uses persistent homology features. Moreover, PATH has the advantage of being interpretable. Finally, we visualize the features captured by persistence fingerprint for variant HIV-1 protease complexes and show that persistence fingerprint captures binding-relevant structural mutations. The source code for PATH is released open-source as part of the osprey protein design software package.
D-peptides, the mirror image of canonical L-peptides, offer numerous biological advantages that make them effective therapeutics. This article details how to use DexDesign, the newest OSPREY-based algorithm, for designing these D-peptides de novo. OSPREY physics-based models precisely mimic energy-equivariant reflection operations, enabling the generation of D-peptide scaffolds from L-peptide templates. Due to the scarcity of D-peptide:L-protein structural data, DexDesign calls a geometric hashing algorithm, Method of Accelerated Search for Tertiary Ensemble Representatives, as a subroutine to produce a synthetic structural dataset. DexDesign enables mixed-chirality designs with a new user interface and also reduces the conformation and sequence search space using three new design techniques: Minimum Flexible Set, Inverse Alanine Scanning, and K*-based Mutational Scanning.
AbstractWith over 270 unique occurrences in the human genome, peptide-recognizing PDZ domains play a central role in modulating polarization, signaling, and trafficking pathways. Mutations in PDZ domains lead to diseases such as cancer and cystic fibrosis, making PDZ domains attractive targets for therapeutic intervention. D-peptide inhibitors offer unique advantages as therapeutics, including increased metabolic stability and low immunogenicity. Here, we introduce DexDesign, a novel OSPREY-based algorithm for computationally designingde novoD-peptide inhibitors. DexDesign leverages three novel techniques that are broadly applicable to computational protein design: the Minimum Flexible Set, K*-based Mutational Scan, and Inverse Alanine Scan, which enable exponential reductions in the size of the peptide sequence search space. We apply these techniques and DexDesign to generate novel D-peptide inhibitors of two biomedically important PDZ domain targets: CAL and MAST2. We introduce a new framework for analyzingde novopeptides—evaluation along a replication/restitution axis—and apply it to the DexDesign-generated D-peptides. Notably, the peptides we generated are predicted to bind their targets tighter than their targets’ endogenous ligands, validating the peptides’ potential as lead therapeutic candidates. We provide an implementation of DexDesign in the free and open source computational protein design software OSPREY.
D-peptide inhibitors offer unique advantages as therapeutics, including increased metabolic stability and low immunogenicity. We introduce DexDesign, an OSPREY-based algorithm for computationally designing de novo D-peptide inhibitors, and use it to design inhibitors of two biomedically important PDZ domains. Novel techniques enabling exponential reductions in peptide search space are presented: Minimum Flexible Set, Inverse Alanine Scans, and K*-based Mutational Scans. Designed D-peptide inhibitors are predicted to significantly improve binding affinity (KD) over the endogenous ligand.
Prospective predictions of drug-resistant protein mutants could improve the design of therapeutics less prone to resistance. Here, we describe RESISTOR, an algorithm that uses structure- and sequence-based criteria to predict resistance mutations. We demonstrate the process of using RESISTOR to predict ERK2 mutants likely to arise in melanoma ablating the efficacy of the ERK1/2 inhibitor SCH772984. RESISTOR is included in the free and open-source computational protein design software OSPREY. For complete details on the use and execution of this protocol, please refer to Guerin et al..1.
We report an Osprey-based computational protocol to prospectively identify oncogenic mutations that act via disruption of molecular interactions. It is applicable to analyse both protein-protein and protein-DNA interfaces and it is validated on a dataset of clinically relevant mutations. In addition, it is used to predict previously uncharacterised patient mutations in CDK6 and p16 genes, which are experimentally confirmed to impair complex formation.
Summary: Broadly neutralizing antibodies (bNAbs) against HIV can reduce viral transmission in humans, but an effective therapeutic will require unusually high breadth and potency of neutralization. We employ the OSPREY computational protein design software to engineer variants of two apex-directed bNAbs, PGT145 and PG9RSH, resulting in increases in potency of over 100-fold against some viruses. The top designed variants improve neutralization breadth from 39% to 54% at clinically relevant concentrations (IC80 < 1 μg/mL) and improve median potency (IC80) by up to 4-fold over a cross-clade panel of 208 strains. To investigate the mechanisms of improvement, we determine cryoelectron microscopy structures of each variant in complex with the HIV envelope trimer. Surprisingly, we find the largest increases in breadth to be a result of optimizing side-chain interactions with highly variable epitope residues. These results provide insight into mechanisms of neutralization breadth and inform strategies for antibody design and improvement.
Podisus maculiventris thanatin has been reported as a potent antimicrobial peptide with antibacterial and antifungal activity. Its antibiotic activity has been most thoroughly characterized against E. coli and shown to interfere with multiple pathways, such as the lipopolysaccharide transport (LPT) pathway comprised of seven different Lpt proteins. Thanatin binds to E. coli LptA and LptD, thus disrupting the LPT complex formation and inhibiting cell wall synthesis and microbial growth. Here, we performed a genomic database search to uncover novel thanatin orthologs, characterized their binding to E. coli LptA using bio-layer interferometry, and assessed their antimicrobial activity against E. coli. We found that thanatins from Chinavia ubica and Murgantia histrionica bound tighter (by 3.6- and 2.2-fold respectively) to LptA and exhibited more potent antibiotic activity (by 2.1- and 2.8-fold respectively) than the canonical thanatin from P. maculiventris. We crystallized and determined the LptA-bound complex structures of thanatins from C. ubica (1.90 Å resolution), M. histrionica (1.80 Å resolution), and P. maculiventris (2.43 Å resolution) to better understand their mechanism of action. Our structural analysis revealed that residues A10 and I21 in C. ubica and M. histrionica thanatin are important for improving the binding interface with LptA, thus overall improving the potency of thanatin against E. coli. We also designed a stapled variant of thanatin that removes the need for a disulfide bond but retains the ability to bind LptA and antibiotic activity. Our discovery presents a library of novel thanatin sequences to serve as starting scaffolds for designing more potent antimicrobial therapeutics.
Computational, in silico prediction of resistance-conferring escape mutations could accelerate the design of therapeutics less prone to resistance. This article describes how to use the Resistor algorithm to predict escape mutations. Resistor employs Pareto optimization on four resistance-conferring criteria-positive and negative design, mutational probability, and hotspot cardinality-to assign a Pareto rank to each prospective mutant. It also predicts the mechanism of resistance, that is, whether a mutant ablates binding to a drug, strengthens binding to the endogenous ligand, or a combination of these two factors, and provides structural models of the mutants. Resistor is part of the free and open-source computational protein design software OSPREY.
Resistance to pharmacological treatments is a major public health challenge. Here, we introduce Resistor—a structure- and sequence-based algorithm that prospectively predicts resistance mutations for drug design. Resistor computes the Pareto frontier of four resistance-causing criteria: the change in binding affinity (ΔKa) of the (1) drug and (2) endogenous ligand upon a protein’s mutation; (3) the probability a mutation will occur based on empirically derived mutational signatures; and (4) the cardinality of mutations comprising a hotspot. For validation, we applied Resistor to EGFR and BRAF kinase inhibitors treating lung adenocarcinoma and melanoma. Resistor correctly identified eight clinically significant EGFR resistance mutations, including the erlotinib and gefitinib “gatekeeper” T790M mutation and five known osimertinib resistance mutations. Furthermore, Resistor predictions are consistent with BRAF inhibitor sensitivity data from both retrospective and prospective experiments using KinCon biosensors. Resistor is available in the open-source protein design software OSPREY.
Antimicrobial resistance presents a significant health care crisis. The mutation F98Y in Staphylococcus aureus dihydrofolate reductase (SaDHFR) confers resistance to the clinically important antifolate trimethoprim (TMP). Propargyl-linked antifolates (PLAs), next generation DHFR inhibitors, are much more resilient than TMP against this F98Y variant, yet this F98Y substitution still reduces efficacy of these agents. Surprisingly, differences in the enantiomeric configuration at the stereogenic center of PLAs influence the isomeric state of the NADPH cofactor. To understand the molecular basis of F98Y-mediated resistance and how PLAs’ inhibition drives NADPH isomeric states, we used protein design algorithms in the osprey protein design software suite to analyze a comprehensive suite of structural, biophysical, biochemical, and computational data. Here, we present a model showing how F98Y SaDHFR exploits a different anomeric configuration of NADPH to evade certain PLAs’ inhibition, while other PLAs remain unaffected by this resistance mechanism.
The K* algorithm provably approximates partition functions for a set of states (e.g., protein, ligand, and protein-ligand complex) to a user-specified accuracy ε. Often, reaching an ε-approximation for a particular set of partition functions takes a prohibitive amount of time and space. To alleviate some of this cost, we introduce two new algorithms into the osprey suite for protein design: fries, a Fast Removal of Inadequately Energied Sequences, and EWAK*, an Energy Window Approximation to K*. fries pre-processes the sequence space to limit a design to only the most stable, energetically favorable sequence possibilities. EWAK* then takes this pruned sequence space as input and, using a user-specified energy window, calculates K* scores using the lowest energy conformations. We expect fries/EWAK* to be most useful in cases where there are many unstable sequences in the design sequence space and when users are satisfied with enumerating the low-energy ensemble of conformations. In combination, these algorithms provably retain calculational accuracy while limiting the input sequence space and the conformations included in each partition function calculation to only the most energetically favorable, effectively reducing runtime while still enriching for desirable sequences. This combined approach led to significant speed-ups compared to the previous state-of-the-art multi-sequence algorithm, BBK*, while maintaining its efficiency and accuracy, which we show across 40 different protein systems and a total of 2,826 protein design problems. Additionally, as a proof of concept, we used these new algorithms to redesign the protein-protein interface (PPI) of the c-Raf-RBD:KRas complex. The Ras-binding domain of the protein kinase c-Raf (c-Raf-RBD) is the tightest known binder of KRas, a protein implicated in difficult-to-treat cancers. fries/EWAK* accurately retrospectively predicted the effect of 41 different sets of mutations in the PPI of the c-Raf-RBD:KRas complex. Notably, these mutations include mutations whose effect had previously been incorrectly predicted using other computational methods. Next, we used fries/EWAK* for prospective design and discovered a novel point mutation that improves binding of c-Raf-RBD to KRas in its active, GTP-bound state (KRasGTP). We combined this new mutation with two previously reported mutations (which were highly-ranked by osprey) to create a new variant of c-Raf-RBD, c-Raf-RBD(RKY). fries/EWAK* in osprey computationally predicted that this new variant binds even more tightly than the previous best-binding variant, c-Raf-RBD(RK). We measured the binding affinity of c-Raf-RBD(RKY) using a bio-layer interferometry (BLI) assay, and found that this new variant exhibits single-digit nanomolar affinity for KRasGTP, confirming the computational predictions made with fries/EWAK*. This new variant binds roughly five times more tightly than the previous best known binder and roughly 36 times more tightly than the design starting point (wild-type c-Raf-RBD). This study steps through the advancement and development of computational protein design by presenting theory, new algorithms, accurate retrospective designs, new prospective designs, and biochemical validation.
The spread of plasmid borne resistance enzymes in clinical Staphylococcus aureus isolates is rendering trimethoprim and iclaprim, both inhibitors of dihydrofolate reductase (DHFR), ineffective. Continued exploitation of these targets will require compounds that can broadly inhibit these resistance-conferring isoforms. Using a structure-based approach, we have developed a novel class of ionized nonclassical antifolates (INCAs) that capture the molecular interactions that have been exclusive to classical antifolates. These modifications allow for a greatly expanded spectrum of activity across these pathogenic DHFR isoforms, while maintaining the ability to penetrate the bacterial cell wall. Using biochemical, structural, and computational methods, we are able to optimize these inhibitors to the conserved active sites of the endogenous and trimethoprim resistant DHFR enzymes. Here, we report a series of INCA compounds that exhibit low nanomolar enzymatic activity and potent cellular activity with human selectivity against a panel of clinically relevant TMP resistant (TMPR) and methicillin resistant Staphylococcus aureus (MRSA) isolates.
The CFTR-associated ligand PDZ domain (CALP) binds to the cystic fibrosis transmembrane conductance regulator (CFTR) and mediates lysosomal degradation of mature CFTR. Inhibition of this interaction has been explored as a therapeutic avenue for cystic fibrosis. Previously, we reported the ensemble-based computational design of a novel peptide inhibitor of CALP, which resulted in the most binding-efficient inhibitor to date. This inhibitor, kCAL01, was designed using osprey and evinced significant biological activity in in vitro cell-based assays. Here, we report a crystal structure of kCAL01 bound to CALP and compare structural features against iCAL36, a previously developed inhibitor of CALP. We compute side-chain energy landscapes for each structure to not only enable approximation of binding thermodynamics but also reveal ensemble features that contribute to the comparatively efficient binding of kCAL01. Finally, we compare the previously reported design ensemble for kCAL01 vs the new crystal structure and show that, despite small differences between the design model and crystal structure, significant biophysical features that enhance inhibitor binding are captured in the design ensemble. This suggests not only that ensemble-based design captured thermodynamically significant features observed in vitro, but also that a design eschewing ensembles would miss the kCAL01 sequence entirely.
We review algorithms for protein design in general. Although these algorithms have a rich combinatorial, geometric, and mathematical structure, they are almost never covered in computer science classes. Furthermore, many of these algorithms admit provable guarantees of accuracy, soundness, complexity, completeness, optimality, and approximation bounds. The algorithms represent a delicate and beautiful balance between discrete and continuous computation and modeling, analogous to that which is seen in robotics, computational geometry, and other fields in computational science. Finally, computer scientists may be unaware of the almost direct impact of these algorithms for predicting and introducing molecular therapies that have gone in a short time from mathematics to algorithms to software to predictions to preclinical testing to clinical trials. Indeed, the overarching goal of these algorithms is to enable the development of new therapeutics that might be impossible or too expensive to discover using experimental methods. Thus the potential impact of these algorithms on individual, community, and global health has the potential to be quite significant.
Computational protein design is a transformative field with exciting prospects for advancing both basic science and translational medical research. New algorithms blend discrete and continuous geometry to address the challenges of creating designer proteins. I will discuss recent progress in this area and some interesting open problems. I will motivate this talk by discussing how, by using continuous geometric representations within a discrete optimization framework, broadly-neutralizing anti-HIV-1 antibodies were computationally designed that are now being tested in humans - the designed antibodies are currently in eight clinical trials, one of which is Phase 2a (NCT03721510). These continuous representations model the flexibility and dynamics of biological macromolecules, which are an important structural determinant of function. However, reconstruction of biomolecular dynamics from experimental observables requires the determination of a conformational probability distribution. These distributions are not fully constrained by the limited geometric information from experiments, making the problem ill-posed in the sense of Hadamard. The ill-posed nature of the problem comes from the fact that it has no unique solution. Multiple or even an infinite number of solutions may exist. To avoid the ill-posed nature, the problem must be regularized by making (hopefully reasonable) assumptions. I will present new ways to both represent and visualize correlated inter-domain protein motions. We use Bingham distributions, based on a quaternion fit to circular moments of a physics-based quadratic form. To find the optimal solution for the distribution, we designed an efficient, provable branch-and-bound algorithm that exploits the structure of analytical solutions to the trigonometric moment problem. Hence, continuous conformational PDFs can be determined directly from NMR measurements. The representation works especially well for multi-domain systems with broad conformational distributions. For more information please see Y. Qi et al. Jour. Mol. Biol. 2018; 430(18 Pt B):3412-3426. doi: 10.1016/j.jmb.2018.06.022. Ultimately, this method has parallels to other branches of geometric computing that balance discrete and continuous representations, including physical geometric algorithms, robotics, computational geometry, and robust optimization. I will advocate for using continuous distributions for protein modeling, and describe future work and open problems.
Protein design algorithms that model continuous sidechain flexibility and conformational ensembles better approximate the in vitro and in vivo behavior of proteins. The previous state of the art, iMinDEE-A*-K*, computes provable ɛ-approximations to partition functions of protein states (e.g., bound vs. unbound) by computing provable, admissible pairwise-minimized energy lower bounds on protein conformations, and using the A* enumeration algorithm to return a gap-free list of lowest-energy conformations. iMinDEE-A*-K* runs in time sublinear in the number of conformations, but can be trapped in loosely-bounded, low-energy conformational wells containing many conformations with highly similar energies. That is, iMinDEE-A*-K* is unable to exploit the correlation between protein conformation and energy: similar conformations often have similar energy. We introduce two new concepts that exploit this correlation: Minimization-Aware Enumeration and Recursive K*. We combine these two insights into a novel algorithm, Minimization-Aware Recursive K* (MARK*), which tightens bounds not on single conformations, but instead on distinct regions of the conformation space. We compare the performance of iMinDEE-A*-K* versus MARK* by running the Branch and Bound over K* (BBK*) algorithm, which provably returns sequences in order of decreasing K* score, using either iMinDEE-A*-K* or MARK* to approximate partition functions. We show on 200 design problems that MARK* not only enumerates and minimizes vastly fewer conformations than the previous state of the art, but also runs up to 2 orders of magnitude faster. Finally, we show that MARK* not only efficiently approximates the partition function, but also provably approximates the energy landscape. To our knowledge, MARK* is the first algorithm to do so. We use MARK* to analyze the change in energy landscape of the bound and unbound states of an HIV-1 capsid protein C-terminal domain in complex with a camelid VHH, and measure the change in conformational entropy induced by binding. Thus, MARK* both accelerates existing designs and offers new capabilities not possible with previous algorithms.