The development of a small-molecule suitable for clinical use involves optimization of ligand binding to an identified target protein but also frequently the reduction of binding to a variety of anti-target proteins known to pose translational risk. These anti-targets, often referred to as ADMET proteins, often contain large, highly flexible binding sites leading to promiscuous binding of small molecules. This property often frustrates efforts to execute structure-based design against these anti-targets. Here we demonstrate the application of a combination of induced fit docking and free energy perturbation to structurally enable several of the most common anti-targets: cytochrome P450 isoforms 3A4 and 2D6, the hERG ion channel, and PXR. Crucially, the enforcement of a consensus binding mode across a congeneric series allowed for the efficient exploration of a wide range of receptor conformation space. Results are presented for both public retrospective datasets and prospective application in active drug discovery programs.
Protein kinases constitute the second largest class of drug targets, most prominently in cancer therapy. In this work, we focus on type II kinase inhibitors that bind to the “classical” inactive conformation. We analyze binding affinities of 50 type II inhibitors across 348 kinases, combining earlier results for 16 inhibitors (the “Davis data set”) with recently obtained measurements for the remaining 34 (the “Schrödinger data set”). Using a Potts statistical energy model, we investigate the role of protein conformational reorganization in kinase selectivity and find that protein reorganization makes a large contribution to the selectivity (ROC AUC ∼0.8). We compare Potts threading predictions for the binding of 50 type II inhibitors to 348 kinases with those of DeepDTAGen, a sequence-based model trained on the Davis data set containing both type I and type II kinase inhibitors. DeepDTAGen performs well for the 16 inhibitors in the Davis data set but poorly for the remaining 34 in the Schrödinger data set, representing unseen data.
Abstract Cytochrome P450 3A4 (CYP3A4) metabolizes roughly half of all marketed drugs, and its inhibition can cause clinically significant drug-drug interactions. The enzyme accommodates chemically diverse ligands, making binding modes and metabolic outcomes difficult to predict. Previous X-ray crystallography efforts have leveraged a truncated construct without the N-terminal segment that tethers CYP3A4 to the membrane. Here we show that the same construct assembles into a symmetric trimer that can be resolved by cryo-EM and determine structures of both unliganded and ligand-bound CYP3A4. Multiple ligands are resolved with density consistent with several mutually exclusive conformations. Protein remodeling to reshape the binding pocket is concentrated in the F/G loop, which is poorly resolved and unmodeled in many X-ray structures. These features likely underlie the poor predictive performance of co-folding methods on this target. The routine use of cryo-EM to resolve CYP3A4 ligand-bound complexes will provide the ground truth data needed to make predictive models of drug metabolism useful in practice.
Abstract The generalizability of co-folding models for protein–ligand structure prediction remains unclear. Here, we benchmark Boltz, a state-of-the-art co-folding model, using a curated set of ligand-bound human G protein-coupled receptors (GPCRs) from families unseen during training. We show that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data when tested with FEP +. We further show that physics‑based refinement of Boltz models can correct ligand poses to near‑experimental accuracy and rescue FEP+ performance to that of the native structure. These results highlight the strengths and limitations of co-folding methods and motivate a workflow that pairs them with physics-based refinement and validation before high-stakes decisions in drug discovery.
Protein-ligand binding affinity prediction is essential for drug discovery and toxicity assessment. While machine learning (ML) promises fast and accurate predictions, its progress is constrained by the availability of reliable data. In contrast, physics-based methods such as absolute binding free energy perturbation (AB-FEP) deliver high accuracy but are computationally prohibitive for high-throughput applications. To bridge this gap, we introduce ToxBench, the first large-scale AB-FEP dataset designed for ML development and focused on a single pharmaceutically critical target, Human Estrogen Receptor Alpha (ERα). ToxBench contains 8,770 ERα-ligand complex structures with binding free energies computed via AB-FEP with a subset validated against experimental affinities at 1.75 kcal/mol RMSE, along with non-overlapping ligand splits to assess model generalizability. Using ToxBench, we further benchmark state-of-the-art ML methods, and notably, our proposed DualBind model, which employs a dual-loss framework to effectively learn the binding energy function. The benchmark results demonstrate the superior performance of DualBind and the potential of ML to approximate AB-FEP at a fraction of the computational cost.
High -quality predicted structures enable structure -based approaches to an expanding number of drug discovery programs. We propose that by utilizing free energy perturbation (FEP), predicted structures can confidently employed to achieve drug design goals. We use structure -based modeling of hERG inhibition to illustrate this value of FEP.
The recently developed AlphaFold2 (AF2) algorithm predicts proteins’ 3D structures from amino acid sequences. The open AlphaFold Protein Structure Database covers the complete human proteome. It shows great potential to provide structural information to enable and enhance existing and new drug discovery projects. Using an industry-leading molecular docking method (Glide), we benchmarked the virtual screening performance of 28 common drug targets each with an AF2 structure and known holo and apo structures from the DUD-E dataset. The AF2 structures show comparable early enrichment of known active compounds (avg. EF 1%: 13.16) to apo structures (avg. EF 1%: 11.56), while falling behind early enrichment of the holo structures (avg. EF 1%: 24.81). We also demonstrated that with the IFD-MD induced-fit docking approach, we can refine the AF2 structures using a known binding ligand to improve the performance in structure-based virtual screening (avg. EF 1%: 19.25). Thus, with proper preparation and refinement, AF2 structures show considerable promise for in silico hit identification.
Free energy perturbation (FEP) remains an indispensable method for computationally assaying prospective compounds in advance of synthesis. But before FEP can be deployed prospectively, it must demonstrate retrospective recapitulation of known experimental data where the subtle details of the atomic ligand-receptor model are consequential. An open question is whether AlphaFold models can serve as useful initial models for FEP in the regime where there exists a congeneric series of known chemical matter but where no experimental structures are available either of the target or of close homologues. As AlphaFold structures are provided without a ligand bound, we employ induced-fit docking to refine the AlphaFold models in the presence of one or more congeneric ligands. In this work, we first validate the performance of our latest induced-fit docking technology, IFD-MD on a retrospective set of public experimental GPCR structures with 95% of crossdocks produc-ing a pose with a ligand RMSD ≤ 2.5 Å in the top 2 predictions. We then apply IFD-MD and FEP on AlphaFold models of the somatostatin receptor family of GPCRs. We use AlphaFold models produced prior to the availability of any experi-mental structure from within this family. We arrive at FEP-validated models for SSTR2, SSTR4, and SSTR5, with RMSE around 1 kcal/mol and explore the challenges of model validation under scenarios of limited ligand-data, ample ligand data, and categorical data.
Accurate prediction of the pKa's of protein residues is crucial to many applications in biological simulation and drug discovery. Here, we present the use of free energy perturbation (FEP) calculations for the prediction of single-protein residue pKa values. We begin with an initial set of 191 residues with experimentally determined pKa values. To isolate sampling limitations from force field inaccuracies, we develop an algorithm to classify residues whose environments are significantly affected by crystal packing effects. We then report an approach to identify buried histidines that require significant sampling beyond what is achieved in typical FEP calculations. We therefore define a clean data set not requiring algorithms capable of predicting major conformational changes on which other pKa prediction methods can be tested. On this data set, we report an RMSE of 0.76 pKa units for 35 ASP residues, 0.51 pKa units for 44 GLU residues, and 0.67 pKa units for 76 HIS residues.
We present a reliable and accurate solution to the induced fit docking problem for protein-ligand binding by combining ligand-based pharmacophore docking (Phase), rigid receptor docking (Glide), and protein structure prediction (Prime) with explicit solvent molecular dynamics simulations. We provide an in-depth description of our novel methodology and present results for 41 targets consisting of 415 cross-docking cases divided amongst a training and test set. For both the training and test-set, we compute binding modes with a ligand-heavy atom RMSD to within 2.5 Å or better in over 90% of cross-docking cases compared to less than 70% of cross-docking cases using our previously published induced-fit docking algorithm and less than 41% using rigid receptor docking. Applications of the predicted ligand-receptor structure in free energy perturbation calculations is demonstrated for both public data and in active drug discovery projects, both retrospectively and prospectively.
Ligand docking is a widely used tool for lead discovery and binding mode prediction based drug discovery. The greatest challenges in docking occur when the receptor significantly reorganizes upon small molecule binding, thereby requiring an induced fit docking (IFD) approach in which the receptor is allowed to move in order to bind to the ligand optimally. IFD methods have had some success but suffer from a lack of reliability. Complementing IFD with all-atom molecular dynamics (MD) is a straightforward solution in principle but not in practice due to the severe time scale limitations of MD. Here we introduce a metadynamics plus IFD strategy for accurate and reliable prediction of the structures of protein-ligand complexes at a practically useful computational cost. Our strategy allows treating this problem in full atomistic detail and in a computationally efficient manner and enhances the predictive power of IFD methods. We significantly increase the accuracy of the underlying IFD protocol across a large data set comprising 42 different ligand-receptor systems. We expect this approach to be of significant value in computationally driven drug design.
Adeno-associated viruses (AAVs) are being development as gene delivery vectors and have shown promise in several clinical trials. And an AAV1-based vector has been approved by the European Commission for the treatment of lipoprotein lipase deficiency disease. However, limitations in tissue transduction specificity as well as neutralization by pre-existing antibodies remain are two major obstacles. To overcome these problems, it is important to characterize functional capsid regions that dictate virus binding to cellular receptors during infection and the potential overlap(s) of these regions with capsid antigenic epitopes. In this study, the sialic acid (SIA) binding site on AAV1 was determined by solving the structure of the AAV1-SIA complex using X-ray crystallography. Residues that form the SIA binding pocket are identical between AAV1 and AAV6; hence, structurally predicted SIA binding sites were mutated on both serotypes to confirm their role in the SIA interaction. Binding and transduction assays with these mutants in CHO cell lines Pro5, Lec2, Lec8, and Lec1, which display different terminal glycans, confirmed the structurally mapped binding pocket. Furthermore, native dot blot showed that several of the AAV1 SIA binding mutants can escape from antibody recognition. This finding is consistent with the overlap between the SIA binding site and the previously mapped AAV1 epitopes. Significantly, a glycan binding mutant with slightly reduced transduction ability is able to escape from antibody recognition. This study identifies an overlap between receptor recognition and antibody reactivity and provides amino acid level information for rational capsid engineering of AAV vectors for improved therapeutic efficacy.
Robust homology modeling to atomic-level accuracy requires in the general case successful prediction of protein loops containing small segments of secondary structure. Further, as loop prediction advances to success with larger loops, the exclusion of loops containing secondary structure becomes awkward. Here, we extend the applicability of the Protein Local Optimization Program (PLOP) to loops up to 17 residues in length that contain either helical or hairpin segments. In general, PLOP hierarchically samples conformational space and ranks candidate loops with a high-quality molecular mechanics force field. For loops identified to possess α-helical segments, we employ an alternative dihedral library composed of (ϕ,ψ) angles commonly found in helices. The alternative library is searched over a user-specified range of residues that define the helical bounds. The source of these helical bounds can be from popular secondary structure prediction software or from analysis of past loop predictions where a propensity to form a helix is observed. Due to the maturity of our energy model, the lowest energy loop across all experiments can be selected with an accuracy of sub-Ångström RMSD in 80% of cases, 1.0 to 1.5 Å RMSD in 14% of cases, and poorer than 1.5 Å RMSD in 6% of cases. The effectiveness of our current methods in predicting hairpin-containing loops is explored with hairpins up to 13 residues in length and again reaching an accuracy of sub-Ångström RMSD in 83% of cases, 1.0 to 1.5 Å RMSD in 10% of cases, and poorer than 1.5 Å RMSD in 7% of cases. Finally, we explore the effect of an imprecise surrounding environment, in which side chains, but not the backbone, are initially in perturbed geometries. In these cases, loops perturbed to 3Å RMSD from the native environment were restored to their native conformation with sub-Ångström RMSD.
Interactions between viruses and the host antibody immune response are critical in the development and control of disease, and antibodies are also known to interfere with the efficacy of viral vector-based gene delivery. The adeno-associated viruses (AAVs) being developed as vectors for corrective human gene delivery have shown promise in clinical trials, but preexisting antibodies are detrimental to successful outcomes. However, the antigenic epitopes on AAV capsids remain poorly characterized. Cryo-electron microscopy and three-dimensional image reconstruction were used to define the locations of epitopes to which monoclonal fragment antibodies (Fabs) against AAV1, AAV2, AAV5, and AAV6 bind. Pseudoatomic modeling showed that, in each serotype, Fabs bound to a limited number of sites near the protrusions surrounding the 3-fold axes of the T=1 icosahedral capsids. For the closely related AAV1 and AAV6, a common Fab exhibited substoichiometric binding, with one Fab bound, on average, between two of the three protrusions as a consequence of steric crowding. The other AAV Fabs saturated the capsid and bound to the walls of all 60 protrusions, with the footprint for the AAV5 antibody extending toward the 5-fold axis. The angle of incidence for each bound Fab on the AAVs varied and resulted in significant differences in how much of each viral capsid surface was occluded beyond the Fab footprints. The AAV-antibody interactions showed a common set of footprints that overlapped some known receptor-binding sites and transduction determinants, thus suggesting potential mechanisms for virus neutralization by the antibodies.
We review advances in implicit solvation and sampling algorithms which have resulted in enhanced capabilities in predicting and refining localized protein structures (e.g. loop regions) to high resolution. Improvements in the generalized Born model and hydrophobicity term yield significantly more accurate energetics; specialized sampling algorithms allow complex local structures, such as a loop-helix-loop region, to be reliably predicted. A novel penalty term is added for loops containing patterns of dihedrals seldom found in experimental structures. We show prediction of diverse sets of large loops, in the native backbone environment, to subångström accuracy. The methodology offers the promise of addressing the refinement problem in homology modeling if an approach can be devised to handle delocalized errors in the structure.