Relative alchemical binding free energy calculations can be used to predict the effect of amino acid mutations on ligand binding affinities. However, these protocols are not well established for proteins containing intrinsically disordered regions (IDRs). In this work, we focus on the development of robust protein-free energy perturbation (FEP) protocols to reproduce experimental binding affinities that have been measured for a panel of mutants of the protein MDM2 against two ligands, AM-7209 and Nutlin-3a. We focus on mutations that occur in the N-terminal IDR lid of MDM2, which is known to undergo ligand-dependent folding upon binding. We systematically assess the effectiveness of both equilibrium and nonequilibrium alchemical protocols in reproducing these experimental binding affinities, in particular for mutations with slowly varying degrees of freedom. We show that the equilibrium protocol outperforms the nonequilibrium protocol in the precision of the free energy estimates obtained. In addition, we demonstrate the effect of the protein force field and the water model used to simulate the highly flexible IDR region. Overall, our findings demonstrate an accurate FEP protocol capable of reproducing these trends and further show the applicability of FEP protocols for elucidating the mutational effects on ligand binding affinity in highly dynamic intrinsically disordered protein regions.
A methodology that combines alchemical free energy calculations (FEP) with machine learning (ML) has been developed to compute accurate absolute hydration free energies. The hybrid FEP/ML methodology was trained on a subset of the FreeSolv database, and retrospectively shown to outperform most submissions from the SAMPL4 competition. Compared to pure machine-learning approaches, FEP/ML yields more precise estimates of free energies of hydration, and requires a fraction of the training set size to outperform standalone FEP calculations. The ML-derived correction terms are further shown to be transferable to a range of related FEP simulation protocols. The approach may be used to inexpensively improve the accuracy of FEP calculations, and to flag molecules which will benefit the most from bespoke forcefield parameterisation efforts.
Bioisostere replacement is a powerful and popular tool used to optimize the potency and selectivity of candidate molecules in drug discovery. Selecting the right bioisosteres to invest resources in for synthesis and subsequent optimization is key to an efficient drug discovery project. Here we demonstrate how 3D-quantitative structure activity relationship (3D-QSAR), and relative binding free energy calculations can be combined into an active learning workflow to prioritize molecules from a pool of hundreds of bioisosteres. We demonstrate on a human aldose reductase test case that the use of this workflow can rapidly locate the strongest-binding bioisosteric replacements with a relatively modest computational cost.
Despite growing accessibility due to advances in computing power, ligand-protein Relative Binding Free Energy (RBFE) calculation remains a resource-intensive activity, which limits its practical application. Improved computational efficiency via bespoke sampling of the alchemical transformation coordinate, λ, is under-explored. We show here that a simple approach for on-the-fly bespoke λ scheduling, named Adaptive Lambda Scheduling (ALS), yields significant reduction in computational cost whilst retaining predictive performance.
Free energy calculations have seen increased usage in structure-based drug design. Despite the rising interest, automation of the complex calculations and subsequent analysis of their results are still hampered by the restricted choice of available tools. In this work, an application for automated setup and processing of free energy calculations is presented. Several sanity checks for assessing the reliability of the calculations were implemented, constituting a distinct advantage over existing open-source tools. The underlying workflow is built on top of the software Sire, SOMD, BioSimSpace and OpenMM and uses the AMBER14SB and GAFF2.1 force fields. It was validated on two datasets originally composed by Schrödinger, consisting of 14 protein structures and 220 ligands. Predicted binding affinities were in good agreement with experimental values. For the larger dataset the average correlation coefficient Rp was 0.70 ± 0.05 and average Kendall’s τ was 0.53 ± 0.05 which is broadly comparable to or better than previously reported results using other methods.
Fibrillar protein aggregates are characteristic of neurodegenerative diseases but represent difficult targets for ligand design, because limited structural information about the binding sites is available. Ligand-based virtual screening has been used to develop a computational method for the selection of new ligands for Aβ(1-42) fibrils, and five new ligands have been experimentally confirmed as nanomolar affinity binders. A database of ligands for Aβ(1-42) fibrils was assembled from the literature and used to train models for the prediction of dissociation constants based on chemical structure. The virtual screening pipeline consists of three steps: a molecular property filter based on charge, molecular weight, and logP; a machine learning model based on simple chemical descriptors; and machine learning models that use field points as a 3D description of shape and surface properties in the Forge software. The three-step pipeline was used to virtually screen 698 million compounds from the ZINC15 database. From the top 100 compounds with the highest predicted affinities, 46 compounds were experimentally investigated by using a thioflavin T fluorescence displacement assay. Five new Aβ(1-42) ligands with dissociation constants in the range 20-600 nM and novel structures were identified, demonstrating the power of this ligand-based approach for discovering new structurally unique, high-affinity amyloid ligands. The experimental hit rate using this virtual screening approach was 10.9%.
A ligand-based virtual screening pipeline was developed to discover structurally novel ligands for Aβ(1-42) fibrils. A database of ligands that bind to Aβ(1-42) fibrils was used to develop a two-step ligand-based virtual screening pipeline. The first step involved developing machine learnings models that use relatively simple chemical descriptors to predict ligand dissociation constants ( K d ). The second step involved constructing 3D models based on field points that describe the surface, shape, and electronic properties of ligands. The combined pipeline was then used to screen 63 million compounds from the ZINC15 database. 1 The most promising compounds were experimentally evaluated using binding assays. For the first chemical descriptor model, a support vector machine model was found to predict the -log( K d /M) of Aβ(1-42) fibril ligands with a mean absolute error of only 0.41. An ensemble of four 3D models were then developed that each predicted -log( K d /M) with a cross-validation regression coefficient of 0.48–0.56. The ZINC15 database compounds were screened through the first chemical descriptor model, and the 10,000 compounds with the highest predicted affinities were then screened through the 3D models. The 46 highest-ranked ligands were then evaluated using Thioflavin T competition assays, resulting in five new Aβ(1-42) ligands with novel structures that exhibited nanomolar binding. A ligand-based virtual screening pipeline has been developed to discover new fibril-binding ligands. Using Aβ(1-42) fibrils as a target, five new structurally diverse ligands were found that exhibit nanomolar binding affinities. 1. Sterling and J. J. Irwin, J. Chem. Inf. Model . 2015 , 55 , 2324-2337.
The development of accurate transferable force fields is key to realizing the full potential of atomistic modeling in the study of biological processes such as protein-ligand binding for drug discovery. State-of-the-art transferable force fields, such as those produced by the Open Force Field Initiative, use modern software engineering and automation techniques to yield accuracy improvements. However, force field torsion parameters, which must account for many stereoelectronic and steric effects, are considered to be less transferable than other force field parameters and are therefore often targets for bespoke parametrization. Here, we present the Open Force Field QCSubmit and BespokeFit software packages that, when combined, facilitate the fitting of torsion parameters to quantum mechanical reference data at scale. We demonstrate the use of QCSubmit for simplifying the process of creating and archiving large numbers of quantum chemical calculations, by generating a dataset of 671 torsion scans for druglike fragments. We use BespokeFit to derive individual torsion parameters for each of these molecules, thereby reducing the root-mean-square error in the potential energy surface from 1.1 kcal/mol, using the original transferable force field, to 0.4 kcal/mol using the bespoke version. Furthermore, we employ the bespoke force fields to compute the relative binding free energies of a congeneric series of inhibitors of the TYK2 protein, and demonstrate further improvements in accuracy, compared to the base force field (MUE reduced from 0.560.390.77 to 0.420.280.59 kcal/mol and R2 correlation improved from 0.720.350.87 to 0.930.840.97).
Relative binding free energy (RBFE) calculations are increasingly used to support the ligand optimisation problem in early-stage drug discovery. Because RBFE calculations frequently rely on alchemical perturbations between ligands in a congeneric series, practitioners are required to estimate an optimal combination of pairwise perturbations for each series. RBFE networks constitute in a collection of edges chosen such that all ligands (nodes) are included in the network, where each edge represents a pairwise RBFE calculation. As there is a vast number of possible configurations it is not trivial to select an optimal perturbation network. Current approaches rely on human intuition and rule-based expert systems for proposing RBFE perturbation networks. This work presents a data-driven alternative to rule-based approaches by using a graph siamese neural network architecture. A novel dataset, RBFE-Space, is presented as a representative and transferable training domain for RBFE machine learning research. The workflow presented in this work matches state-of-the-art programmatic RBFE network generation performance with several key benefits. The workflow provides full transferability of the network generator because RBFE-Space is open-sourced and ready to be applied to other RBFE software. Additionally, the deep learning model represents the first machine-learned predictor of perturbation reliability in RBFE calculations.
Free energy calculations have seen increased usage in structure-based drug design. Despite the rising interest, automation of the complex calculations and subsequent analysis of their results are still hampered by the restricted choice of available tools. In this work, an application for automated setup and processing of free energy calculations is presented. Several sanity checks for assessing the reliability of the calculations were implemented, constituting a distinct advantage over existing open-source tools. The underlying workflow is built on top of the software Sire, SOMD, BioSimSpace and OpenMM and uses the AMBER14SB and GAFF2.1 force fields. It was validated on two datasets originally composed by Schrödinger, consisting of 14 protein structures and 220 ligands. Predicted binding affinities were in good agreement with experimental values. For the larger dataset the average correlation coefficient Rp was 0.70 ± 0.05 and average Kendall's τ was 0.53 ± 0.05 which is broadly comparable to or better than previously reported results using other methods.
A methodology that combines alchemical free energy calculations (FEP) with machine learning (ML) has been developed to compute accurate absolute hydration free energies. The hybrid FEP/ML methodology was trained on a subset of the FreeSolv database, and retrospectively shown to outperform most submissions from the SAMPL4 competition. Compared to pure machine-learning approaches, FEP/ML yields more precise estimates of free energies of hydration, and requires a fraction of the training set size to outperform standalone FEP calculations. The ML-derived correction terms are further shown to be transferable to a range of related FEP simulation protocols. The approach may be used to inexpensively improve the accuracy of FEP calculations, and to flag molecules which will benefit the most from bespoke forcefield parameterisation efforts.
The analysis of activity landscapes andactivity cliffs is a widely used method to locate critical regions of SAR.Knowledge of what changes in a series of molecules caused unexpectedly largechanges in affinity allows the chemist to focus on the molecular features whichare crucial for activity. We examine the usefulness of activity cliff analysiswith a metric based on 3D shape and electrostatic similarity, utilizing aligand-based alignment method. We demonstrate that 3D activity cliff analysisis complementary to the more usual 2D fingerprint-based methods, in that eachfinds cliffs that the other misses. Moreover, we show that analysis of theactivity landscape in the context of a consensus 3D alignment allows the sourceof the activity cliff to be investigated in terms of the effect that astructural change has on the steric and electrostatic properties of a molecule.The technique is illustrated with two set of compounds with activity againstacetylcholinesterase and dipeptidyl peptidase.
Electrostatic interactions between small molecules and their respective receptors are essential for molecular recognition and are also key contributors to the binding free energy. Assessing the electrostatic match of protein-ligand complexes therefore provides important insights into why ligands bind and what can be changed to improve binding. Ideally, the ligand and protein electrostatic potentials at the protein-ligand interaction interface should maximize their complementarity while minimizing desolvation penalties. In this work, we present a fast and efficient tool to calculate and visualize the electrostatic complementarity (EC) of protein-ligand complexes. We compiled benchmark sets demonstrating electrostatically driven structure-activity relationships (SAR) from literature data, including kinase, protein-protein interaction, and GPCR targets, and used these to demonstrate that the EC method can visualize, rationalize, and predict electrostatically driven ligand affinity changes and help to predict compound selectivity. The methodology presented here for the analysis of EC is a powerful and versatile tool for drug design.