Identifying cryptic binding sites in proteins remains a challenge in structure-based drug discovery because these sites are often not apparent in apo structures. Here, we developed and validated a novel "induce-and-identify" workflow that integrates mixed solvent molecular dynamics (MxMD) simulations with SiteMap. This approach leverages MxMD to sample protein conformations to expose hidden pockets, which are then effectively identified and ranked by SiteMap. Using a challenging data set of 65 cryptic binding sites, the developed workflow identified the cryptic binding site within the top 5 predictions in 78.5% of cases. These results suggest that the proposed MxMD + SiteMap workflow provides a robust and valuable tool for early phase drug discovery, enabling the exploration of a broader range of druggable targets by effectively inducing and identifying cryptic binding sites.
The development of a small-molecule suitable for clinical use involves optimization of ligand binding to an identified target protein but also frequently the reduction of binding to a variety of anti-target proteins known to pose translational risk. These anti-targets, often referred to as ADMET proteins, often contain large, highly flexible binding sites leading to promiscuous binding of small molecules. This property often frustrates efforts to execute structure-based design against these anti-targets. Here we demonstrate the application of a combination of induced fit docking and free energy perturbation to structurally enable several of the most common anti-targets: cytochrome P450 isoforms 3A4 and 2D6, the hERG ion channel, and PXR. Crucially, the enforcement of a consensus binding mode across a congeneric series allowed for the efficient exploration of a wide range of receptor conformation space. Results are presented for both public retrospective datasets and prospective application in active drug discovery programs.
A comprehensive study was conducted to address the challenge of identifying cryptic binding sites in proteins, which are crucial for structure-based drug discovery but are often not apparent in apo protein structures. We developed and validated a novel "induce and identify" workflow that integrates mixed solvent molecular dynamics (MxMD) simulations with SiteMap. This approach leverages MxMD sampling of protein conformations to expose hidden pockets, which are then effectively identified and ranked by SiteMap. Using a challenging dataset of 65 cryptic binding sites, the developed workflow identified the cryptic binding site within the top 5 predictions in 78.5% of cases. We believe the proposed MxMD + SiteMap workflow offers a robust and valuable tool for early-phase drug discovery, enabling the exploration of a broader range of druggable targets by effectively inducing and identifying cryptic binding sites.
Ligand optimization is central to drug discovery as hundreds of analogs might be designed and synthesized between an initial hit and a therapeutic candidate. The efficiency of this process is unclear, at least partly because there is no random background for optimization against which to compare. Such a random background might emerge from synthetically accessible but otherwise systematic random small substitutions across starting ligands, measuring likelihood of achieving a substantial improvement in affinity/potency or other property by any single perturbation. Recent literature and ligand-affinity/potency databases suggest that perhaps 10% of analogs with minor modifications improve upon a parent's potency substantially (by ≥10-fold), but this number is clouded by reporting bias, intentional improvement, and inter-group reproducibility. To begin to establish a background expectation for ligand optimization, we comprehensively and systematically modified 18 lead molecules across six targets with single atom changes; 257 compounds were synthesized. Unexpectedly, 11.2% of these random small perturbation analogs improved potency by ≥10-fold over their parents. Conversely, these more potent analogs typically had worse in vitro pharmacokinetics (e.g. reduced metabolic stability, lower plasma free fraction). While it was possible to find analogs where the potency increase compensated for inferior exposure and half-life, resulting in more potent compounds in vivo, overall a frustrated landscape for ligand optimization is revealed. This study begins to establish a background expectation for ligand potency optimization and offers a simple strategy to do so. It also begins to quantify the challenges confronting the field in moving beyond in vitro potency.
Drug resistance is a critical challenge in treating diseases like cancer and infectious disease. This study presents a novel computational workflow for predicting on-target resistance mutations to small molecule inhibitors (SMIs). The approach integrates genetic models with alchemical free energy perturbation (FEP+) calculations to identify likely resistance mutations. Specifically, a genetic model, RECODE, leverages cancer-specific mutation patterns to prioritize probable amino acid changes. Physics-based calculations assess the impact of these mutations on protein stability, endogenous substrate binding, and inhibitor binding. We apply this approach retrospectively to gefitinib and osimertinib, two clinical epidermal growth factor receptor (EGFR) inhibitors used to treat non-small cell lung cancer (NSCLC). Among hundreds of possible mutations, the pipeline accurately predicted 4 out of 11 and 7 out of 19 known binding site mutations for gefitinib and osimertinib, respectively, including the clinically relevant T790M and C797S resistance mutations. This study demonstrates the potential of integrating genetic models and physics-based calculations to predict SMI resistance mutations. This approach can be applied to other kinases and target classes, potentially enabling the design of next-generation inhibitors with improved durability of response in patients.
MALT1 is a key component of the CARD11-BCL10-MALT1 (CBM) complex downstream from BTK on the B-cell receptor signaling pathway. It is a key mediator of NF-κB signaling and considered a potential therapeutic target for several subtypes of non-Hodgkin's B-cell lymphomas. By applying advanced physics-based modeling techniques, including combining free energy calculations with machine learning methods and a chemistry-aware compound enumeration workflow, extensive sets of de novo design ideas were explored to quickly identify a novel hit series. Multiparameter optimization allowed efficient prioritization of molecules with good potency and drug-like properties during lead optimization, which led to the discovery of a highly potent MALT1 inhibitor, SGR-1505, with a well-balanced property profile. It demonstrated strong antitumor activity alone and in combination with BTK inhibitor in multiple in vivo B-cell lymphoma xenograft models and progressed to a phase 1 clinical trial in patients with mature B-cell neoplasms.
Machine learning force fields (MLFFs) have emerged as a sophisticated tool for cost-efficient atomistic simulations approaching DFT accuracy, with recent message passing MLFFs able to cover the entire periodic table. We present an invariant message passing MLFF architecture (MPNICE) which iteratively predicts atomic partial charges, including long-range interactions, enabling the prediction of charge-dependent properties while achieving 5-20x faster inference versus models with comparable accuracy. We train direct and delta-learned MPNICE models for organic systems, and benchmark against experimental properties of liquid and solid systems. We also benchmark the energetics of finite systems, contributing a new set of torsion scans with charged species and a new set of DLPNO-CCSD(T) references for the TorsionNet500 benchmark. We additionally train and benchmark MPNICE models for bulk inorganic crystals, focusing on structural ranking and mechanical properties. Finally, we explore multi-task models for both inorganic and organic systems, which exhibit slightly decreased performance on domain-specific tasks but surprising generalization, stably predicting the gas phase structure of ≃500 Pt/Ir organometallic complexes despite never training to organometallic complexes of any kind.
Homology models have been used for virtual screening and to understand the binding mode of a known active, however rare-ly have the models been shown to be of sufficient accuracy, comparable to crystal structures, to support free-energy perturba-tion (FEP) calculations. We demonstrate here that the use of an advanced induced-fit docking methodology reliably enables predictive FEP calculations on congeneric series across homology models ≥ 30% sequence identity. Further, we show that retrospective FEP calculations on a congeneric series of drug-like ligands is sufficient to discriminate between predicted binding modes. Results are presented for a total of 29 homology models for 14 protein targets, showing FEP results compa-rable to those obtained using experimentally determined crystal structures for 86% of homology models with template struc-ture sequence identities ranging from 30% to 50%. Implications for the use and validation of homology models in drug dis-covery projects are discussed, including the use of AlphaFold2 de novo structures.
In the hit identification stage of drug discovery, a diverse chemical space needs to be explored to identify initial hits. Contrary to empirical scoring functions, absolute protein-ligand binding free energy perturbation (ABFEP) provides a theoretically more rigorous and accurate description of protein-ligand binding thermodynamics and could in principle greatly improve the hit rates in virtual screening. In this work, we describe an implementation of an accurate and reliable ABFEP method in FEP+. We validated the ABFEP method on eight congeneric compound series binding to eight protein receptors including both neutral and charged ligands. For ligands with net charges, the alchemical ion approach is adopted to avoid artifacts in electrostatic potential energy calculations. The calculated binding free energies are highly correlated with experimental results with the weighted average of R2 of 0.55 for the entire dataset and an overall RMSE of 1.1 kcal/mol when protein reorganization effect upon ligand binding was accounted for. Through ABFEP calculations using apo versus holo protein structures, we demonstrated that the protein conformational and protonation state changes between the apo and holo proteins are the main physical factors contributing to the protein reorganization free energy manifested by the overestimation of raw ABFEP calculated binding free energies using the holo structures of the proteins. Furthermore, we performed ABFEP calculations in three virtual screening applications for hit enrichment. ABFEP greatly improves the hit rates as compared to docking scores or other methods like metadynamics. The highly accurate ABFEP results demonstrated in this work position it as a useful tool to improve the hit rates in virtual screening, thus facilitate hit discovery.
Crystal polymorphism is an important and fascinating aspect of solid state chemistry with far reaching implications in the pharmaceuticals, agrisciences, nutraceuticals, battery and aviation industries. Late appearing more stable polymorphs have caused numerous issues in the pharmaceutical industry. Experimental polymorph screening can be very expensive and time consuming, and sometimes may miss important low energy polymorphs due to an inability to exhaust all crystallization conditions. In this paper, we report a crystal structure prediction (CSP) method with state of the art accuracy and efficiency, validated on a large and diverse dataset including 66 molecules with 137 experimentally known polymorphic forms. The method combines a novel systematic crystal packing search algorithm and the use of machine learning force fields in a hierarchical crystal energy ranking. Our method not only reproduces all the experimentally known polymorphs, but also suggests new low energy polymorphs yet to be discovered by experiment that might pose potential risks to development of the currently known forms of these compounds. In addition, we report the prediction results of a blinded study, results for Target XXXI from the seventh CSP blind test, and demonstrate how the method can be used to accelerate clinical formulation design and derisk downstream processing.
Optimizing both on-target and off-target potencies is essential for developing effective and selective small-molecule therapeutics. Free energy calculations offer rapid potency predictions, usually within hours and with experimental accuracy and thus enables efficient identification of promising compounds for synthesis, accelerating early-stage drug discovery campaigns. While free energy predictions are routinely applied to individual proteins, here, we present a free energy framework for efficiently achieving kinome-wide selectivity that led to the discovery of selective Wee1 kinase inhibitors. Ligand-based relative binding free energy calculations rapidly identified multiple novel potent chemical scaffolds. Subsequent protein residue mutation free energy calculations that modified the Wee1 gatekeeper residue, significantly reduced their off-target liabilities across the kinome. Thus, with judicious use of this gatekeeper residue selectivity handle, applying this computational strategy streamlined the optimization of both on-target and off-target potencies, offering a roadmap to expedite drug discovery timelines by decreasing unanticipated off-target toxicities.
The hit identification stage of a drug discovery program generally involves the design of novel chemical scaffolds with desired biological activity against the target(s) of interest. One common approach is scaffold hopping, which is the manual design of novel scaffolds based on known chemical matter. One major limitation of this approach is narrow chemical space exploration, which can lead to difficulties in maintaining or improving biological activity, selectivity, and favorable property space. Another limitation is the lack of preliminary structure-activity relationship (SAR) data around these designs, which could lead to selecting suboptimal scaffolds to advance lead optimization. To address these limitations, we propose AutoDesigner - Core Design (CoreDesign), a de novo scaffold design algorithm. Our approach is a cloud-integrated, de novo design algorithm for systematically exploring and refining chemical scaffolds against biological targets of interest. The algorithm designs, evaluates, and optimizes a vast range - from millions to billions - of molecules in silico, following defined project parameters encompassing structural novelty, physicochemical attributes, potency, and selectivity. In this manner, CoreDesign can generate novel scaffolds and also explore preliminary SAR around each scaffold using FEP+ potency predictions. CoreDesign requires only a single ligand with quantifiable binding affinity and an initial binding hypothesis, making it especially suited for the hit-identification stage where experimental data is often limited. To validate CoreDesign in a real-world drug discovery setting, we applied it to the design of novel, potent Wee1 inhibitors with improved selectivity over PLK1. Starting from a single known ligand, CoreDesign rapidly explored over 23 billion molecules to identify 1,342 novel chemical series with a mean of 4 compounds per scaffold. Importantly, all chemical series met the predefined property space requirements. To rapidly analyze this large amount of data and prioritize chemical scaffolds for synthesis, we utilize t-Distributed Stochastic Neighbor Embedding (t-SNE) plots of in silico properties. The chemical space projections allowed us to rapidly identify a structurally novel 5-5 fused core meeting all the hit-identification requirements. Several compounds were synthesized and assayed from the scaffold, displaying good potency against Wee1 and excellent PLK1 selectivity. Our results suggest that CoreDesign can significantly speed up the hit-identification process and increase the probability of success of drug discovery campaigns by allowing teams to bring forward high-quality chemical scaffolds de-risked by the availability of preliminary SAR.
We report on the development and validation of the OPLS5 force field. OPLS5 further extends the accuracy of our previous model (OPLS4) with the addition of explicit polarization to improve model accuracy for molecular ions and cation-pi interactions. OPLS5 also includes advances to the functional form for metals achieving significant improvements across benchmarks assessing the structure and energetics of metal-organic complexes. Together these advances lead to improved accuracy on our protein-ligand binding benchmarks.
Free energy calculations are revolutionizing early-stage drug-discovery campaigns. Robust free energy methods can rapidly provide accurate on-target and off-target potency predictions to identify promising chemical matter for synthesis, thus, inspiring further rounds of ideation and optimization. Here, we present a free energy framework for efficiently achieving kinome-wide selectivity that led to the discovery of novel selective Wee1 kinase inhibitors. With ligand-based relative binding free energy calculations, multiple novel promising scaffolds were rapidly identified. With protein residue mutation free energy calculations that perturbed the Wee1 gatekeeper residue, off-target liabilities across the kinome of these promising series were efficiently reduced. With judicious identification of a selectivity handle, applying this computational strategy could effectively streamline the optimization of on-target and off-target profiles, thereby accelerating drug discovery timelines and decreasing unanticipated off-target toxicities.
High -quality predicted structures enable structure -based approaches to an expanding number of drug discovery programs. We propose that by utilizing free energy perturbation (FEP), predicted structures can confidently employed to achieve drug design goals. We use structure -based modeling of hERG inhibition to illustrate this value of FEP.
Significant improvements have been made in the past decade to methods that rapidly and accurately predict binding affinity through free energy perturbation (FEP) calculations. This has been driven by recent advances in small-molecule force fields and sampling algorithms combined with the availability of low-cost parallel computing. Predictive accuracies of ∼1 kcal mol-1 have been regularly achieved, which are sufficient to drive potency optimization in modern drug discovery campaigns. Despite the robustness of these FEP approaches across multiple target classes, there are invariably target systems that do not display expected performance with default FEP settings. Traditionally, these systems required labor-intensive manual protocol development to arrive at parameter settings that produce a predictive FEP model. Due to the (a) relatively large parameter space to be explored, (b) significant compute requirements, and (c) limited understanding of how combinations of parameters can affect FEP performance, manual FEP protocol optimization can take weeks to months to complete, and often does not involve rigorous train-test set splits, resulting in potential overfitting. These manual FEP protocol development timelines do not coincide with tight drug discovery project timelines, essentially preventing the use of FEP calculations for these target systems. Here, we describe an automated workflow termed FEP Protocol Builder (FEP-PB) to rapidly generate accurate FEP protocols for systems that do not perform well with default settings. FEP-PB uses an active-learning workflow to iteratively search the protocol parameter space to develop accurate FEP protocols. To validate this approach, we applied it to pharmaceutically relevant systems where default FEP settings could not produce predictive models. We demonstrate that FEP-PB can rapidly generate accurate FEP protocols for the previously challenging MCL1 system with limited human intervention. We also apply FEP-PB in a real-world drug discovery setting to generate an accurate FEP protocol for the p97 system. FEP-PB is able to generate a more accurate protocol than the expert user, rapidly validating p97 as amenable to free energy calculations. Additionally, through the active-learning workflow, we are able to gain insight into which parameters are most important for a given system. These results suggest that FEP-PB is a robust tool that can aid in rapidly developing accurate FEP protocols and increasing the number of targets that are amenable to the technology.
Structure-based drug design frequently operates under the assumption that a single holo structure is relevant. However, a large number of crystallographic examples clearly show that multiple conformations are possible. In those cases, the protein reorganization free energy must be known to accurately predict binding free energies for ligands. Only then can the energetic preference among these multiple protein conformations be utilized to design ligands with stronger binding potency and selectivity. Here, we present a computational method to quantify these protein reorganization free energies. We test it on two retrospective drug design cases, Abl kinase and HSP90, and illustrate how alternative holo conformations can be derisked and lead to large boosts in affinity. This method will allow computer-aided drug design to better support complex protein targets.
Drug discovery is a very expensive and time-consuming process, with the average cost to bring a drug molecule into market reaching $2.5B in 2016. Computational modeling can significantly accelerate the preclinical discovery process and reduce the cost. In this presentation, I am going to discuss how computational modeling, particularly the accurate and reliable protein-ligand binding free energy calculations through FEP+, can help to address some of the challenges in preclinical drug discovery, and present a few real-life drug discovery projects demonstrating the various ways these computational models were able to impact the discovery process.