Epik version 7 is a software program that uses machine learning for predicting the pKa values and protonation state distribution of complex, drug-like molecules. Using an ensemble of atomic graph convolutional neural networks (GCNNs) trained on over 42,000 pKa values across broad chemical space from both experimental and computed origins, the model predicts pKa values with 0.42 and 0.72 log unit median absolute and RMS errors, respectively, across seven test sets. Epik version 7 also generates protonation states and recovers 95% of the most populated protonation states compared to previous versions. Requiring on average only 47 ms per ligand, Epik version 7 is rapid and accurate enough to evaluate protonation states for crucial molecules and prepare ultra-large libraries of compounds to explore vast regions of chemical space. The simplicity of and time required for the training allows for the generation of highly accurate models customized to a program’s specific chemistry.
The recently developed AlphaFold2 (AF2) algorithm predicts proteins’ 3D structures from amino acid sequences. The open AlphaFold Protein Structure Database covers the complete human proteome. It shows great potential to provide structural information to enable and enhance existing and new drug discovery projects. Using an industry-leading molecular docking method (Glide), we benchmarked the virtual screening performance of 28 common drug targets each with an AF2 structure and known holo and apo structures from the DUD-E dataset. The AF2 structures show comparable early enrichment of known active compounds (avg. EF 1%: 13.16) to apo structures (avg. EF 1%: 11.56), while falling behind early enrichment of the holo structures (avg. EF 1%: 24.81). We also demonstrated that with the IFD-MD induced-fit docking approach, we can refine the AF2 structures using a known binding ligand to improve the performance in structure-based virtual screening (avg. EF 1%: 19.25). Thus, with proper preparation and refinement, AF2 structures show considerable promise for in silico hit identification.
With the advent of make-on-demand commercial libraries, the number of purchasable compounds available for virtual screening and assay has grown explosively in recent years, with several libraries eclipsing one billion compounds. Today’s screening libraries are larger and more diverse, enabling discovery of more potent hit compounds and unlocking new areas of chemical space, represented by new core scaffolds. Applying physics-based in-silico screening methods in an exhaustive manner, where every molecule in the library must be enumerated and evaluated independently, is increasingly cost-prohibitive. Here, we introduce a protocol for machine learning-enhanced molecular docking based on active learning to dramatically increase throughput over traditional docking. We leverage a novel selection protocol that strikes a balance between two objectives: (1) Identifying the best scoring compounds and (2) exploring a large region of chemical space, demonstrating superior performance compared to a purely greedy approach. Together with automated redocking of the top compounds, this method captures nearly all the high scoring scaffolds in the library found by exhaustive docking. This protocol is applied to our recent virtual screening campaigns against the D4 and AMPC targets that produced dozens of highly potent, novel inhibitors, and a blinded test against the MT1 target. Our protocol recovers more than 80% of the experimentally confirmed hits with a 14-fold reduction in compute cost, and more than 90% of the hit scaffolds in the top 5% of model predictions, preserving the diversity of the experimentally confirmed hit compounds.
Mechanisms of protein-carbohydrate recognition attract a lot of interest due to their roles in various cellular processes and metabolism disorders. We have performed a large-scale analysis of protein structures solved in complex with glucose, galactose and their substituted analogues. We found that, on average, sugar molecules establish five hydrogen bonds (HBs) in the binding site, including one to three HBs with bridging water molecules. The free energy contribution of bridging and direct HBs was estimated using the free energy perturbation (FEP+) methodology for mono- and disaccharides that bind to l-ABP, ttGBP, TrmB, hGalectin-1 and hGalectin-3. We show that removing hydroxy groups that are engaged in direct HBs with the charged groups of Asp, Arg and Glu residues, protein backbone amide or buried water dramatically decreases binding affinity. In contrast, all solvent-exposed hydroxy groups and hydroxy groups engaged in HBs with the solvent-exposed bridging water molecules contribute weakly to binding affinity and so can be replaced to optimize ligand potency. Finally, we rationalize an effect of binding site water replacement on the binding affinity to l-ABP.
We have developed a new methodology for protein ligand docking and scoring, WScore, incorporating a flexible description of explicit water molecules. The locations and thermodynamics of the waters are derived from a WaterMap molecular dynamics simulation. The water structure is employed to provide an atomic level description of ligand and protein desolvation. WScore also contains a detailed model for localized ligand and protein strain energy and integrates an MM-GBSA scoring component with these terms to assess delocalized strain of the complex. Ensemble docking is used to take into account induced fit effects on the receptor conformation, and protein reorganization free energies are assigned via fitting to experimental data. The performance of the method is evaluated for pose prediction, rank ordering of self-docked complexes, and enrichment in virtual screening, using a large data set of PDB complexes and compared with the Glide SP and Glide XP models; significant improvements are obtained.
Aim: We introduce AutoQSAR, an automated machine-learning application to build, validate and deploy quantitative structure-activity relationship (QSAR) models. Methodology/results: The process of descriptor generation, feature selection and the creation of a large number of QSAR models has been automated into a single workflow within AutoQSAR. The models are built using a variety of machine-learning methods, and each model is scored using a novel approach. Effectiveness of the method is demonstrated through comparison with literature QSAR models using identical datasets for six end points: protein-ligand binding affinity, solubility, blood-brain barrier permeability, carcinogenicity, mutagenicity and bioaccumulation in fish. Conclusion: AutoQSAR demonstrates similar or better predictive performance as compared with published results for four of the six endpoints while requiring minimal human time and expertise.
Progress in structure determination of G protein-coupled receptors (GPCRs) has made it possible to apply structure-based drug design (SBDD) methods to this pharmaceutically important target class. The quality of GPCR structures available for SBDD projects fall on a spectrum ranging from high resolution crystal structures (<2 Å), where all water molecules in the binding pocket are resolved, to lower resolution (>3 Å) where some protein residues are not resolved, and finally to homology models that are built using distantly related templates. Each GPCR project involves a distinct set of opportunities and challenges, and requires different approaches to model the interaction between the receptor and the ligands. In this review we will discuss docking and virtual screening to GPCRs, and highlight several refinement and post-processing steps that can be used to improve the accuracy of these calculations. Several examples are discussed that illustrate specific steps that can be taken to improve upon the docking and virtual screening accuracy. While GPCRs are a unique target class, many of the methods and strategies outlined in this review are general and therefore applicable to other protein families.