Inhibitors of the enzyme adenosine monophosphate deaminase (AMPD) show interesting levels of herbicidal activity. An enzyme mechanism-based approach has been used to design new inhibitors of AMPD starting from nebularine (6) and resulting in the synthesis of 2-deoxy isonebularine (16). This compound is a potent inhibitor of the related enzyme adenosine deaminase (ADA; IC50 16 nM), binding over 5000 times more strongly than nebularine. It is proposed that the herbicidal activity of compound 16 is due to 5́-phosphorylation in planta to give an inhibitor of AMPD. Subsequently, an enzyme structure-based approach was used to design new non-ribosyl AMPD inhibitors. The initial lead structure was discovered by in silico screening of a virtual library against plant AMPD. In a second step, binding to AMPD was further optimised via more detailed molecular modeling leading to 2-(benzyloxy)-5-(imidazo[2,1-f][1,2,4]triazin-7-yl)benzoic acid (36) (IC50 300 nM). This compound does not inhibit ADA and shows excellent selectivity for plant over human AMPD.
Water molecules are of great importance for the correct representation of ligand binding interactions. Throughout the last years, water molecules and their integration into drug design strategies have received increasing attention. Nowadays a variety of tools are available to place and score water molecules. However, the most frequently applied software solutions require substantial computational resources. In addition, none of the existing methods has been rigorously evaluated on the basis of a large number of diverse protein complexes. Therefore, we present a novel method for placing water molecules, called WarPP, based on interaction geometries previously derived from protein crystal structures. Using a large, previously compiled, high-quality validation set of almost 1500 protein-ligand complexes containing almost 20 000 crystallographically observed water molecules in their active sites, we validated our placement strategy. We correctly placed 80% of the water molecules within 1.0 Å of a crystallographically observed one.
Macromolecular structures resolved by X-ray crystallography are essential for life science research. While some methods exist to automatically quantify the quality of the electron density fit, none of them is without flaws. Especially the question of how well individual parts like atoms, small fragments, or molecules are supported by electron density is difficult to quantify. While taking experimental uncertainties correctly into account, they do not offer an answer on how reliable an individual atom position is. A rapid quantification of this atomic position reliability would be highly valuable in structure-based molecular design. To overcome this limitation, we introduce the electron density score EDIA for individual atoms and molecular fragments. EDIA assesses rapidly, automatically, and intuitively the fit of individual as well as multiple atoms (EDIAm) into electron density accompanied by an integrated error analysis. The computation is based on the standard 2fo - fc electron density map in combination with the model of the molecular structure. For evaluating partial structures, EDIAm shows significant advantages compared to the real-space R correlation coefficient (RSCC) and the real-space difference density Z score (RSZD) from the molecular modeler's point of view. Thus, EDIA abolishes the time-consuming step of visually inspecting the electron density during structure selection and curation. It supports daily modeling tasks of medicinal and computational chemists and enables a fully automated assembly of large-scale, high-quality structure data sets. Furthermore, EDIA scores can be applied for model validation and method development in computer-aided molecular design. In contrast to measuring the deviation from the structure model by root-mean-squared deviation, EDIA scores allow comparison to the underlying experimental data taking its uncertainty into account.
ABSTRACTReliable computational prediction of protein side chain conformations and the energetic impact of amino acid mutations are the key aspects for the optimization of biotechnologically relevant enzymatic reactions using structure‐based design. By improving the protein stability, higher yields can be achieved. In addition, tuning the substrate selectivity of an enzymatic reaction by directed mutagenesis can lead to higher turnover rates. This work presents a novel approach to predict the conformation of a side chain mutation along with the energetic effect on the protein structure. The HYDE scoring concept applied here describes the molecular interactions primarily by evaluating the effect of dehydration and hydrogen bonding on molecular structures in aqueous solution. Here, we evaluate its capability of side‐chain conformation prediction in classic remutation experiments. Furthermore, we present a new data set for evaluating “cross‐mutations,” a new experiment that resembles real‐world application scenarios more closely. This data set consists of protein pairs with up to five point mutations. Thus, structural changes are attributed to point mutations only. In the cross‐mutation experiment, the original protein structure is mutated with the aim to predict the structure of the side chain as in the paired mutated structure. The comparison of side chain conformation prediction (“remutation”) showed that the performance of HYDEprotein is qualitatively comparable to state‐of‐the art methods. The ability of HYDEprotein to predict the energetic effect of a mutation is evaluated in the third experiment. Herein, the effect on protein stability is predicted correctly in 70% of the evaluated cases. Proteins 2017; 85:1550–1566. © 2017 Wiley Periodicals, Inc.
Protein ligand interactions are the fundamental basis for molecular design in pharmaceutical research, biocatalysis, and agrochemical development. Especially hydrogen bonds are known to have special geometric requirements and therefore deserve a detailed analysis. In modeling approaches a more general description of hydrogen bond geometries, using distance and directionality, is applied. A first study of their geometries was performed based on 15 protein structures in 1982. Currently there are about 95 000 protein ligand structures available in the PDB, providing a solid foundation for a new large-scale statistical analysis. Here, we report a comprehensive investigation of geometric and functional properties of hydrogen bonds. Out of 22 defined functional groups, eight are fully in accordance with theoretical predictions while 14 show variations from expected values. On the basis of these results, we derived interaction geometries to improve current computational models. It is expected that these observations will be useful in designing new chemical structures for biological applications.
The estimation of free energy of binding is a key problem in structure-based design. We developed the scoring function HYDE based on a consistent description of HYdrogen bond and DEhydration energies in protein–ligand complexes. HYDE is applicable to all types of protein targets since it is not calibrated on experimental binding affinity data or protein–ligand complexes. The comprehensible atom-based score of HYDE is visualized by applying a very intuitive coloring scheme, thereby facilitating the analysis of protein–ligand complexes in the lead optimization process. In this paper, we have revised several aspects of the former version of HYDE which was described in detail previously. The revised HYDE version was already validated in large-scale redocking and screening experiments which were performed in the course of the Docking and Scoring Symposium at 241st ACS National Meeting. In this study, we additionally evaluate the ability of the revised HYDE version to predict binding affinities. On the PDBbind 2007 coreset, HYDE achieves a correlation coefficient of 0.62 between the experimental binding constants and the predicted binding energy, performing second best on this dataset compared to 17 other well-established scoring functions. Further, we show that the performance of HYDE in large-scale redocking and virtual screening experiments on the Astex diverse set and the DUD dataset respectively, is comparable to the best methods in this field.
Combinatorial and parallel chemistry concepts and techniques have a large impact upon the way in which the search for new biologically active lead structures is conducted in modern laboratories. Typically, a library of compounds, synthesized using these techniques, is screened against a biological target to identify small molecule hits which inhibit the target. These are then further elaborated to optimize biological potency and generate lead structures. The way in which the compound library is designed or selected is a key criterion for determining the eventual success of such a research process. Although most libraries will give biological hits, not all small molecule hit structures are suitable for further optimization into leads. The chapter discusses the important concepts, including bioavailability, chemical space, and privileged structures, which are important in successful library design. Modern fragment-, ligand-, and structure-based design approaches are compared and the differences and synergies, strengths and weaknesses are analyzed. Throughout the chapter the most important concepts and methodologies are illustrated using examples taken from the recent literature.
Usually based on molecular mechanics force fields, the post-optimization of ligand poses is typically the most time-consuming step in protein-ligand docking procedures. In return, it bears the potential to overcome the limitations of discretized conformation models. Because of the parallel nature of the problem, recent graphics processing units (GPUs) can be applied to address this dilemma. We present a novel algorithmic approach for parallelizing and thus massively speeding up protein-ligand complex optimizations with GPUs. The method, customized to pose-optimization, performs at least 100 times faster than widely used CPU-based optimization tools. An improvement in Root-Mean-Square Distance (RMSD) compared to the original docking pose of up to 42% can be achieved.
The HYDE scoring function consistently describes hydrogen bonding, the hydrophobic effect and desolvation. It relies on HYdration and DEsolvation terms which are calibrated using octanol/water partition coefficients of small molecules. We do not use affinity data for calibration, therefore HYDE is generally applicable to all protein targets. HYDE reflects the Gibbs free energy of binding while only considering the essential interactions of protein–ligand complexes. The greatest benefit of HYDE is that it yields a very intuitive atom-based score, which can be mapped onto the ligand and protein atoms. This allows the direct visualization of the score and consequently facilitates analysis of protein–ligand complexes during the lead optimization process. In this study, we validated our new scoring function by applying it in large-scale docking experiments. We could successfully predict the correct binding mode in 93% of complexes in redocking calculations on the Astex diverse set, while our performance in virtual screening experiments using the DUD dataset showed significant enrichment values with a mean AUC of 0.77 across all protein targets with little or no structural defects. As part of these studies, we also carried out a very detailed analysis of the data that revealed interesting pitfalls, which we highlight here and which should be addressed in future benchmark datasets.
Corwin Hansch is well-known as the father, inventor, and promoter of quantitative structure-activity relationships. Usually, QSAR is seen as a ligand-based design method correlating molecular structure or property descriptors to biological activity. QSAR is seldom mentioned in relation to structure-based approaches, although it is the centerpiece of nearly every empirical scoring function. QSAR techniques are applied on various levels, from the fitting of scoring terms to biological affinity data in empirical scoring functions up to the fine-tuning of individual aspects of protein-ligand interactions. In the following, we report on current findings for our scoring approach HYDE, which are based upon both the idea of QSAR and Hansch's historical logP data. We relate the molecular surface area of 594 diverse compounds to their experimental octanol/water partition coefficients aiming at new insights in hydrogen bonding and dehydration energies of solutes. Donors and acceptors which are far from each other contribute with nearly constant increments to the logP value. The solubility in the aqueous phase is however not increased with the number of hydrogen bonds a polar group is able to form. Although signs are found that these facts have been known for many years, they have implications for modern scoring function design.
It is highly desirable to have a scoring function that provides guidance for the design of compounds with optimized bioactivity. HYDE [1] is such a scoring function, considering the essential interactions in protein-ligand complexes. HYDE describes consistently hydrogen bonds, the hydrophobic effect and desolvation. Its basic principle is a well balanced assessment of the energetics of desolvation. Compared to most other scoring functions HYDE is not calibrated on affinity data, we use octanol/water partition data of small molecules for calibration. Only three major factors are taken into consideration: (a) local hydrophobicity, (b) solvent accessible surface, and (c) contact surface area. Based on these, energetically favorable and unfavorable contributions to the binding affinity can be assessed on an atomic level. Atomic contributions can be visualized, which turns out to be particularly helpful in a lead-optimization setup. One may immediately identify energetically unfavorable arrangements, like a polar group without a counter-part in an otherwise hydrophobic pocket or two hydrogen bond acceptors facing each other. Medicinal chemists will immediately have ideas how to alter a given structure in order to gain activity. It is demonstrated that HYDE is able to distinguish between strong binders, weak binders, and non-binders considering several p38 MAP kinase inhibitors [2]. In a congeneric series of thrombin inhibitors [3] it is shown that HYDE is able to score and rank single atom exchanges correctly.
We developed a new empirical scoring function, HYDE, for the evaluation of protein-ligand complexes. HYDE estimates binding free energy based on two terms for dehydration and hydrogen bonding only. The essential feature of this scoring function is the integrated use of log P-derived atomic increments for the prediction of free dehydration energy and hydrogen bonding energy. Taking the dehydration of atoms within the interface into account shows that some atoms contribute favorably to the overall score, while others contribute unfavorably. For instance, hydrogen bond functions are penalized if they are dehydrated unless they can overcompensate this loss by forming a hydrogen bond with excellent geometry. The main stabilizing contribution represents the removal of apolar groups from the water: the hydrophobic effect. Initial studies using the DUD dataset show that with HYDE, there is a significant decrease in false positives, a reasonable categorization of compounds as either non-binders, weak, medium or strong binders, and in particular, there is a generally applicable and thermodynamically sensible cutoff score below which there is a high likelihood that the compound is indeed a binder.