
Methods for the calculation of two properties of interest in drug design, namely free energy of aqueous solvation and lipophilicity (log P), using fragmental methods are reviewed here. Though aqueous solvation free energies are commonly estimated using `whole molecule' methods such as GB/SA and AMSOL, we have recently shown that fragmental approaches can offer high quality predictions as well (for molecules of the size of 20 atoms or less). In the case of log P predictions, the more commonly used ALOGP and CLOGP approaches represent the two extremes of the fragmental constant approach: ALOGP uses atom-sized fragments and no correction factors; CLOGP uses larger fragments and correction factors, which are typically obtained for each series of molecules separately. Anew approach (HLOGP) that uses both smaller (atom-sized) and larger fragments is shown to offer better performance than the other two widely used methods for the prediction of lipophilicity. In this approach, an automated `inventory' of fragments (bonded atom combinations) within a molecule, known as molecular hologram, is used as a composite descriptor and it is used in conjunction with partial least squares for the prediction of aqueous solvation or lipophilicity. It is emphasized that these different methods are useful in different types of drug design applications involving small organic molecules.
Knowledge-based scoring functions have recently emerged as an alternative and very promising way of ranking protein-ligand complexes with known 3D structure according to their binding affinities. These simplified potential-based approaches use the structural information stored in databases of protein-ligand complexes to derive atom pair interaction potentials also known as potentials of mean force (PMF). The derived PMF depend on the definition of a suitable reference state. The reference states vary among suggested knowledge — based scoring functions. Therefore, we attempt here to shed some light on the influence of different reference state definitions on the predictive power of a knowledge-based scoring function that has been introduced by us very recently [J. Med. Chem., 42 (1999) 791]. It is shown that a reference state that implicitly and more comprehensively accounts for protein and ligand solvation gives the most consistent scoring results for four test sets of diverse protein-ligand complexes taken from the Brookhaven Protein Data Bank. It is also shown that a reference sphere radius of at least 7–8 Å is needed to effectively capture solvation effects that are treated implicitly in the scoring function.
We describe a ‘virtual NMR screening’ method to assist in the design of inhibitors that occupy different sites within a target. We dock small molecules into the active site of an enzyme and score them. Keeping the tightest-binding lead fixed in space, we dock and score other small molecules in its presence. Using this approach, linker groups are used to join the compounds together to form a high-affinity inhibitor. We present validation of our computational approach by reproducing experimental results for FKBP and stromelysin. Docking simulations are not subject to experimental problems such as proteolysis, protein or compound insolubility, or enzyme size. Because docking is fast and our scoring method can distinguish between high- and low-affinity inhibitors, this docking procedure shows promise as integral part of a drug-design strategy.
A modification of the hydrogen bond score in the docking program FlexX is presented. Hydrogen bonds formed in inaccessible regions of protein cavities thereby gain larger weight than others formed at the protein surface. The modified scoring function is tested with thrombin as a target. Secondly, a recently published knowledge-based scoring function is compared to the FlexX scoring function in several database ranking experiments.
We have investigated the relationship between drug retention in immobilized liposome partitioning chromatography and liposome partitioning and found a strong linear correlation. Separate linear relationships were found depending on the charge of the compound when liposome chromatographic measurements were related to the octanol/water partition coefficients. We have also investigated the importance of the water/octanol partition coefficient in quantitative structure–property relationships related to drug transport properties. The studies show that the inclusion of a parameter related to lipophilicity causes only, at best, a marginal increase in internal predictivity and, at worst, a decrease in external predictivity. The studies also show that parameters related to hydrogen bonding, polarizability and size are important properties that need to be included in quantitative models for drug transport processes. We believe that the use of multivariate characterizations of compounds based on non-composite parameters may result in better and more predictive models compared with models based on parameters of a more composite nature when investigating the possibilities to establish quantitative structure–property relationships.
In this review, we first examine the contextual background of structure–pharmacokinetic relationships. Some concepts in drug disposition are briefly recalled, and inherent difficulties instructure–pharmacokinetic relationships are outlined. Lipophilicity is then investigated in the light of the intermolecular and intramolecular interactions it encodes. In the main body of the review, a number of pharmacokinetic processes are examined for their relations with lipophilicity. These processes are taken in a logical sequence of permeation, absorption (intestine, skin, cornea, brain), plasma protein binding, tissue distribution, volume of distribution and renal clearance. Relations between metabolism and lipophilicity are more complex, since biotransformation involves both low-energy (enzyme binding) and high-energy (catalysis) processes. Only the former may be related to lipophilicity. The conclusion argues against faulty statistics and over-interpretation.
Lipophilicity is a prime physicochemical parameter in describing both pharmacodynamic and pharmacokinetic aspects of drug action. Its quantitative descriptor (log P) is measured in or calculated for different model solvent systems depending on the respective scientific interest. The first part of this article focuses on the hydrophobic fragmental constant approach (Σ f system) for octanol/water partitioning. The original approach and its revisions are presented, followed by recipes for the practical procedure ofΣ f calculations and a detailed description of how to apply the correction factor C M .A Σ f system for application in the study of aliphatic hydrocarbon/water (ahc/w) partitioning is described in the second part of the article. The current version is based on an optimized coupling of the ahc/w system with f constants for octanol/water partitioning. A total of 70 f ahc values are now available.
The large numbers of compounds that are now available in drug discovery programmes have resulted in the need for methods to select compounds, both from external suppliers and in combinatorial library design procedures. In this article we describe a method that has been developed for scoring and ranking compounds according to their likelihood of exhibiting activity. The method can be used to determine the order in which compounds should be screened as well as to guide compound acquisition programmes. We then describe a series of experiments we have conducted that explore the benefits of designing combinatorial libraries through an analysis of product space rather than reactant space. The experiments are based on several different diversity metrics and molecular descriptors. We also show how product-based selection allows multi-objectives to be optimised simultaneously, forexample, diversity and physicochemical properties,allowing the design of diverse and drug-like libraries.
Many different measures of structural similarity have been suggested for matching chemical structures, each such measure focusing upon some particular type of molecular characteristic. The multi-faceted nature of biological activity suggests that an appropriate similarity measure should encompass many different types of characteristic, and this article discusses the use of data fusion methods to combine the results of searches based on multiple similarity measures. Experiments with several different types of dataset and activity suggest that data fusion provides a simple, but effective, approach to the combination of individual similarity measures. The best results were generally obtained with a fusion rule that sums the rank positions achieved by each molecule in searches using individual measures.
In this article, we review the use of in vitro and in silico affinity fingerprints as novel descriptors for similarity searches in molecular databases and QSAR analyses. An affinity fingerprint for a particular molecule is constructed as a vector of either its binding affinities, docking scores or superpositioning pseudo energies against a reference panel of proteins or small molecules. In contrast to most other molecular descriptors, affinity fingerprints are not directly derived from molecular structures. As such, they offer the possibility to detect similarities amongst molecules independent of their structural scaffolds. In this report we introduce the Flexsim-S method, an extension of our previous work on virtual affinity fingerprints. Moreover, we demonstrate that virtual affinity fingerprint methods are comparable to some popular two-dimensional descriptors in terms of correctly classifying compounds, but complementary with respect to the particular search results (hit lists).
We present our database-screening tool SLIDE, which is capable of screening large data sets of organic compounds for potential ligands to a given binding site of a target protein. Its main feature is the modeling of induced complementarity by making adjustments in the protein side chains and ligand upon binding. Mean-field theory is used to balance the conformational changes in both molecules in order to generate a shape-complementary interface. Solvation is considered by prediction of water molecules likely to be conserved from the crystal structure of the ligand-free protein, and allowing them to mediate ligand interactions, if possible, or including a desolvation penalty when they are displaced by ligand atoms that do not replace the lost hydrogen bonds.A data set of over 175 000 organic molecules was screened for potential ligands to the progesterone receptor, dihydrofolate reductase, and a DNA-repair enzyme. In all cases the screening time was less than a day on a Pentium II processor, and known ligands as well as highly complementary new potential ligands were found.
Two methods for structure-based computational ligand design are reviewed. Hydrophobicity maps allow to quantitatively estimate and graphically display the propensity of nonpolar groups to bind at the surface of a protein target [Scarsi et al., Proteins Struct. Funct. Genet., 37 (1999) 565]. The program SEED (Solvation Energy for Exhaustive Docking) finds optimal positions and orientations of nonpolar fragments using the hydrophobicity maps, while polar fragments are docked with at least one hydrogen bond with the protein [Majeux et al., Proteins Struct. Funct. Genet., 37 (1999) 88]. An efficient evaluation of the binding energy, including continuum electrostatic solvation, allows to dock a library of 100 fragments into a 25-residue binding site in about five hours on a personal computer. Applications to thrombin, a key enzyme in the blood coagulation cascade, and the p38 mitogen-activated protein kinase, which is a target for the treatment of inflammatory and neurodegenerative diseases, are presented. The role of the hydrophobicity maps and structure-based docking of a fragment library in exploiting genomes to design drugs is addressed.
Molecular superpositioning is an important task in rational drug design. Usually it is the key step in a comparative analysis of molecules by 3D QSAR methods. Also it is helpful for the elucidation ofa pharmacophore and crucial in the attempt to derive a receptor model. Generally speaking, molecular superpositioning can be seen as the analog of molecular docking if the receptor structure is not available, and direct methods are not applicable. Virtual database screening is the computational counterpart to modern experimental techniques like high throughput screening and assaying ofcombinatorial libraries. Both screening techniques have the commongoal to detect active molecules in a large selection of compounds. Usually hundreds of thousands of candidates are to be tested, hence, time is the limiting factor and rapid processing of utmost importance. Descriptor-based methods that usually provide a simple linear encoding of the molecules meet the demands of computational speed and have been used predominantly for the task of virtual screening, for a long time. However, more powerful superposition methods have been developed during the past few years and now begin also to be applicable to screening large databases. Especially incombination with the faster methods, molecular superpositioning as the final step of a filtering protocol provides a powerful tool for virtual database screening. The present work reports on our latest developments of molecular superpositioning techniques and assessing their applicability to virtual database screening.
A recently introduced molecular size-based model that allows a unified description of enthalpies of vaporization, boiling points, gas–liquid solubilities, and vapor pressures for simple organic liquids using a free energy expression obtained from molecular-level assumptions is summarized. By changing the interaction-related constant ω used by the model when water is the solvent, the model can be extended to describe alkane–water partition,octanol–water partition, and water solubility of solutes that have no hydrogen-bonding or strongly polar substituents. Here, it is shown that thisΔω change, which is most likely related to the changes that the solute produces in the hydrogen-bonded structure of water, agrees very well with the value that can be derived from the modified hydration-shell hydrogen-bond model of Muller. By combining the present molecular size-based model with this hydrogen-bonding model, a simplified but consistent description is obtained for the properties of water and for the hydrophobic effect. This indicates that many unusual properties of water may be accounted for by a proper combination of the nonspecific interactions as extrapolated from other liquids, the unusually small size of its molecules, and an adequate model of hydrogen-bonding. A fully computerized method (QLogP) that can estimate octanol–water log P for a large variety of organic solutes also fits within this unified approach. Despite using only two parameters (molecular volume and a novel, quantified parameter that is probably hydrogen-bonding-related), the predictive power of this method is similar to that of the considerably more complex fragment-contribution methods often used by medicinal chemists (ACD/LogP, AFC, CLOGP, KLogP, MLogP, Rekker), as illustrated by a comparison based on various structures.
A new atom-additive method is presented for calculating octanol/water partition coefficient (log P) of organic compounds. The method, XLOGP v2.0, gives log P values by summing the contributions of component atoms and correction factors. Altogether 90 atom types are used to classify carbon, nitrogen, oxygen, sulfur, phosphorus and halogen atoms, and 10 correction factors are used for some special substructures. The contributions of each atom type and correction factor are derived by multivariate regression analysis of 1853 organic compounds with known experimental log P values. The correlation coefficient (r) for fitting the whole set is 0.973 and the standard deviation (s) is 0.349 log units. Comparison of various log P calculation procedures demonstrates that our method gives much better results than other atom-additive approaches and is at least comparable to fragmental approaches. Because of the simple methodology, the `missing fragment' problem does not occur in our method.
The development of reliable, transferable methods that can compute the energy of interaction between protein sand ligands is a major challenge for computational chemistry. Understanding the energetics ofprotein-ligand interactions would not only provide powerful tools for prediction in structure-assisted ligand and library design, but also enrich our appreciation of the subtleties of structure that underlie molecular recognition in biological systems. One of the central problems in developing effective models is the quality and quantity of experimental data on the structure and thermodynamics of protein-ligand complexes. In this article we discuss some of the issues and some of the experimental programmes of research we have initiated to provide such data. We summarise the characteristics necessary for a model system and the experimental techniques available. This includes a discussion of calorimetry, inhibition assays and crystallographic results on series of complexes in our laboratory, including penicillin acylase, thrombin, sialidase and inparticular the oligopeptide binding protein, OppA. Aswell as discussing the lessons we have learnt about the characteristics of an ideal model system, we also present some preliminary analyses of what our combined structural and thermodynamic data have told us.
The performances of a log P model designed from EVA descriptors based on theoretically derived normal coordinate frequencies and using the classical PLS analysis as statistical engine were compared to those provided by a neural network model employing various autocorrelation vectors for describing the molecules. The superiority of the latter is clearly demonstrated for simulating the lipophilicity of simple chemical structures.
Following the preliminary discussion outlining the relative difficulty of the experimental determination of solubilities versus the lack of general applicability and sound theoretical basis of most current predictive approaches for solubility, we look closely at the recently developed pure thermodynamic model for solubility in real solutions: that derived from mobile order and disorder theory. With successful estimates of the solubility of 62hydroxysteroids and related drugs in common organic solvents of differing polarities and H-bonding capacity, the model has proved to be a valuable tool in predicting the solubility of complex solutes such as polyfunctional drugs. Free of any adjusted regression coefficients and based on a limited number of readily available parameters, the proposed model is a time-saving alternative procedure to experimentation. By properly quantifying the enthalpic and entropic contributions involved in the overall solubility process, the model furthermore assesses the factors that determine solubility differences between steroids and solubility changes upon solvent properties. Therefore, the poor solubility of hydroxysteroids in aliphatic hydrocarbons results from the negative effects due to the change in non-specific cohesion forces upon mixing and due to steroid self-association in solution. In water , the low solubilities are mainly due to the large negative value of the hydrophobic effect which cannot be overcome by steroid–water functional group associations, i.e., the solvation effect. The relatively good solubility of hydroxysteroids in polar non-associated solvent s (ketones, ethers, esters) and in alcohols is explained by the fact that, in both kinds of solvents, steroid self-association is rather well counterbalanced by the formation of a more or less important number of steroid–solvent interactions without being penalized by a strong negative hydrophobic effect in the case of alcohols. Some practical rules regarding how some parameters like the molar volume or the substitution may affect solubility are finally derived, which might help the pharmaceutical scientist to orient the choice of a solvent for liquid pharmaceutical forms.
This study describes the development of the ACD/Log P calculation method. Analysis of 14 calculation methods revealed that the most accurate calculations are obtained when correction factors are used. We evaluated the correction factors used by Hansch and Leo in CLOGP in order to simplify their method. Most of the CLOGP structural factors are included in our fragmental increments. Aliphatic and aromatic factors are replaced with additive interfragmental increments. Missing increments are estimated by two empirical equations with simple physical interpretation. The final method uses three simple equations with several types of parameters. The training set included 3601 compounds and the correlation between experimental and calculated Log P values gave R = 0.992, S = 0.21. The method was validated by comparing it with 17 other methods on various data sets of independently selected drugs and other compounds. In all cases, our method produced the best results. The weakness of this method is that it uses a large number of individual increments for aromatic interactions. Each increment represents a combination of several effects which presently cannot be separated.
Both proton transfer and hydrogen bonding play important roles in biological systems. In order to measure hydrogen bond basicity, we are building a new scale that differs significantly from the pKa scale of proton transfer basicity. The strength of hydrogen bond acceptors (HBAs) is measured from the Gibbs energy change ΔGHB for the formation of 1:1 hydrogen bonding complexes between hydrogen bond acceptors (bases) and a reference hydrogen bond donor (4-fluorophenol) in tetrachloromethane at 298 K. The pKHB database (1.364 pKHB =–ΔGHB (kcal mol-1)) comprises ca. 1000 hydrogen bond acceptors. The HBA strength depends on (i) the position of the acceptor atom in the periodic table, (ii) polarizability, field/inductive and resonance effects of substituents around the acceptor atom, and (iii) proximity effects including steric hindrance of the acceptor site, intramolecular hydrogen bonding and lone-pair–lone-pair repulsions. The ranking of oxygen and sp nitrogen bases does not depend very much on the solvent and the reference hydrogen bond donor, but sp2 and sp3 nitrogen bases gain strength in solvents of higher reaction field than CCl4 and lose strength toward CH and weak NH donors. The complete scatter pattern exhibited by the pKa versus pKHB plot demonstrates the non-equivalence of the two scales. The HBA strength scale is applied to the prediction of the hydrogen bonding site in polybasic drugs (e.g strychnine and carbimazole), and to the calculation of octanol–water partition coefficients. A possible relationship between HBA strength and antihistaminic activity is studied for the `push–pull' drugs cimetidine, ranitidine and famotidine.