Virtual screening and high-throughput screening are two major components of lead discovery within the pharmaceutical industry. In this paper we describe improvements to previously published methods for similarity searching with reduced graphs, with a particular focus on ligand-based virtual screening, and describe a novel use of reduced graphs in the clustering of high-throughput screening data. Literature methods for reduced graph similarity searching encode the reduced graphs as binary fingerprints, which has a number of issues. In this paper we extend the definition of the reduced graph to include positively and negatively ionizable groups and introduce a new method for measuring the similarity of reduced graphs based on a weighted edit distance. Moving beyond simple similarity searching, we show how more flexible queries can be built using reduced graphs and describe a database system that allows iterative querying with multiple representations. Reduced graphs capture many important features of ligand-receptor interactions and, in conjunction with other whole molecule descriptors, provide an informative way to review HTS data. We describe a novel use of reduced graphs in this context, introducing a method we have termed data-driven clustering, that identifies clusters of molecules represented by a particular whole molecule descriptor and enriched in active compounds.
The introduction of combinatorial chemistry groups into pharmaceutical companies provoked a desire for efficient and effective methods for library design and optimisation. This, in turn, has resulted in a large number of scientific publications, detailing a variety of approaches to the problem. This review attempts to describe the major works in the literature, to set them in context both chronologically and scientifically, and to identify the outstanding challenges that must be addressed, if this area of research is to maintain the rapid progress seen hitherto.
In this paper we introduce a quantitative model that relates chemical structural similarity to biological activity, and in particular to the activity of lead series of compounds in high-throughput assays. From this model we derive the optimal screening collection make up for a given fixed size of screening collection, and identify the conditions under which a diverse collection of compounds or a collection focusing on particular regions of chemical space are appropriate strategies. We derive from the model a diversity function that may be used to assess compounds for acquisition or libraries for combinatorial synthesis by their ability to complement an existing screening collection. The diversity function is linked directly through the model to the goal of more frequent discovery of lead series from high-throughput screening. We show how the model may also be used to derive relationships between collection size and probabilities of lead discovery in high-throughput screening, and to guide the judicious application of structural filters.
This paper addresses a major issue in library design, namely how to efficiently optimize the library size (number of products) and configuration (number of reagents at each position) simultaneously with other properties such as diversity, cost, and drug-like physicochemical property profiles. These objectives are often in competition, for example, minimizing the number of reactants while simultaneously maximizing diversity, and thus present difficulties for traditional optimization methods such as genetic algorithms and simulated annealing. Here, a multiobjective genetic algorithm (MOGA) is used to vary library size and configuration simultaneously with other library properties. The result is a family of solutions that explores the tradeoffs in the objectives. This is achieved without the need to assign relative weights to the objectives. The user is then able to make an informed choice on an appropriate compromise solution. The method has been applied to two different virtual libraries: a two-component aminothiazole library and a four-component benzodiazepine library.
Deriving quantitative structure-activity relationship (QSAR) models that are accurate, reliable, and easily interpretable is a difficult task. In this study, two new methods have been developed that aim to find useful QSAR models that represent an appropriate balance between model accuracy and complexity. Both methods are based on genetic programming (GP). The first method, referred to as genetic QSAR (or GPQSAR), uses a penalty function to control model complexity. GPQSAR is designed to derive a single linear model that represents an appropriate balance between the variance and the number of descriptors selected for the model. The second method, referred to as multiobjective genetic QSAR (MoQSAR), is based on multiobjective GP and represents a new way of thinking of QSAR. Specifically, QSAR is considered as a multiobjective optimization problem that comprises a number of competitive objectives. Typical objectives include model fitting, the total number of terms, and the occurrence of nonlinear terms. MoQSAR results in a family of equivalent QSAR models where each QSAR represents a different tradeoff in the objectives. A practical consideration often overlooked in QSAR studies is the need for the model to promote an understanding of the biochemical response under investigation. To accomplish this, chemically intuitive descriptors are needed but do not always give rise to statistically robust models. This problem is addressed by the addition of a further objective, called chemical desirability, that aims to reward models that consist of descriptors that are easily interpretable by chemists. GPQSAR and MoQSAR have been tested on various data sets including the Selwood data set and two different solubility data sets. The study demonstrates that the MoQSAR method is able to find models that are at least as good as models derived using standard statistical approaches and also yields models that allow a medicinal chemist to trade statistical robustness for chemical interpretability.
Early results from screening combinatorial libraries have been disappointing with libraries either failing to deliver the improved hit rates that were expected or resulting in hits with characteristics that make them undesirable as lead compounds. Consequently, the focus in library design has shifted toward designing libraries that are optimized on multiple properties simultaneously, for example, diversity and "druglike" physicochemical properties. Here we describe the program MoSELECT that is based on a multiobjective genetic algorithm and which is able to suggest a family of solutions to multiobjective library design where all the solutions are equally valid and each represents a different compromise between the objectives. MoSELECT also allows the relationships between the different objectives to be explored with competing objectives easily identified. The library designer can then make an informed choice on which solution(s) to explore. Various performance characteristics of MoSELECT are reported based on a number of different combinatorial libraries.
We describe a variety of the computational techniques which we use in the drug discovery and design process. Some of these computational methods are designed to support the new experimental technologies of high-throughput screening and combinatorial chemistry. We also consider some new approaches to problems of long-standing interest such as protein-ligand docking and the prediction of free energies of binding.
High-throughput screening has made a significant impact on drug discovery, but there is an acknowledged need for quantitative methods to analyze screening results and predict the activity of further compounds. In this paper we introduce one such method, binary kernel discrimination, and investigate its performance on two datasets; the first is a set of 1650 monoamine oxidase inhibitors, and the second a set of 101 437 compounds from an in-house enzyme assay. We compare the performance of binary kernel discrimination with a simple procedure which we call "merged similarity search", and also with a feedforward neural network. Binary kernel discrimination is shown to perform robustly with varying quantities of training data and also in the presence of noisy data. We conclude by highlighting the importance of the judicious use of general pattern recognition techniques for compound selection.
PLUMS is a new method to perform rational monomer selection for combinatorial chemistry libraries. The algorithm has been developed to optimize focused libraries with specific two-dimensional and/or three-dimensional properties. A preliminary step is the identification of those molecules in the initial virtual library which satisfy the imposed property constraints; we define these molecules as the virtual hits. From the virtual hits, PLUMS generates a starting library, which is the true combinatorial library that includes all the virtual hits. Monomers are then removed in an iterative fashion, thus reducing the size of the library. At each iteration, the worst monomer is removed. Each sublibrary is selected using a global scoring function, which balances effectiveness and efficiency. The iterative process continues until one is left with a library that consists entirely of virtual hits. The optimal library, which is the best compromise between effectiveness and efficiency, can then be selected according to the score. During the iterative process, equivalent solutions may well occur and are taken into account by the algorithm, according to a user-defined parameter. The number of monomers for each substitution site and the size of the library are parameters that can be either optimized or used to constrain the selection. The results obtained on two test libraries are presented. PLUMS was compared with genetic algorithms (GA) and monomer frequency analysis (MFA), which are widely used for monomer selection. For the two test libraries, PLUMS and GA gave equivalent results. MFA is the fastest method, but it can give misleading solutions. Possible advantages and disadvantages of the different methods are discussed.
Gridding and partitioning (GaP) is a computational method for the classification and selection of monomers for combinatorial libraries. The molecules are described in terms of the pharmacophoric groups they contain and where those pharmacophoric groups can be located in three-dimensional space. The approach involves a detailed conformational analysis of each molecule. This conformational analysis is done within a common coordinate frame, thus enabling the monomers to be compared. The use of a partitioned space is central to this particular application as it facilitates the identification of regions of space which are not well represented by existing compounds. Several ways to extend the use of partitioned pharmacophore spaces are described. Applications of the approach in monomer acquisition and in library design are outlined.
The program SELECT is presented for the design of combinatorial libraries. SELECT is based on a genetic algorithm with a multi-objective fitness function. Any number of objectives can be included, provided that they can be readily calculated. Typically, the objectives would be to maximize structural diversity while ensuring that the compounds in the library have "drug-like" properties. In the examples given, structural diversity is measured using Daylight fingerprints as descriptors and either the normalized sum of pairwise dissimilarities, calculated with the cosine coefficient, or the average nearest neighbor distance, calculated with the Tanimoto coefficient, as the measure of diversity. The objectives are specified at run time. Combinatorial libraries are selected by analyzing product space, which gives significant advantages over methods that are based on analyzing reactant space. SELECT can also be used to choose an optimal configuration for a multicomponent library. The performance of SELECT is demonstrated by its application to the design of a two-component amide library and to the design of a three-component thiazoline-2-imine library.
We describe an integrated suite of computational tools which are used to assist in the selection of compounds for biological assays and the design of combinatorial libraries. These functions are delivered in a platform-independent manner via a corporate intranet and are used by computational experts and nonexperts alike. While the system was primarily designed to be used prior to synthesis, it can also be used to provide structural information for library registration and for decoding beads in tagged libraries. We describe a simple statistical method for monomer selection and compare it to computationally more demanding approaches.
A series of benzophenone derivatives has been synthesized and evaluated as inhibitors of HIV-1 reverse transcriptase (RT) and the growth of HIV-1 in MT-4 cells. Through the use of the structure-activity relationships within this series of compounds and computational chemistry techniques, a binding conformation is proposed. The SAR also indicated that the major interactions of 1h with the RT enzyme are through hydrogen bonding of the amide and benzophenone carbonyls and pi-orbital interactions with the benzophenone nucleus and an aromatic function separated from the benzophenone by a suitable spacer group. The crystal structure of compound 1h has been determined. A number of compounds with potent inhibitory activity against HIV-1 RT and HIV in cellular assays at levels comparable with AZT and our efforts to identify a metabolically stable analogue are described.
A series of substituted imidazo[1,5-b]pyridazines have been prepared and tested for inhibitory activity against the reverse transcriptase of HIV-1 (RT) and their ability to inhibit the growth of infected MT-4 cells. Crystal data are reported on two compounds, 15c and 33. From the structure-activity relationships developed within this and other series, it is proposed that key features of the interaction with RT include hydrogen-bond acceptor and aromatic pi-orbital bonding with the imidazopyridazine nucleus and a benzoyl function separated from the heterocycle by a suitable spacer group. Exceptional activity against the reverse transcriptase of HIV-1 (IC50 = 0.65 nM) was obtained with a 2-imidazolyl-substituted derivative, 7-[2-(1H-imidazol-1-yl)-5-methylimidazo-[1,5-b]pyridazin-7-yl]-1-phenyl-1-heptanone (33) which is attributed to additional binding of the imidazole sp(2) nitrogen atom. A number of the compounds in this series also inhibit the replication of HIV-1 in vitro in MT-4 and C8166 cells at levels observed with the nucleoside AZT.