A new method is presented on how to calculate molecular complexity for chemical reactions by the fractal dimension of educts and products. Two pathways for the total synthesis of strychnine were compared. Significant differences in the two synthesis pathways were reflected by reaction complexity. These results demonstrate that reaction complexity is a powerful measure to group chemical reactions beyond substructural changes.
We present an efficient algorithm for substructure search in combinatorial libraries defined by synthons, i.e. substructures with connection points. Our method improves on existing approaches by introducing powerful heuristics and fast fingerprint screening to quickly eliminate branches of non matching combinations of synthons. With this we achieve typical response times of a few seconds on a standard desktop computer for searches in large combinatorial libraries like the Enamine REAL space. We published the Java source as part of the OpenChemLib under the BSD license, and we implemented tools to enable substructure search in custom combinatorial libraries.
Knowledge about the 3-dimensional structure, orientation and interaction of chemical compounds is important in many areas of science and technology. X-ray crystallography is one of the experimental techniques capable of providing a large amount of structural information for a given compound, and it is widely used for characterisation of organic and metal-organic molecules. The method provides precise 3D coordinates of atoms inside crystals, however, it does not directly deliver information about certain chemical characteristics such as bond orders, delocalization, charges, lone electron pairs or lone electrons. These aspects of a molecular model have to be derived from crystallographic data using refined information about interatomic distances and atom types as well as employing general chemical knowledge. This publication describes a curated automatic pipeline for the derivation of chemical attributes of molecules from crystallographic models. The method is applied to build a catalogue of chemical entities in an open-access crystallographic database, the Crystallography Open Database (COD). The catalogue of such chemical entities is provided openly as a derived database. The content of this catalogue and the problems arising in the fully automated pipeline are discussed, along with the possibilities to introduce manual data curation into the process.
Synthetically accessible chemical spaces provide a valuable source to search for small-molecule analogues or new starting points in drug discovery projects. Having a toolbox at hand that can automatically create searchable representations of such spaces using reaction definitions and building blocks as inputs is a prerequisite to put this approach into practice. Herein, we present a tool kit to create such virtual chemical spaces. It is part of the OpenChemLib, an open-source Cheminformatics tool kit. Furthermore, we demonstrate the creation of a several billion molecules large chemical space from commercial building blocks and a list of common organic chemistry reactions.
In drug discovery, molecules are optimized towards desired properties. In this context, machine learning is used for extrapolation in drug discovery projects. The limits of extrapolation for regression models are known. However, a systematic analysis of the effectiveness of extrapolation in drug discovery has not yet been performed. In response, this study examined the capabilities of six machine learning algorithms to extrapolate from 243 datasets. The response values calculated from the molecules in the datasets were molecular weight, cLogP, and the number of sp3-atoms. Three experimental set ups were chosen for response values. Shuffled data were used for interpolation, whereas data for extrapolation were sorted from high to low values, and the reverse. Extrapolation with sorted data resulted in much larger prediction errors than extrapolation with shuffled data. Additionally, this study demonstrated that linear machine learning methods are preferable for extrapolation.
There has been decades of research on determining and predicting acid dissociation constants (pKa) and the tautomer ratios both experimentally and theoretically. However, the lack of an extensive publicly available database of measured tautomeric ratios in water and nonaqueous solvents poses a challenge for the researchers interested in theoretical studies related to tautomers. Hereby, we present Tautobase, to date and to the best of our knowledge, the first extensive open-source tautomer database of measured and estimated tautomer ratios mainly in water, containing 1680 unique tautomer pairs.
Molecular complexity is an important characteristic of organic molecules for drug discovery. How to calculate molecular complexity has been discussed in the scientific literature for decades. It was known from early on that the numbers of substructures that can be cut out of a molecular graph are of importance for this task. However, it was never realized that the cut-out substructures show self-similarity to the parent structures. A successive removal of one bond and one atom returns a series of fragments with decreasing size. Such a series shows self-similarity similar to fractal objects. Here we used the number of distinct fragments to calculate the fractal dimension of the molecule. The fractal dimension of a molecule is a new matter constant that incorporates all features that are currently known to be important for describing molecular complexity. Furthermore, this is the first work that reveals the fractal nature of organic molecules.
Four datasets measuring DMPK (drug metabolism and pharmacokinetics) parameters, and one target protein-specific dataset were analyzed by machine learning methods. Parameters measured for the five compound sets were biological activity data, plasma protein binding, permeability in MDCK I cell layers, intrinsic clearance by human liver microsomes, and plasma exposure in orally dosed rats. The measured data were sorted chronologically, reflecting the order in which they had been obtained in the discovery project. Subsets of the chronologically sorted data that appeared early in the project were used as training datasets to build predictive models for subsequent compounds based on kNN, partial least squares regression (PLSR), nonlinear PLSR, random forest regression, and support vector regression. A median model was used as a baseline to assess the machine learning model prediction quality. Data sets sorted in order of increasing test set prediction error: intrinsic clearance, plasma protein binding, cell layer permeability, biological activity on target protein, and bioavailability as AUC in rats. Our results give a first estimation of the power of machine learning to predict DMPK properties of compounds in an ongoing drug discovery project.
The Platinum dataset of protein-bound ligand conformations was used to benchmark the ability of the MMFF94s force field to generate bioactive conformations by minimization of randomly generated conformers. Torsion angle parameters that generally caused wrong geometries were reparameterized by conducting dihedral scans using ab initio calculations at the MP2 level. This reparameterization resulted in a systematic improvement of generated conformations.
A new computational method is presented to extract disease patterns from heterogeneous and text-based data. For this study, 22 million PubMed records were mined for co-occurrences of gene name synonyms and disease MeSH terms. The resulting publication counts were transferred into a matrix Mdata. In this matrix, a disease was represented by a row and a gene by a column. Each field in the matrix represented the publication count for a co-occurring disease-gene pair. A second matrix with identical dimensions Mrelevance was derived from Mdata. To create Mrelevance the values from Mdata were normalized. The normalized values were multiplied by the column-wise calculated Gini coefficient. This multiplication resulted in a relevance estimator for every gene in relation to a disease. From Mrelevance the similarities between all row vectors were calculated. The resulting similarity matrix Srelevance related 5,000 diseases by the relevance estimators calculated for 15,000 genes. Three diseases were analyzed in detail for the validation of the disease patterns and the relevant genes. Cytoscape was used to visualize and to analyze Mrelevance and Srelevance together with the genes and diseases. Summarizing the results, it can be stated that the relevance estimator introduced here was able to detect valid disease patterns and to identify genes that encoded key proteins and potential targets for drug discovery projects.
In this case study on an essential instrument of modern drug discovery, we summarize our successful efforts in the last four years toward enhancing the Actelion screening compound collection. A key organizational step was the establishment of the Compound Library Committee (CLC) in September 2013. This cross-functional team consisting of computational scientists, medicinal chemists and a biologist was endowed with a significant annual budget for regular new compound purchases. Based on an initial library analysis performed in 2013, the CLC developed a New Library Strategy. The established continuous library turn-over mode, and the screening library size of 300'000 compounds were maintained, while the structural library quality was increased. This was achieved by shifting the selection criteria from 'druglike' to 'leadlike' structures, enriching for non-flat structures, aiming for compound novelty, and increasing the ratio of higher cost 'Premium Compounds'. Novel chemical space was gained by adding natural compounds, macrocycles, designed and focused libraries to the collection, and through mutual exchanges of proprietary compounds with agrochemical companies. A comparative analysis in 2016 provided evidence for the positive impact of these measures. Screening the improved library has provided several highly promising hits, including a macrocyclic compound, that are currently followed up in different Hit-to-Lead and Lead Optimization programs. It is important to state that the goal of the CLC was not to achieve higher HTS hit rates, but to increase the chances of identified hits to serve as the basis of successful early drug discovery programs. The experience gathered so far legitimates the New Library Strategy.
BACKGROUND:Wikipedia, the world's largest and most popular encyclopedia is an indispensable source of chemistry information. It contains among others also entries for over 15,000 chemicals including metabolites, drugs, agrochemicals and industrial chemicals. To provide an easy access to this wealth of information we decided to develop a substructure and similarity search tool for chemical structures referenced in Wikipedia.RESULTS:We extracted chemical structures from entries in Wikipedia and implemented a web system allowing structure and similarity searching on these data. The whole search as well as visualization system is written in JavaScript and therefore can run locally within a web page and does not require a central server. The Wikipedia Chemical Structure Explorer is accessible on-line at www.cheminfo.org/wikipedia and is available also as an open source project from GitHub for local installation.CONCLUSIONS:The web-based Wikipedia Chemical Structure Explorer provides a useful resource for research as well as for chemical education enabling both researchers and students easy and user friendly chemistry searching and identification of relevant information in Wikipedia. The tool can also help to improve quality of chemical entries in Wikipedia by providing potential contributors regularly updated list of entries with problematic structures. And last but not least this search system is a nice example of how the modern web technology can be applied in the field of cheminformatics. Graphical abstractWikipedia Chemical Structure Explorer allows substructure and similarity searches on molecules referenced in Wikipedia.
Drug discovery projects in the pharmaceutical industry accumulate thousands of chemical structures and ten-thousands of data points from a dozen or more biological and pharmacological assays. A sufficient interpretation of the data requires understanding, which molecular families are present, which structural motifs correlate with measured properties, and which tiny structural changes cause large property changes. Data visualization and analysis software with sufficient chemical intelligence to support chemists in this task is rare. In an attempt to contribute to filling the gap, we released our in-house developed chemistry aware data analysis program DataWarrior for free public use. This paper gives an overview of DataWarrior's functionality and architecture. Exemplarily, a new unsupervised, 2-dimensional scaling algorithm is presented, which employs vector-based or nonvector-based descriptors to visualize the chemical or pharmacophore space of even large data sets. DataWarrior uses this method to interactively explore chemical space, activity landscapes, and activity cliffs.
DDMiner, a new method for mining disease-disease associations in MEDLINE, is presented together with its first results. DDMiner searches for co-occurrences of gene names and disease terms, and finds relationships between diseases by word vector-similarity calculations. All records in PubMed were labeled with around 40,000 gene and protein names, and around 4,000 disease terms. Each disease term was described by a word vector from which the length equals the number of gene names. Each field in the vector represented a gene or a protein. The value in the field was derived from the number of publications in which this gene occurred together with the disease term. Disease-disease associations were calculated by vector-similarity calculation. Five diseases were examined together with their closest neighbor diseases to show the validity of our approach. All five examples showed only disease-disease associations that could be validated by medical literature. These results show that mining for disease-disease associations by second order co-occurrence is a powerful tool for medical science.
This paper reports the computational evaluation and experimental verification of 7-hydroxy-3-(1-phenyl-3-aryl-1H-pyrazol-5-yl)-4H-chromen-4-ones 3 and their o-β-d-glucopyranosides 5 for their antimicrobial and antioxidant activity. The prepared compounds were tested against various Gram-positive and Gram-negative bacteria species. Some of the synthesized compounds have shown potential antimicrobial and antioxidant activity. This POM bioinformatic study could greatly help to pharmacomodulate the potential antibiotics and antioxidants.
A new method is introduced to calculate the complexity of organic molecules in drug discovery. The complexity is calculated by taking the number of unique connected subgraphs u as basis c = f(a, b, p, u). With a and b are the number of atoms and bonds, respectively and p is the ratio of covered bonds by redundant fragments. A set of five datasets with 50 molecules each was analyzed. The datasets were compiled from bioactive natural products, approved drugs, highly bioactive molecules, commercially available compounds for high throughput screening and artificial generated molecules. Comparing the median of c for the five datasets showed a significant increase in the following order: commercially available compounds < bioactive molecules < approved drugs < natural products < artificial molecules. With the introduced complexity value c a meaningful figure of merit was developed to assess automatically the complexity of single compounds and compound libraries in drug discovery.
A new subpharmacophore-based virtual screening method is introduced. Subpharmacophores are derived from large active molecules to detect small bioactive molecules as seeds for starting points in medicinal chemistry programs. A large data set was assembled from the ChEMBL database to check the validity of this approach. Molecules for 133 targets with molecular weights: between 450 and 850 were selected as queries. For the query molecules, the pharmacophore descriptors were calculated. Up to 56 000 subpharmacophore descriptors with five to seven pharmacophore points were derived from the query pharmacophores. The subpharmacophore descriptors were used as queries to screen 1079 test data sets, containing decoys and spike molecules. A maximum upper molecular weight limit of 400 Da was set for the test molecules. Three different chemical fingerprint descriptors were used for comparison purposes. The subpharmacophore approach detected active molecules for 85 out of 133 targets and outperformed the chemical fingerprints. This ligand-based virtual screening experiment was triggered by the needs of medicinal chemistry. Applying the subpharmacophore method in a medicinal chemistry program, where a lead molecule with a molecular weight of 800 Da was available, resulted in a new series of molecules with molecular weights below 400.
A new machine vision method is presented to detect micro tubes with colored solutions or undissolved compounds in solvents on screening plates used in drug discovery. The method presented herein takes an image of a 96 tubes micro tube rack as input. After applying edge detection on the input image, the circles that characterize the borders of the micro tubes are detected via shape matching. A gradient search is used to extract the colors from the micro tubes. Automatic calibration is applied to adapt to changing light conditions. Experiments with more than two thousand micro tube racks (approx. 170,000 compounds) compiled over three years showed that the method is fast and robust.
A new software tool (Gene2diseaseMapper) is presented that takes the HUGO symbol of a gene as input and delivers a ranked list of diseases, considering microarray experiments. Gene name synonym expansion was used to generate a query for MEDLINE. Retrieved article records were filtered by a disease stoplist which was created from Medical Subject Headings (MeSH). From the ArrayExpress microarray database, all experiments were retrieved in which the gene was differentially expressed. The experiments were searched for disease terms extracted from the MEDLINE articles. A similarity function was developed to compare the MeSH terms and the terms in ArrayExpress. A scoring function was implemented which ranked the disease MeSH terms according similarity and frequency. The method was explored with 12 genes for whose corresponding protein a drug was approved or is under development. Gene2diseaseMapper was able to find the diseases of the approved drugs in 11 out of 12 cases.