Recent technological advances enable genomics of individual cells, the building blocks of all living organisms. Single cell data characteristics differ from those of bulk data, which led to a plethora of new analytical strategies. However, solutions are only useful for experts and currently, there are no widely accepted gold standards for single cell data analysis. To meet the requirements of analytical flexibility, ease of use and data security, we developed FASTGenomics ( https://fastgenomics.org ) as a powerful, efficient, versatile, robust, safe and intuitive analytical ecosystem for single-cell transcriptomics.
Support vector machines are a popular machine learning method for many classification tasks in biology and chemistry. In addition, the support vector regression (SVR) variant is widely used for numerical property predictions. In chemoinformatics and pharmaceutical research, SVR has become the probably most popular approach for modeling of non-linear structure-activity relationships (SARs) and predicting compound potency values. Herein, we have systematically generated and analyzed SVR prediction models for a variety of compound data sets with different SAR characteristics. Although these SVR models were accurate on the basis of global prediction statistics and not prone to overfitting, they were found to consistently mispredict highly potent compounds. Hence, in regions of local SAR discontinuity, SVR prediction models displayed clear limitations. Compared to observed activity landscapes of compound data sets, landscapes generated on the basis of SVR potency predictions were partly flattened and activity cliff information was lost. Taken together, these findings have implications for practical SVR applications. In particular, prospective SVR-based potency predictions should be considered with caution because artificially low predictions are very likely for highly potent candidate compounds, the most important prediction targets.
Active compounds can participate in different local structure-activity relationship (SAR) environments and introduce different degrees of local SAR discontinuity, depending on their structural and potency relationships in data sets. Such SAR features have thus far mostly been analyzed using descriptive approaches, in particular, on the basis of activity landscape modeling. However, compounds in different local SAR environments have not yet been predicted. Herein, we adapt the emerging chemical patterns (ECP) method, a machine learning approach for compound classification, to systematically predict compounds with different local SAR characteristics. ECP analysis is shown to accurately assign many compounds to different local SAR environments across a variety of activity classes covering the entire range of observed local SARs. Control calculations using random forests and multiclass support vector machines were carried out and a variety of statistical performance measures were applied. In all instances, ECP calculations yielded comparable or better performance than controls. The approach presented herein can be applied to predict compounds that complement local SARs or prioritize compounds with different SAR characteristics.
A new methodology for activity prediction of compounds from SAR matrices is introduced that is based upon conditional probabilities of activity. The approach has low computational complexity, is primarily designed for hit expansion from biological screening data, and accurately predicts both active and inactive compounds. Its performance is comparable to state-of-the-art machine learning methods such as support vector machines or Bayesian classification. Matrix-based activity prediction of virtual compounds further extends the spectrum of computational methods for compound design.
Support vector machines (SVMs) are among the preferred machine learning algorithms for virtual compound screening and activity prediction because of their frequently observed high performance levels. However, a well-known conundrum of SVMs (and other supervised learning methods) is the black box character of their predictions, which makes it difficult to understand why models succeed or fail. Herein we introduce an approach to rationalize the performance of SVM models based upon the Tanimoto kernel compared with the linear kernel. Model comparison and interpretation are facilitated by a visualization technique, making it possible to identify descriptor features that determine compound activity predictions. An implementation of the methodology has been made freely available.
Support vector machines (SVMs) are among the most popular machine learning methods for compound classification and other chemoinformatics tasks such as, for example, the prediction of ligand-target pairs or compound activity profiles. Depending on the specific applications, different SVM strategies can be used. For example, in the context of potency-directed virtual screening, linear combinations of multiple SVM models have been shown to enrich database selection sets with potent compounds compared to individual models. An open question concerning the use of SVM linear combinations (SVM-LCs) is how to best weight the models on a relative scale. Typically, linear weights are subjectively set. Herein, preferred weighting factors for SVM-LC were systematically determined. Therefore, weights were treated as meta-parameters and optimized by machine learning to enrich data set rankings with highly active compounds. The meta-parameter approach has been applied to 10 screening data sets and found to further improve SVM performance over other SVM-LCs and support vector regression (SVR) models. The results show that optimal weights depend on data set characteristics and chosen molecular representations. In addition, individual models often do not contribute to the performance of SVM-LCs. Taken together, these findings emphasize the need for systematic meta-parameter estimation.
Supervised machine learning models are widely used in chemoinformatics, especially for the prediction of new active compounds or targets of known actives. Bayesian classification methods are among the most popular machine learning approaches for the prediction of activity from chemical structure. Much work has focused on predicting structure-activity relationships (SARs) on the basis of experimental training data. By contrast, only a few efforts have thus far been made to rationalize the performance of Bayesian or other supervised machine learning models and better understand why they might succeed or fail. In this study, we introduce an intuitive approach for the visualization and graphical interpretation of naïve Bayesian classification models. Parameters derived during supervised learning are visualized and interactively analyzed to gain insights into model performance and identify features that determine predictions. The methodology is introduced in detail and applied to assess Bayesian modeling efforts and predictions on compound data sets of varying structural complexity. Different classification models and features determining their performance are characterized in detail. A prototypic implementation of the approach is provided.
Profiling of compound libraries against arrays of targets has become an important approach in pharmaceutical research. The prediction of multi‐target compound activities also represents an attractive task for machine learning with potential for drug discovery applications. Herein, we have explored activity prediction in high‐dimensional target space. Different types of models were derived to predict multi‐target activities. The models included naïve Bayesian (NB) and support vector machine (SVM) classifiers based upon compound structure information and NB models derived on the basis of activity profiles, without considering compound structure. Because the latter approach can be applied to incomplete training data and principally depends on the feature independence assumption, SVM modeling was not applicable in this case. Furthermore, iterative hybrid NB models making use of both activity profiles and compound structure information were built. In high‐dimensional target space, NB models utilizing activity profile data were found to yield more accurate activity predictions than structure‐based NB and SVM models or hybrid models. An in‐depth analysis of activity profile‐based models revealed the presence of correlation effects across different targets and rationalized prediction accuracy. Taken together, the results indicate that activity profile information can be effectively used to predict the activity of test compounds against novel targets.
Profiling of compounds against target families has become an important approach in pharmaceutical research for the identification of hits and analysis of selectivity and promiscuity patterns. We report on modeling of profiling experiments involving 429 potential inhibitors and a panel of 24 different kinases using support vector machine ( SVM ) techniques and naïve Bayesian classification. The experimental matrix contained many different activity profiles. SVM predictions achieved overall high accuracy due to consistently low false‐positive and consistently high true‐negative rates. However, predictions for promiscuous inhibitors were affected by false‐negative rates. Combined target‐based SVM classifiers reached or exceeded the performance of SVM profile prediction methods and were superior to Bayesian classification. The classifiers displayed different prediction characteristics including diverse combinations of false‐positive and true‐negative rates. Predicted and experimentally observed compound activity profiles were compared in detail, revealing activity patterns modeled with different accuracy.
We present an automated procedure for the creation of architectural plant models. It uses an algorithm for the computation of skeletons from sensor data. The skeletons are annotated with semantic labels for the extraction of architectural parameters. The values of several samples are averaged and serve as the basis for the model, which is implemented as a Relational Growth Grammar.
Supervised machine learning approaches, including support vector machines, random forests, Bayesian classifiers, nearest-neighbor similarity searching, and a conceptually distinct mapping algorithm termed DynaMAD, have been investigated for their ability to detect structurally related ligands of a given receptor with different mechanisms of action. For this purpose, a large number of simulated virtual screening trials were carried out with models trained on mechanistic subsets of different classes of receptor ligands. The results revealed that ligands with the desired mechanism of action were frequently contained in database selection sets of limited size. All machine learning approaches successfully detected mechanistic subsets of ligands in a large background database of druglike compounds. However, the early enrichment characteristics considerably differed. Overall, random forests of relatively simple design and support vector machines with Gaussian kernels (Gaussian SVMs) displayed the highest search performance. In addition, DynaMAD was found to yield very small selection sets comprising only ~10 compounds that also contained ligands with the desired mechanism of action. Random forest, Gaussian SVM, and DynaMAD calculations revealed an enrichment of compounds with the desired mechanism over other mechanistic subsets.
The emerging chemical patterns (ECP) approach has been introduced for compound classification. Thus far, only very few ECP applications have been reported. Here, we further investigate the ECP methodology by studying complex classification problems. The analysis involves multi-target data sets with systematically organized subsets of compounds having distinct or overlapping target activities and, in addition, data sets containing classes of specifically active compounds with different mechanism-of-action. In systematic classification trials focusing on individual compound subsets or mechanistic classes, ECP calculations utilizing numerical descriptors achieve moderate to high sensitivity, dependent on the data set, and consistently high specificity. Accurate ECP predictions are already obtained on the basis of very small learning sets with only three positive training instances, which distinguishes the ECP approach from many other machine learning techniques.
Breeding pest resistant crops can reduce the application of plant protection products. This in turn will significantly reduce biotical stress of the plants, and damage to the environment. To enable greater efficiency in crop breeding and to optimize decision making in crop management, there is an increasing demand for non-biased and faster assessment of plant traits in lab and field. Within a subproject of CROP.SENSe.net, an interdisciplinary research network of Bonn University and the research centre Jülich, we are working on a non-destructive solution to analyze and screen plant phenotypes. We aim for complete and precise 3D reconstructions of plants based on sensor data. To meet the challenge of occlusions and self-occlusions in sensor-based measurements, we employ a model-based approach "to make hidden object parts visible." First results show that important breeding characteristics can be derived by combining the visible exterior plant components with the model-induced interior plant components.
In drug discovery, domain experts from different fields such as medicinal chemistry, biology, and computer science often collaborate to develop novel pharmaceutical agents. Computational models developed in this process must be correct and reliable, but at the same time interpretable. Their findings have to be accessible by experts from other fields than computer science to validate and improve them with domain knowledge. Only if this is the case, the interdisciplinary teams are able to communicate their scientific results both precisely and intuitively. This work is concerned with the development and interpretation of machine learning models for drug discovery. To this end, it describes the design and application of computational models for specialized use cases, such as compound profiling and hit expansion. Novel insights into machine learning for ligand-based virtual screening are presented, and limitations in the modeling of compound potency values are highlighted. It is shown that compound activity can be predicted based on high-dimensional target profiles, without the presence of molecular structures. Moreover, support vector regression for potency prediction is carefully analyzed, and a systematic misprediction of highly potent ligands is discovered. Furthermore, a key aspect is the interpretation and chemically accessible representation of the models. Therefore, this thesis focuses especially on methods to better understand and communicate modeling results. To this end, two interactive visualizations for the assessment of näıve Bayes and support vector machine models on molecular fingerprints are presented. These visual representations of virtual screening models are designed to provide an intuitive chemical interpretation of the results.
Computational plant modeling from 3D sensor data is crucial for the early assessment of plant traits. Semantic modeling enables the incorporation of knowledge about the plant species, leading to an improvement of purely geometrical skeletonization approaches. Structural plant features can thereby robustly be extracted from the sensor data.