The development of in silico tools able to predict bioactivity and toxicity of chemical substances is a powerful solution envisioned to assess toxicity as early as possible. To enable the development of such tools, the ToxCast program has generated and made publicly available in vitro bioactivity data for thousands of compounds. The goal of the present study is to characterize and explore the data from ToxCast in terms of Machine Learning capability. For this, a large scale analysis on the entire database has been performed to build models to predict bioactivities measured in in vitro assays. Simple classical QSAR algorithms (ANN, SVM, LDA, random forest, and Bayesian) were first applied on the data, and the results of these algorithms suggested that they do not seem to be well-suited for data sets with a high proportion of inactive compounds. The study then showed for the first time that the use of an ensemble method named "Stacked generalization" could improve the model performance on this type of data. Indeed, for 61% of 483 models, the Stacked method led to models with higher performance. Moreover, the combination of this ensemble method with an applicability domain filter allows one to assess the reliability of the predictions for further compound prioritization. In particular we showed that for 50% of the models, the ROC score is better if we do not consider the compounds that are not within the applicability domain.
In many ways, a living cell can be compared to a complex factory animated by molecular nanomachines, mainly proteins complexes. Hence it is easy to conceive that the expression of proteins, which are cellular effectors, cannot be constant. On the contrary, it is highly dependent on the general context; environmental conditions (pH, temperature, oxygenation, nutrient availability), developmental stage of an organism (fetal spectrum of proteins differ from adult proteins in mammals), response to a stress (UV irradiation, presence of a chemical toxic, osmotic pressure alteration) and even diseases (cancer, attack of a pathogen) are examples of contextual changes in the level of protein expression. In order to understand this cellular state plasticity, a simplified view of this machinery, following general transfers of information according to the central dogma of molecular biology, is the sequence of events: (1) stimulation via a signaling pathway (e.g. presence of an environmental stimulation, followed by internal transduction of the signal), (2) effective stimulation of a transcription factor, (3) activation of the transcription of a particular gene, (4) production of messenger RNA (mRNA) (see Fig. 2.1), (5) translation of mRNA, i.e. production of a functional