We explore a "best-of-both" approach to modeling molecular properties by combining learned molecular descriptors from a graph neural network (GNN) with general-purpose descriptors and a mixed ensemble of machine learning (ML) models. We introduce a MetaModel framework to aggregate predictions from a diverse set of leading ML models. We present a featurization scheme for combining task-specific GNN-derived features with conventional molecular descriptors. We demonstrate that our framework outperforms the cutting-edge ChemProp model on all regression data sets tested and 6 of 9 classification data sets. We further show that including the GNN features derived from ChemProp boosts the ensemble model's performance on several data sets where it otherwise would have underperformed. We conclude that to achieve optimal performance across a wide set of problems, it is vital to combine general-purpose descriptors with task-specific learned features and to use a diverse set of ML models to make the predictions.
Animal pharmacokinetic (PK) data as well as human and animal in vitro systems are utilized in drug discovery to define the rate and route of drug elimination. Accurate prediction and mechanistic understanding of drug clearance and disposition in animals provide a degree of confidence for extrapolation to humans. In addition, prediction of in vivo properties can be used to improve design during drug discovery, help select compounds with better properties, and reduce the number of in vivo experiments. In this study, we generated machine learning models able to predict rat in vivo PK parameters and concentration-time PK profiles based on the molecular chemical structure and either measured or predicted in vitro parameters. The models were trained on internal in vivo rat PK data for over 3000 diverse compounds from multiple projects and therapeutic areas, and the predicted endpoints include clearance and oral bioavailability. We compared the performance of various traditional machine learning algorithms and deep learning approaches, including graph convolutional neural networks. The best models for PK parameters achieved R2 = 0.63 [root mean squared error (RMSE) = 0.26] for clearance and R2 = 0.55 (RMSE = 0.46) for bioavailability. The models provide a fast and cost-efficient way to guide the design of molecules with optimal PK profiles, to enable the prediction of virtual compounds at the point of design, and to drive prioritization of compounds for in vivo assays.
Predicting the sensory properties of compounds is challenging due to the subjective nature of the experimental measurements. This testing relies on a panel of human participants and is therefore also expensive and time-consuming. We describe the application of a state-of-the-art deep learning method, Alchemite™, to the imputation of sparse physicochemical and sensory data and compare the results with conventional quantitative structure–activity relationship methods and a multi-target graph convolutional neural network. The imputation model achieved a substantially higher accuracy of prediction, with improvements in R2 between 0.26 and 0.45 over the next best method for each sensory property. We also demonstrate that robust uncertainty estimates generated by the imputation model enable the most accurate predictions to be identified and that imputation also more accurately predicts activity cliffs, where small changes in compound structure result in large changes in sensory properties. In combination, these results demonstrate that the use of imputation, based on data from less expensive, early experiments, enables better selection of compounds for more costly studies, saving experimental time and resources.
More accurate predictions of the biological properties of chemical compounds would guide the selection and design of new compounds in drug discovery and help to address the enormous cost and low success-rate of pharmaceutical R&D. However this domain presents a significant challenge for AI methods due to the sparsity of compound data and the noise inherent in results from biological experiments. In this paper, we demonstrate how data imputation using deep learning provides substantial improvements over quantitative structure-activity relationship (QSAR) machine learning models that are widely applied in drug discovery. We present the largest-to-date successful application of deep-learning imputation to datasets which are comparable in size to the corporate data repository of a pharmaceutical company (678,994 compounds by 1166 endpoints). We demonstrate this improvement for three areas of practical application linked to distinct use cases; i) target activity data compiled from a range of drug discovery projects, ii) a high value and heterogeneous dataset covering complex absorption, distribution, metabolism and elimination properties and, iii) high throughput screening data, testing the algorithm’s limits on early-stage noisy and very sparse data. Achieving median coefficients of determination, R, of 0.69, 0.36 and 0.43 respectively across these applications, the deep learning imputation method offers an unambiguous improvement over random forest QSAR methods, which achieve median R values of 0.28, 0.19 and 0.23 respectively. We also demonstrate that robust estimates of the uncertainties in the predicted values correlate strongly with the accuracies in prediction, enabling greater confidence in decision-making based on the imputed values.
Imputation is a powerful statistical method that is distinct from the predictive modelling techniques more commonly used in drug discovery. Imputation uses sparse experimental data in an incomplete dataset to predict missing values by leveraging correlations between experimental assays. This contrasts with quantitative structure–activity relationship methods that use only descriptor – assay correlations. We summarize three recent imputation strategies – heterogeneous deep imputation, assay profile methods and matrix factorization – and compare these with quantitative structure–activity relationship methods, including deep learning, in drug discovery settings. We comment on the value added by imputation methods when used in an ongoing project and find that imputation produces stronger models, earlier in the project, over activity and absorption, distribution, metabolism and elimination end points.
Current in vitro models for hepatotoxicity commonly suffer from low detection rates due to incomplete coverage of bioactivity space. Additionally, in vivo exposure measures such as Cmax are used for hepatotoxicity screening and are unavailable early on. Here we propose a novel rule-based framework to extract interpretable and biologically meaningful multiconditional associations to prioritize in vitro end points for hepatotoxicity and understand the associated physicochemical conditions. The data used in this study were derived for 673 compounds from 361 ToxCast bioactivity measurements and 29 calculated physicochemical properties against two lowest effective levels (LEL) of rodent hepatotoxicity from ToxRefDB, namely 15 mg/kg/day and 500 mg/kg/day. To achieve 80% coverage of toxic compounds, 35 rules with accuracies ranging from 96% to 73% using 39 unique ToxCast assays are needed at a threshold level of 500 mg/kg/day, whereas to describe the same coverage at a threshold of 15 mg/kg/day, 20 rules with accuracies of between 98% and 81% were needed, comprising 24 unique assays. Despite the 33-fold difference in dose levels, we found relative consistency in the key mechanistic groups in rule clusters, namely (i) activities against Cytochrome P, (ii) immunological responses, and (iii) nuclear receptor activities. Less specific effects, such as oxidative stress and cell cycle arrest, were used more by rules to describe toxicity at the level of 500 mg/kg/day. Although the endocrine disruption through nuclear receptor activity formulated an essential cluster of rules, this bioactivity was not covered in four commercial assay setups for hepatotoxicity. Using an external set of 29 drugs with drug-induced liver injury (DILI) labels, we found that promiscuity over important assays discriminates between compounds with different levels of liver injury. In vitro-in vivo associations were also improved by incorporating physicochemical properties especially for the potent, 15 mg/kg/day toxicity level as well for assays describing nuclear receptor activity and phenotypic changes. The most frequently used physicochemical properties, predictive for hepatotoxicity in combination with assay activities, are linked to bioavailability, which were the number of rotatable bonds (less than 7) at a of level of 15 mg/kg/day and the number of rings (of less than 3) at level of 500 mg/kg/day. In summary, hepatotoxicity cannot very well be captured by single assay end points, but better by a combination of bioactivities in relevant assays, with the likelihood of hepatotoxicity increasing with assay promiscuity. Together, these findings can be used to prioritize assay combinations that are appropriate to assess potential hepatotoxicity.
Despite the increasing knowledge in both the chemical and biological domains the assimilation and exploration of heterogeneous datasets, encoding information about the chemical, bioactivity and phenotypic properties of compounds, remains a challenge due to requirement for overlap between chemicals assayed across the spaces. Here, we have constructed a novel dataset, larger than we have used in prior work, comprising 579 acute oral toxic compounds and 1427 non-toxic compounds derived from regulatory GHS information, along with their corresponding molecular and protein target descriptors and qHTS in vitro assay readouts from the Tox21 project. We found no clear association between the results of a FAFDrugs4 toxicophore screen and the acute oral toxicity classifications for our compound set; and a screen using a subset of the ToxAlerts toxicophores was also of limited utility, with only slight enrichment toward the toxic set (odds ratio of 1.48). We then investigated to what degree toxic and non-toxic compounds could be separated in each of the spaces, to compare their potential contribution to further analyses. Using an LDA projection, we found the largest degree of separation using chemical descriptors (Cohen's d of 1.95) and the lowest degree of separation between toxicity classes using qHTS descriptors (Cohen's d of 0.67). To compare the predictivity of the feature spaces for the toxicity endpoint, we next trained Random Forest (RF) acute oral toxicity classifiers on either molecular, protein target and qHTS descriptors. RFs trained on molecular and protein target descriptors were most predictive, with ROC AUC values of 0.80-0.92 and 0.70-0.85, respectively, across three test sets. RFs trained on both chemical and protein target descriptors combined exhibited similar predictive performance to the single-domain models (ROC AUC of 0.80-0.91). Model interpretability was improved by the inclusion of protein target descriptors, which allow the identification of specific targets (e.g. Retinal dehydrogenase) with literature links to toxic modes of action (e.g. oxidative stress). The dataset compiled in this study has been made available for future application.
These abstracts were presented at the 2017 annual meeting of the UK In Vitro Toxicology Society (IVTS). The meeting was hosted at the Senate House in London, UK on November 23–24, 2017. The main session topics included hepatotoxicity; dermal and barrier toxicity; IVIVE, exposure and non-mammalian; cardiotoxicity; neurotoxicity; and genotoxicity.
Adverse events resulting from drug therapy can be a cause of drug withdrawal, reduced and or restricted clinical use, as well as a major economic burden for society. To increase the safety of new drugs, there is a need to better understand the mechanisms causing the adverse events. One way to derive new mechanistic hypotheses is by linking data on drug adverse events with the drugs' biological targets. In this study, we have used data mining techniques and mutual information statistical approaches to find associations between reported adverse events collected from the FDA Adverse Event Reporting System and assay outcomes from ToxCast, with the aim to generate mechanistic hypotheses related to structural cardiotoxicity (morphological damage to cardiomyocytes and/or loss of viability). Our workflow identified 22 adverse event-assay outcome associations. From these associations, 10 implicated targets could be substantiated with evidence from previous studies reported in the literature. For two of the identified targets, we also describe a more detailed mechanism, forming putative adverse outcome pathways associated with structural cardiotoxicity. Our study also highlights the difficulties deriving these type of associations from the very limited amount of data available.