Advances in machine learning and artificial intelligence have recently extended to the quantitative prediction of drug-drug interaction (DDI). Because DDIs arise from diverse mechanisms and the required level of predictive accuracy varies with both the endpoint and the stage of drug development, evaluating their significance and deciding what is needed demand unusually broad expertise-ranging from fundamental biology all the way to state-of-the-art machine-learning methods. In this review, DMPK scientists with expertise in machine learning survey and critique the most recent literature covering the following DDI categories: Cytochrome P450 (CYP) substrates, CYP competitive and time-dependent inhibition, CYP induction, non-CYP substrates, non-CYP inhibition, transporter substrates, transporter inhibition, and cutting-edge predictive algorithms based on deep learning applied for the task of DDIs. For each category we summarize current in silico methodologies and their performance, and we provide expert opinions on how these tools can be optimally incorporated into contemporary drug-discovery workflows.
In research focused on protein-protein interaction (PPI) inhibitors, the optimization process to achieve both high inhibitory activity and favorable physicochemical properties remains challenging. Our previous study reported the discovery of novel and bioavailable Keap1-Nrf2 PPI inhibitor 8 which exhibited moderate in vivo activity in rats. In this work, we present our subsequent efforts to optimize this compound. Two distinct approaches were employed, targeting high energy water molecules and Ser602 as "hot spots" from the anchor with good aqueous solubility, metabolic stability, and membrane permeability. Through ligand efficiency (LE)-guided exploration, we identified two novel inhibitors 22 and 33 with good pharmacokinetics (PK) profiles and more potent in vivo activities, which appear to be promising chemical probes among the existing inhibitors.
Various toxicity and pharmacokinetic evaluations as screening experiments are needed at the drug discovery stage. Currently, to reduce the use of animal experiments and developmental expenses, the development of high-performance predictive models based on quantitative structure-activity relationship analysis is desired. From these evaluation targets, we selected 50% lethal dose (LD50), blood-brain barrier penetration (BBBP), and the clearance (CL) pathway for this investigation and constructed predictive models for each target using 636-11,886 compounds. First, we constructed predictive models using the DeepSnap-deep learning (DL) method and images of compounds as features. The calculated area under the curve (AUC) and balanced accuracy (BAC) were, respectively, 0.887 and 0.818 for LD50, 0.893 and 0.824 for BBBP, and 0.883 and 0.763 for the CL pathway. Next, molecular descriptors (MDs) of compounds were calculated using Molecular Operating Environment, alvaDesc, and ADMET Predictor to construct predictive models using the MD-based method. Using these MDs, we constructed predictive models using DataRobot. The calculated AUC and BAC were, respectively, 0.931 and 0.805 for LD50, 0.919 and 0.849 for BBBP, and 0.900 and 0.807 for the CL pathway. In this investigation, we constructed predictive models combining the DeepSnap-DL and MD-based methods. In ensemble models using the mean predictive probability of the DeepSnap-DL and MD-based methods, the calculated AUC and BAC were, respectively, 0.942 and 0.842 for LD50, 0.936 and 0.853 for BBBP, and 0.908 and 0.832 for the CL pathway, with improved predictive performance observed for all variables compared with either single method alone. Moreover, in consensus models that adopted only compounds for which the results of the two methods agreed, the calculated BAC for LD50, BBBP, and the CL pathway were 0.916, 0.918, and 0.847, respectively, indicating higher predictive performance than the ensemble models for all three variables. The predictive models combining the DeepSnap-DL and MD-based methods displayed high predictive performance for LD50, BBBP, and the CL pathway. Therefore, the application of this approach to prediction targets in various drug discovery screenings is expected to accelerate drug discovery.
Oxidative stress is one of the causes of progression of chronic kidney disease (CKD). Activation of the antioxidant protein regulator Nrf2 by inhibition of the Keap1-Nrf2 protein-protein interaction (PPI) is of interest as a potential treatment for CKD. We report the identification of the novel and weak PPI inhibitor 7 with good physical properties by a high throughput screening (HTS) campaign, followed by structural and computational analysis. The installation of only methyl and fluorine groups successfully provided the lead compound 25, which showed more than 400-fold stronger activity. Furthermore, these dramatic substituent effects can be explained by the analysis of using isothermal titration calorimetry (ITC). Thus, the resulting 25, which exhibited high oral absorption and durability, would be a CKD therapeutic agent because of the dose-dependent manner for up-regulation of the antioxidant protein heme oxigenase-1 (HO-1) in rat kidneys.
Pharmacokinetic research plays an important role in the development of new drugs. Accurate predictions of human pharmacokinetic parameters are essential for the success of clinical trials. Clearance (CL) and volume of distribution (Vd) are important factors for evaluating pharmacokinetic properties, and many previous studies have attempted to use computational methods to extrapolate these values from nonclinical laboratory animal models to human subjects. However, it is difficult to obtain sufficient, comprehensive experimental data from these animal models, and many studies are missing critical values. This means that studies using nonclinical data as explanatory variables can only apply a small number of compounds to their model training. In this study, we perform missing-value imputation and feature selection on nonclinical data to increase the number of training compounds and nonclinical datasets available for these kinds of studies. We could obtain novel models for total body clearance (CLtot) and steady-state Vd (Vdss) (CLtot: geometric mean fold error [GMFE], 1.92; percentage within 2-fold error, 66.5%; Vdss: GMFE, 1.64; percentage within 2-fold error, 71.1%). These accuracies were comparable to the conventional animal scale-up models. Then, this method differs from animal scale-up methods because it does not require animal experiments, which continue to become more strictly regulated as time passes.
The toxicity, absorption, distribution, metabolism, and excretion properties of some targets are difficult to predict by quantitative structure-activity relationship analysis. Therefore, there is a need for a new prediction method that performs well for these targets. The aim of this study was to develop a new regression model of rat clearance (CL). We constructed a regression model using 1545 in-house compounds for which we had rat CL data. Molecular descriptors were calculated using molecular operating environment, alvaDesc, and ADMET Predictor software. The classification model of DeepSnap and Deep Learning (DeepSnap-DL) with images of the three-dimensional chemical structures of compounds as features was constructed, and the prediction probabilities for each compound were calculated. For molecular descriptor-based methods that use molecular descriptors and conventional machine learning algorithms selected by DataRobot, the correlation coefficient (R2) and root mean square error (RMSE) were 0.625-0.669 and 0.295-0.318, respectively. We combined molecular descriptors and prediction probability of DeepSnap-DL as features and developed a novel regression method we called the combination model. In the combination model with these two types of features and conventional algorithms selected by DataRobot, R2 and RMSE were 0.710-0.769 and 0.247-0.278, respectively. This finding shows that the combination model performed better than molecular descriptor-based methods. Our combination model will contribute to the design of more rational compounds for drug discovery. This method may be applicable not only to rat CL but also to other pharmacokinetic and pharmacological activity and toxicity parameters; therefore, applying it to other parameters may help to accelerate drug discovery.
Some targets predicted by machine learning (ML) in drug discovery remain a challenge because of poor prediction. In this study, a new prediction model was developed and rat clearance (CL) was selected as a target because it is difficult to predict. A classification model was constructed using 1545 in-house compounds with rat CL data. The molecular descriptors calculated by Molecular Operating Environment (MOE), alvaDesc, and ADMET Predictor software were used to construct the prediction model. In conventional ML using 100 descriptors and random forest selected by DataRobot, the area under the curve (AUC) and accuracy (ACC) were 0.883 and 0.825, respectively. Conversely, the prediction model using DeepSnap and Deep Learning (DeepSnap-DL) with compound features as images had AUC and ACC of 0.905 and 0.832, respectively. We combined the two models (conventional ML and DeepSnap-DL) to develop a novel prediction model. Using the ensemble model with the mean of the predicted probabilities from each model improved the evaluation metrics (AUC = 0.943 and ACC = 0.874). In addition, a consensus model using the results of the agreement between classifications had an increased ACC (0.959). These combination models with a high level of predictive performance can be applied to rat CL as well as other pharmacokinetic parameters, pharmacological activity, and toxicity prediction. Therefore, these models will aid in the design of more rational compounds for the development of drugs.
Research into pharmacokinetics plays an important role in the development process of new drugs. Accurately predicting human pharmacokinetic parameters from preclinical data can increase the success rate of clinical trials. Since clearance (CL) which indicates the capacity of the entire body to process a drug is one of the most important parameters, many methods have been developed. However, there are still rooms to be improved for practical use in drug discovery research; "improving CL prediction accuracy" and "understanding the chemical structure of compounds in terms of pharmacokinetics". To improve those, this research proposes a multimodal learning method based on deep learning that takes not only the chemical structure of a drug but also rat CL as inputs. Good results were obtained compared with the conventional animal scale-up method; the geometric mean fold error was 2.68 and the proportion of compounds with prediction errors of 2-fold or less was 48.5%. Furthermore, it was found to be possible to infer the partial structure useful for CL prediction by a structure contributing factor inference method. The validity of these results of structural interpretation of metabolic stability was confirmed by chemists.