Nucleophilicity and electrophilicity are important properties for evaluating the reactivity and selectivity of chemical reactions. It allows the ranking of nucleophiles and electrophiles on reactivity scales, enabling a better understanding and prediction of reaction outcomes. Building upon our recent work (Digit. Discov., 2024, 3, 347-354), we introduce an atom-based machine learning (ML) approach for predicting methyl cation affinities (MCAs) and methyl anion affinities (MAAs) to estimate nucleophilicity and electrophilicity, respectively. The ML models are trained and validated on QM-derived data from around 50,000 neutral drug-like molecules, achieving Pearson correlation coefficients of 0.97 for MCA and 0.95 for MAA on the held-out test sets. In addition, we demonstrate the ML approach on two different applications: first, as a general tool for filtering retrosynthetic routes based on chemical selectivity predictions, and second, as a tool for assessing the chemical stability of esters and carbamates towards hydrolysis reactions. The code is freely available on GitHub under the MIT open source license and as a web application at www.esnuel.org.
Reactivity scales such as nucleophilicity and electrophilicity are valuable tools for de- termining chemical reactivity and selectivity. However, prior attempts to predict or calculate nucleophilicity and electrophilicity are either not capable of generalizing well to unseen molecular structures or require substantial computing resources. We present a fully automated quantum chemistry (QM)-based workflow that automatically identi- fies nucleophilic and electrophilic sites and computes methyl cation affinities and methyl anion affinities to quantify nucleophilicity and electrophilicity, respectively. The calcu- lations are based on r2SCAN-3c SMD(DMSO) single-point calculations on GFN1-xTB ALPB(DMSO) geometries that, in turn, derive from a GFNFF-xTB ALPB(DMSO) conformational search. The workflow is validated against both experimental and higher- level QM-derived data resulting in very strong correlations while having a median wall time of less than two minutes per molecule. Additionally, we demonstrate the workflow on two different applications: first, as a general tool for filtering retrosynthetic routes based on chemical selectivity predictions, and second, as a tool for determining the relative reactivity of covalent inhibitors. The code is freely available on GitHub under the MIT open source license and as a web application at www.esnuel.org.
An important aspect in the development of small molecules as drugs or agro-chemicals is their systemic availability after intravenous and oral administration. The prediction of the systemic availability from the chemical structure of a potential candidate is highly desirable, as it allows to focus the drug or agrochemical development on compounds with a favorable kinetic profile. However, such pre-dictions are challenging as the availability is the result of the complex interplay between molecular properties, biology and physiology and training data is rare. In this work we improve the hybrid model developed earlier [1]. We reduce the median fold change error for the total oral exposure from 2.85 to 2.35 and for intravenous administration from 1.95 to 1.62. This is achieved by training on a larger data set, improving the neural network architecture as well as the parametrization of mechanistic model. Further, we extend our approach to predict additional endpoints and to handle different covariates, like sex and dosage form. In contrast to a pure machine learning model, our model is able to predict new end points on which it has not been trained. We demonstrate this feature by predicting the exposure over the first 24h, while the model has only been trained on the total exposure.
N-nitrosamines and nitrosamine drug substance related impurities (NDSRIs) became a critical topic for the development and safety of small molecule medicines following the withdrawal of various pharmaceutical products from the market. To assess the mutagenic and carcinogenic potential of different N-nitrosamines lacking robust carcinogenicity data, several approaches are in use including the published carcinogenic potency categorization approach (CPCA), the Enhanced Ames Test (EAT), in vivo mutagenicity studies as well as read-across to analogue molecules with robust carcinogenicity data. We employ quantum chemical calculations as a pivotal tool providing insights into the likelihood of reactive ion formation and subsequent DNA alkylation for a selection of molecules including e.g., carcinogenic N-nitrosopiperazine (NPZ), N-nitrosopiperidine (NPIP), together with N-nitrosodimethylamine (NDMA) as well as non-carcinogenic N-nitrosomethyl-tert-butylamine (NTBA) and bis (butan-2-yl) (nitros)amine (BBNA). In addition, a series of nitroso-methylaminopyridines is compared side-by-side. We draw comparisons between calculated reaction profiles for structures representing motifs common to NDSRIs and those of confirmed carcinogenic and non-carcinogenic molecules with in vivo data from cancer bioassays. Furthermore, our approach enables insights into reactivity and relative stability of intermediate species that can be formed upon activation of several nitrosamines. Most notably, we reveal consistent differences between the free energy profiles of carcinogenic and non-carcinogenic molecules. For the former, the intermediate diazonium ions mostly react, kinetically controlled, to the more stable DNA adducts and less to the water adducts via transition-states of similar heights. Non-carcinogenic molecules yield stable carbocations as intermediates that, thermodynamically controlled, more likely form the statistically preferred water adducts. In conclusion, our data confirm that quantum chemical calculations can contribute to a weight of evidence approach for the risk assessment of nitrosamines.
Pharmacokinetics (PK) is the result of a complex interplay between compound properties and physiology, and a detailed characterization of a molecule's PK during preclinical research is key to understanding the relationship between applied dose, exposure, and pharmacological effect. Predictions of human PK based on the chemical structure of a compound are highly desirable to avoid advancing compounds with unfavorable properties early on and to reduce animal testing, but data to train such models are scarce. To address this problem, we combine well-established Physiologically Based Pharmacokinetic models with Deep Learning models for molecular property prediction into a hybrid model to predict PK parameters for small molecules directly from chemical structure. Our model predicts exposure after oral and intravenous administration with fold change errors of 1.87 and 1.86, respectively, in healthy subjects and 2.32 and 2.23, respectively, in patients with various diseases. Unlike pure Deep Learning models, the hybrid model can predict endpoints on which it was not trained. We validate this extrapolation capability by predicting full concentration-time profiles for compounds with published PK data. Our model enables early selection and prioritization of the most promising drug candidates, which can lead to a reduction in animal testing during drug discovery and development.
Relative solubilities, i.e. whether a given molecule is more soluble in one solvent compared to others, is a critical parameter for pharmaceutical and agricultural formulation development and chemical synthesis, material science, and environmental chemistry. In silico predictions of this crucial variable can help reducing experiments, waste of solvents and synthesis optimization. In this study, we evaluate the performance of different physics-based methods for predicting relative solubilities. Our assessment involves quantum mechanics-based COSMO-RS and molecular dynamics-based free energy methods using OPLS4, the open-source openFF Sage, and GAFF force fields, spanning over 200 solvent-solute combinations. Our investigation highlights the important role of compound multimerization, an effect which must be accounted for to obtain accurate relative solubility predictions. The performance landscape of these methods is varied, with significant differences in precision depending on both the method used and the solute considered, thereby offering an improved understanding of the predictive power of physics-based methods in chemical research.
Aqueous solubility is the most important physicochemical property for agrochemical and drug candidates and a prerequisite for uptake, distribution, transport, and finally the bioavailability in living species. We here present the first-ever direct machine learning models for pH-dependent solubility in water. For this, we combined almost 300000 data points from 11 solubility assays performed over 24 years and over one million data points from lipophilicity and melting point experiments. Data were split into three pH-classes − acidic, neutral and basic − , representing the conditions of stomach and intestinal tract for animals and humans, and phloem and xylem for plants. We find that multi-task neural networks using ECFP-6 fingerprints outperform baseline random forests and single-task neural networks on the individual tasks. Our final model with three solubility tasks using the pH-class combined data from different assays and five helper tasks results in root mean square errors of 0.56 log units overall (acidic 0.61; neutral 0.52; basic 0.54) and Spearman rank correlations of 0.83 (acidic 0.78; neutral 0.86; basic 0.86), making it a valuable tool for profiling of compounds in pharmaceutical and agrochemical research. The model allows for the prediction of compound pH profiles with mean and median RMSE per molecule of 0.62 and 0.56 log units.
Approaches for predicting proteolysis targeting chimera (PROTAC) cell permeability are of major interest to reduce resource-demanding synthesis and testing of low-permeable PROTACs. We report a comprehensive investigation of the scope and limitations of machine learning-based binary classification models developed using 17 simple descriptors for large and structurally diverse sets of cereblon (CRBN) and von Hippel-Lindau (VHL) PROTACs. For the VHL PROTAC set, kappa nearest neighbor and random forest models performed best and predicted the permeability of a blinded test set with >80% accuracy (k ≥ 0.57). Models retrained by combining the original training and the blinded test set performed equally well for a second blinded VHL set. However, models for CRBN PROTACs were less successful, mainly due to the imbalanced nature of the CRBN datasets. All descriptors contributed to the models, but size and lipophilicity were the most important. We conclude that properly trained machine learning models can be integrated as effective filters in the PROTAC design process.
Federated multipartner machine learning has been touted as an appealing and efficient method to increase the effective training data volume and thereby the predictivity of models, particularly when the generation of training data is resource-intensive. In the landmark MELLODDY project, indeed, each of ten pharmaceutical companies realized aggregated improvements on its own classification or regression models through federated learning. To this end, they leveraged a novel implementation extending multitask learning across partners, on a platform audited for privacy and security. The experiments involved an unprecedented cross-pharma data set of 2.6+ billion confidential experimental activity data points, documenting 21+ million physical small molecules and 40+ thousand assays in on-target and secondary pharmacodynamics and pharmacokinetics. Appropriate complementary metrics were developed to evaluate the predictive performance in the federated setting. In addition to predictive performance increases in labeled space, the results point toward an extended applicability domain in federated learning. Increases in collective training data volume, including by means of auxiliary data resulting from single concentration high-throughput and imaging assays, continued to boost predictive performance, albeit with a saturating return. Markedly higher improvements were observed for the pharmacokinetics and safety panel assay-based task subsets.
In this study, we use machine learning algorithms with QM-derived COSMO-RS descriptors, along with Morgan fingerprints, to predict the absolute solubility of drug-like compounds. The QM-derived descriptors account for the molecular properties of the solute, i.e., the solute–solute interactions in an artificial-liquid-state (super-cooled liquid), and the solute–solvent interactions in solution. We employ two main approaches to predict solubility: (i) a hypothetical pathway that involves melting the solute at room temperature T = T¯ ( Δ_fusG_A^⊖ ) and mixing the artificially liquid solute into the solvent ( Δ_mG_(A:B)^⊖ ). In this approach Δ_fusG_A^⊖ is predicted using machine learning models, and the Δ_mG_(A:B)^⊖ is obtained from COSMO-RS calculations; (ii) direct solubility prediction using machine learning algorithms. The models were trained on a large number of Bayer in-house compounds for which water solubility data is available at physiological pH of 6.5 and ambient temperature. We also evaluated our models using external datasets from a solubility challenge. Our models present great improvements compared to the absolute solubility prediction with the QSAR model for the artificial liquid state as implemented in the COSMO therm software, for both in-house and external datasets. We are furthermore able to demonstrate the superiority of QM-derived descriptors compared to cheminformatics descriptors. We finally present low-cost alternative models using fragment-based COSMO quick calculations with only marginal reduction in the quality of predicted solubility.
We present a quantum chemistry (QM)-based method that computes the relative energies of intermediates in the Heck reaction that relate to the regioselective reaction outcome: branched (α), linear (β), or a mix of the two. The calculations are done for two different reaction pathways (neutral and cationic) and are based on r 2SCAN-3c single-point calculations on GFN2-xTB geometries that, in turn, derive from a GFNFF-xTB conformational search. The method is completely automated and is sufficiently efficient to allow for the calculation of thousands of reaction outcomes. The method can mostly reproduce systematic experimental studies where the ratios of regioisomers are carefully determined. For a larger dataset extracted from Reaxys, the results are somewhat worse with accuracies of 63% for β-selectivity using the neutral pathway and 29% for α-selectivity using the cationic pathway. Our analysis of the dataset suggests that only the major or desired regioisomer is reported in the literature in many cases, which makes accurate comparisons difficult. The code is freely available on GitHub under the MIT open-source license: https://github.com/jensengroup/HeckQM.
The well-known concept of quantitative structure-activity relationships (QSAR) has been gaining significant interest in the recent years. Data, descriptors, and algorithms are the main pillars to build useful models that support more efficient drug discovery processes with in silico methods. Significant advances in all three areas are the reason for the regained interest in these models. In this book chapter we review various machine learning (ML) approaches that make use of measured in vitro/in vivo data of many compounds. We put these in context with other digital drug discovery methods and present some application examples.
We present RegioML, an atom-based machine learning model for predicting the regioselectivities of electrophilic aromatic substitution reactions. The model relies on CM5 atomic charges computed using semiempirical tight binding (GFN1-xTB) combined with the ensemble decision tree variant light gradient boosting machine (LightGBM). The model is trained and tested on 21,201 bromination reactions with 101K reaction centers, which is split into a training, test, and out-of-sample datasets with 58K, 15K, and 27K reaction centers, respectively. The accuracy is 93% for the test set and 90% for the out-of-sample set, while the precision (the percentage of positive predictions that are correct) is 88% and 80%, respectively. The test-set performance is very similar to the graph-based WLN method developed by Struble et al. (React. Chem. Eng. 2020, 5, 896) though the comparison is complicated by the possibility that some of the test and out-of-sample molecules are used to train WLN. RegioML out-performs our physics-based RegioSQM20 method (J. Cheminform. 2021, 13:10) where the precision is only 75%. Even for the out-of-sample dataset, RegioML slightly outperforms RegioSQM20. The good performance of RegioML and WLN is in large part due to the large datasets available for this type of reaction. However, for reactions where there is little experimental data, physics-based approaches like RegioSQM20 can be used to generate synthetic data for model training. We demonstrate this by showing that the performance of RegioSQM20 can be reproduced by a ML-model trained on RegioSQM20-generated data.
Accurate calculation of relative tautomer energies in different environments is a prerequisite to many parameters of relevance in drug discovery. This work provides a thorough benchmark of the semiempirical methods AM1, PM3 and GFN2-xTB, the force-field OPLS4, Hartree–Fock and HF-3c, the density functionals PBEh-3c, B97-3c, r2SCAN-3c, PBE, PBE0, TPSS, r2SCAN, ω-B97X-V, M06-2X, B3LYP, B2PLYP, and second-order perturbation theory MP2 versus the gold-standard coupled-cluster DLPNO-CCSD(T) using the def2-QZVPP basis set. The outperforming method identified is M06-2X, whereas r2SCAN-3c is the best-perfoming one in the set of cost-optimized methods. Application of the two methods on a challenging subset from the SAMPL2 challenge provides evidence that deviations from experiment are caused by deficiencies of current continuum solvation methods.
We present RegioSQM20, a new version of RegioSQM (Chem. Sci. 2018, 9, 660),which predicts the regioselectivities of electrophilic aromatic substitution (EAS) re-actions from the calculation of proton affinities. The following improvements havebeen made: The open source semiempirical tight binding program xtb is used insteadof the closed source MOPAC program. Any low energy tautomeric forms of the inputmolecule are identified and regioselectivity predictions are made for each form. Finally,RegioSQM20 offers a qualitative prediction of the reactivity of each tautomer (low,medium, or high) based on the reaction center with the highest proton affinity. Theinclusion of tautomers increases the success rate from 90.7% to 92.7%. RegioSQM20is compared to two machine learning based models: one developed by Struble et al.(React. Chem. Eng. 2020, 5, 896) specifically for regioselectivity predictions of EASreactions (WLN) and a more generally applicable reactivity predictor (IBM RXN) de-veloped by Schwaller et al. (ACS Cent. Sci. 2019, 5, 1572). RegioSQM20 and WLNoffers roughly the same success rates for the entire data sets (without considering tau-tomers), while WLN is many orders of magnitude faster. The accuracy of the moregeneral IBM RXN approach is somewhat lower: 76.3%-85.0%, depending on the dataset. The code is freely available under the MIT open source license and will be madeavailable as a webservice (regiosqm.org) in the near future.
In this study we compare the three algorithms for the generation of conformer ensembles Biovia BEST, Schrödinger Prime macrocycle sampling (PMM) and Conformator (CONF) form the University of Hamburg, with ensembles derived for exhaustive molecular dynamics simulations applied to a dataset of 7 small macrocycles in two charge states and three solvents. Ensemble completeness is a prerequisite to allow for the selection of relevant diverse conformers for many applications in computational chemistry. We apply conformation maps using principal component analysis based on ring torsions. Our major finding critical for all applications of conformer ensembles in any computational study is that maps derived from MD with explicit solvent are significantly distinct between macrocycles, charge states and solvents, whereas the maps for post-optimized conformers using implicit solvent models from all generator algorithms are very similar independent of the solvent. We apply three metrics for the quantification of the relative covered ensemble space, namely cluster overlap, variance statistics, and a novel metric, Mahalanobis distance, showing that post-optimized MD ensembles cover a significantly larger conformational space than the generator ensembles, with the ranking PMM > BEST >> CONF. Furthermore, we find that the distributions of 3D polar surface areas are very similar for all macrocycles independent of charge state and solvent, except for the smaller and more strained compound 7, and that there is also no obvious correlation between 3D PSA and intramolecular hydrogen bond count distributions.
Over the past two decades, an in silico absorption, distribution, metabolism, and excretion (ADMET) platform has been created at Bayer Pharma with the goal to generate models for a variety of pharmacokinetic and physicochemical endpoints in early drug discovery. These tools are accessible to all scientists within the company and can be a useful in assisting with the selection and design of novel leads, as well as the process of lead optimization. Here. we discuss the development of machine-learning (ML) approaches with special emphasis on data, descriptors, and algorithms. We show that high company internal data quality and tailored descriptors, as well as a thorough understanding of the experimental endpoints, are essential to the utility of our models. We discuss the recent impact of deep neural networks and show selected application examples.
We herein report the first thorough analysis of the structure-permeability relationship of semipeptidic macrocycles. In total, 47 macrocycles were synthesized using a hybrid solid-phase/solution strategy, and then their passive and cellular permeability was assessed using the parallel artificial membrane permeability assay (PAMPA) and Caco-2 assay, respectively. The results indicate that semipeptidic macrocycles generally possess high passive permeability based on the PAMPA, yet their cellular permeability is governed by efflux, as reported in the Caco-2 assay. Structural variations led to tractable structure-permeability and structure-efflux relationships, wherein the linker length, stereoinversion, N-methylation, and peptoids site-specifically impact the permeability and efflux. Extensive nuclear magnetic resonance, molecular dynamics, and ensemble-based three-dimensional polar surface area (3D-PSA) studies showed that ensemble-based 3D-PSA is a good predictor of passive permeability.
Selective progesterone receptor modulators are promising therapeutic options for the treatment of uterine fibroids. Vilaprisan, a new chemical entity that was discovered at Bayer is currently in clinical development. In this study we provide a combined experimental and quantum chemical approach providing the data that allowed to present hydroxyestradienone as an acceptable starting material for drug substance synthesis. Hydroxyestradienone has four stereogenic centers leading to 8 diastereomers and 16 enantiomers of which only six diastereomers were synthetically accessible but two not. A computational multistep protocol resulting in density functional P2PLYP-D3(BJ)/dev2-TZVPP Gibbs free energies and SMD solvation free energies led to a clear separation between the existing and the synthetically not accessible enantiomers, whereas multiple geometry-based and cheminformatic descriptors were not able to explain experimental findings.