The efficient separation of chiral molecules is a fundamental challenge in the manufacture of pharmaceuticals and light-polarising materials. We developed an approach that combines machine learning with a physics-based representation to predict resolving agents for chiral molecules, using a transformer-based neural network. In retrospective tests, our approach reaches a four to six-fold improvement over the historical - trial and error based - hit rate. We further validate the model in a prospective experiment, where we use the model to design a resolution screen for six unseen racemates. We successfully resolved three of the six mixtures in a single round of experiments and obtained an overall 8-to-1 true positive to false negative ratio. Together with this study, we release a previously proprietary dataset of over 6000 resolution experiments, the largest diastereomeric salt crystallisation dataset to date. More broadly, our approach and open crystallisation data lay the foundation for accelerating and reducing the costs of chiral resolutions.
High-throughput experimentation (HTE) has the potential to improve our understanding of organic chemistry by systematically interrogating reactivity across diverse chemical spaces. Notable bottlenecks include few publicly available large-scale datasets and the need for facile interpretation of these data’s hidden chemical insights. Here we report the development of a high-throughput experimentation analyser, a robust and statistically rigorous framework, which is applicable to any HTE dataset regardless of size, scope or target reaction outcome, which yields interpretable correlations between starting material(s), reagents and outcomes. We improve the HTE data landscape with the disclosure of 39,000+ previously proprietary HTE reactions that cover a breadth of chemistry, including cross-coupling reactions and chiral salt resolutions. The high-throughput experimentation analyser was validated on cross-coupling and hydrogenation datasets, showcasing the elucidation of statistically significant hidden relationships between reaction components and outcomes, as well as highlighting areas of dataset bias and the specific reaction spaces that necessitate further investigation.
Biomolecular simulations have become an essential tool in contemporary drug discovery and molecular mechanics force fields constitute its cornerstone. Developing a high quality and broad coverage general force field is a significant undertaking that requires substantial expert knowledge and computing resources, which is beyond the scope of general practitioners. Existing force fields originate from only a limited number of groups and organizations, and they either suffer from limited numbers of training sets, lower than desired quality because of oversimplified representations, or are costly for the molecular modeling community to access. To address these issues, in this work, we developed an AMBER-consistent small molecule force field with extensive chemical space coverage to be open-sourced to the entire modeling community. To validate our force field, we carried out benchmarks of quantum mechanics/molecular mechanics conformer comparison and free energy perturbation calculations on several benchmark data sets. Our force field achieves a higher level performance at reproducing quantum mechanics energies and geometries than two popular open-source force fields, OpenFF2 and GAFF2. In relative binding free energy calculations for 31 protein-ligand data sets, comprising 1079 pairs of ligands, the new force field achieves an overall accuracy of 1.19 kcal/mol for ∆∆G and 0.95 kcal/mol for ∆G on a subset of 463 ligands without bespoke fitting to the data sets. The results are on par with the leading commercial series of OPLS force field.
Predicting binding free energy of ligand-protein complexes has been a grand challenge in the field of computational chemistry since the early days of molecular modeling. Multiple computational methodologies exist to predict ligand binding affinities. Pathway-based Free Energy Perturbation (FEP), Thermodynamic Integration (TI) as well as Linear Interaction Energy (LIE), and Molecular Mechanics-Poisson Boltzmann/Generalized Born Surface Area (MM-PBSA/GBSA) have been applied to a variety of biologically relevant problems and achieved different levels of predictive accuracy. Recent advancements in computer hardware and simulation algorithms of molecular dynamics and Monte Carlo sampling, as well as improved general force field parameters, have made FEP a principal approach for calculating the free energy differences, especially when calculating the host-guest binding affinity differences upon chemical modification.Since the FEP-calculated binding free energy difference, denoted ddGFEP only characterizes the difference in free energy between pairs of ligands or complexes, not the absolute binding free energy value of each individual host-guest system, denoted dG, we examine two rarely asked questions in FEP application:1) Which values would be more appropriate as the prediction to assess the ligands prospectively: the calculated pairwise free energy differences, ddGFEP, or the estimated absolute binding energies, d^G, transformed from ddGFEP?2) In the situation where only a limited number of ligand pairs can be calculated in FEP, can the perturbation pairs be optimally selected with respect to the reference ligand(s) to maximize the prediction precision?These two questions underline the viability of an often-neglected assumption in pairwise comparisons: that the pairwise value is sufficient to make a quantitative and reliable characterization of an individual ligand's properties or activities. This implicit assumption would be true if there was no error in each pairwise calculation. Recently pair designs such as multiple pathways or cycle closure analyses provided calculation error estimation but did not address the statistical impact of the two questions above. The error impact is fully minimized by conducting an exhaustive study that obtains all NC2 = N(N-1)/2 pairs for a set N molecules; more if there is directionality (dGi,j != dGj,i). Obviously, that study design is impractical and unnecessary. Thus, we desire to collect the right amount of data that is 1) feasibly attainable, 2) topologically sufficient, and 3) mathematically synthesizable so that we can mitigate inherent calculation errors and have higher confidence in our conclusions.The significance of above questions can be illustrated by a motivating example shown in Figure 1 and Table 1, which considers two different perturbation graph designs for 20 ligands with the same number of FEP perturbation pairs, 19, and the same reference, Ligand 1. These two designs reached different conclusions in rank ordering ligand potencies due to errors inherent in the FEP derived estimates. Based on design A, ligands 5, 7, 14, 15 would be selected as the best four (20%) picks since those d^G estimates are the most favorable. Design B would yield ligands 5, 12, 18, 19 as best for the same reason. Without knowing the true value, dGTrue of the other 19 ligands, we lack a prospective metric to assess which design could be more precise even though, retrospectively, we know that both designs had reasonably good agreement with the true values, as measured through correlation and error metrics. However, the top picks from neither design were consistent with the true top four ligands, which are ligands 7, 10, 12, 18. Yet, if all of the 20C2 =190 pairs could have been calculated as listed in the last column of Table 1, the best four ligands would have been correctly identified. Additionally, the other metrics included in Table 1 were significantly improved. However, as mentioned above, calculating all possible pairs, or even a significant fraction of all possible pairs, is unlikely in practice, especially when number of molecules are large. Given this restriction, is it possible to objectively determine whether design A or B will give more precise predictions?In this report, we investigated the performance of the calculated ddGFEP values compared to the pairwise differences in least squares derived d^G estimates both analytically and through simulations. Based on our findings, we recommend applying weighted least squares to transforming ddGFEP values into d^G estimates. Second, we investigated the factors that contribute to the precision of the d^G estimates, such as the total number of computed pairs, the selection of computed pairs, and the uncertainty in the computed ddGFEP values. The mean squared error, denoted MSE and Spearman's rank correlation, are used as performance metrics.To illustrate, we demonstrated how the structural similarity can be included in design and its potential impact on prediction precision. As in the majority of reported FEP studies on binding affinity prediction, the ddGFEP pairs were selected based on chemical structure similarity. Pairs with small chemical differences are assumed to be more likely to have smaller errors in ddGFEP calculation. Together using the constructed mathematic system and literature examples, we demonstrate that some of pair-selection schemes (designs) are better than the others. To minimize the prediction uncertainty, it is recommended to wisely select design optimality criterion to suitpractical applications accordingly.
Artificial intelligence and machine learning have demonstrated their potential role in predictive chemistry and synthetic planning of small molecules; there are at least a few reports of companies employing in silico synthetic planning into their overall approach to accessing target molecules. A data-driven synthesis planning program is one component being developed and evaluated by the Machine Learning for Pharmaceutical Discovery and Synthesis (MLPDS) consortium, comprising MIT and 13 chemical and pharmaceutical company members. Together, we wrote this perspective to share how we think predictive models can be integrated into medicinal chemistry synthesis workflows, how they are currently used within MLPDS member companies, and the outlook for this field.
Predicting ligand biological activity is a key challenge in drug discovery. Ligand-based statistical approaches are often hampered by noise due to undersampling: The number of molecules known to be active or inactive is vastly less than the number of possible chemical features that might determine binding. We derive a statistical framework inspired by random matrix theory and combine the framework with high-quality negative data to discover important chemical differences between active and inactive molecules by disentangling undersampling noise. Our model outperforms standard benchmarks when tested against a set of challenging retrospective tests. We prospectively apply our model to the human muscarinic acetylcholine receptor M1, finding four experimentally confirmed agonists that are chemically dissimilar to all known ligands. The hit rate of our model is significantly higher than the state of the art. Our model can be interpreted and visualized to offer chemical insights about the molecular motifs that are synergistic or antagonistic to M1 agonism, which we have prospectively experimentally verified.
A correction to this article has been published and is linked from the HTML and PDF versions of this paper. The error has not been fixed in the paper.
It is of great challenge to predict human brain penetration for substrates of multidrug resistance protein 1 (MDR1) and breast cancer resistance protein (BCRP), 2 major efflux transporters at blood-brain barrier. Thus, a physiologically based pharmacokinetic (PBPK) model with the incorporation of in vitro MDR1 and BCRP transporter function data and transporter protein expression levels has been developed. As such, it is crucial to generate MDR1 and BCRP substrate data with a high fidelity. In this study, 2 widely used human MDR1 cell lines from Borst and National Institutes of Health laboratories were evaluated using rodent brain penetration data, and the study suggested that the MDR1 expressed in Madin-Darby canine kidney (MDCK) cell line from National Institutes of Health laboratory predicted brain penetration better, particularly for compounds with a high passive permeability. In addition, human BCRP-MDCK cell line with 1 μM PSC833, a specific MDR1 inhibitor, demonstrated the ability to identify BCRP substrates without the confounding of endogenous canine Mdr1. Comparison of human BCRP and mouse Bcrp transporter functions revealed that the functional differences of BCRP between the 2 species is minimal. The incorporation of both the validated MDR1 and BCRP assays into our brain PBPK model has significantly improved the prediction for the brain penetration of MDR1 and BCRP substrates across species.
Predicting how a complex molecule reacts with different reagents, and how to synthesise complex molecules from simpler starting materials, are fundamental to organic chemistry. We show that an attention-based machine translation model - Molecular Transformer - tackles both reaction prediction and retrosynthesis by learning from the same dataset. Reagents, reactants and products are represented as SMILES text strings. For reaction prediction, the model "translates" the SMILES of reactants and reagents to product SMILES, and the converse for retrosynthesis. Moreover, a model trained on publicly available data is able to make accurate predictions on proprietary molecules extracted from pharma electronic lab notebooks, demonstrating generalisability across chemical space. We expect our versatile framework to be broadly applicable to problems such as reaction condition prediction, reagent prediction and yield prediction.
A major challenge in the development of β-site amyloid precursor protein cleaving enzyme 1 (BACE1) inhibitors for the treatment of Alzheimer's disease is the alignment of potency, drug-like properties, and selectivity over related aspartyl proteases such as Cathepsin D (CatD) and BACE2. The potential liabilities of inhibiting BACE2 chronically have only recently begun to emerge as BACE2 impacts the processing of the premelanosome protein (PMEL17) and disrupts melanosome morphology resulting in a depigmentation phenotype. Herein, we describe the identification of clinical candidate PF-06751979 (64), which displays excellent brain penetration, potent in vivo efficacy, and broad selectivity over related aspartyl proteases including BACE2. Chronic dosing of 64 for up to 9 months in dog did not reveal any observation of hair coat color (pigmentation) changes and suggests a key differentiator over current BACE1 inhibitors that are nonselective against BACE2 in later stage clinical development.
The recent increase in the number of X-ray crystal structures of G-protein coupled receptors (GPCRs) has been enabling for structure-based drug design (SBDD) efforts. These structures have revealed that GPCRs are highly dynamic macromolecules whose function is dependent on their intrinsic flexibility. Unfortunately, the use of static structures to understand ligand binding can potentially be misleading, especially in systems with an inherently high degree of conformational flexibility. Here, we show that docking a set of dopamine D3 receptor compounds into the existing eticlopride-bound dopamine D3 receptor (D3R) X-ray crystal structure resulted in poses that were not consistent with results obtained from site-directed mutagenesis experiments. We overcame the limitations of static docking by using large-scale high-throughput molecular dynamics (MD) simulations and Markov state models (MSMs) to determine an alternative pose consistent with the mutation data. The new pose maintains critical interactions observed in the D3R/eticlopride X-ray crystal structure and suggests that a cryptic pocket forms due to the shift of a highly conserved residue, F6.52. Our study highlights the importance of GPCR dynamics to understand ligand binding and provides new opportunities for drug discovery.
We previously observed a cutaneous type IV immune response in nonhuman primates (NHP) with the mGlu5 negative allosteric modulator (NAM) 7. To determine if this adverse event was chemotype- or mechanism-based, we evaluated a distinct series of mGlu5 NAMs. Increasing the sp3 character of high-throughput screening hit 40 afforded a novel morpholinopyrimidone mGlu5 NAM series. Its prototype, (R)-6-neopentyl-2-(pyridin-2-ylmethoxy)-6,7-dihydropyrimido[2,1-c][1,4]oxazin-4(9H)-one (PF-06462894, 8), possessed favorable properties and a predicted low clinical dose (2 mg twice daily). Compound 8 did not show any evidence of immune activation in a mouse drug allergy model. Additionally, plasma samples from toxicology studies confirmed that 8 did not form any reactive metabolites. However, 8 caused the identical microscopic skin lesions in NHPs found with 7, albeit with lower severity. Holistically, this work supports the hypothesis that this unique toxicity may be mechanism-based although additional work is required to confirm this and determine clinical relevance.
An entry from the Cambridge Structural Database, the world’s repository for small molecule crystal structures. The entry contains experimental data from a crystal diffraction study. The deposited dataset for this entry is freely available from the CCDC and typically includes 3D coordinates, cell parameters, space group, experimental conditions and quality measures.
Macrocycles pose challenges for computer-aided drug design due to their conformational complexity. One fundamental challenge is identifying all low-energy conformations of the macrocyclic ring, which is important for modeling target binding, passive membrane permeation, and other conformation-dependent properties. Macrocyclic polyketides are medically and biologically important natural products characterized by structural and functional diversity. Advances in synthetic biology and semisynthetic methods may enable creation of an even more diverse set of non-natural product polyketides for drug discovery and other applications. However, the conformational sampling of these flexible compounds remains demanding. We developed and optimized a dihedral angle-based macrocycle conformational sampling method for macrocycles of arbitrary structure, and here we apply it to diverse polyketide natural products. First, we evaluated its performance using a data set of 37 polyketides with available crystal structures, with 9-22 rotatable bonds in the macrocyclic ring. Our optimized protocol was able to reproduce the crystal structure of polyketides' aglycone backbone within 0.50 angstrom RMSD for 31 out of 37 polyketides. Consistent with prior structural studies, our analysis suggests that polyketides tend to have multiple distinct low-energy structures, including the bioactive (target-bound) conformation as well as others of unknown significance. For this reason, we also introduce a strategy to improve both efficiency and accuracy of the conformational search by utilizing torsional restraints derived from NMR vicinal proton couplings to restrict the conformational search. Finally, as a first application of the method, we made blinded predictions of the passive membrane permeability of a diverse set of polyketides, based on their predicted structures in low- and high-dielectric media.
Inhibition of β-secretase BACE1 is considered one of the most promising approaches for treating Alzheimer's disease. Several structurally distinct BACE1 inhibitors have been withdrawn from development after inducing ocular toxicity in animal models, but the target mediating this toxicity has not been identified. Here we use a clickable photoaffinity probe to identify cathepsin D (CatD) as a principal off-target of BACE1 inhibitors in human cells. We find that several BACE1 inhibitors blocked CatD activity in cells with much greater potency than that displayed in cell-free assays with purified protein. Through a series of exploratory toxicology studies, we show that quantifying CatD target engagement in cells with the probe is predictive of ocular toxicity in vivo. Taken together, our findings designate off-target inhibition of CatD as a principal driver of ocular toxicity for BACE1 inhibitors and more generally underscore the power of chemical proteomics for discerning mechanisms of drug action.