ABSTRACT Predicting pharmacokinetic (PK) profiles from molecular structures represents a significant advancement in pharmaceutical research with substantial implications for expediting drug discovery processes. We evaluated five approaches to systematically compare five distinct methodological frameworks for predicting rat plasma concentration‐time profiles directly from molecular structures using a consistent dataset and evaluation framework: (1) NCA‐ML, predicted non‐compartmental analysis parameters with one compartmental PK modeling; (2) PBPK‐ML, utilizing ML‐predicted in vitro characteristics in physiologically based PK models; (3) CMT‐ML, neural networks predicting compartmental PK model parameters with two or three compartmental PK modeling; (4) CMT‐PINN, employing physics‐informed neural networks trained on concentration‐time profiles predicting compartmental PK model parameters with two or three compartmental PK modeling; and (5) PURE‐ML, using decision trees to predict concentration values at specific time points. The CMT‐PINN approach achieved highest predictive performance closely followed by PURE‐ML (R2‐log: 0.854 vs. 0.789, Spearman: 0.933 vs. 0.896), with 65.9% versus 61.0% of predictions within a twofold error of the observed concentrations. The other three approaches showed substantially lower performance metrices and higher prediction error margins. Models trained directly on concentration‐time profiles outperformed those trained using derived PK parameters, particularly with limited training datasets. Our findings confirm the viability of predicting PK behavior from molecular structures prior to synthesis. The implementation of these computational approaches enables informed compound selection early in discovery, concentrating resources on promising candidates, and potentially reducing animal studies while accelerating development timelines.
Factor VIIa (FVIIa) catalyzes the first step of the blood coagulation cascade. The expected wide therapeutic window between antithrombotic efficacy and bleeding risk makes FVIIa an attractive drug target. However, no FVIIa inhibitors have reached the market so far, mostly due to poor oral bioavailability. To date, in all ligand-bound X-ray crystal structures of FVIIa, the binding pocket of FVIIa is in an active, open form. Here, we present an X-ray crystal structure of the FVIIa-tissue factor complex with a bound oxazole-based inhibitor at 1.9 Å resolution, with an extensively remodeled active site and a collapsed S1 pocket. Using collectively 0.17 ms of atomistic molecular dynamics simulations, we observed conformational transitions between the collapsed and open forms of the S1 pockets of FVIIa and 12 other serine peptidases out of 16 studied, indicating an equilibrium of open and collapsed states of the S1 pocket in FVIIa and the majority of the serine proteases studied. Therefore, our results point to a general S1 pocket plasticity, which provides the basis for a completely new way of inhibiting FVIIa and other serine proteases.
To increase the chemical space around the well-known GalNAc-ligand as ASGPR-binder, a high-throughput screening campaign was performed, testing approximately 550,000 compounds. After evaluation of the potential screening hits, only one compound, which showed high similarity with guanosine nucleosides, was chosen for further profiling. Crystal structure analysis revealed the coordination of the Ca2+-ion within the ASGPR-binding site by the cis-diol motif of the ribose unit as well as an additional π-π-interaction of the purine heterocycle to tryptophan-243. Based on these findings, guanosine was attached via the 5'-OH group to a recently described morpholino-based nucleotide using two different linker units. The resulting morpholino-guanosine building blocks were conjugated to the 5'-end of a literature-known transthyretin targeting small interfering RNA (siRNA), leading to trivalent siRNA-guanosine conjugates, which were tested for their TTR knockdown and exhibited similar potencies as the analogous GalNAc-conjugates in vitro and in vivo.
Reliable methods to quantify the predictive uncertainty of machine learning (ML) models can significantly increase the impact of molecular property prediction and are routinely used in applications like active learning and ML-guided property optimization. Poor predictive accuracy of ML models is often related to (i) regions of the chemical space, which are characterized by large property differences for structurally similar molecules, and (ii) a lack of representation of test molecules in the training data. Here, we analyze the relationship between these error sources and the predictive uncertainty of popular uncertainty quantification (UQ) methods on molecular activity data sets. We find that several UQ methods struggle to identify poorly predicted compounds in regions of steep structure-activity relationships (SAR). We also demonstrate that the evaluation scenario, as defined by data splitting into training and test sets, significantly impacts observed UQ performance. Based on our findings we introduce a simple but strong and very robust method for UQ that offers significant improvements over previous approaches in several evaluation scenarios and demonstrate its usefulness in an exploratory active learning setting.
Large Language Models (LLMs) have demonstrated great performance in few-shot In-Context Learning (ICL) for a variety of generative and discriminative chemical design tasks. The newly expanded context windows of LLMs can further improve ICL capabilities for molecular inverse design and lead optimization. To take full advantage of these capabilities we developed a new semi-supervised learning method that overcomes the lack of experimental data available for many-shot ICL. Our approach involves iterative inclusion of LLM generated molecules with high predicted performance, along with experimental data. We further integrated our method in a multi-modal LLM which allows for the interactive modification of generated molecular structures using text instructions. As we show, the new method greatly improves upon existing ICL methods for molecular design while being accessible and easy to use for scientists.
Machine learning models support computer-aided molecular design and compound optimization. However, the initial phases of drug discovery often face a scarcity of training data for these models. Meta-learning has emerged as a potentially promising strategy, harnessing the wealth of structure-activity data available for known targets to facilitate efficient few-shot model training for the specific target of interest. In this study, we assessed the effectiveness of two different meta-learning methods, namely model-agnostic meta-learning (MAML) and adaptive deep kernel fitting (ADKF), specifically in the regression setting. We investigated how factors such as dataset size and the similarity of training tasks impact predictability. The results indicate that ADKF significantly outperformed both MAML and a single-task baseline model on the inhibition data. However, the performance of ADKF varied across different test tasks. Our findings suggest that considerable enhancements in performance can be anticipated primarily when the task of interest is similar to the tasks incorporated in the meta-learning process. Meta-learning has emerged as a promising strategy to facilitate efficient few-shot model training for a specific target of interest. Here, we show that the performance of the meta-learned model critically depends on the correlation - captured in terms of molecular similarity and target activity - between the specific target of interest and the training data used to meta-learn the model. image
Lead optimization supported by artificial intelligence (AI)-based generative models has become increasingly important in drug design. Success factors are reagent availability, novelty, and the optimization of multiple properties. Directed fragment-replacement is particularly attractive, as it mimics medicinal chemistry tactics. Here, we present variations of fragment-based reinforcement learning using an actor-critic model. Novel features include freezing fragments and using reagents as the fragment source. Splitting molecules according to reaction schemes improves synthesizability, while tuning network output probabilities allows us to balance novelty versus diversity. Combining fragment-based optimization with virtual library encodings allows the exploration of large chemical spaces with synthesizable ideas. Collectively, these enhancements influence design toward high-quality molecules with favorable profiles. A validation study using 15 pharmaceutically relevant targets reveals that novel structures are obtained for most cases, which are identical or related to independent validation sets for each target. Hence, these modifications significantly increase the value of fragment-based reinforcement learning for drug design. The code is available on GitHub: https://github.com/Sanofi-Public/IDD-papers-fragrl
Amylin receptors (AMYRs), heterodimers of the calcitonin receptor (CTR) and one of three receptor activity-modifying proteins, are promising obesity targets. A hallmark of AMYR activation by Amy is the formation of a 'bypass' secondary structural motif (residues S19-P25). This study explored potential tuning of peptide selectivity through modification to residues 19-22, resulting in a selective AMYR agonist, San385, as well as nonselective dual amylin and calcitonin receptor agonists (DACRAs), with San45 being an exemplar. We determined the structure and dynamics of San385-bound AMY3R, and San45 bound to AMY3R or CTR. San45, via its conjugated lipid at position 21, was anchored at the edge of the receptor bundle, enabling a stable, alternative binding mode when bound to the CTR, in addition to the bypass mode of binding to AMY3R. Targeted lipid modification may provide a single intervention strategy for design of long-acting, nonselective, Amy-based DACRAs with potential anti-obesity effects.
A key challenge in drug discovery is to optimize, in silico, various absorption and affinity properties of small molecules. One strategy that was proposed for such optimization process is active learning. In active learning molecules are selected for testing based on their likelihood of improving model performance. To enable the use of active learning with advanced neural network models we developed two novel active learning batch selection methods. These methods were tested on several public datasets for different optimization goals and with different sizes. We have also curated new affinity datasets that provide chronological information on state-of-the-art experimental strategy. As we show, for all datasets the new active learning methods greatly improved on existing and current batch selection methods leading to significant potential saving in the number of experiments needed to reach the same model performance. Our methods are general and can be used with any package including the popular DeepChem library.
Molecular generative artificial intelligence is drawing significant attention in the drug design community, with several experimentally validated proof of concepts already published. Nevertheless, generative models are known for sometimes generating unrealistic, unstable, unsynthesizable, or uninteresting structures. This calls for methods to constrain those algorithms to generate structures in drug-like portions of the chemical space. While the concept of applicability domains for predictive models is well studied, its counterpart for generative models is not yet well-defined. In this work, we empirically examine various possibilities and propose applicability domains suited for generative models. Using both public and internal data sets, we use generative methods to generate novel structures that are predicted to be actives by a corresponding quantitative structure-activity relationships model while constraining the generative model to stay within a given applicability domain. Our work looks at several applicability domain definitions, combining various criteria, such as structural similarity to the training set, similarity of physicochemical properties, unwanted substructures, and quantitative estimate of drug-likeness. We assess the structures generated from both qualitative and quantitative points of view and find that the applicability domain definitions have a strong influence on the drug-likeness of generated molecules. An extensive analysis of our results allows us to identify applicability domain definitions that are best suited for generating drug-like molecules with generative models. We anticipate that this work will help foster the adoption of generative models in an industrial context.
Human glucose transporters (GLUTs) are responsible for cellular uptake of hexoses. Elevated expression of GLUTs, particularly GLUT1 and GLUT3, is required to fuel the hyperproliferation of cancer cells, making GLUT inhibitors potential anticancer therapeutics. Meanwhile, GLUT inhibitor-conjugated insulin is being explored to mitigate the hypoglycemia side effect of insulin therapy in type 1 diabetes. Reasoning that exofacial inhibitors of GLUT1/3 may be favored for therapeutic applications, we report here the engineering of a GLUT3 variant, designated GLUT3exo, that can be probed for screening and validating exofacial inhibitors. We identify an exofacial GLUT3 inhibitor SA47 and elucidate its mode of action by a 2.3 Å resolution crystal structure of SA47-bound GLUT3. Our studies serve as a framework for the discovery of GLUTs exofacial inhibitors for therapeutic development.
The identification and optimization of promising lead molecules is essential for drug discovery. Recently, artificial intelligence (AI) based generative methods provided complementary approaches for generating molecules under specific design constraints of relevance in drug design. The goal of our study is to incorporate protein 3D information directly into generative design by flexible docking plus an adapted protein-ligand scoring function, thereby moving towards automated structure-based design. First, the protein-ligand scoring function RFXscore integrating individual scoring terms, ligand descriptors, and combined terms was derived using the PDBbind database and internal data. Next, design results for different workflows are compared to solely ligand-based reward schemes. Our newly proposed, optimal workflow for structure-based generative design is shown to produce promising results, especially for those exploration scenarios, where diverse structures fitting to a protein binding site are requested. Best results are obtained using docking followed by RFXscore, while, depending on the exact application scenario, it was also found useful to combine this approach with other metrics that bias structure generation into "drug-like" chemical space, such as target-activity machine learning models, respectively.
The accurate prediction of protein-ligand binding affinity belongs to one of the central goals in computer-based drug design. Molecular dynamics (MD)-based free energy calculations have become increasingly popular in this respect due to their accuracy and solid theoretical basis. Here, we present a combined study which encompasses experimental and computational studies on two series of factor Xa ligands, which enclose a broad chemical space including large modifications of the central scaffold. Using this integrated approach, we identified several new ligands with different heterocyclic scaffolds different from the previously identified indole-2-carboxamides that show superior or similar affinity. Furthermore, the so far underexplored terminal alkyne moiety proved to be a suitable non-classical bioisosteric replacement for the higher halogen-π aryl interactions. With this challenging example, we demonstrated the ability of the MD-based non-equilibrium free energy calculation approach for guiding crucial modifications in the lead optimization process, such as scaffold replacement and single-site modifications at molecular interaction hot spots.
In silico models based on Deep Neural Networks (DNNs) are promising for predicting activities and properties of new molecules. Unfortunately, their inherent black-box character hinders our understanding, as to which structural features are important for activity. However, this information is crucial for capturing the underlying structure-activity relationships (SARs) to guide further optimization. To address this interpretation gap, "Explainable Artificial Intelligence" (XAI) methods recently became popular. Herein, we apply and compare multiple XAI methods to projects of lead optimization data sets with well-established SARs and available X-ray crystal structures. As we can show, easily understandable and comprehensive interpretations are obtained by combining DNN models with some powerful interpretation methods. In particular, SHAP-based methods are promising for this task. A novel visualization scheme using atom-based heatmaps provides useful insights into the underlying SAR. It is important to note that all interpretations are only meaningful in the context of the underlying models and associated data.
Abstract Targeting amylin (Amy) receptors (AMYRs) can reduce body weight with additional benefits to other anti-obesity treatments such as glucagon-like peptide-1 receptor (GLP-1R) agonists. AMYRs are heterodimers of the calcitonin receptor (CTR) and one of three receptor activity-modifying proteins (RAMPs), yielding AMY1R, AMY2R and AMY3R, respectively. A hallmark of AMYR activation by Amy is the formation of a secondary structural motif, termed a “bypass motif” (residues S19-P25) that partly contributes to selective activation of cAMP responses at AMYRs over CTR. This study explored the feasibility of tuning the selectivity of Amy analogues by modifying the residues (19-22) located within the bypass motif, resulting in a selective AMYR agonist, San385, as well as a series of non-selective dual amylin and calcitonin receptor agonists (DACRAs), with San45 being an exemplar. We determined the structure and dynamics of San385-bound AMY3R, as well as San45-bound AMY3R and CTR, decoding the structure-activity relationship (SAR) of these peptides. In particular, San45 is conjugated at position 19 with a lipid modification that anchors the peptide at the edge of receptor bundle and enables an alternate binding mode when bound to the CTR, in addition to the bypass mode of binding to AMY3R. This unique mechanism provides a single intervention strategy through targeted lipid modification to the structure-based design of long-acting, non-selective, Amy-based DACRAs with potential anti-obesity effects.
Artificial intelligence has seen an incredibly fast development in recent years. Many novel technologies for property prediction of drug molecules as well as for the design of novel molecules were introduced by different research groups. These artificial intelligence-based design methods can be applied for suggesting novel chemical motifs in lead generation or scaffold hopping as well as for optimization of desired property profiles during lead optimization. In lead generation, broad sampling of the chemical space for identification of novel motifs is required, while in the lead optimization phase, a detailed exploration of the chemical neighborhood of a current lead series is advantageous. These different requirements for successful design outcomes render different combinations of artificial intelligence technologies useful. Overall, we observe that a combination of different approaches with tailored scoring and evaluation schemes appears beneficial for efficient artificial intelligence-based compound design.
In silico driven optimization of compound properties related to pharmacokinetics, pharmacodynamics, and safety is a key requirement in modern drug discovery. Nowadays, large and harmonized datasets allow to implement deep neural networks (DNNs) as a framework for leveraging predictive models. Nevertheless, various available model architectures differ in their global applicability and performance in lead optimization projects, such as stability over time and interpretability of the results. Here, we describe and compare the value of established DNN‐based methods for the prediction of key ADME property trends and biological activity in an industrial drug discovery environment, represented by microsomal lability, CYP3A4 inhibition and factor Xa inhibition. Three architectures are exemplified, our earlier described multilayer perceptron approach (MLP), graph convolutional network‐based models (GCN) and a vector representation approach, Mol2Vec. From a statistical perspective, MLP and GCN were found to perform superior over Mol2Vec, when applied to external validation sets. Interestingly, GCN‐based predictions are most stable over a longer period in a time series validation study. Apart from those statistical observations, DNN prove of value to guide local SAR. To illustrate this important aspect in pharmaceutical research projects, we discuss challenging applications in medicinal chemistry towards a more realistic picture of artificial intelligence in drug discovery.
A detailed understanding of ligand–protein interactions is a key element in the search for novel drugs. The activity profile of novel compounds is essential to understand their mechanism of action as well as potential side effects. In phenotypic screening, compounds are identified due to a beneficial effect in a cellular or organism model, but the protein targets are typically not known. Various in silico methods have been established, which are capable to suggest the putative target or targets with good success rates. Here, we give an overview of ligand- and structure-based in silico methods for target prediction. These methods clearly profit from the large volume of ligand–protein data in public and corporate databases. The benefits, potential opportunities, and challenges for different methods are discussed. The number of available validation studies as well as prospective applications underline the importance of in silico methods as an integrated part of target prediction to help enable the discovery of novel targets.
Physiological processes rely on initial recognition events between cellular components and other molecules or modalities. Biomolecules can have multiple sites or mode of interaction with other molecular entities, so that a resolution of the individual binding events in terms of spatial localization as well as association and dissociation kinetics is required for a meaningful description. Here we describe a trichromatic fluorescent binding- and displacement assay for simultaneous monitoring of three individual binding sites in the important transporter and binding protein human serum albumin. Independent investigations of binding events by X-ray crystallography and time-resolved dynamics measurements (switchSENSE technology) confirm the validity of the assay, the localization of binding sites and furthermore reveal conformational changes associated with ligand binding. The described assay system allows for the detailed characterization of albumin-binding drugs and is therefore well-suited for prediction of drug-drug and drug-food interactions. Moreover, conformational changes, usually associated with binding events, can also be analyzed.