Machine learning has a long history in the life and natural sciences. In the age of artificial intelligence (AI), both the opportunities and challenges associated with predictive and generative modeling are increasing. To exert significant influence on experimental programs in interdisciplinary research, AI models and their predictions must be demystified as far as possible, and misconceptions regarding what highly complex models can and cannot do must be avoided. Furthermore, it is necessary to evaluate best practices for the use of AI in life science research including drug discovery. This contribution discusses the evolution of AI models in the life and natural sciences, as well as current developments and remaining challenges in this rapidly evolving field. It is dedicated to the memory of Terry R. Stouch, a valued colleague and dear friend whom I had the privilege of knowing for more than 30 years and who left us far too soon.
We present a protocol to generate dual-target compounds (DT-CPDs) and triple-target compounds (TT-CPDs) interacting with two and three target proteins, respectively, using transformer-based chemical language models. We describe steps for installing software, preparing data, and pre-training the model on pairs of single-target compounds (ST-CPDs), which bind to an individual protein, and DT-CPDs. We then detail procedures for assembling ST-CPD and corresponding DT- or TT-CPD data for fine-tuning on specific protein pairs or triplets and for evaluating model performance on hold-out test sets.For complete details on the use and execution of this protocol, please refer to Srinivasan et al. 1 for DT-CPDs and to Srinivasan et al.2 for TT-CPDs. This protocol is an update to Srinivasan et al.3
Machine learning (ML) is widely used in medicinal chemistry, but accurate predictions alone are insufficient. Researchers need insight into which molecular features determine compound properties. We present an application-oriented case study that analyzes a trained model for compound potency as a source of domain knowledge. The model is converted into decision rules, and topic-guided visual analytics is used to identify co-occurring feature conditions associated with high predicted potency. These patterns are then mapped back to molecular substructures, yielding chemically interpretable motifs and testable hypotheses about structure-activity relationships. The study demonstrates how combining rule-based representations, topic modeling, and visual exploration can turn potency predictions into mechanistic insight, and outlines a reusable workflow for interpreting ML models of molecular properties.
Applications of machine learning (ML) in the natural sciences and interdisciplinary research typically benefit from downstream analyses using explainable artificial intelligence (XAI) to rationalize predictions. We have designed and implemented an AI agent system powered by a large language model (LLM) to seamlessly link various ML and XAI methods for applications in chemistry. To this end, we focused on chemically intuitive XAI concepts. The agent system was carefully evaluated using different prediction tasks combining (deep) ML and XAI that also required discussion and interpretation of the results, without intermediate user intervention. In these proof-of-concept applications, the agent produced accurate results and showed a partly surprising capacity for chemical interpretation, although the central LLM did not undergo specific chemical training. We also detected limitations of the system’s ability to evaluate results from a chemical perspective, for instance, when challenged with comparing learning characteristics of different models. These findings demonstrated that the agent was not “thinking” independently. However, in terms of technical performance and efficiency, adaptability, and the capacity for chemical interpretation, the agent system significantly advances the use of XAI in molecular ML.
Screening of drugs on cancer cell lines is a first step in the search for compounds leading to cancer cell death, providing the basis for follow-up investigations. Cell line-based drug sensitivity data can be used for predictive modeling, for example, in the context of drug repurposing. In this work, we report a machine learning framework for the systematic prediction of drug sensitivity based on transcriptomic data from cancer cell line screening. Three different categories of complementary classification models are introduced to predict cell line sensitivities to known drugs based on encoded gene expression profiles, predict activity of new drugs for given cell lines based on compound fingerprints, and predict responses of new cell lines to new drugs based on combined molecular representations and gene expression profiles. The models are found to have adequate predictive performance for cell-based data and are shown to be applicable to predict drug sensitivity for sarcoma cell lines, a rare form of cancer, for which data are limited. For the subset of sarcoma cell lines, tested drugs were also ranked based on cell line sensitivity rates. Taken together, our results suggest that the complementary machine learning models have potential for practical applications to search for new compounds for the treatment of different types of cancers including sarcoma.
The definition of activity cliffs (ACs) depends on compound similarity and activity difference criteria and on activity data types. ACs are usually defined as pairs or groups of structurally similar compounds or structural analogues that are active against the same target, but have large differences in potency (requiring numerical potency values, preferably equilibrium constants). In addition, ACs have also been defined as pairs of structural analogues that are active or inactive in screening assays. In medicinal chemistry, ACs are of particular interest because they often reveal structure-activity relationship (SAR) determinants during compound optimization. In cheminformatics, ACs present challenging test cases for machine learning (ML) and activity predictions because they represent an extreme form of SAR discontinuity in compound data sets. Given their composition, ACs are notoriously difficult to predict based on chemical structure representations. Various attempts have been made to predict compound pairs forming ACs or the activity of AC compounds via ML, often reporting high accuracy. However, in the absence of data leakage from training to test sets, AC prediction accuracy based on chemical structure is low or moderate at best. As an alternative to structural representations, biological/functional compound descriptors might be considered such as biological assay profiles, which have been investigated for other compound activity prediction. In this work, we report the prediction of AC compounds based on bioactivity profiles derived from a compound profiling matrix using data partitioning schemes designed to control information or data leakage during ML. Under these stringent conditions, AC compound predictions based on bioactivity profiles often failed. However, we also observed subsets of highly accurate predictions and explored in detail why these predictions succeeded, but others failed. The analysis revealed a critically important role of assay similarity for successful AC compound predictions. Most profile assays did not measurably influence the predictions. By contrast, accurate predictions mostly depended on the presence of one or at most a few profile assays that were similar to test assays. In most cases, these profile assays could be identified and exploited for predictions by nearest neighbor searching, thus putting ML performance into perspective.Scientific contribution As an alternative to the generally difficult prediction of ACs based on chemical structure, we introduce prediction of AC compounds based on bioactivity profiles. We show that accurate activity predictions of AC compounds do not depend on the global information content of assay profiles, but are largely determined by nearest neighbor relationships between profile and test assays.
Protein kinases (PKs) are among the most intensely investigated drug targets. Small molecular PK inhibitors (PKIs) are preferred agents for therapeutic intervention. Late in 2025, the 100th PKI reached drug approval by the U.S. Food and Drug Administration. However, current PKI drugs are directed against a limited number of primary PK targets. Thus, numerous PKs comprising the human kinome remain available as potential pharmaceutical targets. In 2019, the U.S. National Institutes of Health reported 162 understudied human kinases including 155 PKs (and seven lipid kinases), referred to as dark kinases or the dark kinome. These kinases were identified primarily considering limited functional information and the lack of detection reagents. Herein, we report a new categorization of the human kinome based on publicly available biological and PKI data. The unified categorization integrates biological and chemical exploration of PKs, introduces new PK categories for target evaluation, and defines the understudied kinome.
Compound optimization is of central relevance in medicinal chemistry. We introduce a new machine learning framework for iterative chemical optimization that integrates compound potency predictions, the explanation of predictions, and generative modeling and that is applicable to individual compounds. The approach identifies substituents in active compounds that limit their potency and iteratively replaces these substituents with others supporting potency increases. In proof-of-concept calculations, the methodology effectively optimizes compound potency. Furthermore, the optimization framework is combined with a large language model via the model concept protocol to generate an AI agent system for interactive optimization. The system is shown to successfully carry out optimization tasks of increasing complexity based on simple prompts, without the need for additional fine-tuning. The interactive computational optimization approach is accessible to non-experts and expected to be of particular interest for practical medicinal chemistry.
Hierarchical pooling is a promising mechanism to enhance graph neural networks (GNNs) by enabling multi-scale representation learning. Rationalization of hierarchical GNN predictions remains an underexplored area. In this work, we investigate the impact of hierarchical pooling on GNNs for molecular property prediction. We designed architectural variants integrating pharmacophore features with pooling GNNs at different levels. GNN models with pharmacophore-based graph reduction or hierarchical pooling achieved comparable compound classification performance. Explainable artificial intelligence (XAI) methods were applied to compare feature importance and substructure attribution for the different model architectures. Qualitative and quantitative analyses of the resulting explanations demonstrated that the GNN variants had different internal learning characteristics. GNN models based on reduced graphs matched the prediction accuracy of models based on complete graph representations following different variant-dependent learning strategies.
Diffusion models have gained considerable importance in many scientific fields. These models perturb the original data distributions by adding noise during the diffusion phase, which is then removed again during the subsequent denoising phase when generating the output. Predictions of diffusion are difficult to rationalize. Current efforts to explain the generative process and outputs of diffusion models are essentially limited to training data attribution for image generation. Herein, we report a feature attribution approach to better understand how diffusion models used in molecular design arrive at their results. As a proof of concept, we analyze an equivariant diffusion model for generating linkers in fragment-based compound design. The analysis yields plausible explanations for linker predictions and uncovers unexpected atomic contributions to generative design. Notably, the equivariant diffusion model does not require learning the underlying chemistry to produce chemically sound compound structures.
Chemical language models (CLMs), particularly encoder-decoder transformers, have advanced generative molecular design. Transformer CLMs are able to learn a variety of molecular mappings for compound design that can be conditioned using context-dependent rules. However, their black-box nature complicates the interpretation of predictions. Current analysis methods mostly focus on attention weights of token relationships or attention flow in encoder and decoder modules and cannot explain predictions at the molecular level. Sequence-based compound design was used as a model system to investigate transformer learning characteristics through systematic control calculations involving modifications of protein sequences and sequence-compound pairs. The analysis revealed that compound reproducibility depended on similarity relationships between training and test data and on compound memorization, while specific sequence information was not learned. These findings indicate that predictions of transformer CLMs are driven by memorization effects and statistical correlations rather than by learning specific chemical or biological information. Understanding this learning behavior aids in avoiding over-interpretation of model outputs and informs the appropriate application of transformer-based CLMs in molecular design.
Polypharmacology-based drug discovery relies on compounds with multi-target activity that are identified in screening and target profiling assays or using computational methods. Contemporary design of multi-target compounds is advanced by deep generative modeling. Dual-target compounds (DT-CPDs) are known for having a large number of target combinations, providing a sound basis for machine learning. By contrast, only comparably small numbers of triple-target compounds (TT-CPDs) are available, covering a very limited target space. Here, we investigate how this data restriction might be overcome to enable generative design of new TT-CPDs. Therefore, a transformer model is pre-trained to generate DT-CPDs from corresponding single-target compounds and used as a base model for triple-target fine-tuning. For different target combinations, the resulting models correctly reproduce known TT-CPDs not encountered during fine-tuning. Feature importance analysis explains the predictions and reveals structural motifs implicated in target selectivity or triple-target activity, thus providing a chemically intuitive rationale for the approach.
The rise of artificial intelligence (AI) has taken machine learning (ML) in molecular design to a new level. As ML increasingly relies on complex deep learning frameworks, the inability to understand predictions of black-box models has become a topical issue. Consequently, there is strong interest in the field of explainable AI (XAI) to bridge the gap between black-box models and the acceptance of their predictions, especially at interfaces with experimental disciplines. Therefore, XAI methods must go beyond extracting learning patterns from ML models and present explanations of predictions in a human-centered, transparent, and interpretable manner. In this Perspective, we examine current challenges and opportunities for XAI in molecular design and evaluate the benefits of incorporating domain-specific knowledge into XAI approaches for model refinement, experimental design, and hypothesis testing. In this context, we also discuss the current limitations in evaluating results from chemical language models that are increasingly used in molecular design and drug discovery.
In drug design, transformer networks adopted from natural language processing are applied in a variety of ways. We have used sequence-based generative compound design as a model system to explore the learning characteristics of transformers and determine if these models learned information relevant for protein-ligand interactions. The analysis reveals that sequence-based predictions of active compounds using transformer models required a proportion of at least ∼60% of the original test sequences. Moreover, predictions depended on sequence and compound similarity of training and test data and on compound memorization effects. The predictions were purely statistically driven by associating sequence patterns with molecular structures, thus rationalizing their strict dependence on detectable similarities. Moreover, the transformer models did not learn target sequence information relevant for ligand binding. While the results do not call sequence-based compound design approaches generally into question, they caution against over-interpretation of transformer models used for such applications.
Transformer-based chemical language models (CLMs) were derived to generate structurally and topologically diverse embeddings of core structure fragments, substituents, or core/substituent combinations in chemically proper compounds, representing a design task that is difficult to address using conventional structure generation methods. To this end, CLM variants were challenged to learn different fragment-to-compound mappings in the absence of structural rules or any other fragment linking or synthetic information. The resulting alternative models were found to have high syntactic fidelity, but displayed notable differences in their ability to generate valid candidate compounds containing test fragments, with a clear preference for a model variant processing core/substituent combinations. However, the majority of valid candidate compounds generated with all models were distinct from training data and structurally novel. In addition, the CLMs exhibited high chemical diversification capacity and often generated structures with new topologies not encountered during training. Furthermore, all models produced large numbers of close structural analogues of known bioactive compounds covering a large target space, thus indicating the relevance of newly generated candidates for pharmaceutical research. As a part of our study, the new methodology and all data are made publicly available.
Protein kinases (PKs) play a central role in cellular signaling. Uncontrolled signaling by deregulated PKs is implicated in a variety of diseases. As a consequence, PKs are among the most popular pharmaceutical targets. The preferred strategy for therapeutic intervention of medical conditions caused by deregulated PKs is the inhibition of their catalytic phosphorylation activity. Accordingly, small-molecular PK inhibitors (PKIs) have become a major drug class in oncology and prime candidates in other therapeutic areas. While cellular functions of many PKs and potential involvement in disease biology have been intensely investigated, others have received comparably little attention, leading to the identification of understudied 162 kinases representing the dark "dark kinome". Dark PKs have for the most part been categorized based on the absence of functional information and lack of reagents. Large-magnitude projects have been initiated to further explore and functionally characterize the dark kinome. In addition, different categories of PKs have also been defined based on their degree of chemical exploration in medicinal chemistry, representing complementary assessments of understudied PKs.