Disease-associated macrophage states contribute to the pathogenesis of numerous inflammatory disorders. While small molecule-mediated macrophage repolarization represents a promising strategy to restore homeostasis, current approaches often rely on predefined molecular markers and may therefore overlook previously unrecognized modulators of macrophage plasticity. To address this limitation, we performed a phenotypic high-content cell painting screen in induced pluripotent stem cell-derived and blood monocyte-derived macrophages. Using our previously established machine learning-based cell painting analysis pipeline, we screened the annotated opnMe and JUMP-CP compound libraries and identified 26 compounds that induced M1- or M2-like polarization phenotypes. Among these hits, we identified WZ4003, an inhibitor of AMPK-related kinases NUAK1 and NUAK2, as novel macrophage-polarizing compound. Combining cell painting feature profiling with functional analyses, we revealed that the NUAK1/2 inhibitor WZ4003 induces a distinct macrophage state characterized by a rounded M1-like morphology, ferroptosis protection, mitochondrial reactive oxygen species increase, antioxidative adaptation, and phagocytosis and efferocytosis impairment. These findings link specific morphological signatures to functional macrophage states. Overall, this study provides a resource of macrophage-polarizing compounds and demonstrates the utility of machine learning-based cell painting for identifying novel modulators of macrophage polarization states.
ABSTRACT A successful drug needs to combine several properties including high potency and good pharmacokinetic (PK) properties to sustain efficacious plasma concentration over time. To estimate required doses for preclinical animal efficacy models or for the clinics, in vivo PK studies need to be conducted. Although the prediction of ADME properties of compounds using machine learning (ML) models based on chemical structures is well established in drug discovery, the prediction of complete plasma concentration–time profiles has only recently gained attention. In this study, we systematically compare various approaches that integrate ML models with empiric or mechanistic PK models to predict PK profiles in rats after intravenous administration prior to synthesis. More specifically, we compare a standard noncompartmental analysis (NCA)‐based approach (prediction of CL and Vss), a pure ML approach (non‐mechanistic PK description), a compartmental modeling approach, and a physiologically based pharmacokinetic (PBPK) approach. Our study based on internal preclinical data shows that the latter three approaches yield PK profile predictions of comparable accuracy across a large data set (evaluated as geometric mean fold errors for each profile of over 1000 small molecules). In summary, we demonstrate the improved ability to prioritize drug candidates with desirable PK properties prior to synthesis with ML predictions.
As training volume increases predictive model quality, leveraging existing external data sources holds the promise of time- and cost-efficiency. In a drug discovery setting, pharmaceutical companies all own substantial but confidential datasets. The MELLODDY project develops a privacy-preserving federated machine learning solution and deploys it at an unprecedented scale (more than 100,000 tasks across ten major pharmaceutical companies), while ensuring the security and privacy of each partner’s sensitive data. Each partner builds models that benefit from a shared representation, for their own private assays. Established predictive performance metrics such as AUC ROC or AUC PR are constrained to unseen labelled chemical space. However, they cannot gauge performance gains in unlabelled chemical space. Federated learning indirectly extends labelled space, but in a privacy-preserving context, a partner cannot use this label extension for performance assessment. Metrics that estimate uncertainty on a prediction can be calculated even where no label is known. Practically, the chemical space covered with predictions of sufficient confidence, reflects the applicability domain of a model. After establishing a link to established performance metrics, we propose the efficiency from the conformal prediction framework (‘conformal efficiency’) as a proxy to the applicability domain size. A documented extension of the applicability domain would qualify as a tangible benefit from federated learning. In interim assessments, MELLODDY partners report a median increase in conformal efficiency of the federated over the single-partner model of 5.5% (with increases up to 9.7%). Subject to distributional conditions, that efficiency increase can be directly interpreted as the expected increase in conformal i.e. high confidence predictions. In conclusion, we present the first evidence that privacy-preserving federated machine learning across massive drug-discovery datasets from ten pharma partners indeed extends the applicability domain of property prediction models.
Macrophage polarization critically contributes to a multitude of human pathologies. Hence, modulating macrophage polarization is a promising approach with enormous therapeutic potential. Macrophages are characterized by a remarkable functional and phenotypic plasticity, with pro-inflammatory (M1) and anti-inflammatory (M2) states at the extremes of a multidimensional polarization spectrum. Cell morphology is a major indicator for macrophage activation, describing M1(-like) (rounded) and M2(-like) (elongated) states by different cell shapes. Here, we introduced cell painting of macrophages to better reflect their multifaceted plasticity and associated phenotypes beyond the rigid dichotomous M1/M2 classification. Using high-content imaging, we established deep learning- and feature-based cell painting image analysis tools to elucidate cellular fingerprints that inform about subtle phenotypes of human blood monocyte-derived and iPSC-derived macrophages that are characterized as screening surrogate. Moreover, we show that cell painting feature profiling is suitable for identifying inter-donor variance to describe the relevance of the morphology feature ‘cell roundness’ and dissect distinct macrophage polarization signatures after stimulation with known biological or small-molecule modulators of macrophage (re-)polarization. Our novel established AI-fueled cell painting analysis tools provide a resource for high-content-based drug screening and candidate profiling, which set the stage for identifying novel modulators for macrophage (re-)polarization in health and disease.
ADME (Absorption, Distribution, Metabolism, Excretion) properties are key parameters to judge whether a drug candidate exhibits a desired pharmacokinetic (PK) profile. In this study, we tested multi-task machine learning (ML) models to predict ADME and animal PK endpoints trained on in-house data generated at Boehringer Ingelheim. Models were evaluated both at the design stage of a compound (i.e., no experimental data of test compounds available) and at testing stage when a particular assay would be conducted (i.e., experimental data of earlier conducted assays may be available). Using realistic time-splits, we found a clear benefit in performance of multi-task graph-based neural network models over single-task models, which was even stronger when experimental data of earlier assays is available. In an attempt to explain the success of multi-task models, we found that especially endpoints with the largest numbers of data points (physicochemical endpoints, clearance in microsomes) are responsible for increased predictivity in more complex ADME and PK endpoints. In summary, our study provides insight into how data for multiple ADME/PK endpoints in a pharmaceutical company can be best leveraged to optimize predictivity of ML models.
A successful drug needs to combine several properties including high potency and good pharmacokinetic (PK) properties to sustain efficacious plasma concentration over time. To estimate required doses for preclinical animal efficacy models or for the clinics, in vivo PK studies need to be conducted. While the prediction of ADME properties of compounds using Machine Learning (ML) models based on chemical structures is well established in drug discovery, the prediction of complete plasma concentration-time profiles has only recently gained attention. In this study, we systematically compare various approaches that integrate ML models with mechanistic PK models to predict PK profiles in rats after i.v. administration prior to synthesis. More specifically, we compare a standard noncompartmental analysis (NCA) based approach (prediction of CL and Vss), a pure ML approach (non-mechanistic PK description), a compartmental modeling approach, and a physiologically based pharmacokinetic (PBPK) approach. Our study based on internal preclinical data shows that the latter three approaches yield PK profile predictions of comparable accuracy (evaluated as geometric mean fold errors for each profile) across a large test set (>1000 small molecules). In summary, we demonstrate the improved ability to prioritize drug candidates with desirable PK properties prior to synthesis with ML predictions. ### Competing Interest Statement The authors have declared no competing interest.
In a drug discovery setting, pharmaceutical companies own substantial but confidential datasets. The MELLODDY project developed a privacy-preserving federated machine learning solution and deployed it at an unprecedented scale. Each partner built models for their own private assays that benefitted from a shared representation. Established predictive performance metrics such as AUC ROC or AUC PR are constrained to unseen labeled chemical space and cannot gage performance gains in unlabeled chemical space. Federated learning indirectly extends labeled space, but in a privacy-preserving context, a partner cannot use this label extension for performance assessment. Metrics that estimate uncertainty on a prediction can be calculated even where no label is known. Practically, the chemical space covered with predictions above an uncertainty threshold, reflects the applicability domain of a model. After establishing a link to established performance metrics, we propose the efficiency from the conformal prediction framework (‘conformal efficiency’) as a proxy to the applicability domain size. A documented extension of the applicability domain would qualify as a tangible benefit from federated learning. In interim assessments, MELLODDY partners reported a median increase in conformal efficiency of the federated over the single-partner model of 5.5% (with increases up to 9.7%). Subject to distributional conditions, that efficiency increase can be directly interpreted as the expected increase in conformal i.e. low uncertainty predictions. In conclusion, we present the first indication that privacy-preserving federated machine learning across massive drug-discovery datasets from ten pharma partners indeed extends the applicability domain of property prediction models.
Federated multipartner machine learning has been touted as an appealing and efficient method to increase the effective training data volume and thereby the predictivity of models, particularly when the generation of training data is resource-intensive. In the landmark MELLODDY project, indeed, each of ten pharmaceutical companies realized aggregated improvements on its own classification or regression models through federated learning. To this end, they leveraged a novel implementation extending multitask learning across partners, on a platform audited for privacy and security. The experiments involved an unprecedented cross-pharma data set of 2.6+ billion confidential experimental activity data points, documenting 21+ million physical small molecules and 40+ thousand assays in on-target and secondary pharmacodynamics and pharmacokinetics. Appropriate complementary metrics were developed to evaluate the predictive performance in the federated setting. In addition to predictive performance increases in labeled space, the results point toward an extended applicability domain in federated learning. Increases in collective training data volume, including by means of auxiliary data resulting from single concentration high-throughput and imaging assays, continued to boost predictive performance, albeit with a saturating return. Markedly higher improvements were observed for the pharmacokinetics and safety panel assay-based task subsets.
AbstractRational drug design deals with computational methods to accelerate the development of new drugs. Among other tasks, it is necessary to analyze huge databases of small molecules. Since a direct relationship between the structure of these molecules and their effect (e.g., toxicity) can be assumed in many cases, a wide set of methods is based on the modeling of the molecules as graphs with attributes.Here, we discuss our results concerning structural molecular similarity searches and molecular clustering and put them into the wider context of graph similarity search. In particular, we discuss algorithms for computing graph similarity w.r.t. maximum common subgraphs and their extension to domain specific requirements.
Knowledge about interrelationships between different proteins is crucial in fundamental research for the elucidation of protein networks and pathways. Furthermore, it is especially critical in chemical biology to identify further key regulators of a disease and to take advantage of polypharmacology effects. A comprehensive scaffold-based analysis uncovered an unexpected relationship between bromodomain-containing protein 4 (BRD4) and peroxisome-proliferator activated receptor gamma (PPARγ). They are both important drug targets for cancer therapy and many more important diseases. Both proteins share binding site similarities near a common hydrophobic subpocket which should allow the design of a polypharmacology-based ligand targeting both proteins. Such a dual-BRD4-PPARγ-modulator could show synergistic effects with a higher efficacy or delayed resistance development in, for example, cancer therapy. Thereon, a complex structure of sulfasalazine was obtained that involves two bromodomains and could be a potential starting point for the design of a bivalent BRD4 inhibitor.
Machine learning models predicting the bioactivity of chemical compounds belong nowadays to the standard tools of cheminformaticians and computational medicinal chemists. Multi-task and federated learning are promising machine learning approaches that allow privacy-preserving usage of large amounts of data from diverse sources, which is crucial for achieving good generalization and high-performance results. Using large, real world data sets from six pharmaceutical companies, here we investigate different strategies for averaging weighted task loss functions to train multi-task bioactivity classification models. The weighting strategies shall be suitable for federated learning and ensure that learning efforts are well distributed even if data are diverse. Comparing several approaches using weights that depend on the number of sub-tasks per assay, task size, and class balance, respectively, we find that a simple sub-task weighting approach leads to robust model performance for all investigated data sets and is especially suited for federated learning.
With the increase in applications of machine learning methods in drug design and related fields, the challenge of designing sound test sets becomes more and more prominent. The goal of this challenge is to have a realistic split of chemical structures (compounds) between training, validation and test set such that the performance on the test set is meaningful to infer the performance in a prospective application. This challenge is by its own very interesting and relevant,but is even more complex in a federated machine learning approach where multiple partners jointly train a model under privacy-preserving conditions where chemical structures must not be shared between the different participating parties in the federated learning. In this work we discuss three methods which provide a splitting of the data set and are applicable in a federated privacy-preserving setting, namely: a. locality-sensitive hashing (LSH), b. sphere exclusion clustering, c. scaffold-based binning (scaffold network). For evaluation of these splitting methods we consider the following quality criteria: bias in prediction performance, label and data imbalance, distance of the test set compounds to the training set and compare them to a random splitting. The main findings of the paper are a. both sphere exclusion clustering and scaffold-based binning result in high quality splitting of the data sets, b. in terms of compute costs sphere exclusion clustering is very expensive in the case of federated privacy-preserving setting.
Chemical similarity between two molecules is a fundamental concept in cheminformatics and structure comparison is therefore an often applied and important task. Structural comparison is used, e.g., to predict biological activities or to analyze molecular datasets. One approach for the identification of chemical similarity is based on a graph representation of molecules, because a molecule can intuitively be interpreted as a graph structure. In this article we focus on algorithms for the calculation of chemical similarity based on a graph representation, which is expressed as the maximum common substructure between two molecules.
Protein ligand interaction fingerprints are a powerful approach for the analysis and assessment of docking poses to improve docking performance in virtual screening. In this study, a novel interaction fingerprint approach (PADIF, protein per atom score contributions derived interaction fingerprint) is presented which was specifically designed for utilising the GOLD scoring functions' atom contributions together with a specific scoring scheme. This allows the incorporation of known protein-ligand complex structures for a target-specific scoring. Unlike many other methods, this approach uses weighting factors reflecting the relative frequency of a specific interaction in the references and penalizes destabilizing interactions. In addition, and for the first time, an exhaustive validation study was performed that assesses the performance of PADIF and two other interaction fingerprints in virtual screening. Here, PADIF shows superior results, and some rules of thumb for a successful use of interaction fingerprints could be identified.
A common issue during drug design and development is the discovery of novel scaffolds for protein targets. On the one hand the chemical space of purchasable compounds is rather limited; on the other hand artificially generated molecules suffer from a grave lack of accessibility in practice. Therefore, we generated a novel virtual library of small molecules which are synthesizable from purchasable educts, called CHIPMUNK (CHemically feasible In silico Public Molecular UNiverse Knowledge base). Altogether, CHIPMUNK covers over 95 million compounds and encompasses regions of the chemical space that are not covered by existing databases. The coverage of CHIPMUNK exceeds the chemical space spanned by the Lipinski rule of five to foster the exploration of novel and difficult target classes. The analysis of the generated property space reveals that CHIPMUNK is well suited for the design of protein–protein interaction inhibitors (PPIIs). Furthermore, a recently developed structural clustering algorithm (StruClus) for big data was used to partition the sub‐libraries into meaningful subsets and assist scientists to process the large amount of data. These clustered subsets also contain the target space based on ChEMBL data which was included during clustering.
The era of big data is influencing the way how rational drug discovery and the development of bioactive molecules is performed and versatile tools are needed to assist in molecular design workflows. Scaffold Hunter is a flexible visual analytics framework for the analysis of chemical compound data and combines techniques from several fields such as data mining and information visualization. The framework allows analyzing high-dimensional chemical compound data in an interactive fashion, combining intuitive visualizations with automated analysis methods including versatile clustering methods. Originally designed to analyze the scaffold tree, Scaffold Hunter is continuously revised and extended. We describe recent extensions that significantly increase the applicability for a variety of tasks.
The ever increasing bioactivity data that are produced nowadays allow exhaustive data mining and knowledge discovery approaches that change chemical biology research. A wealth of chemoinformatics tools, web services and applications therefore exists that supports a careful evaluation and analysis of experimental data to draw conclusions that can influence the further development of chemical probes and potential lead structures. This review focuses on open-source approaches that can be handled by scientists who are not familiar with computational methods having no expert knowledge in chemoinformatics and modeling. Our aim is to present an easily manageable toolbox for support of every day laboratory work. This includes, among other things, the available bioactivity and related molecule databases as well as tools to handle and analyze in-house data.
The ever increasing bioactivity data that are produced nowadays allow exhaustive data mining and knowledge discovery approaches that change chemical biology research. A wealth of chemoinformatics tools, web services, and applications therefore exists that supports a careful evaluation and analysis of experimental data to draw conclusions that can influence the further development of chemical probes and potential lead structures. This review focuses on open-source approaches that can be handled by scientists who are not familiar with computational methods having no expert knowledge in chemoinformatics and modeling. Our aim is to present an easily manageable toolbox for support of every day laboratory work. This includes, among other things, the available bioactivity and related molecule databases as well as tools to handle and analyze in-house data.