Abstract We present a data-driven and unsupervised approach for extracting 3D pharmacophore hypotheses, without prior ligand selection, from a chemogenomic kinase dataset (406 kinases and 645 compounds, PKIS2). A metric called NEM for Normalized Enrichment Measure is introduced for each pharmacophore, which quantifies change in the proportion of active compounds consistent with the pharmacophore compared to the original dataset. Based on this metric, we can identify pharmacophores associated with specific kinases and, conversely, determine all kinases that share similar metric values. This approach enables the characterization of polypharmacological profiles linked to individual pharmacophore hypotheses. We further evaluate the consistency of our results with various biological datasets, including ChEMBL, DrugBank, LINCS, KINOMEscan, and Kinobeads, showing agreement across representative case studies. This study highlights the potential of our approach for elucidating relationships between pharmacophores and kinase selectivity profiles, providing a scalable framework for exploring kinase–ligand interaction landscapes.
Summary Tuberculosis is characterized by broad clinical heterogeneity that hinders infection control, with differences in lesion development, progression, and treatment outcomes. This complexity is likely associated with Mycobacterium tuberculosis inherent phenotypic variation and its capacity to diversify under host microenvironmental and antimicrobial stressors. Here, we analyze M. tuberculosis at the single-cell and subpopulation level using fluorescent reporters, imaging, transcriptomic, and functional assays. We identify RNA signatures specific to stress-responsive bacilli with translational potential. Focusing on the clinically validated chaperone GroEL2, we find that it correlates with M. tuberculosis growth rate and stress tolerance in vitro and intracellularly. Furthermore, GroEL2 phenotypic diversity influences innate responses in macrophages, which experience different polarization, in turn affecting GroEL2 expression. We also show that targeting GroEL2 impairs pathogen survival and dampens inflammation. This study provides a link between pathogen phenotypic variation and macrophage fates, with implications for early infection outcomes, local disease progression, and subpopulation-targeted interventions.
Membrane type 5-matrix metalloproteinase (MT5-MMP, MMP-24), an η-secretase involved in amyloid precursor protein processing, is a promising but unexplored target in Alzheimer's disease. We report here the identification of a first nonpeptidic hit for MT5-MMP through a structure-guided approach. A homology model of the MT5-MMP catalytic domain was built from the MT3-MMP/batimastat structure and validated by both docking and experimental inhibition data obtained with batimastat (IC50 = 3 nM). To account for binding-site plasticity, especially that of the S1' pocket, molecular dynamics and ensemble docking were applied to a zinc-binding group (ZBG)-focused library of 3851 compounds. Although the initial screening campaign yielded only weakly active candidates, analysis of docking poses identified a relevant scaffold for optimization. Replacement of a carboxylic acid ZBG by a hydroxamic acid led to compound 17, which inhibited MT5-MMP with an IC50 of 6 μM and established a first nonpeptidic hit for future optimization.
The French Chemoinformatics Society (SFCi) is a learned society that unites French academics, students and industrial scientists in chemoinformatics, promoting this discipline at the interface of chemistry, computer science, data science and AI through conferences, training initiatives, networking and community-building.
Understanding ligand selectivity and efficacy at melatonin (MT1, MT2) and serotonin (5-HT2C) receptors remains a key challenge in central nervous system (CNS) drug discovery. Ligands combining MT1/MT2 agonism with 5-HT2C antagonism are of therapeutic relevance in depression and circadian rhythm disorders. This review provides a comprehensive overview of ligands reported for these receptors in ChEMBL, with an emphasis on compounds evaluated across multiple targets. A pharmacophore-based approach was applied to group ligands by shared features, revealing pharmacomodulations influencing receptor selectivity and polypharmacology. Together, these insights can inspire the rational design of selective or polypharmacological ligands, offering new opportunities for multitarget drug discovery.
Research fields that leverage relational data, like many others, have been significantly impacted by Deep Learning (DL) techniques, particularly Graph Neural Networks (GNNs). Among these fields, drug design, which aims to create new molecules with optimal affinities for specific targets, is a crucial step in the development of new medicinal drugs. In silico approaches in this area often rely on molecular graphs that encode the atoms and bonds of a molecule, without prior knowledge of the biological properties to be predicted. To address this limitation, pharmacophoric features are essential, as they contain structural information that captures important biological properties. These features have proven effective in tasks involving protein-ligand interactions. In this context, we propose the MCP-GNN model, which combines molecular representations with complete graphs of pharmacophoric features, both based on 2D information, to classify biological activity. Our experimental results demonstrate that this approach, using simple yet efficient techniques, achieves better performance than more complex architectures.
Drug-recalcitrant infections are a leading global-health concern. Bacterial cells benefit from phenotypic variation, which can suggest effective antimicrobial strategies. However, probing phenotypic variation entails spatiotemporal analysis of individual cells that is technically challenging, and hard to integrate into drug discovery. In this work, we develop a multi-condition microfluidic platform suitable for imaging two-dimensional growth of bacterial cells during transitions between separate environmental conditions. With this platform, we implement a dynamic single-cell screening for pheno-tuning compounds, which induce a phenotypic change and decrease cell-to-cell variation, aiming to undermine the entire bacterial population and make it more vulnerable to other drugs. We apply this strategy to mycobacteria, as tuberculosis poses a major public-health threat. Our lead compound impairs Mycobacterium tuberculosis via a peculiar mode of action and enhances other anti-tubercular drugs. This work proves that harnessing phenotypic variation represents a successful approach to tackle pathogens that are increasingly difficult to treat.
The exploration of chemical space is a fundamental aspect of chemoinformatics, particularly when one explores a large compound data set to relate chemical structures with molecular properties. In this study, we extend our previous work on chemical space visualization at the pharmacophoric level. Instead of using conventional binary classification of affinity (active vs inactive), we introduce a refined approach that categorizes compounds into four distinct classes based on their activity levels: super active, very active, active, and inactive. This classification enriches the color scheme applied to pharmacophore space, where the color representation of a pharmacophore hypothesis is driven by the associated compounds. Using the BCR-ABL tyrosine kinase as a case study, we identified intriguing regions corresponding to pharmacophore activity discontinuities, providing valuable insights for structure-activity relationships analysis. image
In this work, we propose to analyze the potential of a new type of pharmacophoric descriptors coupled to a novel feature transformation technique, called Weight-Matrix Learning (WML, based on a feed-forward neural network). The application concerns virtual screening on a tyrosine kinase named BCR-ABL. First, the compounds were described using three different families of descriptors: our new pharmacophoric descriptors, and two circular fingerprints, ECFP4 and FCFP4. Afterwards, each of these original molecular representations were transformed using either an unsupervised WML method or a supervised one. Finally, using these transformed representations, K-Means clustering algorithm was applied to automatically partition the molecules. Combining our pharmacophoric descriptors with supervised Weight-Matrix Learning (SWMLR ) leads to clearly superior results in terms of several quality measures.
Maximum common substructures (MCS) have received a lot of attention in the chemoinformatics community. They are typically used as a similarity measure between molecules, showing high predictive performance when used in classification tasks, while being easily explainable substructures. In the present work, we applied the Pairwise Maximum Common Subgraph Feature Generation (PMCSFG) algorithm to automatically detect toxicophores (structural alerts) and to compute fingerprints based on MCS. We present a comparison between our MCS-based fingerprints and 12 well-known chemical fingerprints when used as features in machine learning models. We provide an experimental evaluation and discuss the usefulness of the different methods on mutagenicity data. The features generated by the MCS method have a state-of-the-art performance when predicting mutagenicity, while they are more interpretable than the traditional chemical fingerprints.
This paper presents a novel approach called Pharmacophore Activity Delta for extracting outstanding pharmacophores from a chemogenomic dataset, with a specific focus on a kinase target known as BCR-ABL. The method involves constructing a Hasse diagram, referred to as the pharmacophore network, by utilizing the subgraph partial order as an initial step, leading to the identification of pharmacophores for further evaluation. A pharmacophore is classified as a ‘Pharmacophore Activity Delta’ if its capability to effectively discriminate between active vs inactive molecules significantly deviates (by at least δ standard deviations) from the mean capability of its related pharmacophores. Among the 1479 molecules associated to BCR-ABL binding data, 130 Pharmacophore Activity Delta were identified. The pharmacophore network reveals distinct regions associated with active and inactive molecules. The study includes a discussion on representative key areas linked to different pharmacophores, emphasizing structure–activity relationships.
The purpose of pattern mining is to help experts understand their data. Following the assumption that an analyst expects neighbouring patterns to show similar behavior, we investigate the interestingness of a pattern given its neighborhood. We define a new way of selecting outstanding patterns, based on an order relation between patterns and a quality score. An outstanding pattern shows only small syntactic variations compared to its neighbors but deviates strongly in quality. Using several supervised quality measures, we show experimentally that only very few patterns turn out to be outstanding. We also illustrate our approach with patterns mined from molecular data.
This paper introduces a general method that can be used to create groups of pharmacophores to support their further in-depth analysis. A BCR-ABL molecular dataset was used to calculate graph edit distances between pharmacophores and led to their organization into a novel pharmacophore network. The application of a graph layout algorithm allowed us to discriminate between the pharmacophores associated with active compounds and those associated with inactive compounds. A clustering approach was used to refine the partitioning by grouping the pharmacophores based on their structures, activities, and binding modes. Analysis of a newly spatialized pharmacophore network provided us with critical insight into structure-activity relationships, most notably those that revealed distinctions between activity classes and chemical families. As shown, this method permits us to identify families of structurally homogeneous pharmacophores.
All along the drug development process, one of the most frequent adverse side effects, leading to the failure of drugs, is the cardiac arrhythmias. Such failure is mostly related to the capacity of the drug to inhibit the human ether-à-go-go-related gene (hERG) cardiac potassium channel. The early identification of hERG inhibition properties of biological active compounds has focused most of attention over the years. In order to prevent the cardiac side effects, a great number of in silico, in vitro and in vivo assays have been performed. The main goal of these studies is to understand the reasons of these effects, and then to give information or instructions to scientists involved in drug development to avoid the cardiac side effects. To evaluate anticipated cardiovascular effects, early evaluation of hERG toxicity has been strongly recommended for instance by the regulatory agencies such as U.S. Food and Drug Administration (FDA) and European Medicines Agency (EMA). Thus, following an initial screening of a collection of compounds to find hits, a great number of pharmacomodulation studies on the novel identified chemical series need to be performed including activity evaluation towards hERG. We provide in this concise review clear guidelines, based on described examples, illustrating successful optimization process to avoid hERG interactions as cases studies and to spur scientists to develop safe drugs.
All along the drug development process, one of the most frequent adverse side effects, leading to the failure of drugs, is the cardiac arrhythmias. Such failure is mostly related to the capacity of the drug to inhibit the human ether-a-go-go-related gene (hERG) cardiac potassium channel. The early identification of hERG inhibition properties of biological active compounds has focused most of attention over the years. In order to prevent the cardiac side effects, a great number of in silico, in vitro and in vivo assays have been performed. The main goal of these studies is to understand the reasons of these effects, and then to give information or instructions to scientists involved in drug development to avoid the cardiac side effects. To evaluate anticipated cardiovascular effects, early evaluation of hERG toxicity has been strongly recommended for instance by the regulatory agencies such as U.S. Food and Drug Administration (FDA) and European Medicines Agency (EMA). Thus, following an initial screening of a collection of compounds to find hits, a great number of pharmacomodulation studies on the novel identified chemical series need to be performed including activity evaluation towards hERG. We provide in this concise review clear guidelines, based on described examples, illustrating successful optimization process to avoid hERG interactions as cases studies and to spur scientists to develop safe drugs. (C) 2020 Elsevier Masson SAS. All rights reserved.