AGAPE (computAtional G-quadruplex Affinitiy PrEdiction) is an innovative machine learning (ML)-based workflow developed to predict the potential stabilization of small molecules for G-quadruplex (G4) structures. G4s, especially prevalent in telomeres and oncogene promoters, represent promising therapeutic targets, but designing selective binders remains challenging. This study addresses this gap by implementing an ML framework in KNIME, specifically designed for ease of use across the scientific community. The AGAPE workflow integrates 5666 classical and quantum chemical (QC) descriptors for G4 ligands, enabling a comprehensive representation of binding-relevant molecular features. Using data from G4LDB and in-house collections, we created a robust dataset of 1217 compounds categorized by Forster Resonance Energy Transfer (FRET) assays as ACTIVE or INACTIVE based on G4 stabilization capacity. Feature selection algorithms and ML models, particularly XGBoost with Naive Bayes and Random Forest classifiers, were employed to achieve an optimized prediction model with an accuracy close to 90%. A consensus voting system among top-performing models further improved classification reliability. AGAPE efficiently predicts G4 stabilization, offers interpretability of key chemical interactions, and provides a scalable, accessible tool for G4-focused drug discovery. This workflow lays the foundation for targeted therapeutic development, enhancing ligand selectivity, and expanding the application of AI in chemoinformatics ### Competing Interest Statement The authors have declared no competing interest.
Background Lysine demethylase enzymes (KDMs) are an emerging class of therapeutic targets, that catalyse the removal of methyl marks from histone lysine residues regulating chromatin structure and gene expression. KDM4A isoform plays an important role in the epigenetic dysregulation in various cancers and is linked to aggressive disease and poor clinical outcomes. Despite several efforts, the KDM4 family lacks successful specific molecular inhibitors. Results Herein, starting from a structure-based fragments virtual screening campaign we developed a synergic framework as a guide to rationally design efficient KDM4A inhibitors. Commercial libraries were used to create a fragments collection and perform a virtual screening campaign combining docking and pharmacophore approaches. The most promising compounds were tested in-vitro by a Homogeneous Time-Resolved Fluorescence-based assay developed for identifying selective substrate-competitive inhibitors by means of inhibition of H3K9me3 peptide demethylation. 2-(methylcarbamoyl)isonicotinic acid was identified as a preliminary active fragment, displaying inhibition of KDM4A enzymatic activity. Its chemical exploration was deeply investigated by computational and experimental approaches which allowed a rational fragment growing process. The in-silico studies guided the development of derivatives designed as expansion of the primary fragment hit and provided further knowledge on the structure–activity relationship. Conclusions Our study describes useful insights into key ligand-KDM4A protein interaction and provides structural features for the development of successful selective KDM4A inhibitors.
The COVID-19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small-molecule drugs that are widely available, including in low- and middle-income countries, is an ongoing challenge. In this work, we report the results of a community effort, the “Billion molecules against Covid-19 challenge”, to identify small-molecule inhibitors against SARS-CoV-2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 potentially active molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (Nsp12 domain), and (alpha) spike protein S. Overall, 27 potential inhibitors were experimentally confirmed by binding-, cleavage-, and/or viral suppression assays and are presented here. All results are freely available and can be taken further downstream without IP restrictions. Overall, we show the effectiveness of computational techniques, community efforts, and communication across research fields (i.e., protein expression and crystallography, in silico modeling, synthesis and biological assays) to accelerate the early phases of drug discovery.
The COVID‐19 pandemic continues to pose a substantial threat to human lives and is likely to do so for years to come. Despite the availability of vaccines, searching for efficient small‐molecule drugs that are widely available, including in low‐ and middle‐income countries, is an ongoing challenge. In this work, we report the results of an open science community effort, the “Billion molecules against COVID‐19 challenge”, to identify small‐molecule inhibitors against SARS‐CoV‐2 or relevant human receptors. Participating teams used a wide variety of computational methods to screen a minimum of 1 billion virtual molecules against 6 protein targets. Overall, 31 teams participated, and they suggested a total of 639,024 molecules, which were subsequently ranked to find ‘consensus compounds’. The organizing team coordinated with various contract research organizations (CROs) and collaborating institutions to synthesize and test 878 compounds for biological activity against proteases (Nsp5, Nsp3, TMPRSS2), nucleocapsid N, RdRP (only the Nsp12 domain), and (alpha) spike protein S. Overall, 27 compounds with weak inhibition/binding were experimentally identified by binding‐, cleavage‐, and/or viral suppression assays and are presented here. Open science approaches such as the one presented here contribute to the knowledge base of future drug discovery efforts in finding better SARS‐CoV‐2 treatments.
The family of protein kinases comprises more than 500 genes involved in numerous functions. Hence, their physiological dysfunction has paved the way toward drug discovery for cancer, cardiovascular, and inflammatory diseases. As a matter of fact, Kinase binding sites high similarity has a double role. On the one hand it is a critical issue for selectivity, on the other hand, according to poly-pharmacology, a synergistic controlled effect on more than one target could be of great pharmacological interest. Another important aspect of binding similarity is the possibility of exploit it for repositioning of drugs on targets of the same family. In this study, we propose our approach called Kinase drUgs mAchine Learning frAmework (KUALA) to automatically identify kinase active ligands by using specific sets of molecular descriptors and provide a multi-target priority score and a repurposing threshold to suggest the best repurposable and non-repurposable molecules. The comprehensive list of all kinase-ligand pairs and their scores can be found at https://github.com/molinfrimed/multi-kinases .
In recent years, the debate in the field of applications of Deep Learning to Virtual Screening has focused on the use of neural embeddings with respect to classical descriptors in order to encode both structural and physical properties of ligands and/or targets. The attention on embeddings with the increasing use of Graph Neural Networks aimed at overcoming molecular fingerprints that are short range embeddings for atomic neighborhoods. Here, we present EMBER, a novel molecular embedding made by seven molecular fingerprints arranged as different “spectra” to describe the same molecule, and we prove its effectiveness by using deep convolutional architecture that assesses ligands’ bioactivity on a data set containing twenty protein kinases with similar binding sites to CDK1. The data set itself is presented, and the architecture is explained in detail along with its training procedure. We report experimental results and an explainability analysis to assess the contribution of each fingerprint to different targets.
In the last decades, HOX proteins have been extensively studied due to their pivotal role in transcriptional events. HOX proteins execute their activity by exploiting a cooperative binding to PBX proteins and DNA. Therefore, an increase or decrease in HOX activity has been associated with both solid and haematological cancer diseases. Thus, inhibiting HOX-PBX interaction represents a potential strategy to prevent these malignancies, as demonstrated by the patented peptide HTL001 that is being studied in clinical trials. In this work, a computational study is described to identify novel potential peptides designed by employing a database of non-natural amino acids. For this purpose, residue scanning of the HOX minimal active sequence was performed to select the mutations to be further processed. According to these results, the peptides were point-mutated and used for Molecular Dynamics (MD) simulations in complex with PBX1 protein and DNA to evaluate complex binding stability. MM-GBSA calculations of the resulting MD trajectories were exploited to guide the selection of the most promising mutations that were exploited to generate twelve combinatorial peptides. Finally, the latter peptides in complex with PBX1 protein and DNA were exploited to run MD simulations and the ΔGbinding average values of the complexes were calculated. Thus, the analysis of the results highlighted eleven combinatorial peptides that will be considered for further assays.
: The Polycomb Repressive complex 2 (PRC2) maintains a repressive chromatin state and silences many genes, acting as methylase on histone tails. This enzyme was found overexpressed in many types of cancer. In this work, we have set up a Computer-Aided Drug Design approach based on the allosteric modulation of PRC2. In order to minimize the possible bias derived from using a single set of coordinates within the protein-ligand complex, a dynamic workflow was developed. In details, molecular dynamic was used as tool to identify the most significant ligand-protein interactions from several crystallized protein structures. The identified features were used for the creation of dynamic pharmacophore models and docking grid constraints for the design of new PRC2 allosteric modulators. Our protocol was retrospectively validated using a dataset of active and inactive compounds, and the results were compared to the classic approaches, through ROC curves and enrichment factor. Our approach suggested some important interaction features to be adopted for virtual screening performance improvement.
Coronavirus disease 2019 (COVID-19) has spread out as a pandemic threat affecting over 2 million people. The infectious process initiates via binding of SARS-CoV-2 Spike (S) glycoprotein to host angiotensin-converting enzyme 2 (ACE2). The interaction is mediated by the receptor-binding domain (RBD) of S glycoprotein, promoting host receptor recognition and binding to ACE2 peptidase domain (PD), thus representing a promising target for therapeutic intervention. Herein, we present a computational study aimed at identifying small molecules potentially able to target RBD. Although targeting PPI remains a challenge in drug discovery, our investigation highlights that interaction between SARS-CoV-2 RBD and ACE2 PD might be prone to small molecule modulation, due to the hydrophilic nature of the bi-molecular recognition process and the presence of druggable hot spots. The fundamental objective is to identify, and provide to the international scientific community, hit molecules potentially suitable to enter the drug discovery process, preclinical validation and development.
NLRP3 (NOD-like receptor family, pyrin domain-containing protein 3) activation has been linked to several chronic pathologies, including atherosclerosis, type-II diabetes, fibrosis, rheumatoid arthritis, and Alzheimer's disease. Therefore, NLRP3 represents an appealing target for the development of innovative therapeutic approaches. A few companies are currently working on the discovery of selective modulators of NLRP3 inflammasome. Unfortunately, limited structural data are available for this target. To date, MCC950 represents one of the most promising noncovalent NLRP3 inhibitors. Recently, a possible region for the binding of MCC950 to the NLRP3 protein was described but no details were disclosed regarding the key interactions. In this communication, we present an in silico multiple approach as an insight useful for the design of novel NLRP3 inhibitors. In detail, combining different computational techniques, we propose consensus-retrieved protein residues that seem to be essential for the binding process and for the stabilization of the protein-ligand complex.