Machine-learning models that predict bioactivity of novel compounds are dependent on the quality, quantity, and relatedness of their training data. For regression modelling, large amounts of publicly available lower-quality data need to be disregarded due to significant noise or incompatible endpoints. In this paper, we implement a pairwise auxiliary ranking task in a message-passing neural network for proteochemometrics regression bioactivity modelling. The auxiliary training task compares a relative ranking of the affinity of different compound pairs and runs in parallel to the standard regression training. This simple modification allows for the extension of training data with percentage-inhibition data points, increasing the chemical space available to regression models. We benchmark models on data sampled from standard sources, as well as on data related to the extended training data to characterize the influence of these modifications. Significant performance increases can be achieved on related data without sacrificing predictive power on standard tests. The pairwise auxiliary ranking task is flexible in the types of endpoints it can support and the regression problems it can address.
Predicting how chemical modifications affect drug binding is central to rational drug design. Free energy perturbation (FEP) calculations provide accurate estimates of these binding affinity changes, but existing methods often require substantial computational resources and expert knowledge. Here, we present QligFEP v2.1.0, a flexible open-source workflow based on a graphical and command-line interface for calculating relative binding free energies using spherical boundary conditions, which dramatically reduces simulation system size by confining simulations to a focused region around the binding site. QligFEP features a configurable restraint algorithm that automatically handles diverse chemical transformations, streamlined setup procedures, and enhanced analysis tools. We validated the method using industry benchmarks comprising 16 protein targets and 639 ligand transformations. Statistical analysis demonstrates that QligFEP achieves comparable accuracy to established commercial and open-source alternatives while requiring only a fraction of the computational resources. The perturbation protocol simulates ∼6250 atoms per perturbation leg and completes transformation replicates in under 2 h on standard computational clusters. Unlike full-system simulations, QligFEP's modest computational requirements make FEP accessible for less than $1 on current AWS spot instances. The combination of accuracy, flexibility, and computational efficiency positions QligFEP as a practical solution for accelerating compound optimization in drug discovery, making rigorous binding affinity predictions accessible for large scale applications and to research groups with limited computational infrastructure.
Current methods for determining the neurotoxic potential of (agro)chemicals are not comprehensive enough, as is suggested by the increased incidence of Parkinson's disease (PD) among people exposed to certain pesticides. Mechanism-based in silico screening can address this shortcoming by predicting molecular initiating event activation, a precursor of the adverse outcome. However, a limited amount of protein-binding data has been collected on pesticides, meaning that for screening, an approach is required that is well suited for extrapolation to a wide variety of chemicals. Here, the group I metabotropic glutamate receptors (mGluRs) were taken as a case study because of their role in chemical-induced PD. Compounds with known activity for these receptors were docked into the allosteric binding site, interaction fingerprints (IFPs) were computed, and used to train classification models. Afterward, model enrichment was evaluated, feature importance was derived, and the applicability domain was analyzed. Both IFP-based mGluR models demonstrated good enrichment with the area under the Receiver Operating Characteristic Curve (ROC AUC) being 0.78 and 0.66. The interactions that were most important for model predictions were hydrogen bond formation with Asn760 for mGluR1 and aromatic interactions with Trp785 for mGluR5. The applicability domain varied depending on the training set but was consistently larger for IFPs than for Morgan fingerprints, a fragment-based descriptor. A virtual screen using the final models identified 132 potential mGluR binders, of which one, bifenthrin, had been previously found to bind in in vitro experiments. To promote the implementation of the screening technique presented here, a platform was created where users can make predictions using our models. All in all, these results highlight the potential of combining IFPs with machine learning and contribute to the shift toward mechanism-based in silico toxicology.
Molecular generators enable the exploration of chemical space to identify novel compounds with desirable properties. However, assessing their performance remains challenging due to the structural diversity and volume of the generated molecules. Commonly used evaluation metrics, focusing on chemical validity and novelty, do not fully align with the primary goal of molecular generation: the discovery of new biologically active compounds. To address this limitation, we introduce scaffold-based metrics that enable a fair comparison by evaluating a generator’s ability to recover biologically relevant scaffolds absent from the input set. We applied the scaffold Recovery Score (RS), SEt scaffold Diversity (SED), and Absolute SEt scaffold Recall (ASER) metrics to compare several molecular generators, including Molpher, DrugEx, REINVENT, and Graph-based genetic algorithm. The proposed scaffold-based metrics provide a realistic framework for evaluating and optimizing molecular generators for their practical use in drug discovery scenarios, particularly in the design of focused virtual chemical libraries. The metrics are available as open-source in a GitHub repository at https://github.com/filvaleriia/scaffold-based-metrics.
Virtual screening (VS) is a powerful approach to exploring a vast chemical space, encompassing libraries of millions to billions of compounds. However, the low hit rates of VS require testing numerous candidates to validate true binders, followed by iterative optimization cycles, which makes experimental validation costly and time-consuming. Here, we report COMBINAUT, an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement. Using a faculty-wide collection of in-house building blocks, the system enables enumeration of over 22.9 million compounds, each designed for parallelized synthesis within 32 h using repurposed solid-phase peptide synthesis equipment. Using this platform, we performed large-scale VS targeting the allosteric pocket of the immuno-oncology target, C-C chemokine receptor 2 (CCR2). Our approach facilitated the rapid synthesis and testing of 100 VS hits spanning diverse molecular architectures. In radioligand binding assays, we successfully validated nine hits with distinct scaffolds, including completely novel CCR2 ligand chemotypes. Iterative hit-to-lead optimization using the automated workflow produced cell-active CCR2 antagonists. This work demonstrates the synergy of automated synthesis and VS, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.
Artificial intelligence (AI) is receiving increasing attention across the entire lifecycle of medicines, from early development to postauthorization use. While various AI tools have been developed in commercial and academic settings, the extent of their use in regulatory contexts within the European Union remains unknown. In this study, we systematically analyzed the use of AI for regulatory evidence by reviewing Public reports and internal development advice documents from the European Medicines Agency (EMA). Of 26,480 documents screened, 52 documents contained AI use, reflecting 43 unique AI tools used across quality, nonclinical, clinical, and pharmacovigilance domains. The majority of AI tools were deployed in clinical applications, and outputs were used either directly as endpoints, in shaping clinical trial methodology, or in ancillary analyses. Comments and advice from the EMA regarding AI use varied according to the context of use and covered aspects of documentation, methodology, validation, and lifecycle management. Our findings indicate that AI use in regulatory evidence is still limited but showing an increasing trend. As the field is still novel, there are limited regulatory precedents, and AI-specific guideline development and harmonization across regulatory jurisdictions are still ongoing. In this light, the quantitative characterization of AI tools and AI-related regulatory comments captured in this analysis provide concrete insights into AI use and frequent regulatory considerations, aiding in both AI tool and AI guideline development.
Virtual screening (VS) is a powerful approach to exploring a vast chemical space, encompassing libraries of millions to billions of compounds. However, the low hit rates of VS require testing numerous candidates to validate true binders, followed by iterative optimization cycles, which makes experimental validation costly and time-consuming. Here, we report COMBINAUT, an automated parallel synthesis platform that generates diverse chemical scaffolds to accelerate hit validation and refinement. Using a faculty-wide collection of in-house building blocks, the system enables enumeration of over 22.9 million compounds, each designed for parallelized synthesis within 32 h using repurposed solid-phase peptide synthesis equipment. Using this platform, we performed large-scale VS targeting the allosteric pocket of the immuno-oncology target, C-C chemokine receptor 2 (CCR2). Our approach facilitated the rapid synthesis and testing of 100 VS hits spanning diverse molecular architectures. In radioligand binding assays, we successfully validated nine hits with distinct scaffolds, including completely novel CCR2 ligand chemotypes. Iterative hit-to-lead optimization using the automated workflow produced cell-active CCR2 antagonists. This work demonstrates the synergy of automated synthesis and VS, enabling the efficient exploration of chemical space and the rapid discovery of novel ligands.
Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here, we apply sparse autoencoders (SAEs) to MolFormer, an encoder-only cLM, to mechanistically examine how molecular representations are built across layers. We discover that early layers rely on position-tracking latents to parse molecular grammar, while later layers encode atom-in-substructure and pharmacologically relevant features. Additionally, we show that non-canonical SMILES produce more disruptive representation shifts than invalid SMILES, driven by position-latent disruption propagating across layers. To support further exploration, we develop InterMol, an interactive visualizer for SAE activations on molecular strings and structures.
MOTIVATION:The generation and analysis of diverse mutants of a protein is a powerful tool for understanding protein function. However, generating such mutants can be time-consuming, while the commercial option of buying a series of mutant plasmids can be expensive. In contrast, the insertion of a synthesized double-stranded DNA (dsDNA) fragment into a plasmid is a fast and low-cost method to generate a large library of mutants with one or more point mutations, insertions, or deletions. RESULTS:To aid in the design of these DNA fragments, we have developed PyVADesign: a Python package that makes the design and ordering of dsDNA fragments straightforward and cost-effective. In PyVADesign, the mutations of interest are clustered in different cloning groups for efficient exchange into the target plasmid. Additionally, primers that prepare the target plasmid for insertion of the dsDNA fragment, as well as primers for sequencing, are automatically designed within the same program. AVAILABILITY AND IMPLEMENTATION:PyVADesign is open source and available at https://github.com/CDDLeiden/PyVADesign and archived via Zenodo (https://doi.org/10.5281/zenodo.15057525).
Drug-induced liver injury (DILI) presents a critical challenge in drug development, often leading to the withdrawal of promising therapeutic candidates. Traditional predictive models for DILI, typically relying on molecular descriptors and pharmacokinetic properties, are insufficient due to the complex and multifactorial nature of liver toxicity. This complexity stems from overlapping biological stress responses activated by both hepatotoxic and non-hepatotoxic compounds, making it difficult to distinguish between them accurately. Additionally, the scarcity of DILI-positive compounds in available datasets results in significant class imbalance, further limiting the efficacy of conventional predictive models. Addressing these challenges requires novel approaches incorporating molecular and bioactivity data to enhance predictive power. In this study, we developed a custom oversampling strategy tailored to handle DILI's biological complexity and class imbalance. We integrated stress pathway activations, particularly focusing on the oxidative, unfolded protein, DNA damage, heat shock, and cytokine signalling stress responses, with molecular descriptors and bioactivity profiles to improve model performance. The custom oversampling technique demonstrated improved specificity and overall predictive accuracy, mitigating the effects of class imbalance without overfitting. Despite these advances, significant challenges remain in refining predictive models, particularly in identifying the most informative biological markers and optimising experimental protocols for better data acquisition. Our results suggest that while incorporating diverse data types and novel oversampling strategies improves DILI prediction, further efforts are required to create robust, generalisable models capable of reliably predicting hepatotoxicity in the drug development process.
Pancreatic ductal adenocarcinoma (PDAC) is an aggressive malignancy with a 5-year survival rate of approximately 5-7%, and complete surgical resection remains the only curative treatment but is often unfeasible. Fluorescence-guided surgery (FGS) using tumor-targeted probes may improve tumor visualization and facilitate complete resection. This study aimed to identify and validate tumor targets for FGS during PDAC resection procedures. RNA expression data from over 4000 cell surface genes, obtained from public genomic databases, were analyzed to identify genes encoding PDAC-associated proteins. Eleven potential tumor targets were identified, including CEACAM5, TMPRSS4, COL17A1, CLDN18, and AQP5. Protein expression was evaluated by immunohistochemistry (IHC) in tissues from 44 PDAC and 7 chronic pancreatitis (CP) patients. All targets, except COL17A1, showed significantly higher expression in PDAC tissue compared to healthy pancreatic, CP, and duodenal tissue (p < 0.001), as well as in tumor-positive versus tumor-negative lymph nodes. Especially CEACAM5, TMPRSS4, and AQP5 were identified as the most promising targets for distinguishing PDAC from healthy tissues and detecting lymph node metastasis during FGS. The development of probes targeting multiple markers, such as AQP5 with CEACAM5 and/or TMPRSS4, may help overcome interpatient variability and enhance detection across patients.
CC chemokine receptor (CCR) 2 and 5 are G protein-coupled receptors that play a crucial role in immunohomeostasis. Accordingly, overactivation of their signaling pathways is involved in various immunopathologies and cancer. Extensive research focusing on discovering CCR2 and CCR5 orthosteric antagonists, ultimately resulted in some clinical success, but the area of intracellular allosteric modulators is still underexplored and the move from orthosteric to allosteric modulation could be an interesting paradigm shift. To this end, we document the development of novel CCR2 and CCR5 intracellular allosteric antagonists through a virtual screen on a small combinatorial library derived from existing CCR2, CCR5, and CCR4 ligands. Using a molecular docking approach, the created library was screened in its entirety utilizing a refined AlphaFold model of CCR5 based on the crystal structure of its close homologue, CCR2. The screening resulted in the identification of several virtual hits, out of which one was developed further by in-house synthesis. In total, 18 analogues were prepared and experimentally evaluated for their binding affinity for CCR2 and functional inhibition on CCR5. This expeditious and simple workflow beginning from docking to compound evaluation identified 3 hits for CCR2 (Ki = 1.3-6 μM) and 1 hit (IC50 = 10.8 μM) for CCR5. The obtained structure-activity relationships were also further rationalized using structural information available for both CCR5 and CCR2 providing valuable insights for future development of intracellular allosteric ligands.
Applications of machine learning in chemistry are often limited by the scarcity and expense of labeled data, restricting traditional supervised methods. In this work, we introduce a framework for molecular reasoning using general-purpose Large Language Models (LLMs) that operates without requiring labeled training data. Our method anchors chain-of-thought reasoning to the molecular structure by using unique atomic identifiers. First, the LLM performs a one-shot task to identify relevant fragments and their associated chemical labels or transformation classes. In an optional second step, this position-aware information is used in a few-shot task with provided class examples to predict the chemical transformation. We apply our framework to single-step retrosynthesis, a task where LLMs have previously underperformed. Across academic benchmarks and expert-validated drug discovery molecules, our work enables LLMs to achieve high success rates in identifying chemically plausible reaction sites (≥90%), named reaction classes (≥40%), and final reactants (≥74%). Beyond solving complex chemical tasks, our work also provides a method to generate theoretically grounded synthetic datasets by mapping chemical knowledge onto the molecular structure and thereby addressing data scarcity.
The unbound brain-to-plasma partition coefficient (Kp,uu,BBB) is an essential parameter for predicting central nervous system (CNS) drug disposition using physiologically-based pharmacokinetic (PBPK) modeling. Kp,uu,BBB values for specific compounds are however often unavailable, and are moreover time consuming to obtain experimentally. The aim of this study was to develop a quantitative structure–property relationship (QSPR) model to predict the Kp,uu,BBB and to demonstrate how QSPR-model predictions can be integrated into a physiologically-based pharmacokinetic model for the CNS. Rat Kp,uu,BBB values were obtained for 98 compounds from literature or in house historical data. For all compounds, 2D and 3D physico-chemical and structural properties were derived using the Molecular Operating Environment (MOE) software. Multiple machine learning (ML) regression models were compared for prediction of the Kp,uu,BBB, including random forest, support vector machines, K-nearest neighbors, and (sparse-) partial least squares. Finally, we demonstrate how the developed QSPR model predictions can be integrated into a CNS PBPK modeling workflow. Among all ML algorithms, a random forest showed the best predictive performance for Kp,uu,BBB on test data with R2 value of 0.61 and 61
Uncertainty quantification (UQ) has been recognized as a prerequisite for reliable and trustworthy computational modeling in drug discovery. Two widely considered paradigms, Bayesian methods (deep ensemble and MC dropout) and evidential learning, differ in their computational demands and expressivity of uncertainties, excelling in complementary settings. Here, we propose hybrid approaches that combine both paradigms and benchmark them on the Papyrus++ data set across two end points (xC50, Kx) and multiple split strategies. Our ensemble of evidential models (EOE) consistently achieves the best overall performance, yielding the lowest RMSE and leading CRPS and interval scores, including under the most challenging distributional shifts. While large ensembles often excel in rejection-based utility, EOE matches or surpasses them at a fraction of the computational cost. Statistical tests confirm its advantage, and a hardware-agnostic compute analysis highlights favorable performance-efficiency trade-offs. These results demonstrate that combining evidential and Bayesian principles yields more accurate and informative uncertainties for bioactivity modeling, with EOE offering a robust─and computationally practical─default for uncertainty-aware decision-making in drug discovery.
The virtually monomorphic antigen presentation molecule HLA-E can present self- and non-self peptides to the NKG2A/CD94 co-receptor inhibitory complex expressed on natural killer (NK) cells and to T cell receptors (TCRs) expressed on T cells. HLA-E presents self-peptides to NKG2A/CD94 to regulate tissue homeostasis, whereas HLA-E restricted T cells mediate regulatory and cytotoxic responses toward pathogen-infected cells. In this study, we directly compared HLA-E/peptide recognition and signaling between NKG2A/CD94 and 2 HLA-E restricted TCRs that can recognize self-peptides or identical peptide mimics from the viral UL40 protein of cytomegalovirus using position substituted peptide variants. We show that position 7 is critical for interaction with NKG2A/CD94, whereas position 8 is important for interaction with the TCRs. The Arginine at position 5 of these peptides is an essential residue for recognition by both receptors. Thus, NKG2A/CD94 and TCRs have different requirements for recognition of peptides presented in HLA-E.
Alchemical free energy calculations are becoming an increasingly prevalent tool in drug discovery efforts. Over the past decade, significant progress has been made in automating various aspects of this technique. However, one aspect hampering wider application is the construction of perturbation networks to connect ligands of interest. More specifically, ligand pairs with large dissimilarities should be avoided since they can lower convergence and decrease accuracy. Here, we propose a technique for automatic generation of intermediate molecules to break up problematic edges─calculations connecting two different ligands or molecules─into smaller perturbations. To this end, a modular tool was developed that generates intermediates for a molecule pair by enumerating R-group combinations called IMERGE-FEP (Intermediate MolEculaR GEnerator for Free Energy Perturbation). Intermediate enumeration of multiple, representative congeneric series showed that intermediates increase similarity regarding shared substructures, geometry, and LOMAP scores. Taken together, this tool eases integration of intermediate steps into free energy calculation protocols.
Background: The escalating global crisis of antibiotic resistance necessitates the discovery of novel antimicrobial agents. Antimicrobial peptides (AMPs) represent a promising alternative to combat multidrug-resistant (MDR) pathogens. Because traditional AMP discovery is labour-intensive and costly, machine learning (ML) is applied to identify AMPs effective against MDR bacteria and skin infections. Methods: The ML-based CalcAMP model predicts the antimicrobial activity of 16,384 unique 14-amino-acid peptide sequences, resulting in a novel Guided Designed Smart antimicrobial Therapeutic (GDST) peptide catalogue. Parent sequences and retro-inverso (RI) variants of two prime GDST peptides undergo extensive testing against MDR bacteria and in skin infection models. Results: GDST-038 and GDST-045, along with their RI variants, show potent antimicrobial activity against Acinetobacter baumannii and Staphylococcus aureus, rapidly depolarizing the cytoplasmic membrane, exhibiting broad-spectrum bactericidal effects against ESKAPE pathogens, and causing minimal haemolysis. RI variants display superior A. baumannii biofilm killing compared to parent sequences, while all GDST peptides achieve >3-log reductions in S. aureus biofilm CFU within 24 h. Potent efficacy is observed in a 3D human skin epidermal infection model, with elimination of S. aureus at ≥15 μM. No resistance develops after 22 passages. Conclusions: ML-driven screening enables rapid identification of two novel candidate AMPs, highlighting the therapeutic potential of GDST peptides for MDR bacterial infections.
Background/Objectives: Preclinical models of liver fibrosis only partially mimic human disease processes. Particularly, traditional transforming growth factor beta 1 (TGFβ1)-induced hepatic stellate cell (HSC) models lack relevant processes, including hypoxia-induced pathways. Here, the ability of a hypoxia-mimicking compound (IOX2) to more accurately reflect the human fibrotic phenotype on a functional level was investigated. Methods: Human primary HSCs were stimulated (TGFβ1 +/- IOX2), and the cell viability and fibrotic phenotype were determined. The latter was assessed as protein levels of fibrosis markers-collagen, TIMP-1, and Fibronectin. Next-generation sequencing (NGS), differential expression analyses (DESeq2), and Ingenuity Pathway Analysis (IPA) were performed for mechanistic evaluation and biological annotation. Results: Stimulation with TGFβ1 + IOX2 significantly increased fibrotic marker levels. Also, fibrosis-related pathways were activated, and hypoxia-related genes and collagen modifications, such as crosslinking, increased dose-dependently. Comparative analysis with human fibrotic DEGs showed improved disease representation in the HSC model in the presence of IOX2. Conclusions: In conclusion, the HSC model better recapitulated liver fibrosis by IOX2 administration. Therefore, hypoxia-mimicking compounds hold promise for enhancing the translational value of in vitro fibrosis models, providing valuable insights in liver fibrosis pathogenesis and potential therapeutic strategies.