Neisseria gonorrhoeae is a common Gram-negative pathogen with increasing resistance to all recommended antibiotics. There is a critical need to improve the efficiency of the antibiotic hit discovery process to replenish the drug development pipeline. Here, we show that deep learning models can augment high-throughput screens to identify readily available molecules with narrow-spectrum activity against difficult-to-treat strains of N. gonorrhoeae. We phenotypically tested 38,650 small molecules for N. gonorrhoeae growth inhibition to train a predictive graph neural network (GNN) model. We benchmarked the model's performance against other architectures, including a large language model, and found that GNNs more accurately identify active, drug-like molecules that are structurally distinct from the training set and known antibiotics. Using the model to virtually screen ~6 million compounds, we identified 213 compounds for experimental validation and found that 83 (39%) inhibited N. gonorrhoeae growth. Two of these compounds were structurally dissimilar to existing antibiotics, maintained potency against multidrug-resistant N. gonorrhoeae strains in vitro, exhibited promising selectivity indices, and were rapidly bactericidal with low frequencies of resistance. Proteomic studies revealed their distinct mechanisms of action, with one compound targeting alanine racemase, an enzyme involved in the essential process of peptidoglycan synthesis. Furthermore, the compounds showed early promise in reducing N. gonorrhoeae titers in a human vagina-on-a-chip infection model and a mouse vaginal infection model. Our work establishes the deep learning-enabled discovery of selective antibacterial compounds against N. gonorrhoeae as a much-needed hit discovery tool to address the growing crisis of antimicrobial resistance for this pathogen.
Abstract Hair thinning arises from multi-faceted dysfunction within the hair follicle, driven by both intrinsic cellular pathways and pathways responding to extrinsic hormonal and microenvironmental cues. Here, we present an AI-enabled discovery framework to discover small molecules that promote hair follicle rejuvenation. This framework integrates graph neural networks trained on phenotypic screening data with structure-based virtual screening to prioritize compounds that modulate complementary biological pathways. Through AI-enabled screening, hit-to-lead optimization, and medicinal chemistry, we identified four compounds that increase follicle dermal papilla cell viability, stabilize hypoxia signaling by inhibiting prolyl hydroxylase domain protein 2 (PHD2), and suppress androgen-mediated follicular miniaturization by inhibiting 5α-reductases (5-ARs). RNA sequencing analyses confirmed pathway engagement, and functional validation across primary cells and a 3D hair follicle organoid model demonstrated high activity and cellular specificity. The lead compounds were incorporated into a water-based formulation, where they demonstrated robust solubility and combinatorial efficacy to increase sprouting length of follicle organoids. These results establish an AI-enabled platform for discovering multi-pathway modulators of hair follicle rejuvenation.
Gastrulation, a critical developmental stage involving germ layer specification and axes formation, is a major point of failure in human development, contributing to pregnancy loss and congenital malformations. However, due to ethical constraints and anatomical differences in animal models, the failure modes underlying human gastrulation remain poorly understood. To elucidate these failure modes, we introduce FATE-MAP (Failure Analysis and Trajectory Evaluation via Mechanistic-AI Prediction), an integrated platform that combines high-throughput perturbations of human 2D gastruloids with quantitative phenotypic mapping, predictive deep learning, and mechanistic morphogen modeling. Analyzing over 2000 drug-treated human 2D gastruloids, we mapped a phenotypic morphospace that separates canonical patterning, in which primitive-streak fates are correctly specified and radially organized, from failure modes, defined as departures from this organization and marked by a loss of a required fate and/or radial symmetry. To predict and interpret patterning outcomes, FATE-MAP combines a transformer linking chemical structure to phenotype with PDE simulations of morphogen transport and cell fate specification, and projects both outputs onto the experimentally defined morphospace. Applying this framework, we flagged two clinical molecules as potential teratogens and identified two parameters, cell density and SOX2 stability, that form orthogonal morphospace axes along which canonically patterned gastruloids systematically vary. FATE-MAP thus provides a roadmap for decoding human developmental trajectories and accelerating safe therapeutic discovery.
The antimicrobial resistance crisis necessitates structurally distinct antibiotics. While deep learning approaches can identify antibacterial compounds from existing libraries, structural novelty remains limited. Here, we developed a generative artificial intelligence framework for designing de novo antibiotics through two approaches: a fragment-based method to comprehensively screen >107 chemical fragments in silico against Neisseria gonorrhoeae or Staphylococcus aureus, subsequently expanding promising fragments, and an unconstrained de novo compound generation, each using genetic algorithms and variational autoencoders. Of 24 synthesized compounds, seven demonstrated selective antibacterial activity. Two lead compounds exhibited bactericidal efficacy against multidrug-resistant isolates with distinct mechanisms of action and reduced bacterial burden in vivo in mouse models of N. gonorrhoeae vaginal infection and methicillin-resistant S. aureus skin infection. We further validated structural analogs for both compound classes as antibacterial. Our approach enables the generative deep-learning-guided design of de novo antibiotics, providing a platform for mapping uncharted regions of chemical space.
Understanding the physiological effects of antibiotics on bacterial cells is important for informing antibiotic use. Bacterial communities treated with antibiotics in a microfluidic device maintain glucose consumption at the community periphery, protecting interior cells from the effects of antibiotics.
RNAs represent a class of programmable biomolecules capable of performing diverse biological functions. Recent studies have developed accurate RNA three-dimensional structure prediction methods, which may enable new RNAs to be designed in a structure-guided manner. Here, we develop a structure-to-sequence deep learning platform for the de novo generative design of RNA aptamers. We show that our approach can design RNA aptamers that are predicted to be structurally similar, yet sequence dissimilar, to known light-up aptamers that fluoresce in the presence of small molecules. We experimentally validate several generated RNA aptamers to have fluorescent activity, show that these aptamers can be optimized for activity in silico, and find that they exhibit a mechanism of fluorescence similar to that of known light-up aptamers. Our results demonstrate how structural predictions can guide the targeted and resource-efficient design of new RNA sequences. A deep learning platform for structure-guided, generative design of RNA sequences is developed and used to discover fluorescent RNA aptamers.
Artificial intelligence (AI) and machine learning (ML) models are being deployed in many domains of society and have recently reached the field of drug discovery. Given the increasing prevalence of antimicrobial resistance, as well as the challenges intrinsic to antibiotic development, there is an urgent need to accelerate the design of new antimicrobial therapies. Antimicrobial peptides (AMPs) are therapeutic agents for treating bacterial infections, but their translation into the clinic has been slow owing to toxicity, poor stability, limited cellular penetration and high cost, among other issues. Recent advances in AI and ML have led to breakthroughs in our abilities to predict biomolecular properties and structures and to generate new molecules. The ML-based modelling of peptides may overcome some of the disadvantages associated with traditional drug discovery and aid the rapid development and translation of AMPs. Here, we provide an introduction to this emerging field and survey ML approaches that can be used to address issues currently hindering AMP development. We also outline important limitations that can be addressed for the broader adoption of AMPs in clinical practice, as well as new opportunities in data-driven peptide design.
Nonlinear biomolecular interactions on membranes drive membrane remodeling crucial for biological processes including chemotaxis, cytokinesis, and endocytosis. The complexity of biomolecular interactions, their redundancy, and the importance of spatiotemporal context in membrane organization impede understanding of the physical principles governing membrane mechanics. Developing a minimal in vitro system that mimics molecular signaling and membrane remodeling while maintaining physiological fidelity poses a major challenge. Inspired by chemotaxis, we reconstructed chemically regulated actin polymerization inside vesicles, guiding membrane self-organization. An external, undirected chemical input induced directed actin polymerization and membrane deformation uncorrelated with upstream biochemical cues, suggesting symmetry breaking. A biophysical model incorporating actin dynamics and membrane mechanics proposes that uneven actin distributions cause nonlinear membrane deformations, consistent with experimental findings. This protocellular system illuminates the interplay between actin dynamics and membrane shape during symmetry breaking, offering insights into chemotaxis and other cell biological processes.
Accurate prediction of RNA three-dimensional (3D) structures remains an unsolved challenge. Determining RNA 3D structures is crucial for understanding their functions and informing RNA-targeting drug development and synthetic biology design. The structural flexibility of RNA, which leads to the scarcity of experimentally determined data, complicates computational prediction efforts. Here we present RhoFold+, an RNA language model-based deep learning method that accurately predicts 3D structures of single-chain RNAs from sequences. By integrating an RNA language model pretrained on ~23.7 million RNA sequences and leveraging techniques to address data scarcity, RhoFold+ offers a fully automated end-to-end pipeline for RNA 3D structure prediction. Retrospective evaluations on RNA-Puzzles and CASP15 natural RNA targets demonstrate the superiority of RhoFold+ over existing methods, including human expert groups. Its efficacy and generalizability are further validated through cross-family and cross-type assessments, as well as time-censored benchmarks. Additionally, RhoFold+ predicts RNA secondary structures and interhelical angles, providing empirically verifiable features that broaden its applicability to RNA structure and function studies.
Deep learning approaches have been increasingly applied to the discovery of novel chemical compounds. These predictive approaches can accurately model compounds and increase true discovery rates, but they are typically black box in nature and do not generate specific chemical insights. Explainable deep learning aims to 'open up' the black box by providing generalizable and human-understandable reasoning for model predictions. These explanations can augment molecular discovery by identifying structural classes of compounds with desired activity in lieu of lone compounds. Additionally, these explanations can guide hypothesis generation and make searching large chemical spaces more efficient. Here we present an explainable deep learning platform that enables vast chemical spaces to be mined and the chemical substructures underlying predicted activity to be identified. The platform relies on Chemprop, a software package implementing graph neural networks as a deep learning model architecture. In contrast to similar approaches, graph neural networks have been shown to be state of the art for molecular property prediction. Focusing on discovering structural classes of antibiotics, this protocol provides guidelines for experimental data generation, model implementation and model explainability and evaluation. This protocol does not require coding proficiency or specialized hardware, and it can be executed in as little as 1-2 weeks, starting from data generation and ending in the testing of model predictions. The platform can be broadly applied to discover structural classes of other small molecules, including anticancer, antiviral and senolytic drugs, as well as to discover structural classes of inorganic molecules with desired physical and chemical properties.
An artificial-intelligence graph neural network was trained on experimental data and used to identify chemical substructures that underlie selective antibiotic activity in more than 12 million compounds. This led to the discovery of a class of antibiotics with in vitro and in vivo activity against Gram-positive bacteria, including Staphylococcus aureus.
The design choices underlying machine-learning (ML) models present important barriers to entry for many biologists who aim to incorporate ML in their research. Automated machine-learning (AutoML) algorithms can address many challenges that come with applying ML to the life sciences. However, these algorithms are rarely used in systems and synthetic biology studies because they typically do not explicitly handle biological sequences (e.g., nucleotide, amino acid, or glycan sequences) and cannot be easily compared with other AutoML algorithms. Here, we present BioAutoMATED, an AutoML platform for biological sequence analysis that integrates multiple AutoML methods into a unified framework. Users are automat-ically provided with relevant techniques for analyzing, interpreting, and designing biological sequences. BioAutoMATED predicts gene regulation, peptide-drug interactions, and glycan annotation, and designs optimized synthetic biology components, revealing salient sequence characteristics. By auto-mating sequence modeling, BioAutoMATED allows life scientists to incorporate ML more readily into their work.
We used graph neural networks trained on experimental data to identify senolytic compounds from vast chemical libraries of over 800,000 compounds and discovered structurally diverse senolytics that have potent in vitro and in vivo activity, as well as favorable medicinal chemistry properties.
The discovery of novel structural classes of antibiotics is urgently needed to address the ongoing antibiotic resistance crisis1-9. Deep learning approaches have aided in exploring chemical spaces1,10-15; these typically use black box models and do not provide chemical insights. Here we reasoned that the chemical substructures associated with antibiotic activity learned by neural network models can be identified and used to predict structural classes of antibiotics. We tested this hypothesis by developing an explainable, substructure-based approach for the efficient, deep learning-guided exploration of chemical spaces. We determined the antibiotic activities and human cell cytotoxicity profiles of 39,312 compounds and applied ensembles of graph neural networks to predict antibiotic activity and cytotoxicity for 12,076,365 compounds. Using explainable graph algorithms, we identified substructure-based rationales for compounds with high predicted antibiotic activity and low predicted cytotoxicity. We empirically tested 283 compounds and found that compounds exhibiting antibiotic activity against Staphylococcus aureus were enriched in putative structural classes arising from rationales. Of these structural classes of compounds, one is selective against methicillin-resistant S. aureus (MRSA) and vancomycin-resistant enterococci, evades substantial resistance, and reduces bacterial titres in mouse models of MRSA skin and systemic thigh infection. Our approach enables the deep learning-guided discovery of structural classes of antibiotics and demonstrates that machine learning models in drug discovery can be explainable, providing insights into the chemical substructures that underlie selective antibiotic activity.
The accumulation of senescent cells is associated with aging, inflammation and cellular dysfunction. Senolytic drugs can alleviate age-related comorbidities by selectively killing senescent cells. Here we screened 2,352 compounds for senolytic activity in a model of etoposide-induced senescence and trained graph neural networks to predict the senolytic activities of >800,000 molecules. Our approach enriched for structurally diverse compounds with senolytic activity; of these, three drug-like compounds selectively target senescent cells across different senescence models, with more favorable medicinal chemistry properties than, and selectivity comparable to, those of a known senolytic, ABT-737. Molecular docking simulations of compound binding to several senolytic protein targets, combined with time-resolved fluorescence energy transfer experiments, indicate that these compounds act in part by inhibiting Bcl-2, a regulator of cellular apoptosis. We tested one compound, BRD-K56819078, in aged mice and found that it significantly decreased senescent cell burden and mRNA expression of senescence-associated genes in the kidneys. Our findings underscore the promise of leveraging deep learning to discover senotherapeutics.
Despite advances in molecular biology, genetics, computation, and medicinal chemistry, infectious disease remains an ominous threat to public health. Addressing the challenges posed by pathogen outbreaks, pandemics, and antimicrobial resistance will require concerted interdisciplinary efforts. In conjunction with systems and synthetic biology, artificial intelligence (AI) is now leading to rapid progress, expanding anti-infective drug discovery, enhancing our understanding of infection biology, and accelerating the development of diagnostics. In this Review, we discuss approaches for detecting, treating, and understanding infectious diseases, underscoring the progress supported by AI in each case. We suggest future applications of AI and how it might be harnessed to help control infectious disease outbreaks and pandemics.
There is a need to discover and develop non-toxic antibiotics that are effective against metabolically dormant bacteria, which underlie chronic infections and promote antibiotic resistance. Traditional antibiotic discovery has historically favored compounds effective against actively metabolizing cells, a property that is not predictive of efficacy in metabolically inactive contexts. Here, we combine a stationary-phase screening method with deep learning-powered virtual screens and toxicity filtering to discover compounds with lethality against metabolically dormant bacteria and favorable toxicity profiles. The most potent and structurally distinct compound without any obvious mechanistic liability was semapimod, an anti-inflammatory drug effective against stationary-phase E . coli and A . baumannii. . Integrating microbiological assays, biochemical measurements, and single-cell microscopy, we show that semapimod selectively disrupts and permeabilizes the bacterial outer membrane by binding lipopolysaccharide. This work illustrates the value of harnessing non-traditional screening methods and deep learning models to identify non-toxic antibacterial compounds that are effective in infection-relevant contexts.
Efficient identification of drug mechanisms of action remains a challenge. Computational docking approaches have been widely used to predict drug binding targets; yet, such approaches depend on existing protein structures, and accurate structural predictions have only recently become available from AlphaFold2. Here, we combine AlphaFold2 with molecular docking simulations to predict protein-ligand interactions between 296 proteins spanning Escherichia coli's essential proteome, and 218 active antibacterial compounds and 100 inactive compounds, respectively, pointing to widespread compound and protein promiscuity. We benchmark model performance by measuring enzymatic activity for 12 essential proteins treated with each antibacterial compound. We confirm extensive promiscuity, but find that the average area under the receiver operating characteristic curve (auROC) is 0.48, indicating weak model performance. We demonstrate that rescoring of docking poses using machine learning-based approaches improves model performance, resulting in average auROCs as large as 0.63, and that ensembles of rescoring functions improve prediction accuracy and the ratio of true-positive rate to false-positive rate. This work indicates that advances in modeling protein-ligand interactions, particularly using machine learning-based approaches, are needed to better harness AlphaFold2 for drug discovery.
β-Lactam antibiotics disrupt the assembly of peptidoglycan (PG) within the bacterial cell wall by inhibiting the enzymatic activity of penicillin-binding proteins (PBPs). It was recently shown that β-lactam treatment initializes a futile cycle of PG synthesis and degradation, highlighting major gaps in our understanding of the lethal effects of PBP inhibition by β-lactam antibiotics. Here, we assess the downstream metabolic consequences of treatment of Escherichia coli with the β-lactam mecillinam and show that lethality from PBP2 inhibition is a specific consequence of toxic metabolic shifts induced by energy demand from multiple catabolic and anabolic processes, including accelerated protein synthesis downstream of PG futile cycling. Resource allocation into these processes is coincident with alterations in ATP synthesis and utilization, as well as a broadly dysregulated cellular redox environment. These results indicate that the disruption of normal anabolic-catabolic homeostasis by PBP inhibition is an essential factor for β-lactam antibiotic lethality.
Understanding how bactericidal antibiotics kill bacteria remains an open question. Previous work has proposed that primary drug-target corruption leads to increased energetic demands, resulting in the generation of reactive metabolic byproducts (RMBs), particularly reactive oxygen species, that contribute to antibiotic-induced cell death. Studies have challenged this hypothesis by pointing to antibiotic lethality under anaerobic conditions. Here, we show that treatment of Escherichia coli with bactericidal antibiotics under anaerobic conditions leads to changes in the intracellular concentrations of central carbon metabolites, as well as the production of RMBs, particularly reactive electrophilic species (RES). We show that antibiotic treatment results in DNA double-strand breaks and membrane damage and demonstrate that antibiotic lethality under anaerobic conditions can be decreased by RMB scavengers, which reduce RES accumulation and mitigate associated macromolecular damage. This work indicates that RMBs, generated in response to antibiotic-induced energetic demands, contribute in part to antibiotic lethality under anaerobic conditions.