RNA helicases are essential, dynamic proteins that consume ATP to unwind and rearrange RNA. These activities place RNA helicases in roles as central mediators of signaling, especially those pathways dependent on RNA metabolism. Their binding of both ATP and RNA, as well as limited literature examples of small molecule ligands, support the tractability of RNA helicases. We employed structure-based virtual screening to rationally identify ligands that occupy the ATP-binding site of three human DExD/H-box RNA helicases: MDA5, LGP2, and DDX1. Following alignment of the well conserved nucleotide binding pocket for these RNA helicases, we docked and refined the list of potential ligands from the MolPort-2022-03 ligand library of ∼3.7 million members. A chemical lead with favorable solubility emerged from the 144 purchased compounds, which were evaluated in MDA5, LGP2, and DDX1 ATPase assays as well as corresponding SPR assays. It was found to be ATP un-competitive for MDA5 and to have similar affinity for the three RNA helicases. Analogs of the lead compound were designed to optimize the potency and selectivity of the scaffold, yielding both pan-helicase inhibitors and other analogs that are biased toward MDA5 inhibition.
Chemical pollution is a global threat to human health, yet the toxicity mechanisms of most contaminants remains unknown. Here, we applied an ultrahigh-throughput affinity selection-mass spectrometry (AS-MS) platform to systematically identify protein targets of prioritized chemical contaminants. After benchmarking the platform, we screened 50 human proteins against 481 prioritized chemicals, including 446 ToxCast chemicals and 35 per- and polyfluoroalkyl substances (PFAS). Among 24,050 interactions assessed, we discovered 35 interactions involving 13 proteins, with fatty acid-binding proteins (FABPs) emerging as the most ligandable protein family. Given this, we selected FABPs for further validation, which revealed a distinct PFAS binding pattern: legacy PFAS selectively bound to FABP1, whereas replacement compounds, perfluoroether carboxylic acids, unexpectedly interacted with all FABPs. X-ray crystallography further revealed that the ether group enhances the molecular flexibility of alternative PFAS to accommodate the binding pockets of FABPs. Our findings demonstrate that AS-MS is a robust platform for the discovery of protein targets beyond the scope of ToxCast and highlight the broader protein-binding spectrum of alternative PFAS as potential regrettable substitutes.
Humans are exposed to complex mixtures of environmental chemicals, raising two questions: how many chemicals may elicit toxicity at realistic exposure levels, and whether toxicity is driven by the mixture effects of numerous chemicals. Here, we propose a probabilistic model addressing both questions using large-scale chemical–protein interaction data from affinity selection mass spectrometry. Across 3.51 million interactions from 407,271 compounds and 110 human proteins, we observe a log-log relationship between binding probability and binding affinity. This model predicts that human toxicity is dominated by a small number (~100) of potent chemicals and their analogues under a Pareto law, rather than by the cumulative contributions of numerous weak compounds. These findings challenge the long-standing emphasis on mixture effects and instead support a ‘key driver’ hypothesis.
Machine learning (ML) is increasingly used in DNA-encoded library (DEL) screening for ligand discovery, but its success depends on access to suitable data sets, which are often proprietary and costly. To overcome this, we present the first fully open, automated DEL-ML framework using public DEL data sets and chemical fingerprints to enable reproducible, accessible drug discovery. Our workflow─from model training to virtual screening and compound selection─requires no human intervention. As a proof of concept, we identified binders for WDR91 by training ML models on the HitGen OpenDEL library (3B molecules) and screening the Enamine REAL Space library (37B molecules), yielding 50 candidates. Experimental testing confirmed seven novel binders with dissociation constants between 2.7-21 μM. Our open-source approach matches the performance of proprietary methods, demonstrating that public DEL data can support robust ML-driven ligand discovery and fostering transparency and broader community participation in drug development.
We report an enantioselective protein affinity selection mass spectrometry screening approach (E-ASMS) that enables the detection of weak binders, informs on selectivity, and generates orthogonal confirmation of binding. After method development with control proteins, we screen 31 human proteins against a designed library of 8,217 chiral compounds. We identify 16 binders to 12 targets, including many proteins predicted to be "challenging to ligand", and confirm their interactions through orthogonal biophysical assays. Seven binders to six targets display enantioselective binding, with KD values ranging from 3 to 20 µM. Binders for four targets (DDB1, WDR91, WDR55, and HAT1) are selected for in-depth characterization using X-ray crystallography. In all four cases, the mechanisms underlying enantioselectivity are readily explained. These results demonstrate that E-ASMS enables the identification and characterization of selective and weakly binding ligands for novel protein targets with unprecedented throughput and sensitivity.
The leucine-rich repeat kinase 2 (LRRK2) is the most mutated gene in familial Parkinson's disease, and its mutations lead to pathogenic hallmarks of the disease. The LRRK2 WDR domain is an understudied drug target for Parkinson's disease, with no known inhibitors prior to the first phase of the Critical Assessment of Computational Hit-Finding Experiments (CACHE) Challenge. A unique advantage of the CACHE Challenge is that the predicted molecules are experimentally validated in-house. Here, we report the design and experimental confirmation of LRRK2 WDR inhibitor molecules. We used an active learning (AL) machine learning (ML) workflow based on optimized free-energy molecular dynamics (MD) simulations utilizing the thermodynamic integration (TI) framework to expand a chemical series around two of our previously confirmed hit molecules. We identified 8 experimentally verified novel inhibitors out of 35 experimentally tested (23% hit rate). These results demonstrate the efficacy of our free-energy-based active learning workflow to explore large chemical spaces quickly and efficiently while minimizing the number and length of expensive simulations. This workflow is widely applicable to screening any chemical space for small-molecule analogs with increased affinity, subject to the general constraints of RBFE calculations. The mean absolute error of the TI MD calculations was 2.69 kcal/mol, with respect to the measured KD of hit compounds.
Protein arginine methyltransferase 9 (PRMT9) is part of the PRMT family, and it is suspected to function in pathways relevant to neurodevelopment. It is thought to participate in alternative splicing through interactions with the splicing factor SF3B2 (SAP145). In this study, we report 26 families (35 individuals) with bi-allelic loss-of-function variants in PRMT9, implicating PRMT9 in an autosomal-recessive human disease. Individuals primarily present with a neurodevelopmental disorder characterized by global developmental delay, learning disabilities, mild to severe intellectual disability, autism spectrum disorder, epilepsy, and hypotonia. The mutation spectrum includes 26 different variants such as frameshifting indels, nonsense variants, missense variants, and two copy-number variants. Mapping of the disease-causing missense variants onto the crystal structure of PRMT9 revealed that several of the variants reside within the catalytically active module of PRMT9, likely impairing its methyltransferase activity and resulting in a loss of function. In skin fibroblasts derived from affected individuals, we observed reduced expression at the RNA and/or protein level and subsequent aberrant methylation activity. Moreover, transcriptomic analysis of fibroblasts from affected individuals indicated differential expression of genes related to intellectual disability, autism, and cilia, suggesting a role of PRMT9 during ciliogenesis. Under ciliogenesis conditions, the skin-derived fibroblasts exhibited anomalies in the length of primary cilia but normal amounts of cilia. In addition, a prmt9 knockout zebrafish model displayed abnormal social preference in adult animals. Altogether, our findings implicate bi-allelic PRMT9 loss-of-function variants as causal for neurodevelopmental disorders.
53BP1 is a DNA damage response protein recruited to sites of double strand breaks through recognition of dimethylated lysine on histone 4 by its tandem Tudor domains. Like 53BP1, BRCA-1 plays a role in the regulation of DNA repair pathways, and BRCA-1 mutations have been strongly linked to breast and ovarian cancer. Interestingly, mice null for 53BP1 and BRCA-1 genes display minimal tumor formation, suggesting that the effects of deleterious BRCA-1 mutations could be prevented with potent 53BP1 small molecule antagonists. Herein, we describe a fragment screen that was used to identify compounds that bind to the 53BP1 Tudor domain and a chemoinformatic workflow to select near-neighbor analogues and establish structure activity relationships for these binders. The marked affinity improvements of the analogues over their parent fragments highlights the developability of these series and the utility of this approach in discovering novel hit compounds for 53BP1 and other methyl-lysine reader proteins.
Critical Assessment of Computational Hit-Finding Experiments (CACHE) Challenges emerged as real-life stress tests for computational hit-finding strategies. In CACHE Challenge #1, 23 participants contributed their original workflows to identify small-molecule ligands for the WD40 repeat (WDR) of LRRK2, a promising Parkinson's target. We applied the FRASE-based hit-finding robot (FRASE-bot), a platform for interaction-based screening allowing a drastic reduction of the explorable chemical space and a concurrent detection of putative ligand-binding sites. In two screening rounds, 84 compounds were procured for experimental testing and 8 were confirmed to bind LRRK2-WDR with dissociation constants (K d) ranging from 3 to 41 μM. To investigate the functional effect of WDR ligands, they were tested for their ability to modify the LRRK2 activity markers in HEK293T cells. Two compounds showed statistically significant increases in the kinase activity of WT LRRK2, and two compounds affected the conformation and kinase activity of major LRRK2 mutants.
Human DExD/H-box RNA helicases are ubiquitous molecular motors that unwind and rearrange RNA secondary structures in an ATP-dependent manner. These enzymes play essential roles in nearly all aspects of RNA metabolism. While their biological functions are well-characterized, the kinetic mechanisms remain relatively understudied in vitro. In this study, we describe the development and optimization of a bioluminescence-based assay to characterize the ATPase activity of three human RNA helicases: MDA5, LGP2, and DDX1. The assays were conducted using annealed 24-mer ds-RNA (blunt-ended double-stranded RNA) or double-stranded RNA with a 25-nt 3ʹ overhang (partial ds-RNA). These findings establish a robust and high-throughput in vitro assay suitable for a 384-well format, enabling the discovery and characterization of inhibitors targeting MDA5, LGP2, and DDX1. This work provides a valuable resource for advancing our understanding of these helicases and their therapeutic potential in Alzheimer's disease.
Human DExD/H-box RNA helicases are ubiquitous molecular motors that unwind and rearrange RNA secondary structures in an ATP-dependent manner. These enzymes play essential roles in nearly all aspects of RNA metabolism. While their biological functions are well-characterized, the kinetic mechanisms remain relatively understudied in vitro. In this study, we describe the development and optimization of a bioluminescence-based assay to kinetically characterize three human RNA helicases: MDA5, LGP2, and DDX1. The assays were conducted using annealed 24-mer RNA (blunt-ended double-stranded RNA) or double-stranded RNA (ds-RNA) with a 25-nt 3' overhang. These findings establish a robust and high-throughput in vitro assay suitable for a 384-well format, enabling the discovery and characterization of inhibitors targeting MDA5, LGP2, and DDX1. This work provides a valuable resource for advancing our understanding of these helicases and their therapeutic potential in Alzheimer's disease.
Damaged DNA-binding protein-1 (DDB1)- and CUL4-associated factor 12 (DCAF12) serves as the substrate recognition component within the Cullin4-RING E3 ligase (CRL4) complex, capable of identifying C-terminal double-glutamic acid degrons to promote the degradation of specific substrates through the ubiquitin proteasome system. Melanoma-associated antigen 3 (MAGEA3) and T-complex protein 1 subunit epsilon (CCT5) proteins have been identified as cellular targets of DCAF12. To further characterize the interactions between DCAF12 and both MAGEA3 and CCT5, we developed a suite of biophysical and proximity-based cellular NanoBRET assays showing that the C-terminal degron peptides of both MAGEA3 and CCT5 form nanomolar affinity interactions with DCAF12 in vitro and in cells. Furthermore, we report here the 3.17 & Aring; cryo-EM structure of DDB1-DCAF12-MAGEA3 complex revealing the key DCAF12 residues responsible for C-terminal degron recognition and binding. Our study provides new insights and tools to enable the discovery of small molecule handles targeting the WD40-repeat domain of DCAF12 for future proteolysis targeting chimera design and development.
Target class-focused drug discovery has a strong track record in pharmaceutical research, yet public domain data indicate that many members of protein families remain unliganded. Here we present a systematic approach to scale up the discovery and characterization of small molecule ligands for the WD40 repeat (WDR) protein family. We developed a comprehensive suite of protocols for protein production, crystallography, and biophysical, biochemical, and cellular assays. A pilot hit-finding campaign using DNA-encoded chemical library selection followed by machine learning (DEL-ML) to predict ligands from virtual libraries yielded first-in-class, drug-like ligands for 7 of the 16 WDR domains screened, thus demonstrating the broader ligandability of WDRs. This study establishes a template for evaluation of protein family wide ligandability and provides an extensive resource of WDR protein biochemical and chemical tools, knowledge, and protocols to discover potential therapeutics for this highly disease-relevant, but underexplored target class.
Protein class-focused drug discovery has a long and successful history in pharmaceutical research, yet most members of druggable protein families remain unliganded, often for practical reasons. Here we combined experiment and computation to enable discovery of ligands for WD40 repeat (WDR) proteins, one of the largest human protein families. This resource includes expression clones, purification protocols, and a comprehensive assessment of the druggability for hundreds of WDR proteins. We solved 21 high resolution crystal structures, and have made available a suite of biophysical, biochemical, and cellular assays to facilitate the discovery and characterization of small molecule ligands. To this end, we use the resource in a hit-finding pilot involving DNA-encoded library (DEL) selection followed by machine learning (ML). This led to the discovery of first-in-class, drug-like ligands for 9 of 20 targets. This result demonstrates the broad ligandability of WDRs. This extensive resource of reagents and knowledge will enable further discovery of chemical tools and potential therapeutics for this important class of proteins.### Competing Interest StatementBLS, BG, JSD, JZ, JWC, MvR, PR, SK, and TK are employees and shareholders of Relay Therapeutics.* BLI : Biolayer Interferometry DEL : DNA Encoded Library DLID : drug-like density DSF : Differential Scanning Fluorimetry FP : Fluorescence Polarization GCNN : Graph Convolutional Neural Network HDX : Hydrogen-Deuterium eXchange HTS : High-Throughput Screening ML : Machine Learning NanoBRET : NanoLuciferase Bioluminescence Resonance Energy Transfer PPI : Protein-Protein Interaction PROTAC : Proteolysis Targeting Chimera SPR : Surface Plasmon Resonance Tm : Protein melting temperature WDR : Tryptophan-Aspartate Repeat
The CACHE challenges are a series of prospective benchmarking exercises to evaluate progress in the field of computational hit-finding. Here we report the results of the inaugural CACHE challenge in which 23 computational teams each selected up to 100 commercially available compounds that they predicted would bind to the WDR domain of the Parkinson's disease target LRRK2, a domain with no known ligand and only an apo structure in the PDB. The lack of known binding data and presumably low druggability of the target is a challenge to computational hit finding methods. Of the 1955 molecules predicted by participants in Round 1 of the challenge, 73 were found to bind to LRRK2 in an SPR assay with a KD lower than 150 μM. These 73 molecules were advanced to the Round 2 hit expansion phase, where computational teams each selected up to 50 analogs. Binding was observed in two orthogonal assays for seven chemically diverse series, with affinities ranging from 18 to 140 μM. The seven successful computational workflows varied in their screening strategies and techniques. Three used molecular dynamics to produce a conformational ensemble of the targeted site, three included a fragment docking step, three implemented a generative design strategy and five used one or more deep learning steps. CACHE #1 reflects a highly exploratory phase in computational drug design where participants adopted strikingly diverging screening strategies. Machine learning-accelerated methods achieved similar results to brute force (e.g., exhaustive) docking. First-in-class, experimentally confirmed compounds were rare and weakly potent, indicating that recent advances are not sufficient to effectively address challenging targets.
Recent advances in DNA-encoded library (DEL) screening have created bioactivity datasets containing billions of molecules, unlocking new opportunities for machine learning (ML) in drug discovery. However, most ultra-large DEL libraries are proprietary, limiting the advancement of ML tools for big chemical data analytics and hindering the democratization of DEL-ML technology. We address this gap by developing an open, end-to-end DEL-ML framework using public datasets, where enriched binders are represented by common chemical fingerprints, ensuring proprietary data protection. We demonstrate that ML models can be built and validated on fingerprinted DEL data and then applied to virtual screening (VS) of billion-sized, publicly accessible chemical libraries. As a proof-of-concept, we screened the human protein WDR91 using the HitGen OpenDEL library (3 billion molecules) and trained ML models, which were used to screen the Enamine REAL Space library (37 billion molecules). Fifty potential binders were identified, 48 of which were tested, and seven were confirmed as novel binders with dissociation constants (KD) from 2.7 to 21 μM that were successfully co-crystalized with WDR91. This fully automated, open-source workflow demonstrates the potential of DEL-ML models in discovering novel binders and promotes the use of open chemical bioactivity datasets and ML to accelerate drug discovery.
We herewith applied a priori a generic hit identification method (POEM) for difficult targets of known three-dimensional structure, relying on the simple knowledge of physicochemical and topological properties of a user-selected cavity. Searching for local similarity to a set of fragment-bound protein microenvironments of known structure, a point cloud registration algorithm is first applied to align known subpockets to the target cavity. The resulting alignment then permits us to directly pose the corresponding seed fragments in a target cavity space not typically amenable to classical docking approaches. Last, linking potentially connectable atoms by a deep generative linker enables full ligand enumeration. When applied to the WD40 repeat (WDR) central cavity of leucine-rich repeat kinase 2 (LRRK2), an unprecedented binding site, POEM was able to quickly propose 94 potential hits, five of which were subsequently confirmed to bind in vitro to LRRK2-WDR.
WD40 repeat-containing protein 91 (WDR91) regulates early-to-late endosome conversion and plays vital roles in endosome fusion, recycling, and transport. WDR91 was recently identified as a potential host factor for viral infection. We employed DNA-encoded chemical library (DEL) selection against the WDR domain of WDR91, followed by machine learning to predict ligands from the synthetically accessible Enamine REAL database. Screening of predicted compounds identified a WDR91 selective compound 1, with a KD of 6 ± 2 μM by surface plasmon resonance. The co-crystal structure confirmed the binding of 1 to the WDR91 side pocket, in proximity to cysteine 487, which led to the discovery of covalent analogues 18 and 19. The covalent adduct formation for 18 and 19 was confirmed by intact mass liquid chromatography-mass spectrometry. The discovery of 1, 18, and 19, accompanying structure-activity relationship, and the co-crystal structures provide valuable insights for designing potent and selective chemical tools against WDR91 to evaluate its therapeutic potential.
ABSTRACT Cbl-b is a RING-type E3 ubiquitin ligase that is expressed in several immune cell lineages, where it negatively regulates the activity of immune cells. Cbl-b has specifically been identified as an attractive target for cancer immunotherapy due to its role in promoting an immunosuppressive tumor environment, and Nx-1607, is in phase I clinical trials for advanced solid tumor malignancies. Using a suite of biophysical and cellular assays, we confirmed potent binding of C7683 (an analogue of Nx-1607) to the full-length Cbl-b and its N-terminal fragment containing the TKBD-LHR-RING domains. To further elucidate its mechanism of inhibition, we determined the co-crystal structure of Cbl-b with C7683, revealing compound interaction with both the TKBD and LHR, but not the ring domain. Here, we provide structural insights into a novel mechanism of Cbl-b inhibition by a small-molecule inhibitor that locks the protein in an inactive conformation by acting as an intramolecular glue.