ABSTRACT Proteolysis Targeting Chimeras (PROTACs) are heterobifunctional molecules that bring a ubiquitin E3 ligase into proximity of a target protein to polyubiquitinate and degrade the target. PROTACs act catalytically, offering distinct advantages over conventional inhibitors and are the subject of intense study. The development of PROTACs involves extensive optimization of the chemical moiety linking two different protein-binding chemotypes, often requiring the synthesis, purification and testing of hundreds of PROTAC candidates. We used this approach to rapidly explore the landscape of targeted degradation of four different targets in parallel, combining and comparing a recently reported FBXO22-recruiting chemical warhead with warheads for the commonly used CRBN and VHL E3 ligases. Using a limited number of compounds (175 compounds in total) we observed no FBXO22-dependent degradation of these four targets. However, our libraries generated potent FBXO22 homo-PROTACs inducing self-degradation, as well as CRBN- and VHL-mediated degraders of FBXO22.
The third Critical Assessment of Computational Hit-finding Experiments (CACHE) challenged computational teams to identify chemically novel ligands targeting the macrodomain 1 of SARS-CoV-2 Nsp3, a promising coronavirus drug target. Twenty-three groups deployed diverse design strategies to collectively select 1739 ligand candidates. While over 85% of the designed molecules were chemically novel, the best experimentally confirmed hits were structurally similar to previously published compounds. Confirming a trend observed in CACHE #1 and #2, two of the best-performing workflows used compounds selected by physics-based computational screening methods to train machine learning models able to rapidly screen large chemical libraries, while four others used exclusively physics-based approaches. Three pharmacophore searches and one fragment growing strategy were also part of the seven winning workflows. While active molecules discovered by CACHE #3 participants largely mimicked the adenine ring of the endogenous substrate, ADP-ribose, preserving the canonical chemotype commonly observed in previously reported Nsp3-Mac1 ligands, they still provide novel structure-activity relationship insights that may inform the development of future antivirals. Collectively, these results show that multiple molecular design strategies can efficiently converge on similar potent molecules.
Hydrogen-Deuterium eXchange Mass Spectrometry (HDX-MS) is a dynamics-sensitive structural method that has rapidly achieved widespread adoption in the biopharmaceuticals industry, owing to its ability to quickly identify binding sites and to provide a molecular mechanism of action for drug candidates. However, a central limitation of conventional HDX-MS is that it has substantially lower spatial resolution than other structural techniques, typically providing information averaged over "segments" of five amino acids or more. Here, we demonstrate a sensitive, broadly applicable method for single amino acid resolution, i.e., site-specific HDX-MS measurements. Using a set of five therapeutic candidates targeting the highly druggable cancer target WDR5, we explore the greatly enhanced analytical power that arises from site specificity, including binding mode characterization, affinity ranking, and the detection of features that are "silent" in conventional peptide-level HDX-MS experiments.
Chemical pollution is a global threat to human health, yet the toxicity mechanisms of most contaminants remains unknown. Here, we applied an ultrahigh-throughput affinity selection-mass spectrometry (AS-MS) platform to systematically identify protein targets of prioritized chemical contaminants. After benchmarking the platform, we screened 50 human proteins against 481 prioritized chemicals, including 446 ToxCast chemicals and 35 per- and polyfluoroalkyl substances (PFAS). Among 24,050 interactions assessed, we discovered 35 interactions involving 13 proteins, with fatty acid-binding proteins (FABPs) emerging as the most ligandable protein family. Given this, we selected FABPs for further validation, which revealed a distinct PFAS binding pattern: legacy PFAS selectively bound to FABP1, whereas replacement compounds, perfluoroether carboxylic acids, unexpectedly interacted with all FABPs. X-ray crystallography further revealed that the ether group enhances the molecular flexibility of alternative PFAS to accommodate the binding pockets of FABPs. Our findings demonstrate that AS-MS is a robust platform for the discovery of protein targets beyond the scope of ToxCast and highlight the broader protein-binding spectrum of alternative PFAS as potential regrettable substitutes.
Targeted protein degradation (TPD) through the ubiquitin-proteasome system is driven by compound-mediated polyubiquitination of a protein-of-interest by an E3 ubiquitin (Ub) ligase. Relatively few E3s have been successfully utilized for TPD and the governing principles of functional ternary complex formation between the E3, degrader, and protein target remain elusive. FBXO22 has recently been harnessed for TPD applications by degraders that covalently modify its cysteine residues. Here, we reveal that the aldehyde derivative of UNC10088 promotes cooperative binding of FBXO22 to NSD2, a histone methyltransferase and oncogenic protein, leading to a cryo-EM structure of the SKP1-CUL1-F-box (SCF)-FBXO22 complex with NSD2. This structure revealed a conformational change in the FBXO22 loop surrounding C326, further exposing the cysteine for covalent recruitment. Additional medicinal chemistry efforts led to the discovery of benzaldehyde-based non-prodrug degraders that similarly engage C326 of FBXO22 and potently degrade NSD2. Unlike many degraders, our molecules recruit NSD2 to a different surface of FBXO22 than the known FBXO22 substrate BACH1, allowing for concurrent complex formation and structural determination of SCFFBXO22 bound to both the neosubstrate NSD2 and native substrate BACH1. Overall, we demonstrate the biochemical and structural basis for NSD2 degradation, revealing key principles for efficient and selective TPD by SCFFBXO22.
Acute myeloid leukemia (AML) is a hematopoietic malignancy caused by abnormal proliferation and differentiation of blasts. PRMT5, a methyltransferase that catalyzes symmetric dimethylation of arginine (SDMA) residues, has been implicated in cancer stem cell homeostasis and shown to be a potential therapeutic target in AML. However, given the toxicity of complete PRMT5 inhibition, there is a need to identify effective synergistic therapies. Through a targeted screen of compounds that inhibit key nodes of PRMT5-regulated pathways, we identified a synthetic lethality between inhibition of PRMT5 and LSD1, a lysine demethylase known to affect AML blast differentiation. The two inhibitors broadly reshape the transcriptome of targeted cells and synergize to promote AML differentiation and eventually growth inhibition and apoptosis, in a p53-dependent manner. To leverage this synthetic lethal interaction, we generated new dual compounds to inhibit both enzymes and recapitulated the effects of the drug combination. Our results uncover an unexpected convergence of PRMT5- and LSD1-regulated targets, paving the way for new therapeutic opportunities.
Short linear motifs (SLiMs) within intrinsically disordered protein regions mediate transient interactions crucial for cell physiology1. However, the global interaction landscape of human SLiMs remains largely uncharted. Here we present the Atlas of SLiM-mediated Human protein-protein Interactions (ASHI), which maps more than 20,000 interactions by screening over 800 human protein domains against a library of one million peptides tiling the human disordered proteome. ASHI expands the SLiM interactome, uncovers novel binding modes for known peptide-binding domains, and reveals unexpected peptide-binding activities in enzymes, chaperones, RNA-binding proteins, and modification-reader domains. Furthermore, intrinsically disordered regions emerge as densely encoded interaction platforms where interaction specificity is governed by diverse mechanisms, including key motif determinants, flanking residues, competition, and multivalency. These data provide an unprecedented foundation for modeling dynamic interaction networks, interpreting disease-associated variants, and decoding the dark proteome.
The MYC family of transforming oncogenes functions as regulators of gene transcription and is composed of three members, MYC, MYCN, and MYCL. As c-MYC (MYC) is deregulated in >50% of human cancers, the role, regulation, and structural features of MYC have been well-studied. By contrast, L-MYC has been understudied, as historically, oncogenic deregulation was evident only in a subset of lung carcinomas. However, recent deep genomic analyses of primary patient samples have demonstrated that L-MYC is deregulated in numerous human cancers. With this revelation it is important to understand how L-MYC compares to MYC at the structural level, particularly for the development of broad-spectrum inhibitors of the MYC family. Here, we show that L-MYC expression is elevated in several primary patient tumor samples, providing further evidence for L-MYC as a driver oncoprotein in primary human cancers. Next, we provide new biophysical insights into an N-terminal region within the L-MYC transactivation domain, which harbors two regions conserved amongst MYC proteins: MYC box 0 (MB0) and MYC box I (MBI). Nuclear magnetic resonance spectroscopy of residues 1-80 of L-MYC confirms that it is largely intrinsically disordered and interacts with the known MYC MB0 interactor, PNUTS (phosphatase 1 nuclear targeting subunit). Comparatively, L-MYC does not interact with the MYC-MBI interactor and tumor suppressor Bin1 (bridging integrator 1). Together, these results further substantiate the oncogenic role of L-MYC in human cancer and enhance our understanding of the biophysical nature of L-MYC to improve strategies for the development of anti-cancer therapeutics targeting MYC family members.
All folded proteins continuously fluctuate between their low-energy native structures and higher energy conformations that can be partially or fully unfolded. These rare states influence protein function, interactions, aggregation, and immunogenicity, yet they remain far less understood than protein native states. Although native protein structures are now often predictable with impressive accuracy, conformational fluctuations and their energies remain largely invisible and unpredictable, and experimental challenges have prevented large-scale measurements that could improve machine learning and physics-based modeling. Here, we introduce a multiplexed experimental approach to analyze the energies of conformational fluctuations for hundreds of protein domains in parallel using intact protein hydrogen-deuterium exchange mass spectrometry. We analyzed 5,778 domains 28-64 amino acids in length, revealing hidden variation in conformational fluctuations even between sequences sharing the same fold and global folding stability. Site-resolved hydrogen exchange NMR analysis of 13 domains showed that these fluctuations often involve entire secondary structural elements with lower stability than the overall fold. Computational modeling of our domains identified structural features that correlated with the experimentally observed fluctuations, enabling us to design mutations that stabilized low-stability structural segments. Our dataset enables new machine learning-based analysis of protein energy landscapes, and our experimental approach promises to reveal these landscapes at unprecedented scale.
Ubiquitin-specific peptidase 16 (USP16) is a deubiquitinase that specifically cleaves ubiquitin from histone H2A and modulates gene expression, cell cycle regulation, and various other cellular processes. The USP16 zinc-finger ubiquitin-binding domain (UBD) binds the free C-terminal end of both ubiquitin and ISG15, two major signaling proteins that mediate many biological pathways, but the precise function of USP16-UBD remains unclear. A small-molecule antagonist targeting USP16-UBD could enable cellular studies to elucidate its biological role. Here we report SGC-UBD1031 (15), a chemical probe targeting USP16-UBD with similar in vitro binding profiles to HDAC6-UBD and selectivity over nine other UBDs. In cellular assays, 15 disrupts the interaction between the C-terminus of ISG15 and USP16-UBD, as well as the interaction between ISG15 and HDAC6-UBD, at a concentration of 1 μM. The corresponding enantiomer SGC-UBD1031N (16) does not interfere with these interactions and thus serves as a negative control.
Hydrogen-Deuterium eXchange Mass Spectrometry (HDX-MS) is a dynamics-sensitive structural method that has rapidly achieved widespread adoption in the biopharmaceuticals industry, owing to its ability to quickly identify binding sites and to provide a molecular mechanism of action for drug candidates. However, a central limitation of conventional HDX-MS is that it has substantially lower spatial resolution than other structural techniques, typically providing information averaged over "segments" of five amino acids or more. Here, we demonstrate a sensitive, broadly applicable method for single amino acid resolution, i.e., site-specific HDX-MS measurements. Using a set of five therapeutic candidates targeting the highly druggable cancer target WDR5, we explore the greatly enhanced analytical power that arises from site specificity, including binding mode characterization, affinity ranking, and the detection of features that are "silent" in conventional peptide-level HDX-MS experiments.
Glioblastoma (GBM) is an aggressive brain cancer with a poor survival rate. Despite hundreds of clinical trials, there is no effective targeted therapy. Glioblastoma stem cells (GSCs) are an important GBM model system. In culture, these cells form spatial structures that share morphological aspects with their source tumors. We collected 17,000 phase-contrast images of 15 patient-derived GSC lines growing to confluence. We find that GSCs grow in characteristic multicellular patterns depending on their transcriptional state. Interpretable computer vision algorithms identified specific image features that predict transcriptional state across multiple cell confluency levels. This relationship will be useful in developing GSC screens where image features can be used to identify how GSC biology changes in response to perturbations simply by imaging cultured cells on plates.
Artificial intelligence (AI) and machine learning (ML) can transform early-stage drug discovery by enabling data-driven hit identification for target proteins at scale. However, progress remains constrained by the limited availability of large, open, standardized, and reusable protein-ligand interaction datasets for training, benchmarking, and reproducibility. Modern screening technologies such as DNA-encoded libraries (DEL) and affinity selection mass spectrometry (AS-MS), including enantioselective AS-MS (E-ASMS), can generate binding data for libraries ranging from hundreds of thousands to billions of compounds, yet these datasets are often fragmented, inconsistently processed, poorly annotated, and difficult to reuse for AI/ML. We introduce AIRCHECK (Artificial Intelligence Ready CHEmiCal Knowledge-base), an open, cloud-based data and computing infrastructure that converts DEL and E-ASMS screening outputs, together with downstream experimental validation assays when available, into curated, standardized, AI-ready datasets. AIRCHECK applies shared standard operating procedures, automated, cloud-native pipelines for data validation, harmonized labelling, and chemical featurization, producing downloadable data packages with rich metadata and multiple molecular fingerprints and descriptors that can be accessed at scale via application programming interfaces. Beyond data access, AIRCHECK provides open-source, containerized AI/ML workflows with persistent experiment tracking and benchmarking, as well as baseline AI/ML models to support training, evaluation, and deployment across all included targets. As of January 2026, AIRCHECK hosts curated screening data for 114 protein targets (~350 GB of screening data and ~125 TB of associated chemical representations). Developed within the Target 2035 open-science framework, AIRCHECK enables collaborative, reproducible virtual screening and experimental follow-up to accelerate and democratize hit discovery across the human proteome.
Small cell lung cancer (SCLC) is a highly aggressive form of cancer, commonly treated with DNA-damaging therapies such as chemotherapy and radiotherapy. Unfortunately, relapse occurs early and frequently, suggesting that epigenetic mechanisms may play a role in this aggressive behavior. Targeting these mechanisms during initial treatment could potentially enhance anti-cancer effects. This study investigated the combination of DNA-damaging treatments with a panel of Epigenetic Chemical Probes (EpiProbes). Among these, MS023, a PRMT inhibitor, showed the greatest synergy with cisplatin and etoposide across various SCLC cell lines. The cytotoxicity of MS023 was correlated with PRMT1 gene expression and protein levels. BioID analysis revealed that many PRMT1 interactors are involved in mRNA splicing. Mechanistic validation demonstrated that MS023 impaired RNA splicing, increased DNA:RNA hybrids, and caused DNA double-strand breaks (DSBs). When combined with ionizing radiation (IR), MS023 significantly increased DSBs, as indicated by γH2AX foci. Additionally, MS023 enhanced the effects of IR and the PARP inhibitor talazoparib, both in vitro and in vivo. Therefore, targeting PRMT1 in combination with DNA-damaging therapies presents a promising strategy to improve treatment outcomes for SCLC.
A critical assessment of computational hit-finding experiments (CACHE) challenge was conducted to predict ligands for the SARS-CoV-2 Nsp13 helicase RNA binding site, a highly conserved COVID-19 target. Twenty-three participating teams comprised of computational chemists and data scientists used protein structure and data from fragment-screening paired with advanced computational and machine learning methods to each predict up to 100 inhibitory ligands. Across all teams, 1957 compounds were predicted and were subsequently procured from commercial catalogs for biophysical assays. Of these compounds, 0.7% were confirmed to bind to Nsp13 in a surface plasmon resonance assay. The six best-performing computational workflows used fragment growing, active learning, or conventional virtual screening with and without complementary deep-learning scoring functions. Follow-up functional assays resulted in identification of two compound scaffolds that bound Nsp13 with a Kd below 10 μM and inhibited in vitro helicase activity. Overall, CACHE #2 participants were successful in identifying hit compound scaffolds targeting Nsp13, a central component of the coronavirus replication-transcription complex. Computational design strategies recurrently successful across the first two CACHE challenges include linking or growing docked or crystallized fragments and docking small and diverse libraries to train ultrafast machine-learning models. The CACHE #2 competition reveals how crowd-sourcing ligand prediction efforts using a distinct array of approaches followed with critical biophysical assays can result in novel lead compounds to advance drug discovery efforts.
Here we describe ProtacID, a flexible BioID (proximity-dependent biotinylation)-based approach to identify PROTAC-proximal proteins in living cells. ProtacID analysis of VHL- and CRBN-recruiting PROTACs targeting a number of different proteins (localized to chromatin or cellular membranes, and tested across six different human cell lines) demonstrates how this technique can be used to validate PROTAC degradation targets and identify non-productive (i.e. non-degraded) PROTAC-interacting proteins, addressing a critical need in the field of PROTAC development. We also demonstrate that ProtacID can be used to characterize native, endogenous multiprotein complexes without the use of antibodies, or modification of the protein of interest with epitope tags or biotin ligase tagging.
SARS-CoV-2 relies on host proteases to prime its spike protein for cell entry through either the endosomal or plasma membrane pathway. Although cysteine cathepsins are known to mediate the endosomal route, the identity of the dominant enzyme has remained unclear. Here, we identify human Cathepsin K (hCatK), a lysosomal cysteine protease, as a previously unrecognized yet functionally important mediator of spike activation. While human Cathepsin L (hCatL) has long been regarded as the principal endosomal protease for spike processing, inhibition of hCatK with the selective inhibitor Odanacatib suppressed viral infection in endothelial cells as effectively as the broad-spectrum cysteine protease inhibitor E-64d, implicating hCatK as a key driver of spike processing during the endosomal viral entry. Comprehensive enzymatic profiling demonstrated that hCatK exhibits 24- to 63-fold higher catalytic efficiency toward the Furin-cleavage site (FCS) sequence than hCatL and displays a distinct substrate-recognition pattern at the Omicron FCS relative to the Wuhan variant. We further demonstrate that hCatK is an off-target of Nirmatrelvir, a clinically approved 3CL-Mpro inhibitor, with a sub-micromolar potency (IC50 = 0.6 ± 0.1 µM). A 1.9 Å crystal structure of the hCatK–Nirmatrelvir complex delineates the molecular basis of inhibitor binding and supports the rational design of dual-acting antivirals. Collectively, these findings redefine the landscape of host proteases involved in SARS-CoV-2 spike activation and establish hCatK as a previously overlooked but strategic target for antiviral intervention. ### Competing Interest Statement The authors have declared no competing interest. The Natural Sciences and Engineering Research Council of Canada, RGPIN 2019-06720 Canadian Institutes of Health Research, PJT 155979
Expansion of the CAG trinucleotide repeat tract in exon 1 of the Huntingtin (HTT) gene causes Huntington's disease (HD) through the expression of a polyglutamine-expanded form of the HTT protein. This mutation triggers cellular and biochemical pathologies, leading to cognitive, motor, and psychiatric symptoms in HD patients. Targeting HTT splicing with small molecule drugs is a compelling approach to lowering HTT protein levels to treat HD, and splice modulators are currently being tested in the clinic. Here, we identify PRMT5 as a novel regulator of HTT messenger RNA (mRNA) splicing and alternative polyadenylation. PRMT5 inhibition disrupts the splicing of HTT introns 9 and 10, leading to the activation of multiple proximal intronic polyadenylation sites within these introns and promoting premature termination, cleavage, and polyadenylation of the HTT mRNA. This suggests that HTT protein levels may be lowered due to this mechanism. We also detected increasing levels of these truncated HTT transcripts across a series of neuronal differentiation samples, which correlated with lower PRMT5 expression. Notably, PRMT5 inhibition in glioblastoma stem cells potently induced neuronal differentiation. We posit that PRMT5-mediated regulation of intronic polyadenylation, premature termination, and cleavage of the HTT mRNA modulates HTT expression and plays an important role during neuronal differentiation.