Ubiquilin-2 (UBQLN2) is a ubiquitin (Ub)-binding shuttle protein that is mutated in X-linked forms of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD). ALS/FTD-linked mutations in UBQLN2 disrupt its conformation, increasing its tendency to form cytoplasmic aggregates that may disrupt cellular regulation through loss-of-function (LOF) and gain-of-function (GOF) effects. Here, we performed quantitative mass spectrometry (MS)-based interactome analysis of wild-type (UBQLN2WT) and ALS-mutant UBQLN2 (UBQLN2ALS) proteins using inducible pluripotent stem cells (iPSCs) and induced motor neurons (iMNs). Proteins showing enhanced association with UBQLN2ALS proteins included PEG10, a known degradation target of UBQLN2, and BAG6, a chaperone involved in the triage of mislocalized proteins (MLPs). BAG6 knockdown inhibited the solubility recovery of both UBQLN2WT and UBQLN2ALS proteins following heat stress (HS), suggesting it functions as a UBQLN2 holdase. In addition, knockdown of BAG6 or knockout of UBQLN2 led to PEG10 accumulation, implicating both in PEG10 turnover; however, neither BAG6 nor UBQLN2 was required for PEG10 degradation in response to HS. A highly aggregation-prone UBQLN24XALS mutant harboring four different ALS-associated mutations showed increased PEG10 binding and modestly delayed PEG10 turnover while PEG10 degradation was not significantly different between UBQLN2WT and iPSCs expressing a UBQLN2P497H clinical mutant. The combined findings implicate BAG6 as a UBQLN2 holdase and identify a suite of proteins whose altered binding may contribute to pathologic changes in UBQLN2-associated ALS/FTD.
The dehydroamino acids (DHAAs) dehydroalanine and dehydrobutyrine are formed in proteins posttranslationally from serine and threonine, respectively. They contain an electrophilic alkene, which reacts readily with nucleophiles by Michael addition, giving rise to a variety of chemical modifications as well as protein crosslinks. Recent work has shown DHAAs and their modifications to be highly prevalent in protein aggregates isolated from brains of Alzheimer's Disease (AD) patients. As there is no human enzyme known to catalyze the formation of DHAAs, this raised the question of how they are produced. We show here that metal oxides, particularly the iron oxides goethite and hematite, efficiently catalyze the formation of DHAAs in model phosphopeptides in vitro under physiologically relevant conditions. Iron oxides are known to accumulate in the human brain with age, and this deposition is enhanced in neurodegenerative disorders like AD. The DHAA formation by heterogeneous catalysis reported here provides a previously unknown mechanism linking iron deposition to protein crosslinking in AD and potentially other neurodegenerative diseases where iron accumulation and protein aggregation co-occur.
Alzheimer's disease (AD) is characterized by the accumulation of protein aggregates, which are thought to be influenced by posttranslational modifications (PTMs). Dehydroamino acids (DHAAs) are rarely observed PTMs that contain an electrophilic alkene capable of forming protein-protein crosslinks, which may lead to protein aggregation. We report here the discovery of DHAAs in the protein aggregates from AD, constituting an unknown and previously unsuspected source of extensive proteomic complexity. We used mass spectrometry-based proteomics to discover 404 sites of DHAA formation in 171 proteins from protein aggregate-enriched human brain samples, 6-fold more sites than observed in the soluble protein fractions. The DHAA modifications are observed both directly and in the form of conjugates after reacting with abundant cellular nucleophiles or crosslinking to nucleophilic amino acid residues. We report 11 such crosslinks, including three in the Tau protein, which are 10-fold more abundant in AD samples compared with age-matched controls. Many of the proteins found to contain DHAAs and their conjugates are involved in protein aggregation or pathways dysregulated in AD. DHAAs are prevalent modifications in the AD brain proteome and give rise to protein crosslinks that may contribute to protein aggregation.
Long noncoding RNAs (lncRNAs) exert regulatory functions in a wide spectrum of biological contexts, and certain regulatory functions involve the formation of RNA-protein complexes. Discovering the structure and function of these complexes may unveil important functional insights. The DDX41 gene encoding the DEAD-box RNA helicase 41 protein (DDX41) is subject to extensive germline genetic variation, and certain variants create a predisposition to develop myelodysplastic syndrome and acute myeloid leukemia. While the importance of DDX41 for the control of hematopoiesis is established, many questions remain regarding the mechanisms of how DDX41 functions in hematopoietic stem and progenitor cells. Previously, we identified a DDX41-regulated lncRNA, growth-arrest-specific 5 (Gas5). As the Gas5 function in hematopoiesis is unknown, we analyzed the protein interactors of Gas5 lncRNA using HyPR-MS (hybridization purification of RNA-protein complexes, followed by mass spectrometry). A total of 303 proteins were identified as Gas5 lncRNA interactors, five of which were experimentally validated as Gas5 lncRNA interactors by RNA immunoprecipitation qPCR (RIP-qPCR) analysis. The identification of protein interactors with a DDX41-regulated lncRNA establishes a foundation on which to guide future mechanistic and biological studies.
How RIF1 (RAP1 interacting factor 1) fulfills its diverse roles in DNA double-strand break repair, DNA replication, and nuclear organization remains elusive. Here, we show that alternative splicing of a cassette exon (Ex32) encoding a Ser/Lys-rich cassette in the RIF1 C-terminal domain (CTD) gives rise to RIF1-Long (RIF1-L) and RIF1-Short (RIF1-S) isoforms with different functional characteristics. We demonstrate that RIF1-Ex32 splice-in is mediated by an exonic splicing enhancer that is recognized by the serine and arginine rich splicing factor 1 (SRSF1) and antagonized by SRSF3 and SRSF7. Exposure to DNA damage inhibited Ex32 splice-in, potentiated the association of SRSF3 and SRSF7 with RIF1 pre-mRNA, and caused an increase in RIF1-S protein expression, which was also observed across a diverse set of primary cancers. Isoform-specific proteomic analyses revealed RIF1-L preferentially associated with mediator of DNA damage checkpoint 1 (MDC1) and sustained MDC1 focus formation to a greater extent than RIF1-S. We further show that the Ser/Lys-rich cassette stabilized a novel phase separation activity of the RIF1 CTD and enhanced RIF1-L chromatin retention, which was reversed by cyclin-dependent kinase 1-dependent phosphorylation of the RIF1 CTD in response to G2 DNA damage checkpoint inhibition. These combined findings suggest DNA damage-dependent RIF1 alternative splicing contributes to RIF1 functional diversification in genome protection.
A critical part of the hepatitis B virus (HBV) life cycle is the packaging of the pregenomic RNA (pgRNA) into nucleocapsids. While this process is known to involve several viral elements, much less is known about the identities and roles of host proteins in this process. To better understand the role of host proteins, we isolated pgRNA and characterized its protein interactome in cells expressing either packaging-competent or packaging-incompetent HBV genomes. We identified over 250 host proteins preferentially associated with pgRNA from the packaging-competent version of the virus. These included proteins already known to support capsid formation, enhance viral gene expression, catalyze nucleocapsid dephosphorylation, and bind to the viral genome, demonstrating the ability of the approach to effectively reveal functionally significant host-virus interactors. Three of these host proteins, AURKA, YTHDF2, and ATR, were selected for follow-up analysis. RNA immunoprecipitation qPCR (RIP-qPCR) confirmed pgRNA-protein association in cells, and siRNA knockdown of the proteins showed decreased encapsidation efficiency. This study provides a template for the use of comparative RNA-protein interactome analysis in conjunction with virus engineering to reveal functionally significant host-virus interactions.
Top-down proteomics, the characterization of intact proteoforms by tandem mass spectrometry, is the principal method for proteoform characterization in complex samples. Top-down proteomics relies on precursor isolation and subsequent gas-phase fragmentation to make proteoform identifications. While this strategy can produce highly detailed molecular information, the reliance on time-intensive tandem MS limits the speed with which proteoforms can be identified. We suggest that once proteoforms have been identified by top-down analysis in a system of interest, and archived in a system-specific Proteoform Atlas, subsequent analyses in that system can utilize the Atlas information to enable simpler and faster MS1-only identifications. We explore this idea here, using the E. coli ribosome as a model system of limited complexity. We used deep top-down analysis to construct an E. coli ribosomal Proteoform Atlas containing 2099 proteoforms from 52 of the 54 proteins that make up the E. coli ribosome. We show that using the Atlas enables confident MS1-only identifications of E. coli ribosomal proteoforms from E. coli that were perturbed by exposure to cold. Furthermore, this Atlas strategy identifies proteoforms up to 77% more rapidly compared to top-down identifications that require acquisition of both MS1 and MS2 spectra.
Background: Alzheimers disease (AD) is characterized by accumulation of two types of protein aggregates, senile plaques and neurofibrillary tangles, which strongly contribute to the pathogenesis of the disease. Senile plaques consist primarily of aggregated amyloid-β, while neurofibrillary tangles form via aggregation of the protein Tau, as well as other microtubule- associated proteins such as CRMP2. Posttranslational modifications have been hypothesized to contribute to the initial aggregation events that lead to SPs and NFTs. Dehydroamino acids (DHAAs) are posttranslational modifications rarely observed in humans and have not previously been reported in AD. DHAAs arise from the eliminylation of serine, threonine, or cysteine, yielding a double bond with distinct molecular geometry and reactivity. Their geometry can produce secondary structure rearrangements, such as those seen in senile plaque and neurofibrillary tangle formation, while their reactivity can cause intramolecular or intermolecular (protein-protein) crosslinking. We hypothesized that this modification might be present in protein aggregation-associated neurodegenerative disorders like AD. Methods: We performed mass spectrometry-based bottom-up proteomics on the sarkosyl-insoluble (protein aggregate-enriched) material from ten AD brains and three age-matched controls. Identifications of DHAA-mediated crosslinked peptides were validated using both an isotopic labeling strategy and spike-in experiments employing synthetic crosslinked peptide standards. Similar findings were obtained in searches of publicly available proteomic datasets from AD and control brains. Results: We identified 412 sites of DHAA modification in 184 proteins, with the highest prevalence in the neurofibrillary tangle-forming proteins Tau and CRMP2. Comparison with results of previous protein aggregate interactomics studies show that proteins containing the DHAA modification are more highly associated with protein aggregates than are proteins containing any other individual posttranslational modification. We further observed 11 protein crosslinks arising from DHAAs, including three from the Tau protein. Label free quantification showed that Tau crosslinks are an order of magnitude more prevalent in AD samples than in age- matched controls. Conclusions: Dehydroamino acids and their derivatives are prevalent modifications in the Alzheimers disease brain proteome. These modifications give rise to protein crosslinks which may contribute to protein aggregation processes. ### Competing Interest Statement The authors have declared no competing interest.
Host RNA binding proteins recognize viral RNA and play key roles in virus replication and antiviral mechanisms. SARS-CoV-2 generates a series of tiered subgenomic RNAs (sgRNAs), each encoding distinct viral protein(s) that regulate different aspects of viral replication. Here, for the first time, we demonstrate the successful isolation of SARS-CoV-2 genomic RNA and three distinct sgRNAs (N, S, and ORF8) from a single population of infected cells and characterize their protein interactomes. Over 500 protein interactors (including 260 previously unknown) were identified as associated with one or more target RNA. These included protein interactors unique to a single RNA pool and others present in multiple pools, highlighting our ability to discriminate between distinct viral RNA interactomes despite high sequence similarity. Individual interactomes indicated viral associations with cell response pathways, including regulation of cytoplasmic ribonucleoprotein granules and posttranscriptional gene silencing. We tested the significance of three protein interactors in these pathways (APOBEC3F, PPP1CC, and MSI2) using siRNA knockdowns, with several knockdowns affecting viral gene expression, most consistently PPP1CC. This study describes a new technology for high-resolution studies of SARS-CoV-2 RNA regulation and reveals a wealth of new viral RNA-associated host factors of potential functional significance to infection.
RNA-protein interactions are key to many aspects of cellular homeostasis and their identification is important to understanding cellular function. Multiple strategies have been developed for the RNA-centric characterization of RNA-protein complexes. However, these studies have all been done in immortalized cell lines that do not capture the complexity of heterogeneous tissue samples. Here, we develop hybridization purification of RNA-protein complexes followed by mass spectrometry (HyPR-MS) for use in tissue samples. We isolated both polyadenylated RNA and the specific long noncoding RNA MALAT1 and characterized their protein interactomes. These results demonstrate the feasibility of HyPR-MS in tissue for the multiplexed characterization of specific RNA-protein complexes.
The rapid and accurate quantification of peptides is a critical element of modern proteomics that has become increasingly challenging as proteomic data sets grow in size and complexity. We present here FlashLFQ, a computer program for high-speed label-free quantification of peptides and proteins following a search of bottom-up mass spectrometry data. FlashLFQ is approximately an order of magnitude faster than established label-free quantification methods and can quantify data-dependent analysis (DDA) search results from any proteomics search program. It is available as a graphical user interface program, a command line tool, a Docker image, and integrated into the MetaMorpheus search software.
Mutations in the ubiquitin (Ub) chaperone Ubiquilin 2 (UBQLN2) cause X-linked forms of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) through unknown mechanisms. Here, we show that aggregation-prone, ALS-associated mutants of UBQLN2 (UBQLN2ALS) trigger heat stress-dependent neurodegeneration in Drosophila. A genetic modifier screen implicated endolysosomal and axon guidance genes, including the netrin receptor, Unc-5, as key modulators of UBQLN2 toxicity. Reduced gene dosage of Unc-5 or its coreceptor Dcc/frazzled diminished neurodegenerative phenotypes, including motor dysfunction, neuromuscular junction defects, and shortened lifespan, in flies expressing UBQLN2ALS alleles. Induced pluripotent stem cells (iPSCs) harboring UBQLN2ALS knockin mutations exhibited lysosomal defects while inducible motor neurons (iMNs) expressing UBQLN2ALS alleles exhibited cytosolic UBQLN2 inclusions, reduced neurite complexity, and growth cone defects that were partially reversed by silencing of UNC5B and DCC. The combined findings suggest that altered growth cone dynamics are a conserved pathomechanism in UBQLN2-associated ALS/FTD.
Interpreting proteomics data remains challenging due to the large number of proteins that are quantified by modern mass spectrometry methods. Weighted gene correlation network analysis (WGCNA) can identify groups of biologically related proteins using only protein intensity values by constructing protein correlation networks. However, WGCNA is not widespread in proteomic analyses due to challenges in implementing workflows. To facilitate the adoption of WGCNA by the proteomics field, we created MetaNetwork, an open-source, R-based application to perform sophisticated WGCNA workflows with no coding skill requirements for the end user. We demonstrate MetaNetwork's utility by employing it to identify groups of proteins associated with prostate cancer from a proteomic analysis of tumor and adjacent normal tissue samples. We found a decrease in cytoskeleton-related protein expression, a known hallmark of prostate tumors. We further identified changes in module eigenproteins indicative of dysregulation in protein translation and trafficking pathways. These results demonstrate the value of using MetaNetwork to improve the biological interpretation of quantitative proteomics experiments with 15 or more samples.
In eukaryotes, splice sites define the introns of pre-mRNAs and must be recognized and excised with nucleotide precision by the spliceosome to make the correct mRNA product. In one of the earliest steps of spliceosome assembly, the U1 small nuclear ribonucleoprotein (snRNP) recognizes the 5' splice site (5' SS) through a combination of base pairing, protein-RNA contacts, and interactions with other splicing factors. Previous studies investigating the mechanisms of 5' SS recognition have largely been done in vivo or in cellular extracts where the U1/5' SS interaction is difficult to deconvolute from the effects of trans-acting factors or RNA structure. In this work we used colocalization single-molecule spectroscopy (CoSMoS) to elucidate the pathway of 5' SS selection by purified yeast U1 snRNP. We determined that U1 reversibly selects 5' SS in a sequence-dependent, two-step mechanism. A kinetic selection scheme enforces pairing at particular positions rather than overall duplex stability to achieve long-lived U1 binding. Our results provide a kinetic basis for how U1 may rapidly surveil nascent transcripts for 5' SS and preferentially accumulate at these sequences rather than on close cognates.
Human immunodeficiency virus type 1 (HIV-1) remains a deadly infectious disease despite existing antiretroviral therapies. A comprehensive understanding of the specific mechanisms of viral infectivity remains elusive and currently limits the development of new and effective therapies. Through in-depth proteomic analysis of HIV-1 virions, we discovered the novel post-translational modification of highly conserved residues within the viral matrix and capsid proteins to the dehydroamino acids, dehydroalanine and dehydrobutyrine. We further confirmed their presence by labeling the reactive alkene, characteristic of dehydroamino acids, with glutathione via Michael addition. Dehydroamino acids are rare, understudied, and have been observed mainly in select bacterial and fungal species. Until now, they have not been observed in HIV proteins. We hypothesize that these residues are important in viral particle maturation and could provide valuable insight into HIV infectivity mechanisms.
RNA-protein interactions are integral to maintaining proper cellular function and homeostasis, and the disruption of key RNA-protein interactions is central to many disease states. HyPR-MS (hybridization purification of RNA-protein complexes followed by mass spectrometry) is a highly versatile and efficient technology which enables multiplexed discovery of specific RNA-protein interactomes. This chapter provides extensive guidance for successful application of HyPR-MS to the system and target RNA(s) of interest, as well as a detailed description of the fundamental HyPR-MS procedure, including: (1) experimental design of controls, capture oligonucleotides, and qPCR assays; (2) formaldehyde cross-linking of cell culture; (3) cell lysis and RNA solubilization; (4) isolation of target RNA(s); (5) RNA purification and RT-qPCR analysis; (6) protein preparation and mass spectrometric analysis; and (7) mass spectrometric data analysis.
RNA-binding proteins are crucial to the function of coding and non-coding RNAs. The disruption of RNA–protein interactions is involved in many different pathological states. Several computational and experimental strategies have been developed to identify protein binders of selected RNA molecules. Amongst these, ‘in cell’ hybridization methods represent the gold standard in the field because they are designed to reveal the proteins bound to specific RNAs in a cellular context. Here, we compare the technical features of different ‘in cell’ hybridization approaches with a focus on their advantages, limitations, and current and potential future applications.
Mutations in the ubiquitin (Ub) chaperone Ubiquilin 2 (UBQLN2) cause X-linked forms of amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (FTD) through unknown mechanisms. Here we show that aggregation-prone, ALS-associated mutants of UBQLN2 (UBQLN2 ALS ) trigger heat stress-dependent neurodegeneration in Drosophila. A genetic modifier screen implicated endolysosomal and axon guidance genes, including the netrin receptor, Unc-5, as key modulators of UBQLN2 toxicity. Reduced gene dosage of Unc-5 or its coreceptor Dcc/frazzled diminished neurodegenerative phenotypes, including motor dysfunction, neuromuscular junction defects, and shortened lifespan, in flies expressing UBQLN2 ALS alleles. Induced pluripotent stem cells (iPSCs) harboring UBQLN2 ALS knockin mutations exhibited lysosomal defects while inducible motor neurons (iMNs) expressing UBQLN2 ALS alleles exhibited cytosolic UBQLN2 inclusions, reduced neurite complexity, and growth cone defects that were partially reversed by silencing of UNC5B and DCC. The combined findings suggest that altered growth cone dynamics are a conserved pathomechanism in UBQLN2-associated ALS/FTD.
HIV-1 generates unspliced (US), partially spliced (PS), and completely spliced (CS) classes of RNAs, each playing distinct roles in viral replication. Elucidating their host protein ‘interactomes’ is crucial to understanding virus-host interplay. Here, we present HyPR-MSSV for isolation of US, PS, and CS transcripts from a single population of infected CD4+ T-cells and mass spectrometric identification of their in vivo protein interactomes. Analysis revealed 212 proteins differentially associated with the unique RNA classes, including preferential association of regulators of RNA stability with US and PS transcripts and, unexpectedly, mitochondria-linked proteins with US transcripts. Remarkably, >80 of these factors screened by siRNA knockdown impacted HIV-1 gene expression. Fluorescence microscopy confirmed several to co-localize with HIV-1 US RNA and exhibit changes in abundance and/or localization over the course of infection. This study validates HyPR-MSSV for discovery of viral splice variant protein interactomes and provides an unprecedented resource of factors and pathways likely important to HIV-1 replication.