The high throughput analysis of proteins with mass spectrometry (MS) is highly valuable for understanding human biology, discovering disease biomarkers, identifying therapeutic targets, and exploring pathogen interactions. To achieve these goals, specialized proteomics subfields, including plasma proteomics, immunopeptidomics, and metaproteomics, must tackle specific analytical challenges, such as an increased identification ambiguity compared to routine proteomics experiments. Technical advancements in MS instrumentation can mitigate these issues by acquiring more discerning information at higher sensitivity levels. This is exemplified by the incorporation of ion mobility and parallel accumulation and serial fragmentation (PASEF) technologies in timsTOF instruments. In addition, AI-based bioinformatics solutions can help overcome ambiguity issues by integrating more data into the identification workflow. Here, we introduce TIMS2Rescore, a data-driven rescoring workflow optimized for DDA-PASEF data from timsTOF instruments. This platform includes new timsTOF MS2PIP spectrum prediction models and IM2Deep, a new deep learning-based peptide ion mobility predictor. Furthermore, to fully streamline data throughput, TIMS2Rescore directly accepts Bruker raw mass spectrometry data and search results from ProteoScape and many other search engines, including Sage and PEAKS. We showcase TIMS2Rescore performance on plasma proteomics, immunopeptidomics (HLA class I and II), and metaproteomics data sets. TIMS2Rescore is open-source and freely available at https://github.com/compomics/tims2rescore.
The high throughput analysis of proteins with mass spectrometry (MS) is highly valuable for understanding human biology, discovering disease biomarkers, identifying therapeutic targets, and exploring pathogen interactions. To achieve these goals, specialized proteomics subfields – such as plasma proteomics, immunopeptidomics, and metaproteomics – must tackle specific analytical challenges, such as an increased identification ambiguity compared to routine proteomics experiments. Technical advancements in MS instrumentation can counter these issues by acquiring more discerning information at higher sensitivity levels, as is exemplified by the incorporation of ion mobility and parallel accumulation - serial fragmentation (PASEF) technologies in timsTOF instruments. In addition, AI-based bioinformatics solutions can help overcome ambiguity issues by integrating more data into the identification workflow. Here, we introduce TIMS2Rescore, a data-driven rescoring workflow optimized for DDA-PASEF data from timsTOF instruments. This platform includes new timsTOF MS2PIP spectrum prediction models and IM2Deep, a new deep learning-based peptide ion mobility predictor. Furthermore, to fully streamline data throughput, TIMS2Rescore directly accepts Bruker raw mass spectrometry data, and search results from ProteoScape and many other search engines, including MS Amanda and PEAKS. We showcase TIMS2Rescore performance on plasma proteomics, immunopeptidomics (HLA class I and II), and metaproteomics data sets. TIMS2Rescore is open-source and freely available at . ### Competing Interest Statement J.K., T.S., G.R., D.T. are employees of subsidiaries of Bruker Corp but will not gain financially from what is disclosed.
Objectives Minimal residual disease (MRD) status in multiple myeloma (MM) is an important prognostic biomarker. Personalized blood-based targeted mass spectrometry detecting M-proteins (MS-MRD) was shown to provide a sensitive and minimally invasive alternative to MRD-assessment in bone marrow. However, MS-MRD still comprises of manual steps that hamper upscaling of MS-MRD testing. Here, we introduce a proof-of-concept for a novel workflow using data independent acquisition-parallel accumulation and serial fragmentation (dia-PASEF) and automated data processing. Methods Using automated data processing of dia-PASEF measurements, we developed a workflow that identified unique targets from MM patient sera and personalized protein sequence databases. We generated patient-specific libraries linked to dia-PASEF methods and subsequently quantitated and reported M-protein concentrations in MM patient follow-up samples. Assay performance of parallel reaction monitoring (prm)-PASEF and dia-PASEF workflows were compared and we tested mixing patient intake sera for multiplexed target selection. Results No significant differences were observed in lowest detectable concentration, linearity, and slope coefficient when comparing prm-PASEF and dia-PASEF measurements of serial dilutions of patient sera. To improve assay development times, we tested multiplexing patient intake sera for target selection which resulted in the selection of identical clonotypic peptides for both simplex and multiplex dia-PASEF. Furthermore, assay development times improved up to 25× when measuring multiplexed samples for peptide selection compared to simplex. Conclusions Dia-PASEF technology combined with automated data processing and multiplexed target selection facilitated the development of a faster MS-MRD workflow which benefits upscaling and is an important step towards the clinical implementation of MS-MRD.
Real-time database searching allows for simpler and automated proteomics workflows as it eliminates technical bottlenecks in high-throughput experiments. Most importantly, it enables results-dependent acquisition (RDA), where search results can be used to guide data acquisition during acquisition. This is especially beneficial for glycoproteomics since the wide range of physicochemical properties of glycopeptides lead to a wide range of optimal acquisition parameters. We established here the GlycoPaSER prototype by extending the Parallel Search Engine in Real-time (PaSER) functionality for real-time glycopeptide identification from fragmentation spectra. Glycopeptide fragmentation spectra were decomposed into peptide and glycan moiety spectra using common N-glycan fragments. Each moiety was subsequently identified by a specialized algorithm running in real-time. GlycoPaSER can keep up with the rate of data acquisition for real-time analysis with similar performance to other glycoproteomics software and produces results that are in line with the literature reference data. The GlycoPaSER prototype presented here provides the first proof-of-concept for real-time glycopeptide identification that unlocks the future development of RDA technology to transcend data acquisition.
Data independent acquisition (DIA) has become the go to method for deep and quantitative proteomic analysis given the ability to sample large m/z windows in a reproducible and non-stochastic manner. Using a method termed dia-PASEF on a TIMS enabled Q-TOF lends additional advantages in both duty cycle and selectivity using the ion mobility space. DIA-PASEF allows for deep proteomes in short gradient times ( , 20 min.) and even deeper proteomes with longer greater times. Despite the rapid advance- ments in hardware and applications of proteomics, which has directly lead tothegenerationofthousandsofproteomicdatasets,theproteomicbottle-neckremainsdataanalysis.ThePaSERplatformprovidesasolution through its GPU-powered ability to perform CCS-enabled DDA and DIA analysis in real time.With the release of PaSER 2022c, we introduced the fi rst vendor integrated version of DIA-NN, namely TIMS DIA-NN: a CCS- enabled analysis tool for the identi fi cation and quanti fi cation of dia-PASEF data. TIMS-DIA-NN takes advantage of the TIMScore model that greatly increases peptide identi fi cations to build more robust libraries.Here we present PaSER 2023 which includes signi fi cant improvements to the peak picking and match between runs algorithms. To show the capability of this platform, we used Human, Yeast and Ecoli (HYE) mixtures at different but known ratios to identify over 140000 precursors and over 12500 proteins excellent with quantitative accuracy.To alleviate the bottleneck of data processing, library based searches of dia-PASEF data can be searched in near real-time, providing answers to biological questions without any necessityofwaitingforanalysistoprocess.
The use of isobaric mass tags in quantitative proteomics is a common practice where ratios of the liberated reporter ions provide quantitative information. Multiplexing more channels within the same experiment has obvious advantages, but neutron encoded mass tags have a mass defect requiring >60,000 mass resolution at 100 m/z for baseline resolution. Here we demonstrate the multiplexing of 9 channels where accurate quantitation is afforded because baseline resolution of each of the reporter ions is achieved. Furthermore, co-fragmentation and underestimated reporter ion intensities is a common problem, we demonstrate how different trapping times in a trapped ion mobility spectrometer (TIMS) operated in varying parallel accumulation serial fragmentation (PASEF) settings mitigates the compression effect common to isobaric mass tagging experiments.
Photosynthetic organisms provide food and energy for nearly all life on Earth, yet half of their protein-coding genes remain uncharacterized1,2. Characterization of these genes could be greatly accelerated by new genetic resources for unicellular organisms. Here we generated a genome-wide, indexed library of mapped insertion mutants for the unicellular alga Chlamydomonas reinhardtii. The 62,389 mutants in the library, covering 83% of nuclear protein-coding genes, are available to the community. Each mutant contains unique DNA barcodes, allowing the collection to be screened as a pool. We performed a genome-wide survey of genes required for photosynthesis, which identified 303 candidate genes. Characterization of one of these genes, the conserved predicted phosphatase-encoding gene CPL3, showed that it is important for accumulation of multiple photosynthetic protein complexes. Notably, 21 of the 43 higher-confidence genes are novel, opening new opportunities for advances in understanding of this biogeochemically fundamental process. This library will accelerate the characterization of thousands of genes in algae, plants, and animals.
Photosynthetic organisms provide food and energy for nearly all life on Earth, yet half of their protein-coding genes remain uncharacterized1,2. Characterization of these genes could be greatly accelerated by new genetic resources for unicellular organisms that complement the use of multicellular plants by enabling higher-throughput studies. Here, we generated a genome-wide, indexed library of mapped insertion mutants for the flagship unicellular algaChlamydomonas reinhardtii(Chlamydomonas hereafter). The 62,389 mutants in the library, covering 83% of nuclear, protein-coding genes, are available to the community. Each mutant contains unique DNA barcodes, allowing the collection to be screened as a pool. We leveraged this feature to perform a genome-wide survey of genes required for photosynthesis, which identified 303 candidate genes. Characterization of one of these genes, the conserved predicted phosphataseCPL3, showed it is important for accumulation of multiple photosynthetic protein complexes. Strikingly, 21 of the 43 highest-confidence genes are novel, opening new opportunities for advances in our understanding of this biogeochemically fundamental process. This library is the first genome-wide mapped mutant resource in any unicellular photosynthetic organism, and will accelerate the characterization of thousands of genes in algae, plants and animals.
ABSTRACT Gram-negative bacteria have an outer membrane (OM) impermeable to many toxic compounds that can be further strengthened during stress. In Enterobacteriaceae, the envelope contains enterobacterial common antigen (ECA), a carbohydrate-derived moiety conserved throughout Enterobacteriaceae, the function of which is poorly understood. Previously, we identified several genes in Escherichia coli K-12 responsible for an RpoS-dependent decrease in envelope permeability during carbon-limited stationary phase. For one of these, yhdP, a gene of unknown function, deletion causes high levels of both vancomycin and detergent sensitivity, independent of growth phase. We isolated spontaneous suppressor mutants of yhdP with loss-of-function mutations in the ECA biosynthesis operon. ECA biosynthesis gene deletions suppressed envelope permeability from yhdP deletion independently of envelope stress responses and interactions with other biosynthesis pathways, demonstrating suppression is caused directly by removing ECA. Furthermore, yhdP deletion changed cellular ECA levels and yhdP was found to co-occur phylogenetically with the ECA biosynthesis operon. Cells make three forms of ECA: ECA lipopolysaccharide (LPS), an ECA chain linked to LPS core; ECA phosphatidylglycerol, a surface-exposed ECA chain linked to phosphatidylglycerol; and cyclic ECA, a cyclized soluble ECA molecule found in the periplasm. We determined that the suppression of envelope permeability with yhdP deletion is caused specifically by the loss of cyclic ECA, despite lowered levels of this molecule found with yhdP deletion. Furthermore, removing cyclic ECA from wild-type cells also caused changes to OM permeability. Our data demonstrate cyclic ECA acts to maintain the OM permeability barrier in a manner controlled by YhdP. IMPORTANCE Enterobacterial common antigen (ECA) is a surface antigen made by all members of Enterobacteriaceae, including many clinically relevant genera (e.g., Escherichia, Klebsiella, Yersinia). Although this surface-exposed molecule is conserved throughout Enterobacteriaceae, very few functions have been ascribed to it. Here, we have determined that the periplasmic form of ECA, cyclic ECA, plays a role in maintaining the outer membrane permeability barrier. This activity is controlled by a protein of unknown function, YhdP, and deletion of yhdP damages the OM permeability barrier in a cyclic ECA-dependent manner, allowing harmful molecules such as antibiotics into the cell. This role in maintenance of the envelope permeability barrier is the first time a phenotype has been described for cyclic ECA. As the Gram-negative envelope is generally impermeable to antibiotics, understanding the mechanisms through which the barrier is maintained and antibiotics are excluded may lead to improved antibiotic delivery.
Post‐transcriptional modifications play an important role in RNA biology. In particular, the addition of small chemical groups to the nucleobases of mRNA can affect how modified transcripts are processed in the cell, thereby impacting gene expression programs. In order to study the molecular mechanisms underlying these modifications, it is necessary to characterize their ‘readers’, that is, proteins that directly bind to these modifications to mediate their functional consequences; this is a major challenge because we lack approaches to precisely manipulate RNA chemistry in the cell and because protein–modified RNA interactions can be low affinity. In this unit, we describe in detail a photocrosslinking‐based RNA chemical proteomics approach to profile the protein–modified RNA interactome modulated by N 6 ‐methyladenosine (m 6 A), the most abundant internal modification in eukaryotic mRNA. First, we present protocols for the synthesis and characterization of short, diazirine‐containing synthetic RNA probes, followed by a description of their use in mass spectrometry–based proteomics with HeLa cell lysate and a short commentary on data analysis and result interpretation. © 2018 by John Wiley & Sons, Inc.
Topoisomerase IIα (TOP2α) is essential for chromosomal condensation and segregation, as well as genomic integrity. Here we report that RNF168, an E3 ligase mutated in the human RIDDLE syndrome, interacts with TOP2α and mediates its ubiquitylation. RNF168 deficiency impairs decatenation activity of TOP2α and promotes mitotic abnormalities and defective chromosomal segregation. Our data also indicate that RNF168 deficiency, including in human breast cancer cell lines, confers resistance to the anti-cancer drug and TOP2 inhibitor etoposide. We also identify USP10 as a deubiquitylase that negatively regulates TOP2α ubiquitylation and restrains its chromatin association. These findings provide a mechanistic link between the RNF168/USP10 axis and TOP2α ubiquitylation and function, and suggest a role for RNF168 in the response to anti-cancer chemotherapeutics that target TOP2.
von Hippel-Lindau (VHL) disease is a rare familial cancer predisposition syndrome caused by a loss or mutation in a single gene, VHL, but it exhibits a wide phenotypic variability that can be categorized into distinct subtypes. The phenotypic variability has been largely argued to be attributable to the extent of deregulation of the α subunit of hypoxia-inducible factor α, a well established target of VHL E3 ubiquitin ligase, ECV (Elongins/Cul2/VHL). Here, we show that erythropoietin receptor (EPOR) is hydroxylated on proline 419 and 426 via prolyl hydroxylase 3. EPOR hydroxylation is required for binding to the β domain of VHL and polyubiquitylation via ECV, leading to increased EPOR turnover. In addition, several type-specific VHL disease-causing mutants, including those that have retained proper binding and regulation of hypoxia-inducible factor α, showed a severe defect in binding prolyl hydroxylated EPOR peptides. These results identify EPOR as the second bona fide hydroxylation-dependent substrate of VHL that potentially influences oxygen homeostasis and contributes to the complex genotype-phenotype correlation in VHL disease.
We generated a global genetic interaction network for Saccharomyces cerevisiae, constructing more than 23 million double mutants, identifying about 550,000 negative and about 350,000 positive genetic interactions. This comprehensive network maps genetic interactions for essential gene pairs, highlighting essential genes as densely connected hubs. Genetic interaction profiles enabled assembly of a hierarchical model of cell function, including modules corresponding to protein complexes and pathways, biological processes, and cellular compartments. Negative interactions connected functionally related genes, mapped core bioprocesses, and identified pleiotropic genes, whereas positive interactions often mapped general regulatory connections among gene pairs, rather than shared functionality. The global network illustrates how coherent sets of genetic interactions connect protein complex and pathway modules to map a functional wiring diagram of the cell.
Oligomeric ubiquitin structures (i.e. ubiquitin “chains”) may be formed through any of seven different lysine residues in the polypeptide, or via the amine group of Met 1. Different types of ubiquitin chains can confer very different biological outcomes to a protein substrate, yet the structural characteristics of E2s and E3s that determine ubiquitin linkage specificity remain poorly understood. In vitro autoubiquitylation assays combined with ubiquitin protein variants bearing individually mutated lysine residues (“K‐to‐R” mutants) have thus been widely used to characterize E2–E3 linkage specificity. However, how this type of assay compares to direct identification of ubiquitin linkage types using mass spectrometry (MS) has not been rigorously tested. Here, we characterize the linkage specificity of 12 different E2–E3 combinations using both approaches. The simple MS‐based method described here is more robust, requires less material and is less prone to bias introduced by, e.g. the use of mutant proteins with unknown effects on E1, E2 or E3 recognition, antibodies with uncharacterized epitopes, the low dynamic range of X‐ray film, and additional sources of experimental error. Indeed, our results suggest that the K‐to‐R assay be approached with some caution.
Abstract MYC activity is regulated by a complex network of signaling cascades and protein interactions, some of which result in post-translation modifications (PTMs). For example, phosphorylation of the regulatory threonine 58 (T58) and serine 62 (S62) plays a pivotal role in regulating MYC stability and activity. Loss of this regulatory pathway in cancer can lead to MYC dysregulation and contribute to tumorigenesis. Despite their important role, the majority of MYC PTMs arising downstream of different signaling cascades and their consequences on regulating MYC activity remain largely unknown. To capture the broad spectrum of MYC PTMs, a mass spectrometry (MS)-based approach is being pursued to attain high sequence coverage and enable the identification of PTMs throughout the protein. MYC was overexpressed in the HEK293T cell line, immunoprecipitated, digested, and analyzed on the Velos Orbitrap MS. This analysis facilitated the identification of a range of MYC PTMs, including phosphorylation, sumoylation, ubiquitination and acetylation. Phosphorylation was readily detectable at residues T58 and S62, consistent with previous reports. Additionally, we also observed a previously reported phosphorylation cluster involving residues T343/S344/S347/S348, suggesting that these sites might play a role in regulating MYC function. Indeed, converting these residues to alanine to prevent phosphorylation resulted in a gain-of-function mutant (See Penn Lab poster Lorenco et al.). Another phosphorylation was observed mapping to residues S71 and/or S81. Mutating these sites to alanine also potentiated MYC transformation, highlighting the role of PTMs in regulating MYC function. A MYC sumoylation site was identified on lysine 326 (K326) and is subject of ongoing investigation (See Penn Lab poster Kalkat et al.). Additionally, we will discuss approaches presently underway to increase sequence coverage and identify additional MYC PTMs. PTMs play an important role in regulating the activity of transcription factors, including that of MYC. Moreover, PTMs can contribute to the dysregulation of its activity in cancer. Thus, mapping the array of PTMs in MYC, understanding their role in regulating MYC function, and elucidating the pathways that converge on these PTMs may pave the way for the development of novel therapeutic strategies aimed at targeting MYC-induced tumorigenesis. Citation Format: Diana Resetca, Manpreet Kalkat, Corey Lourenco, Pak-Kei Chan, Tharan Srikumar, Brian Raught, Linda Penn. Identifying MYC post-translational modifications using a mass spectrometry-based approach. [abstract]. In: Proceedings of the AACR Special Conference on Myc: From Biology to Therapy; Jan 7-10, 2015; La Jolla, CA. Philadelphia (PA): AACR; Mol Cancer Res 2015;13(10 Suppl):Abstract nr A12.
Small ubiquitin-like modifier-1 (SUMO1) plays a number of roles in cellular events and recent evidence has given momentum for its contributions to neuronal development and function. Here, we have generated a SUMO1 transgenic mouse model with exclusive overexpression in neurons in an effort to identify in vivo conjugation targets and the functional consequences of their SUMOylation. A high-expressing line was examined which displayed elevated levels of mono-SUMO1 and increased high molecular weight conjugates in all brain regions. Immunoprecipitation of SUMOylated proteins from total brain extract and proteomic analysis revealed ~95 candidate proteins from a variety of functional classes, including a number of synaptic and cytoskeletal proteins. SUMO1 modification of synaptotagmin-1 was found to be elevated as compared to non-transgenic mice. This observation was associated with an age-dependent reduction in basal synaptic transmission and impaired presynaptic function as shown by altered paired pulse facilitation, as well as a decrease in spine density. The changes in neuronal function and morphology were also associated with a specific impairment in learning and memory while other behavioral features remained unchanged. These findings point to a significant contribution of SUMO1 modification on neuronal function which may have implications for mechanisms involved in mental retardation and neurodegeneration.