Non-canonical (i.e., unannotated) open reading frames (ncORFs) have until recently been omitted from reference genome annotations, despite evidence of their translation, limiting their incorporation into biomedical research. To address this, in 2022, we initiated the TransCODE consortium and built the first community-driven consensus catalog of human ncORFs, which was openly distributed to the research community via Ensembl-GENCODE. While this catalog represented a starting point for reference ncORF annotation, major technical and scientific issues remained. In particular, this initial catalogue had no standardized framework to judge the evidence of translation for individual ncORFs. Here, we present an expanded and refined catalog of the human reference annotation of ncORFs. By incorporating more datasets and by lifting constraints on ORF length and start-codon, we define a comprehensive set of 28,359 ncORFs that is nearly four times the size of the previous catalog. Furthermore, to aid users who wish to work with ncORFs with the strongest and most reproducible signals of translation, we utilized a data-driven framework (i.e. translation signature scores) to assess the accumulated evidence for any individual ncORF. Using this approach, we derive a subset of 7,888 ncORFs with translation evidence on par with canonical protein-coding genes, which we refer to as the Primary set. This set can serve as a reliable reference for downstream analyses and validation, with a particular emphasis on high quality. Overall, this update reflects continual community-driven efforts to make ncORFs accessible and actionable to the broader research public and further iterations of the catalog will continue to expand and refine this resource.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
Proteomics data-dependent acquisition data sets collected with high-resolution mass-spectrometry (MS) can achieve very high-quality results, but nearly every analysis yields results that are thresholded at some accepted false discovery rate, meaning that a substantial number of results are incorrect. For study conclusions that rely on a small number of peptide-spectrum matches being correct, it is thus important to examine at least some crucial spectra to ensure that they are not one of the incorrect identifications. We present Quetzal, a peptide fragment ion spectrum annotation tool to assist researchers in annotating and examining such spectra to ensure that they correctly support study conclusions. We describe how Quetzal annotates spectra using the new Human Proteome Organization (HUPO) Proteomics Standards Initiative (PSI) mzPAF standard for fragment ion peak annotation, including the Python-based code, a web-service end point that provides annotation services, and a web-based application for annotating spectra and producing publication-quality figures. We illustrate its functionality with several annotated spectra of varying complexity. Quetzal provides easily accessible functionality that can assist in the effort to ensure and demonstrate that crucial spectra support study conclusions. Quetzal is publicly available at https://proteomecentral.proteomexchange.org/quetzal/.
We benchmarked eight software tools in crosslinking mass spectrometry (MS) to evaluate their accuracy in detecting protein-protein interactions. Whereas about half performed reliably, others reported up to 31-fold more false interactions than indicated by their reported false discovery rates. During this community-driven effort, we helped correct one of the poorer-performing tools. Our findings demonstrate that, with robust algorithms, crosslinking MS can deliver highly dependable insights into biological systems. ### Competing Interest Statement The authors have declared no competing interest. Deutsche ForschungsgemeinschaftDeutsche Forschungsgemeinschaft, https://ror.org/018mejw64, Excellence Strategy EXC 2008 390540038 UniSysCat National Institutes of HealthNational Institutes of Health, https://ror.org/01cwqze88, R01 GM087221, S10 OD026936 National Science FoundationNational Science Foundation, https://ror.org/021nxhr62, MRI-1920268
A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.
Malaria parasites must respond quickly to environmental changes, including during their transmission between mammalian and mosquito hosts. Therefore, female gametocytes proactively produce and translationally repress mRNAs that encode essential proteins that the zygote requires to establish a new infection. While the release of translational repression of individual mRNAs has been documented, the details of the global release of translational repression have not. Moreover, changes in the spatial arrangement and composition of the DOZI/CITH/ALBA complex that contribute to translational control are also not known. Therefore, we have conducted the first quantitative, comparative transcriptomics and DIA-MS proteomics of Plasmodium parasites across the host-to-vector transmission event to document the global release of translational repression. Using female gametocytes and zygotes of P. yoelii, we found that ~200 transcripts are released for translation soon after fertilization, including those encoding essential functions. Moreover, we identified that many transcripts remain repressed beyond this point. TurboID-based proximity proteomics of the DOZI/CITH/ALBA regulatory complex revealed substantial spatial and/or compositional changes across this transmission event, which are consistent with recent, paradigm-shifting models of translational control. Together, these data provide a model for the essential translational control mechanisms that promote Plasmodium’s efficient transmission from mammalian host to mosquito vector.
We describe a new release of the Candida albicans PeptideAtlas proteomics spectral resource (build 2024-03), providing a sequence coverage of 79.5% at the canonical protein level, matched mass spectrometry spectra, and experimental evidence identifying 3382 and 536 phosphorylated serine and threonine sites with false localization rates of 1% and 5.3%, respectively. We provide a tutorial on how to use the PeptideAtlas and associated tools to access this information. The C. albicans PeptideAtlas summary web page provides “Build overview”, “PTM coverage”, “Experiment contribution”, and “Data set contribution” information. The protein and peptide information can also be accessed via the Candida Genome Database via hyperlinks on each protein page. This allows users to peruse identified peptides, protein coverage, post-translational modifications (PTMs), and experiments that identify each protein. Given the value of understanding the PTM landscape in the sequence of each protein, a more detailed explanation of how to interpret and analyze PTM results is provided in the PeptideAtlas of this important pathogen. Candida albicans PeptideAtlas web page: https://db.systemsbiology.net/sbeams/cgi/PeptideAtlas/buildDetails?atlas_build_id=578.
Advances in mass spectrometry (MS) instrumentation, including higher resolution, faster scan speeds, and improved sensitivity, have dramatically increased the data volume and complexity. The adoption of imaging and ion mobility further amplifies these challenges in proteomics, metabolomics, and lipidomics. Current open formats such as mzML and imzML struggle to keep pace due to large file sizes, slow data access, and limited metadata support. Vendor-specific formats offer faster access but lack interoperability and long-term archival guarantees. We here lay the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multidimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human- and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will support complex workflows and regulatory compliance through faster access, improved metadata, and interoperability. For vendors, it offers a streamlined, open alternative to proprietary formats. mzPeak aims to become a cornerstone of MS data management, enabling sustainable, high-performance solutions for future data types and fostering collaboration across the mass spectrometry community.
Liquid chromatography-tandem mass spectrometry employing data-dependent acquisition (DDA) is a mature, widely used proteomics technique routinely applied to proteome profiling, protein-protein interaction studies, biomarker discovery, and protein modification analysis. Numerous tools exist for searching DDA data and myriad file formats are output as results. While some search and post processing tools include data visualization features to aid biological interpretation, they are often limited or tied to specific software pipelines. This restricts the accessibility, sharing and interpretation of data, and hinders comparison of results between different software pipelines. We developed Limelight, an easy-to-use, open-source, freely available tool that provides data sharing, analysis and visualization and is not tied to any specific software pipeline. Limelight is a data visualization tool specifically designed to provide access to the whole "data stack", from raw and annotated scan data to peptide-spectrum matches, quality control, peptides, proteins, and modifications. Limelight is designed from the ground up for sharing and collaboration and to support data from any DDA workflow. We provide tools to import data from many widely used open-mass and closed-mass search software workflows. Limelight helps maximize the utility of data by providing an easy-to-use interface for finding and interpreting data, all using the native scores from respective workflows.
There is a growing interest in serum-based biomarkers that reveal the mechanisms of healthy aging. The Longevity Consortium generated untargeted mass spectrometry-based proteomic and metabolomic datasets from serum samples across four cohorts; the Osteoporotic Fractures in Men (MrOS) Study, the Study of Osteoporotic Fractures (SOF), the Health, Aging, and Body Composition (Health ABC) Study, and the Cardiovascular Health Study (CHS) totaling 3380 total participants with a 1:3 case-cohort design. This dataset offers a unique opportunity to identify a robust prospective signature of Longevity, defined as survival to the 98th percentile based on sex- and birth cohort-specific survival distributions. We applied multiblock sparse partial least squares discriminant analysis and systems biology approaches to integrate and construct multi-omic signatures predictive of longevity. From approximately 5000 metabolites and 500 proteins, we identified biomarkers associated with coagulation and complement and cholesterol transport. Interestingly, the cholesterol transport signature (apolipoproteins) was only enriched in females, suggesting sex-specific mechanisms of longevity. Furthermore, we validate these findings by comparing omics data from human centenarian and mouse longevity studies that also capture molecular signatures induced by life-extending interventions. This study underscores the power of integrative systems biology methods in characterizing the heterogeneity of molecular aging phenotypes, ultimately enabling the development of robust longevity signatures. The identified biomarker signatures have potential implications for personalized interventions aimed at promoting healthy aging and mitigating age-related diseases.
Johanson-Blizzard Syndrome (JBS) is an autosomal recessive spectrum disorder associated with the UBR-1 ubiquitin ligase that features developmental delay including motor abnormalities. Here, we demonstrate that C. elegans UBR-1 regulates high-intensity locomotor behavior and developmental viability via both ubiquitin ligase and scaffolding mechanisms. Super-resolution imaging with CRISPR-engineered UBR-1 and genetic results demonstrated that UBR-1 is expressed and functions in the nervous system including in pre-motor interneurons. To decipher mechanisms of UBR-1 function, we deployed CRISPR-based proteomics using C. elegans which identified a cadre of glutamate metabolic enzymes physically associated with UBR-1 including GLN-3, GOT-2.2, GFAT-1 and GDH-1. Similar to UBR-1, all four glutamate enzymes are genetically linked to human developmental and neurological deficits. Proteomics, multi-gene interaction studies, and pharmacological findings indicated that UBR-1, GLN-3 and GOT-2.2 form a signaling axis that regulates glutamate homeostasis. Developmentally, UBR-1 is expressed in embryos and functions with GLN-3 to regulate viability. Overall, our results suggest UBR-1 is an enzyme hub in a GOT-2.2/UBR-1/GLN-3 axis that maintains glutamate homeostasis required for efficient locomotion and organismal viability. Given the prominent role of glutamate within and outside the nervous system, the UBR-1 glutamate homeostatic network we have identified could contribute to JBS etiology.
The Human Proteome Project (HPP), the flagship initiative of the Human Proteome Organization (HUPO), has pursued two goals: (1) to credibly identify at least one isoform of every protein-coding gene and (2) to make proteomics an integral part of multiomics studies of human health and disease. The past year has seen major transitions for the HPP. neXtProt was retired as the official HPP knowledge base, UniProtKB became the reference proteome knowledge base, and Ensembl-GENCODE provides the reference protein target list. A function evidence FE1-5 scoring system has been developed for functional annotation of proteins, parallel to the PE1-5 UniProtKB/neXtProt scheme for evidence of protein expression. This report includes updates from neXtProt (version 2023-09) and UniProtKB release 2024_04, with protein expression detected (PE1) for 18138 of the 19411 GENCODE protein-coding genes (93%). The number of non-PE1 proteins ("missing proteins") is now 1273. The transition to GENCODE is a net reduction of 367 proteins (19,411 PE1-5 instead of 19,778 PE1-4 last year in neXtProt). We include reports from the Biology and Disease-driven HPP, the Human Protein Atlas, and the HPP Grand Challenge Project. We expect the new Functional Evidence FE1-5 scheme to energize the Grand Challenge Project for functional annotation of human proteins throughout the global proteomics community, including π-HuB in China.
Covalent protein adducts formed by drugs or their reactive metabolites are risk factors for adverse reactions, and inactivation of cytochrome P450 (CYP) enzymes. Characterization of drug-protein adducts is limited due to lack of methods identifying and quantifying covalent adducts in complex matrices. This study presents a workflow that combines data-dependent and data-independent acquisition (DDA and DIA) based liquid chromatography with tandem mass spectrometry (LC-MS/MS) to detect very low abundance adducts resulting from CYP mediated drug metabolism in human liver microsomes (HLMs). HLMs were incubated with raloxifene as a model compound and adducts were detected in 78 proteins, including CYP3A and CYP2C family enzymes. Experiments with recombinant CYP3A and CYP2C enzymes confirmed adduct formation in all CYPs tested, including CYPs not subject to time-dependent inhibition by raloxifene. These data suggest adducts can be benign. DIA analysis showed variable adduct abundance in many peptides between livers, but no concomitant decrease of unadducted peptides. This study sets a new standard for adduct detection in complex samples, offering insights into the human adductome resulting from reactive metabolite exposure. The methodology presented will aid mechanistic studies to identify, quantify and differentiate between adducts that result in adverse drug reactions and those that are benign.
Glycosylphosphatidylinositol (GPI) anchor protein modification in Plasmodium species is well known and represents the principal form of glycosylation in these organisms. The structure and biosynthesis of GPI anchors of Plasmodium spp. has been primarily studied in the asexual blood stage of P. falciparum and is known to contain the typical conserved GPI structure of EtN-P-Man3GlcN-PI. Here, we have investigated the circumsporozoite protein (CSP) for the presence of a GPI-anchor. CSP is the major surface protein of Plasmodium sporozoites, the infective stage of the malaria parasite. While it is widely assumed that CSP is a GPI-anchored cell surface protein, compelling biochemical evidence for this supposition is absent. Here, we employed metabolic labeling and mass-spectrometry based approaches to confirm the presence of a GPI anchor in CSP. Biosynthetic radiolabeling of CSP with [3H]-palmitic acid and [3H]-ethanolamine, with the former being base-labile and therefore ester-linked, provided strong evidence for the presence of a GPI anchor on CSP, but these data alone were not definitive. To provide further evidence, immunoprecipitated CSP was analyzed for presence of myo-inositol (a characteristic component of GPI anchor) using strong acid hydrolysis and GC-MS for a highly sensitive and quantitative detection. The single ion monitoring (SIM) method for GC-MS analysis confirmed the presence of the myo-inositol component in CSP. Taken together, these data provide confidence that the long-assumed presence of a GPI anchor on this important parasite protein is correct.
Malaria parasites must be able to respond quickly to changes in their environment, including during their transmission between mammalian hosts and mosquito vectors. Therefore, before transmission, female gametocytes proactively produce and translationally repress mRNAs that encode essential proteins that the zygote requires to establish a new infection. This essential regulatory control requires the orthologues of DDX6 (DOZI), LSM14a (CITH), and ALBA proteins to form a translationally repressive complex in female gametocytes that associates with many of the affected mRNAs. However, while the release of translational repression of individual mRNAs has been documented, the details of the global release of translational repression have not. Moreover, the changes in spatial arrangement and composition of the DOZI/CITH/ALBA complex that contribute to translational control are also not known. Therefore, we have conducted the first quantitative, comparative transcriptomics and DIA-MS proteomics of Plasmodium parasites across the host-to-vector transmission event to document the global release of translational repression. Using female gametocytes and zygotes of P. yoelii , we found that nearly 200 transcripts are released for translation soon after fertilization, including those with essential functions for the zygote. However, we also observed that some transcripts remain repressed beyond this point. In addition, we have used TurboID-based proximity proteomics to interrogate the spatial and compositional changes in the DOZI/CITH/ALBA complex across this transmission event. Consistent with recent models of translational control, proteins that associate with either the 5’ or 3’ end of mRNAs are in close proximity to one another during translational repression in female gametocytes and then dissociate upon release of repression in zygotes. This observation is cross-validated for several protein colocalizations in female gametocytes via ultrastructure expansion microscopy and structured illumination microscopy. Moreover, DOZI exchanges its interaction from NOT1-G in female gametocytes to the canonical NOT1 in zygotes, providing a model for a trigger for the release of mRNAs from DOZI. Finally, unenriched phosphoproteomics revealed the modification of key translational control proteins in the zygote. Together, these data provide a model for the essential translational control mechanisms used by malaria parasites to promote their efficient transmission from their mammalian host to their mosquito vector.
ABSTRACTThe complex life cycle of Plasmodium parasites, the eukaryotic pathogens that cause malaria, features three distinct invasive forms tailored specifically to the equally distinct host environment they must navigate and invade for progression of the life cycle. One conserved feature of all these invasive forms is the presence of micronemes, apically oriented secretory organelles involved in egress, motility, adhesion and invasion. Micronemes are tailored to their specific host environment and feature stage specific contents. Here we investigate the role of GPI-anchored micronemal antigen (GAMA), which shows a micronemal localization in all zoite forms of the rodent infecting species Plasmodium berghei. While GAMA is dispensable during asexual blood stages, GAMA knock out parasites are severely defective for invasion of the mosquito midgut, resulting in reduced numbers of oocysts. Once formed, oocysts develop normally, however sporozoites are unable to egress and these sporozoites exhibit defective motility. Epitope-tagging of GAMA revealed tight temporal expression late during sporogony and showed that GAMA is shed during sporozoite gliding motility in a similar manner to circumsporozoite protein. Complementation of P. berghei knock out parasites with full length P. falciparum GAMA partially restored infectivity to mosquitoes, indicating a conservation of function across Plasmodium species. A suite of parasites with GAMA expressed under the promoters of the known ookinete-to-sporozoite stage-specific genes: CTRP, CAP380 and TRAP, further confirmed the involvement of GAMA in midgut infection, motility and infection of the mammalian host and revealed a lethal consequence to overexpression of GAMA during oocyst development. Combined, the research suggest that GAMA plays independent roles in sporozoite motility, egress and invasion, possibly implicating GAMA as a regulator of microneme function.AUTHOR SUMMARYMalaria remains a major source of morbidity and mortality across the globe. Completion of a complex life cycle between vertebrates and mosquitoes is required for the maintenance of parasite populations and the persistence of malaria disease and death. Three invasive forms across the complex lifecycle of the parasite must successfully egress and invade specific cell types within the vertebrate and mosquito hosts to maintain parasite populations and consequently disease and suffering. A conserved feature of all invasive forms are the micronemes, apically oriented secretory organelles which contain proteins required for motility, egress and invasion. Few proteins are expressed in the micronemes of all three invasive forms. One such protein is GPI-anchored micronemal antigen (GAMA). Here we reveal that GAMA is required for the invasion of the mosquito midgut, egress of sporozoites from oocysts and invasion of the vertebrate host. Our finding indicate that while GAMA is essential for sporozoite motility, the defects in oocyst egress and hepatocyte invasion occur independently of the motility defect, implicating the requirement of GAMA in all three processes.