Although ribosome profiling (Ribo-seq) has revealed widespread translation of non-canonical open reading frames (ncORFs), integrating this information into routine mass spectrometry (MS) workflows remains challenging due to biochemical constraints and large search databases. We retrained TIS Transformer, a transcriptomic language model, by incorporating 40,832 Ribo-seq-derived non-canonical translation initiation sites, enabling the prediction of 48,265 ncORFs, including non-AUG starts, while preserving canonical protein detection. We combined these predictions with Swiss-Prot sequences to create Swiss-Prot/MicroProt, a size-controlled database suitable for standard proteomic analysis. Reanalysis of a comprehensive HeLa MS dataset using this database maintained over 98% concordance with Swiss-Prot searches, recovering 13,128 canonical proteins and identifying 348 microproteins, 270 of which had orthogonal evidence of their existence. This unified framework facilitates systematic microprotein detection in conventional proteomics workflows, bridging the gap between canonical and non-canonical proteome analyses without requiring specialized experimental protocols.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
Non-canonical (i.e., unannotated) open reading frames (ncORFs) have until recently been omitted from reference genome annotations, despite evidence of their translation, limiting their incorporation into biomedical research. To address this, in 2022, we initiated the TransCODE consortium and built the first community-driven consensus catalog of human ncORFs, which was openly distributed to the research community via Ensembl-GENCODE. While this catalog represented a starting point for reference ncORF annotation, major technical and scientific issues remained. In particular, this initial catalogue had no standardized framework to judge the evidence of translation for individual ncORFs. Here, we present an expanded and refined catalog of the human reference annotation of ncORFs. By incorporating more datasets and by lifting constraints on ORF length and start-codon, we define a comprehensive set of 28,359 ncORFs that is nearly four times the size of the previous catalog. Furthermore, to aid users who wish to work with ncORFs with the strongest and most reproducible signals of translation, we utilized a data-driven framework (i.e. translation signature scores) to assess the accumulated evidence for any individual ncORF. Using this approach, we derive a subset of 7,888 ncORFs with translation evidence on par with canonical protein-coding genes, which we refer to as the Primary set. This set can serve as a reliable reference for downstream analyses and validation, with a particular emphasis on high quality. Overall, this update reflects continual community-driven efforts to make ncORFs accessible and actionable to the broader research public and further iterations of the catalog will continue to expand and refine this resource.
A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1–3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of ‘peptideins’ as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome. A large-scale proteomics analysis of the dark proteome by the TransCODE Consortium reveals many translated non-canonical open reading frames to encode microproteins and peptideins.
Mitochondrial alternative open reading frames (ORFs) substantially broaden the functional scope traditionally attributed to mitochondrial DNA, encoding peptides and proteins that participate in diverse cellular processes. These newly identified ORFs are embedded within annotated sequences, both coding and non-coding, and reveal layers of overlapping genetic information. We report the discovery of MTALTCO1, a 259 amino-acid protein, the longest mitochondrial alternative protein identified to date, encoded by an ORF located within the human cytochrome oxidase 1 gene, in the +3 reading frame. We confirm the expression and mitochondrial origin of MTALTCO1 through multiple independent lines of evidence, including a custom-designed antibody, mass spectrometry-derived peptides, sequence analysis and inhibitors of mitochondrial expression. Despite encoding AGR codons as arginine, contrary to the prevailing view that these function invariably as stop codons in the vertebrate mitochondrial genetic code, MTALTCO1 shows strong evidence of mitochondrial translation, challenging established models of mitochondrial codon usage and gene expression. Co-immunoprecipitations and pulldown assays delineate MTALTCO1's interaction landscape across major cellular pathways. Finally, we present the first in-depth analysis of conservation for a mitochondrial alternative ORF overlapping a reference protein-coding gene and discuss the results in light of MTALTCO1's suggested role in protein scaffolding. This article is part of the theme issue 'Evolutionary genetics of mitochondria: on diverse and common evolutionary constraints across eukarya'.
SUMMARY Non-canonical open reading frames (ncORFs) are an emerging area of research that is quickly gaining momentum. Many peptides and proteins missed in initial annotation efforts (ncProts) were subsequently shown to be crucial for a wide range of biological processes. The discovery of ncORFs continues to improve the accuracy of loss-of-function studies because they often occupy the same genomic spaces as annotated ORFs. While databases of mutant phenotypes linked to genomic loci exist in a few species, none of these databases integrate the information on ncORFs present in already characterized loci. In this study, we introduce a nearly comprehensive loss-of-function phenomics dataset of Medicago truncatula (673 loci characterized over the past 30 years), which was integrated as a new track into the genome browser of this organism. This dataset helped critically analyze the potential contribution of ncORFs to published phenotypes. We detected mass spectrometry (MS)-validated ncORFs in 10 characterized genes, including major regulators of development and symbiotic relationships. We also found conserved ncORFs in 113 characterized genes, including four genes with highly conserved ncORFs. In some studies, the contribution of these ncORFs can be ruled out, while in others it cannot. Using real examples, we systematized ambiguities associated with ncORFs. Furthermore, we highlighted little-known trans effects of insertional mutagenesis on splicing as contributors to that ambiguity. Finally, our meta-analysis of published phenotypes revealed that different protein classes have significantly different (unique) proportions of unconditional, conditional, and neutral phenotypes, potentially reflecting their relative functional importance. Significance statement This study is the first to merge a nearly comprehensive inventory of loss-of-function studies in a eukaryotic organism with the information on novel MS-validated and conserved ncORFs.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
BACKGROUND:Most mutations in the COL6A3 gene lead to collagen VI-related myopathies. This is due to a reduced expression or mislocalization of the COL6A3 protein. Therefore, studying the consequence of knocking out the Col6a3 gene in mouse models is relevant, but the Col6a3 mouse models reported so far do not entirely abolish COL6A3 protein expression. METHODS:Here, we present the development, validation and preliminary phenotypic characterization of a novel CRISPR-based knockout mouse model targeting Col6a3 exon 3 (Col6a3d3/d3). RESULTS:In this mouse model, Col6a3 mRNA is still expressed at a similar level to wild-type littermates, although the expected protein is undetectable by mass spectrometry. Histological analysis of Col6a3d3/d3 quadriceps revealed an abnormally high frequency of muscle cells with internally nucleated muscle cells, consistent with a myopathy phenotype. Interestingly, Col6a3d3/d3 mice are smaller in size, with their fat, muscle, and bone kept proportional compared to wild-type littermates. CONCLUSIONS:In summary, we performed the validation and preliminary phenotypic characterization of a novel Col6a3 knockout mouse model that could be further characterized and used to study COL6A3 biology and model collagen VI-associated diseases.
The high complexity of eukaryotic organisms enabled their evolutionary success, driven by the diversification of their proteomes. Various mechanisms contributed to this process. Alternative splicing had the largest known impact among these mechanisms. Earlier, we hypothesized that along with alternative splicing, a different but conceptually similar mechanism creates novel versions of existing proteins in all eukaryotes. However, this mechanism operates at the level of translation, where amino acid sequence novelty arises through multiple programmed ribosomal frameshifting events occurring within the same transcript. This mechanism, which is termed mosaic translation, is very difficult to demonstrate even with the most up-to-date molecular tools. Thus, it remained unnoticed so far. Using a subset of mass spectrometry proteomic data from various organs of the model plant Medicago truncatula, we took the first step toward experimental validation of this hypothesis. Our original in silico approach resulted in the discovery of two candidates for mosaic proteins (homologs of EF1α and RuBisCo) and 154 candidates for chimeric peptides. Chimeric peptides and polypeptides are produced in the course of one ribosomal frameshifting event and may correspond to parts of mosaic proteins. In addition, our analysis reveals the possibility of translation of chimeric peptides from five ribosomal RNA transcripts, ten long non-coding RNA transcripts, and one transfer RNA transcript. These findings are novel and will form the basis for future experimental validation. We also present multiple lines of indirect evidence supporting the validity of our in silico data.
Recently, proteomics analyses using databases of unannotated ORFs revealed that ubiquitin (Ub) variants can be encoded and expressed from pseudogenes. One such pseudogene, UBBP4, produces UbKEKS, which contains four substitutions (Q2K, K33E, Q49K, and N60S) relative to canonical Ub. Unlike Ub, UbKEKS does not promote proteasomal degradation through K48 linkages and instead modifies a distinct set of proteins. To elucidate the structural basis of this divergence, we solved the NMR solution structure of UbKEKS and characterized its backbone dynamics by 15N-relaxation. While UbKEKS retains the overall helix-grip fold, we observed significant rearrangements and amplified motions in residues governing the Ub pincer mode, a conformational switch that determines whether UIMs engage the canonical I44 interface or the α1–β3 edge. Specifically, Q2K and K33E cooperate to enhance motions on both fast (ps–ns) and slow (µs–ms) timescales within α1, the β1–β2 loop, and β5—regions central to pincer mode regulation. In addition, Q49K, adjacent to I44, perturbs UIM recognition and likely interferes with K48 chain formation and binding to the proteasomal receptor S5a. Collectively, our findings identify structural and dynamical determinants that explain UbKEKS's distinct substrate profile and inability to target proteins for degradation.
A major scientific drive is to characterize the protein-coding genome as it provides the primary basis for the study of human health. But the fundamental question remains: what has been missed in prior genomic analyses? Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states, with major implications for proteomics, genomics, and clinical science. However, the impact of ncORFs has been limited by the absence of a large-scale understanding of their contribution to the human proteome. Here, we report the collaborative efforts of stakeholders in proteomics, immunopeptidomics, Ribo-seq ORF discovery, and gene annotation, to produce a consensus landscape of protein-level evidence for ncORFs. We show that at least 25% of a set of 7,264 ncORFs give rise to translated gene products, yielding over 3,000 peptides in a pan-proteome analysis encompassing 3.8 billion mass spectra from 95,520 experiments. With these data, we developed an annotation framework for ncORFs and created public tools for researchers through GENCODE and PeptideAtlas. This work will provide a platform to advance ncORF-derived proteins in biomedical discovery and, beyond humans, diverse animals and plants where ncORFs are similarly observed.
Pseudogenes, traditionally considered non-functional gene copies, have garnered attention due to emerging evidence of their transcription and translation. Ubiquitin is canonically expressed from UBA52 and RPS27A genes as fusion proteins, with additional polyubiquitin precursors encoded by UBB and UBC. Several pseudogenes of these loci are annotated as non-functional. Here, we report that the RPS27A pseudogene, RPS27AP5, expresses two proteins: a ubiquitin variant (UbP5) and a ribosomal protein variant (S27aP5). These proteins mature through cleavage and exhibit localization and biochemical characteristics similar to their parental counterparts. S27aP5 integrates into ribosomes, and its overexpression leads to an increased 80S monosome fraction. Using affinity purification and polysome profiling, we show that S27aP5-containing ribosomes exhibit altered mRNA associations. The findings suggest that RPS27A, a processed pseudogene, can give rise to a ribosomal protein variant capable of integrating into monosomes and influencing mRNA association aligns with growing evidence that ribosomes may exhibit functional diversity.
Pseudogenes, traditionally considered non-functional gene copies resulting from evolutionary mutations, have garnered attention due to recent transcriptomics and proteomics revealing their unexpected expressions and consequential cellular functions. Ubiquitin, transcribed from UBA52 and RPS27A genes, fused to ribosomal proteins eL40 and eS31, and polyubiquitin precursors encoded by UBB and UBC genes, has additional pseudogenes labeled as non-functional. However, recent evidence challenges this notion, demonstrating that these pseudogenes produce ubiquitin variants with minimal differences from the canonical sequence, suggesting a new regulatory dimension in ubiquitin-mediated cellular processes. To systematically catalogue possible Ubiquitin (Ub) and Ubiquitin-like (Ubl) variants from pseudogenes, expression data was compiled, identifying potential functional variants. Among these pseudogenes, RPS27AP5 expresses both Ubiquitin variant (UbP5) and ribosomal protein variant (S27aP5), with precursor proteins maturing through cleavage and exhibiting behavior similar to their counterparts post-translation. Notably, S27aP5 integrates into translating ribosomes, increasing the 80S monosomal ribosomal fraction and indirectly influencing p16INK4A transcriptional activation. The discovery of a functional S27a pseudogene supports the concept that a subset of ribosomes may incorporate diverse subunits for specific translational functions.
Proteogenomics is becoming a powerful tool in personalized medicine by linking genomics, transcriptomics and mass spectrometry (MS)-based proteomics. Due to increasing evidence of alternative open reading frame-encoded proteins (AltProts), proteogenomics has a high potential to unravel the characteristics, variants, expression levels of the alternative proteome, in addition to already annotated proteins (RefProts). To obtain a broader view of the proteome of ovarian cancer cells compared to ovarian epithelial cells, cell-specific total RNA-sequencing profiles and customized protein databases were generated. In total, 128 RefProts and 30 AltProts were identified exclusively in SKOV-3 and PEO-4 cells. Among them, an AltProt variant of IP_715944, translated from DHX8, was found mutated (p.Leu44Pro). We show high variation in protein expression levels of RefProts and AltProts in different subcellular compartments. The presence of 117 RefProt and two AltProt variants was described, along with their possible implications in the different physiological/pathological characteristics. To identify the possible involvement of AltProts in cellular processes, cross-linking-MS (XL-MS) was performed in each cell line to identify AltProt-RefProt interactions. This approach revealed an interaction between POLD3 and the AltProt IP_183088, which after molecular docking, was placed between POLD3-POLD2 binding sites, highlighting its possibility of the involvement in DNA replication and repair.
The OpenProt proteogenomic resource (https://www.openprot.org/) provides users with a complete and freely accessible set of non-canonical or alternative open reading frames (AltORFs) within the transcriptome of various species, as well as functional annotations of the corresponding protein sequences not found in standard databases. Enhancements in this update are largely the result of user feedback and include the prediction of structure, subcellular localization, and intrinsic disorder, using cutting-edge algorithms based on machine learning techniques. The mass spectrometry pipeline now integrates a machine learning-based peptide rescoring method to improve peptide identification. We continue to help users explore this cryptic proteome by providing OpenCustomDB, a tool that enables users to build their own customized protein databases, and OpenVar, a genomic annotator including genetic variants within AltORFs and protein sequences. A new interface improves the visualization of all functional annotations, including a spectral viewer and the prediction of multicoding genes. All data on OpenProt are freely available and downloadable. Overall, OpenProt continues to establish itself as an important resource for the exploration and study of new proteins.
Mitochondrial derived peptides and proteins significantly expand the coding potential of the human mitogenome. Here, we report the discovery of MTALTCO1, a 259 amino acid protein encoded by a mitochondrial alternative open reading frame (mtaltORF) found in the +3 reading frame of the cytochrome oxidase 1 (CO1) gene. Using custom antibodies, we confirmed the mitochondrial expression of MTALTCO1 in human cell lines. Sequence analysis revealed high arginine content and an elevated isoelectric point that were not contingent on CO1's amino acid sequence, suggesting selective pressures acting on this protein. MTALTCO1 displays extensive fusion-fission dynamics at the interspecies level, yet produces a full-length protein throughout human haplogroups. Our findings highlight the importance of identifying novel mtaltORFs in expanding our understanding of the mitochondrial proteome. ### Competing Interest Statement The authors have declared no competing interest.
Proteogenomics has enabled the detection of novel proteins encoded in noncanonical or alternative open reading frames (altORFs) in genes already coding a reference protein. Reanalysis of proteomic and ribo-seq data revealed that the p53-induced death domain-containing protein (orPIDD1) gene encodes a second 171 amino acid protein, altPIDD1, in addition to the known 910-amino acid-long PIDD1 protein. The two ORFs overlap almost completely, and the translation initiation site of altPIDD1 is located upstream of PIDD1. AltPIDD1 has more translational and protein level evidence than PIDD1 across various cell lines and tissues. In HEK293 cells, the altPIDD1 to PIDD1 ratio is 40 to 1, as measured with isotope-labeled (heavy) peptides and targeted proteomics. AltPIDD1 localizes to cytoskeletal structures labeled with phalloidin and interacts with cytoskeletal proteins. Unlike most noncanonical proteins, altPIDD1 is not evolutionarily young but emerged in placental mammals. Overall, we identifyPIDD1as a dual-coding gene, with altPIDD1, not the annotated protein, being the primary product of translation.
Proteogenomics has revealed the translation of unannotated open reading frames (ORFs) present in mRNAs and in noncoding RNAs (ncRNAs). OpenProt annotates all ORFs with a minimum of 30 codons in the transcriptome of several species and displays many functional features associated with the corresponding proteins. Two types of proteins are annotated: reference or canonical proteins which are proteins already annotated in UniProt, RefSeq, or Ensembl and noncanonical proteins. Noncanonical proteins form two groups: predicted novel isoforms that display a significant level of homology with a reference protein and alternative proteins that are new proteins with no significant homology to known proteins. This chapter describes how to check whether a gene and/or transcript contains multiple open reading frames and how to use OpenProt databases for the detection of alternative proteins and novel isoforms by mass spectrometry-based proteomics.
Proteogenomics has enabled the detection of novel proteins encoded in noncanonical or alternative open reading frames (altORFs) in genes already coding a reference protein. Reanalysis of proteomic and ribo-seq data revealed that the p53-induced death domain-containing protein (or PIDD1) gene encodes a second 171 amino acid protein, altPIDD1, in addition to the known 910-amino acid-long PIDD1 protein. The two ORFs overlap almost completely, and the translation initiation site of altPIDD1 is located upstream of PIDD1. AltPIDD1 has more translational and protein level evidence than PIDD1 across various cell lines and tissues. In HEK293 cells, the altPIDD1 to PIDD1 ratio is 40 to 1, as measured with isotope-labeled (heavy) peptides and targeted proteomics. AltPIDD1 localizes to cytoskeletal structures labeled with phalloidin and interacts with cytoskeletal proteins. Unlike most noncanonical proteins, altPIDD1 is not evolutionarily young but emerged in placental mammals. Overall, we identify PIDD1 as a dual-coding gene, with altPIDD1, not the annotated protein, being the primary product of translation.