AMD1 encodes adenosylmethionine decarboxylase 1 (AMD1), a key enzyme in polyamine biosynthesis. A subset of ribosomes translating the AMD1 coding sequence read through the stop codon and pause at a second in-frame stop 384 nucleotides downstream, producing a conserved C-terminal extension (C-tail). Despite growing evidence that such cis-acting elements regulate translation of their genes, the molecular mechanism by which the C-tail mediates ribosome stalling remains unclear. Here, we determined the structure of the ribosome nascent chain complex paused by the AMD1 C-tail which traps eukaryotic release factor 1 (eRF1) with the ATP-binding cassette subfamily E member 1 (ABCE1). The nascent chain forms a molecular clamp that positions an arginine hook in the peptidyl-transferase center, occluding the accommodation of the eRF1 GGQ motif thereby hampering translation termination. Analysis of aggregated ribosome profiling data revealed several genes with a pattern of stop codon readthrough followed by ribosome stalling at a specific location, suggesting that regulatory readthrough-stall mechanisms may not be limited to AMD1.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
Non-canonical (i.e., unannotated) open reading frames (ncORFs) have until recently been omitted from reference genome annotations, despite evidence of their translation, limiting their incorporation into biomedical research. To address this, in 2022, we initiated the TransCODE consortium and built the first community-driven consensus catalog of human ncORFs, which was openly distributed to the research community via Ensembl-GENCODE. While this catalog represented a starting point for reference ncORF annotation, major technical and scientific issues remained. In particular, this initial catalogue had no standardized framework to judge the evidence of translation for individual ncORFs. Here, we present an expanded and refined catalog of the human reference annotation of ncORFs. By incorporating more datasets and by lifting constraints on ORF length and start-codon, we define a comprehensive set of 28,359 ncORFs that is nearly four times the size of the previous catalog. Furthermore, to aid users who wish to work with ncORFs with the strongest and most reproducible signals of translation, we utilized a data-driven framework (i.e. translation signature scores) to assess the accumulated evidence for any individual ncORF. Using this approach, we derive a subset of 7,888 ncORFs with translation evidence on par with canonical protein-coding genes, which we refer to as the Primary set. This set can serve as a reliable reference for downstream analyses and validation, with a particular emphasis on high quality. Overall, this update reflects continual community-driven efforts to make ncORFs accessible and actionable to the broader research public and further iterations of the catalog will continue to expand and refine this resource.
A major scientific drive is to characterize the protein-coding genome, which is a primary basis for studying human health. But the fundamental question remains of what has been missed in previous analyses. Over the past decade, the translation of non-canonical open reading frames (ncORFs) has been observed across human cell types and disease states1–3, with major implications for biomedical science. However, a key gap in knowledge has been which ncORFs produce small microproteins or alternative protein molecules that contribute to the human proteome. Here we report the collaborative efforts of the TransCODE Consortium4 to produce a consensus landscape of protein-level evidence for ncORFs. We show that about 25% of a set of 7,264 ncORFs gives rise to detectable peptides in a large-scale analysis of 95,520 proteomics experiments. We develop an annotation framework for ncORF-encoded microproteins as human proteins and codify the new conceptual model of ‘peptideins’ as microproteins that have indeterminate potential as functional proteins. To probe the biological implications of peptideins, we create an evolutionary analysis approach, termed ORF relative branch length (ORBL), and determine that evolutionary constraint is common and associates with observation of ncORF-derived peptides. We then characterize a pan-essential cellular phenotype for one peptidein from the OLMALINC long non-coding RNA. Overall, we generate public research tools supported by GENCODE and PeptideAtlas and advance biomedical discovery for understudied components of the human proteome. A large-scale proteomics analysis of the dark proteome by the TransCODE Consortium reveals many translated non-canonical open reading frames to encode microproteins and peptideins.
We show that chronic impairment of mitochondrial respiration is associated with marked accumulation of cytochrome c (Cytc) protein. Using SCO2-deficient HCT116 cells lacking functional cytochrome c oxidase and wild-type cells exposed to sustained hypoxia, we found that substantial mitochondrial Cytc accumulation parallels reduced electron flux through Cytc. SCO2-deficient cells exhibited equally elevated Cytc levels under normoxia (19% O2) and hypoxia (0.1-3% O2). Wild-type cells under sustained hypoxia accumulated Cytc, reaching levels comparable to those in SCO2-deficient cells. This effect was reversible upon reoxygenation. Increased Cytc protein levels were also observed in other cell models, including primary cortical neurons cultured under chronic hypoxia and in cerebral cortex tissue from hypoxia-exposed mice. Cytc accumulation occurred independently of CYCS transcription, mRNA translation, HIF activation, ROS production and changes in mitochondrial network. Pharmacological inhibition of complex III was likewise accompanied by increased Cytc levels, whereas mitochondrial uncoupling had no effect, suggesting that impaired electron transfer rather than membrane depolarisation per se underlies this association. Raman spectroscopy revealed enrichment of reduced Cytc and an increased Cytc-to-cytochrome b ratio in respiration-deficient cells. Further supporting a stabilisation-based mechanism, the fraction of membrane-unbound ferro-Cytc was decreased in SCO2-deficient cells, consistent with moderate cardiolipin enrichment, which is known to enhance retention of Cytc at the inner mitochondrial membrane. Despite elevated mitochondrial Cytc content, SCO2-deficient cells were less susceptible to apoptosis induced by intermittent hypoxia or dichloroacetate. Together, these findings indicate that reduced electron flux through complex IV is associated with Cytc accumulation through increased protein stability and membrane retention without enhancing apoptotic sensitivity.
BACKGROUND:Nucleotide sequence can be translated in three reading frames producing distinct protein products. Many examples of RNA translation in two reading frames (dual coding) have been identified so far. RESULTS:We report translation of mRNA transcripts derived from SRD5A1 locus in all three reading frames that result in the synthesis of long polypeptides. This occurs due to initiation at three nearby AUG codons occurring in all three reading frames. Only one of the three proteoforms contains the conserved catalytical domain of SRD5A1 produced either from the second or the third AUG codon depending on the transcript. Paradoxically, ribosome profiling data and expression reporters indicate that the most efficient translation would produce catalytically inactive polypeptide. While phylogenetic analysis suggests that the long triple decoding region is specific to primates, occurrence of nearby AUGs in all three reading frames is ancestral to placental mammals. This suggests that their evolutionary significance belongs to regulation of translation rather than biological role of their products. By analysing multiple publicly available ribosome profiling data and with gene expression assays carried out in different cellular environments, we show that relative expression of these proteoforms is mutually dependent and varies across environments supporting this conjecture. We show that a remarkable feature of triple decoding is its resistance to frameshift causing variants with apparent implications to clinical interpretation of genomic sequence variants. CONCLUSIONS:We argue for the importance of identification, characterisation and annotation of productive RNA translation irrespective of the presumed biological roles of its products.
Thousands of short open reading frames (sORFs) are translated outside of annotated coding sequences. Recent studies have pioneered searching for sORF-encoded microproteins in mass spectrometry (MS)-based proteomics and peptidomics datasets. Here, we assessed literature-reported MS-based identifications of unannotated human proteins. We find that studies vary by three orders of magnitude in the number of unannotated proteins they report. Of nearly 10,000 reported sORF-encoded peptides, 96% were unique to a single study, and 12% mapped to annotated proteins or proteoforms. Manual curation of a benchmark dataset of 406 manually evaluated spectra from 204 sORF-encoded proteins revealed large variation in peptide-spectrum match (PSM) quality between studies, with immunopeptidomics studies generally reporting higher quality PSMs than conventional enzymatic digests of whole cell lysates. We estimate that 65% of predicted sORF-encoded protein detections in immunopeptidomics studies were supported by high-quality PSMs versus 7.8% in non-immunopeptidomics datasets. Our work stresses the need for standardized protocols and analysis workflows to guide future advancements in microprotein detection by MS towards uncovering how many human microproteins exist.
Translational control shapes the proteome and is particularly important in regulating gene expression under stress. A key source of endothelial stress is treatment with tyrosine kinase inhibitors (TKIs), which lowers cancer mortality but increases cardiovascular mortality. Using a human induced pluripotent stem cell-derived endothelial cell (hiPSC-EC) model of sunitinibinduced vascular dysfunction combined with ribosome profiling, we assessed the role of translational control in hiPSC-ECs in response to stress. We identified staphylococcal nuclease and tudor domain-containing protein 1 (SND1) as a sunitinibdependent translationally repressed gene. SND1 translational repression was mediated by the mTORC1/4E-BP1 pathway. SND1 inhibition led to endothelial dysfunction, whereas SND1 OE protected against sunitinib-induced endothelial dysfunction. Mechanistically, SND1 transcriptionally regulated UBE2N, an E2-conjugating enzyme that mediates K63-linked ubiquitination. UBE2N along with the E3 ligases RNF8 and RNF168 regulated the DNA damage repair response pathway to mitigate the deleterious effects of sunitinib. In silico analysis of FDA-approved drugs led to the identification of an ACE inhibitor, ramipril, that protected against sunitinib-induced vascular dysfunction in vitro and in vivo, all while preserving the efficacy of cancer therapy. Our study established a central role fortranslational control of SND1 in sunitinib-induced endothelial dysfunction that could potentially be therapeutically targeted to reduce sunitinib-induced vascular toxicity.
Ribosomal frameshifting is an important, albeit rare, mRNA decoding mechanism that generally allows the synthesis of a single protein from two different reading frames. +1 frameshifting is commonly presumed to involve re-pairing of the P-site tRNA with the +1 codon. However, in several occurrences in the yeast Saccharomyces cerevisiae, P-site tRNA re-pairing with the +1 codon is impossible. In one model, +1 frameshifting occurs according to a common mechanism involving P-site tRNA movement without re-pairing with the +1 codon. The alternative is a distinct mechanism allowing A-site tRNA acceptance at the +1 codon in the absence of P-site tRNA movement. Here, we experimentally compared all known +1 ribosomal frameshifting sites in S. cerevisiae, including a novel case discovered during this study in LLP1. We identified a conserved RNA secondary structure upstream of the ABP140 frameshifting site that increases frameshifting efficiency. The location of the structure suggests that it creates an mRNA-pulling effect favouring +1 codon in the P-site. Placing the stimulator upstream of various known frameshifting sites revealed that its stimulatory action is selective to those frameshifting sites where P-site tRNA re-pairing is possible, reinforcing the idea of two distinct mechanisms of +1 ribosomal frameshifting.
Dual reporters encoding two distinct proteins within the same mRNA have had a crucial role in identifying and characterizing unconventional mechanisms of eukaryotic translation. These mechanisms include initiation via internal ribosomal entry sites (IRESs), ribosomal frameshifting, stop codon readthrough and reinitiation. This design enables the expression of one reporter to be influenced by the specific mechanism under investigation, while the other reporter serves as an internal control. However, challenges arise when intervening test sequences are placed between these two reporters. Such sequences can inadvertently impact the expression or function of either reporter, independent of translation-related changes, potentially biasing the results. These effects may occur due to cryptic regulatory elements inducing or affecting transcription initiation, splicing, polyadenylation and antisense transcription as well as unpredictable effects of the translated test sequences on the stability and activity of the reporters. Unfortunately, these unintended effects may lead to misinterpretation of data and the publication of incorrect conclusions in the scientific literature. To address this issue and to assist the scientific community in accurately interpreting dual-reporter experiments, we have developed comprehensive guidelines. These guidelines cover experimental design, interpretation and the minimal requirements for reporting results. They are designed to aid researchers conducting these experiments as well as reviewers, editors and other investigators who seek to evaluate published data.
Programmed ribosomal frameshifting is a process where a proportion of ribosomes change their reading frame on an mRNA. While frameshifting is commonly used by viruses, very few phylogenetically conserved examples are known in nuclear encoded genes. Here, we report a +1 frameshifting event during decoding of the human gene PLEKHM2 that provides access to a second internally overlapping ORF. The new carboxyl-terminal domain of this frameshift protein forms an α helix, which relieves PLEKHM2 from autoinhibition and allows it to move to the tips of cells without activation by ARL8. Reintroducing both the canonically translated and frameshifted protein are necessary to restore normal contractile function of PLEKHM2 knockout cardiomyocytes, demonstrating the necessity of frameshifting for normal cardiac activity.
Upstream open reading frames (uORFs) are a widespread class of translated regions (translons) occurring in 5′ leaders of mRNAs and serving critical roles in post-transcriptional regulation. However, their specific biological activities in human cells remains to be fully elucidated. Here, we conducted a genome-wide CRISPR-Cas9 loss-of-function screen of 978 uORFs identified with ribosome profiling, across human cell lines of distinct origin (HAP1, A549 and HEK293T). A total of 155 uORFs were identified as being essential for cell proliferation. These uORFs showed a high cell-type specificity, with only a few being universally essential. Subsequent analysis has revealed that the primary reason underlying the uORF essentiality is not encoded micropeptides, but rather cis -regulatory mechanisms. Moreover, uORFs located within short 5′ UTRs were disproportionately sensitive to frameshift-inducing indels, which frequently lead to the uORF extension and overlap with the coding region (CDS), resulting in translational repression. Finally, by intersecting regions of essential uORF with ClinVar and dbSNP datasets, we identified naturally occurring variants with the potential to disrupt their function and contribute to disease phenotypes. These findings highlight a pervasive and underappreciated layer of translational control in human cells and establish uORFs as critical cis -regulatory elements with potential relevance to human health. ### Competing Interest Statement Pavel V Baranov is a co-founder and a shareholder of Eirna Bio. The remaining authors declare no competing interests. Russian Science Foundation, https://ror.org/03y2gwe85, 23-14-00058
Ribosome profiling is a powerful technique used to study gene expression on a transcriptome-wide scale. It involves sequencing of mRNA fragments protected by ribosomes from ribonuclease digestion. The initial steps commonly involve cell lysis followed by centrifugation and ribonuclease digestion. We find that centrifugation depletes 329 translated mRNAs in HEK293T cells. Many of these mRNAs encode cytoskeleton proteins. This suggests that the expression of a subset of mRNAs may be significantly underestimated in most ribosome profiling experiments. We show that omitting the centrifugation step after cell lysis can resolve this issue.
Stimulation of resting T cells triggers a rapid transition into an activated state, characterized by significant changes in phenotype, metabolism and secretory function. To explore the earliest responses, we examined transcriptional and translational changes in isolated human T cells within the first four hours post-stimulation. Polyclonal stimulation initially upregulates cytokine genes while downregulating certain transcription factors and regulators. Subsequently, selective mRNA translation occurs on 5⍰-terminal oligopyrimidine (TOP) motif-containing mRNAs and ATF4 , alongside global transcriptional reprogramming involving alternative transcription and splicing. Notably, both stimulated and unstimulated T cells undergo dynamic gene expression changes over time, reflecting adaptation to in vitro conditions and the loss of in vivo homeostatic signals. This includes heightened translation initiation stringency, particularly in non-stimulated cells. Thus, stress responses induced by standard T cell isolation protocols must be considered when investigating T cell signaling and activation. ### Competing Interest Statement P.V.B, A.M.M. and G.L. are co-founders and shareholders of EIRNAbio Ltd. Russian Science Foundation, 24-14-00213 to D.E.A. (data analysis), 24-15-00332 to Y.P.R. (experiments with T cells) Research Ireland Centre for Research Training in Genomics Data Science, 18/CRT/6214 to P.O’B Irish Research Council Enterprise Partnership Studentship Award, EPSPG/2020/461 to A.O’C Taighde Éireann – Research Ireland Frontiers for the Future Award, 20/FFP-A/8929 to P.V.B
The application of ribosome profiling has revealed an unexpected abundance of translation in addition to that responsible for the synthesis of previously annotated protein-coding regions. Multiple short sequences have been found to be translated within single RNA molecules, within both annotated protein-coding and noncoding regions. The biological significance of this translation is a matter of intensive investigation. However, current schematic or annotation-based representations of mRNA translation generally do not account for the apparent multitude of translated regions within the same molecules. They also do not take into account the stochasticity of the process that allows alternative translations of the same RNA molecules by different ribosomes. There is a need for formal representations of mRNA complexity that would enable the analysis of quantitative information on translation and more accurate models for predicting the phenotypic effects of genetic variants affecting translation. To address this, we developed a conceptually novel abstraction that we term ribosome decision graphs (RDGs). RDGs represent translation as multiple ribosome paths through untranslated and translated mRNA segments. We termed the latter “translons.” Nondeterministic events, such as initiation, reinitiation, selenocysteine insertion, or ribosomal frameshifting, are then represented as branching points. This representation allows for an adequate representation of eukaryotic translation complexity and focuses on locations critical for translation regulation. We show how RDGs can be used for depicting translated regions and for analyzing genetic variation and quantitative genome-wide data on translation for characterization of regulatory modulators of translation.
ABSTRACT Bacteria have evolved diverse defense mechanisms to counter bacteriophage attacks. Genetic programs activated upon infection characterize phage–host molecular interactions and ultimately determine the outcome of the infection. In this study, we applied ribosome profiling to monitor protein synthesis during the early stages of sk1 bacteriophage infection in Lactococcus cremoris . Our analysis revealed major changes in gene expression within 5 minutes of sk1 infection. Notably, we observed a specific and severe downregulation of several pyr operons which encode enzymes required for uridine monophosphate biosynthesis. Consistent with previous findings, this is likely an attempt of the host to starve the phage of nucleotides it requires for propagation. We also observed a gene expression response that we expect to benefit the phage. This included the upregulation of 40 ribosome proteins that likely increased the host’s translational capacity, concurrent with a downregulation of genes that promote translational fidelity ( lepA and raiA ). In addition to the characterization of host–phage gene expression responses, the obtained ribosome profiling data enabled us to identify two putative recoding events as well as dozens of loci currently annotated as pseudogenes that are actively translated. Furthermore, our study elucidated alterations in the dynamics of the translation process, as indicated by time-dependent changes in the metagene profile, suggesting global shifts in translation rates upon infection. Additionally, we observed consistent modifications in the ribosome profiles of individual genes, which were apparent as early as 2 minutes post-infection. The study emphasizes our ability to capture rapid alterations of gene expression during phage infection through ribosome profiling. IMPORTANCE The ribosome profiling technology has provided invaluable insights for understanding cellular translation and eukaryotic viral infections. However, its potential for investigating host–phage interactions remains largely untapped. Here, we applied ribosome profiling to Lactococcus cremoris cultures infected with sk1, a major infectious agent in dairy fermentation processes. This revealed a profound downregulation of genes involved in pyrimidine nucleotide synthesis at an early stage of phage infection, suggesting an anti-phage program aimed at restricting nucleotide availability and, consequently, phage propagation. This is consistent with recent findings and contributes to our growing appreciation for the role of nucleotide limitation as an anti-viral strategy. In addition to capturing rapid alterations in gene expression levels, we identified translation occurring outside annotated regions, as well as signatures of non-standard translation mechanisms. The gene profiles revealed specific changes in ribosomal densities upon infection, reflecting alterations in the dynamics of the translation process.
Upstream open reading frames (uORFs) are a class of translated regions (translons) in mRNA 5' leaders. uORFs are believed to be pervasive regulators of the translation of mammalian mRNAs. Some uORFs are highly repressive but others have little or no impact on downstream mRNA translation either due to inefficient recognition of their start codon(s) or/and due to efficient reinitiation after uORF translation. While experiments with uORF reporter constructs proved to be instrumental in the investigation of uORF-mediated mechanisms of translation control, they can have serious limitations as manipulations with uORF sequences can yield various artefacts. Here we propose a general approach for using translation complex profiling (TCP-seq) data for exploring uORF regulatory characteristics. Using several examples, we show how TCP-seq could be used to estimate both repressiveness and modes of action of individual uORFs. We demonstrate how this approach could be used to assess the mechanisms of uORF-mediated translation control in the mRNA of several human genes, including EIF5, IFRD1, MDM2, MIEF1, PPP1R15B, TAF7, and UCP2.
Ribosome profiling (Ribo-Seq) has revolutionised our understanding of translation, but the increasing complexity and volume of Ribo-Seq data present challenges for its reuse. Here, we formally introduce RiboSeq.Org, an integrated suite of resources designed to facilitate Ribo-Seq data analysis and visualisation within a web browser. RiboSeq.Org comprises several interconnected tools: GWIPS-viz for genome-wide visualisation, Trips-Viz for transcriptome-centric analysis, RiboGalaxy for data processing and the newly developed RiboSeq data portal (RDP) for centralised dataset identification and access. The RDP currently hosts preprocessed datasets corresponding to 14840 sequence libraries (samples) from 969 studies across 96 species, in various file formats along with standardised metadata. RiboSeq.Org addresses key challenges in Ribo-Seq data reuse through standardised sample preprocessing, semi-automated metadata curation and programmatic information access via a REST API and command-line utilities. RiboSeq.Org enhances the accessibility and utility of public Ribo-Seq data, enabling researchers to gain new insights into translational regulation and protein synthesis across diverse organisms and conditions. By providing these integrated, user-friendly resources, RiboSeq.Org aims to lower the barrier to reproducible research in the field of translatomics and promote more efficient utilisation of the wealth of available Ribo-Seq data. [GRAPHICS] .
Summary: The integrated stress response (ISR) is critical for cell survival under stress. In response to diverse environmental cues, eIF2α becomes phosphorylated, engendering a dramatic change in mRNA translation. The activation of ISR plays a pivotal role in the early embryogenesis, but the eIF2-dependent translational landscape in pluripotent embryonic stem cells (ESCs) is largely unexplored. We employ a multi-omics approach consisting of ribosome profiling, proteomics, and metabolomics in wild-type (eIF2α+/+) and phosphorylation-deficient mutant eIF2α (eIF2αA/A) mouse ESCs (mESCs) to investigate phosphorylated (p)-eIF2α-dependent translational control of naive pluripotency. We show a transient increase in p-eIF2α in the naive epiblast layer of E4.5 embryos. Absence of eIF2α phosphorylation engenders an exit from naive pluripotency following 2i (two chemical inhibitors of MEK1/2 and GSK3α/β) withdrawal. p-eIF2α controls translation of mRNAs encoding proteins that govern pluripotency, chromatin organization, and glutathione synthesis. Thus, p-eIF2α acts as a key regulator of the naive pluripotency gene regulatory network.