m6A is the most widespread mRNA modification and is primarily implicated in controlling mRNA stability. Fundamental questions pertaining to m6A are the extent to which it is dynamically modulated within cells and across stimuli, and the forces underlying such modulation. Prior work has focused on investigating active mechanisms governing m6A levels, such as recruitment of m6A writers or erasers leading to either global or site-specific modulation. Here, we propose that changes in m6A levels across subcellular compartments and biological trajectories may result from passive changes in gene-level mRNA metabolism. To predict the intricate interdependencies between m6A levels, mRNA localization, and mRNA decay, we establish a differential model m6ADyn encompassing mRNA transcription, methylation, export, and m6A-dependent and independent degradation. We validate the predictions of m6ADyn in the context of intracellular m6A dynamics, where m6ADyn predicts associations between relative mRNA localization and m6A levels, which we experimentally confirm. We further explore m6ADyn predictions pertaining to changes in m6A levels upon controlled perturbations of mRNA metabolism, which we also experimentally confirm. Finally, we demonstrate the relevance of m6ADyn in the context of cellular heat stress response, where genes subjected to altered mRNA product and export also display predictable changes in m6A levels, consistent with m6ADyn predictions. Our findings establish a framework for dissecting m6A dynamics and suggest the role of passive dynamics in shaping m6A levels in mammalian systems.
Ribosomal RNA (rRNA) constitutes the core of ribosomes and is extensively chemically modified. Technical challenges have precluded systematically dissecting rRNA modifications and their dynamics. We develop Pan-Mod-seq, permitting inference of 16 distinct modifications across dozens of samples in parallel. We applied Pan-Mod-seq to RNA from 14 species spanning all domains of life, cultured under highly diverse conditions. While dynamic modifications are rare in mesophiles, in extreme hyperthermophiles, ∼50% of modifications are dynamic. We dissect the biogenesis and function of a conserved module of tandem m5C-ac4C modifications, co-induced at high temperatures, via enzymes intrinsically regulated by temperature and required for growth at higher temperatures. Cryo-electron microscopy (cryo-EM) structures of ribosomes from wild-type (WT) and enzyme-deficient archaea reveal recurrent molecular interactions through which they confer structural stability, and biophysical studies demonstrate their synergistic thermostabilizing role. Our findings systematically dissect rRNA modification plasticity and pave the way for surveying the rRNA epitranscriptome in health and disease.
N4-acetylcytidine (ac4C) is a ubiquitous RNA modification incorporated by cytidine acetyltransferase enzymes. Here, we report the biochemical characterization of Thermococcus kodakarensis Nat10 (TkNat10), an RNA acetyltransferase involved in archaeal thermotolerance. We demonstrate that TkNat10's catalytic activity is critical for T. kodakarensis fitness at elevated temperatures. Unlike eukaryotic homologs, TkNat10 exhibits robust stand-alone activity, modifying diverse RNA substrates in a temperature, ATP, and acetyl-CoA-dependent manner. Transcriptome-wide analysis reveals TkNat10 preferentially modifies unstructured RNAs containing a 5'-CCG-3' consensus sequence. Using a high-throughput mutagenesis approach, we define sequence and structural determinants of TkNat10 substrate recognition. We find TkNat10 can be engineered to facilitate use of propionyl-CoA, providing insight into its cofactor specificity. Finally, we demonstrate TkNat10's utility for site-specific acetylation of RNA oligonucleotides, enabling analysis of ac4C-dependent RNA-protein interactions. Our findings establish a framework for understanding archaeal RNA acetylation and a new tool for studying the functional consequences of ac4C in diverse RNA contexts.
N6-methyladenosine (m6A), a widespread destabilizing mark on mRNA, is non-uniformly distributed across the transcriptome, yet the basis for its selective deposition is unknown. Here, we propose that m6A deposition is not selective. Instead, it is exclusion based: m6A consensus motifs are methylated by default, unless they are within a window of ∼100 nt from a splice junction. A simple model which we extensively validate, relying exclusively on presence of m6A motifs and exon-intron architecture, allows in silico recapitulation of experimentally measured m6A profiles. We provide evidence that exclusion from splice junctions is mediated by the exon junction complex (EJC), potentially via physical occlusion, and that previously observed associations between exon-intron architecture and mRNA decay are mechanistically mediated via m6A. Our findings establish a mechanism coupling nuclear mRNA splicing and packaging with the covalent installation of m6A, in turn controlling cytoplasmic decay.
Millions of adenosines are deaminated throughout the transcriptome by ADAR1 and/or ADAR2 at varying levels, raising the question of what are the determinants guiding substrate specificity and how these differ between the two enzymes. We monitor how secondary structure modulates ADAR2 vs ADAR1 substrate selectivity, on the basis of systematic probing of thousands of synthetic sequences transfected into cell lines expressing exclusively ADAR1 or ADAR2. Both enzymes induce symmetric, strand-specific editing, yet with distinct offsets with respect to structural disruptions: −26 nt for ADAR2 and −35 nt for ADAR1. We unravel the basis for these differences in offsets through mutants, domain-swaps, and ADAR homologs, and find it to be encoded by the differential RNA binding domain (RBD) architecture. Finally, we demonstrate that this offset-enhanced editing can allow an improved design of ADAR2-recruiting therapeutics, with proof-of-concept experiments demonstrating increased on-target and potentially decreased off-target editing.
AbstractRNA can be extensively modified post-transcriptionally with >170 covalent modifications, expanding its functional and structural repertoire. Pseudouridine (Ψ), the most abundant modified nucleoside in rRNA and tRNA, has recently been found within mRNA molecules. It remains unclear whether pseudouridylation of mRNA can be snoRNA-guided, bearing important implications for understanding the physiological target spectrum of snoRNAs and for their potential therapeutic exploitation in genetic diseases. Here, using a massively parallel reporter based strategy we simultaneously interrogate Ψ levels across hundreds of synthetic constructs with predesigned complementarity against endogenous snoRNAs. Our results demonstrate that snoRNA-mediated pseudouridylation can occur on mRNA targets. However, this is typically achieved at relatively low efficiencies, and is constrained by mRNA localization, snoRNA expression levels and the length of the snoRNA:mRNA complementarity stretches. We exploited these insights for the design of snoRNAs targeting pseudouridylation at premature termination codons, which was previously shown to suppress translational termination. However, in this and follow-up experiments in human cells we observe no evidence for significant levels of readthrough of pseudouridylated stop codons. Our study enhances our understanding of the scope, ‘design rules’, constraints and consequences of snoRNA-mediated pseudouridylation.
Oligo library pools are powerful tools for systematic investigation of genetic and transcriptomic machinery such as promoter function and gene regulation, non-coding RNAs, or RNA modifications. Here, we provide a detailed protocol for cloning DNA oligo pools made up of tens of thousands of different constructs, aiming to preserve the complexity of the pools. This system would be suitable for expression in cell lines and can be followed up by next-generation sequencing analysis. For complete details on the use and execution of this profile, please refer to Uzonyi et al. (2021).
Adenosine-to-inosine editing is catalyzed by ADAR1 at thousands of sites transcriptome-wide. Despite intense interest in ADAR1 from physiological, bioengineering, and therapeutic perspectives, the rules of ADAR1 substrate selection are poorly understood. Here, we used large-scale systematic probing of ∼2,000 synthetic constructs to explore the structure and sequence context determining editability. We uncover two structural layers determining the formation and propagation of A-to-I editing, independent of sequence. First, editing is robustly induced at fixed intervals of 35 bp upstream and 30 bp downstream of structural disruptions. Second, editing is symmetrically introduced on opposite sites on a double-stranded structure. Our findings suggest a recursive model for RNA editing, whereby the structural alteration induced by the editing at one site iteratively gives rise to the formation of an additional editing site at a fixed periodicity, serving as a basis for the propagation of editing along and across both strands of double-stranded RNA structures.
RNA modifications are present in most cellular RNAs and are formed posttranscriptionally by enzymatic machineries that involve hundreds of enzymes and cofactors. RNA modifications impact the life cycle of the RNA, its stability, folding, cellular localization, as well as interactions with RNA and protein partners. RNA modifications are important for mitochondrial function and are required for proper processing and function of mitochondrial (mt) tRNA and rRNA. Underscoring their importance, several mitochondrial diseases are caused by defects in mt-RNA modifications, stemming from mutations in mtDNA at or near mt-RNA modification sites or in nuclear-encoded mt-RNA modifying enzymes. A highly abundant RNA modification, involved in mitochondrial physiology and pathology is pseudouridylation (Ψ), which is catalyzed by enzymes of the Pseudouridine Synthase (PUS) family. Although some Ψ sites in mt-rRNA and mt-tRNA have been identified, little is known about the functional role of these modifications. Furthermore, it is unknown which enzyme facilitates the modification of each site and it is likely that many yet undiscovered mt-RNA modifications exist, as is evidenced by recent work showing some Ψ sites on mRNA. Here, we present mito-Ψ-Seq, a high-throughput method for semiquantitative mapping of Ψ in mt-RNA.
N4-acetylcytidine (ac4C) is an ancient and highly conserved RNA modification that is present on tRNA and rRNA and has recently been investigated in eukaryotic mRNA1–3. However, the distribution, dynamics and functions of cytidine acetylation have yet to be fully elucidated. Here we report ac4C-seq, a chemical genomic method for the transcriptome-wide quantitative mapping of ac4C at single-nucleotide resolution. In human and yeast mRNAs, ac4C sites are not detected but can be induced—at a conserved sequence motif—via the ectopic overexpression of eukaryotic acetyltransferase complexes. By contrast, cross-evolutionary profiling revealed unprecedented levels of ac4C across hundreds of residues in rRNA, tRNA, non-coding RNA and mRNA from hyperthermophilic archaea. Ac4C is markedly induced in response to increases in temperature, and acetyltransferase-deficient archaeal strains exhibit temperature-dependent growth defects. Visualization of wild-type and acetyltransferase-deficient archaeal ribosomes by cryo-electron microscopy provided structural insights into the temperature-dependent distribution of ac4C and its potential thermoadaptive role. Our studies quantitatively define the ac4C landscape, providing a technical and conceptual foundation for elucidating the role of this modification in biology and disease4–6. A method termed ac4C-seq is introduced for the transcriptome-wide mapping of the RNA modification N4-acetylcytidine, revealing widespread temperature-dependent acetylation that facilitates thermoadaptation in hyperthermophilic archaea.
Despite much research, our understanding of the architecture and cis-regulatory elements of human promoters is still lacking. Here, we devised a high-throughput assay to quantify the activity of approximately 15,000 fully designed sequences that we integrated and expressed from a fixed location within the human genome. We used this method to investigate thousands of native promoters and preinitiation complex (PIC) binding regions followed by in-depth characterization of the sequence motifs underlying promoter activity, including core promoter elements and TF binding sites. We find that core promoters drive transcription mostly unidirectionally and that sequences originating from promoters exhibit stronger activity than those originating from enhancers. By testing multiple synthetic configurations of core promoter elements, we dissect the motifs that positively and negatively regulate transcription as well as the effect of their combinations and distances, including a 10-bp periodicity in the optimal distance between the TATA and the initiator. By comprehensively screening 133 TF binding sites, we find that in contrast to core promoters, TF binding sites maintain similar activity levels in both orientations, supporting a model by which divergent transcription is driven by two distinct unidirectional core promoters sharing bidirectional TF binding sites. Finally, we find a striking agreement between the effect of binding site multiplicity of individual TFs in our assay and their tendency to appear in homotypic clusters throughout the genome. Overall, our study systematically assays the elements that drive expression in core and proximal promoter regions and sheds light on organization principles of regulatory regions in the human genome.
N6-methyladenosine (m6A) is the most abundant modification on mRNA, and is implicated in critical roles in development, physiology and disease. The ability to map m6A using immunoprecipitation-based approaches has played a critical role in dissecting m6A functions and mechanisms of action. Yet, these approaches are of limited specificity, unknown sensitivity, and unable to quantify m6A stoichiometry. These limitations have severely hampered our ability to unravel the factors determining where m6A will be deposited, to which levels (the ‘m6A code’), and to quantitatively profile m6A dynamics across biological systems. Here, we used the RNase MazF, which cleaves specifically at unmethylated RNA sites, to develop MASTER-seq for systematic quantitative profiling of m6A sites at 16-25% of all m6A sites at single nucleotide resolution. We established MASTER-seq for orthogonal validation and de novo detection of m6A sites, and for tracking of m6A dynamics in yeast gametogenesis and in early mammalian differentiation. We discover that antibody-based approaches severely underestimate the number of m6A sites, and that both the presence of m6A and its stoichiometry are ‘hard-coded’ via a simple and predictable code within the extended sequence composition at the methylation sites. This code accounts for ~50% of the variability in methylation levels across sites, allows excellent de novo prediction of methylation sites, and predicts methylation acquisition and loss across evolution. We anticipate that MASTER-seq will pave the path towards a more quantitative investigation of m6A biogenesis and regulation in a wide variety of systems, including diverse cell types, stimuli, subcellular components, and disease states.
N6-methyladenosine (m6A) is the most abundant modification on mRNA and is implicated in critical roles in development, physiology, and disease. A major limitation has been the inability to quantify m6A stoichiometry and the lack of antibody-independent methodologies for interrogating m6A. Here, we develop MAZTER-seq for systematic quantitative profiling of m6A at single-nucleotide resolution at 16%-25% of expressed sites, building on differential cleavage by an RNase. MAZTER-seq permits validation and de novo discovery of m6A sites, calibration of the performance of antibody-based approaches, and quantitative tracking of m6A dynamics in yeast gametogenesis and mammalian differentiation. We discover that m6A stoichiometry is "hard coded" in cis via a simple and predictable code, accounting for 33%-46% of the variability in methylation levels and allowing accurate prediction of m6A loss and acquisition events across evolution. MAZTER-seq allows quantitative investigation of m6A regulation in subcellular fractions, diverse cell types, and disease states.
Despite extensive research, the sequence features affecting microRNA-mediated regulation are not well understood, limiting our ability to predict gene expression levels in both native and synthetic sequences. Here we employed a massively parallel reporter assay to investigate the effect of over 14,000 rationally designed 3' UTR sequences on reporter construct repression. We found that multiple factors, including microRNA identity, hybridization energy, target accessibility, and target multiplicity, can be manipulated to achieve a predictable, up to 57-fold, change in protein repression. Moreover, we predict protein repression and RNA levels with high accuracy (R = 0.84 and R = 0.80, respectively) using only 3' UTR sequence, as well as the effect of mutation in native 3' UTRs on protein repression (R = 0.63). Taken together, our results elucidate the effect of different sequence features on miRNA-mediated regulation and demonstrate the predictability of their effect on gene expression with applications in regulatory genomics and synthetic biology.
Following synthesis, RNA can be modified with over 100 chemically distinct modifications, which can potentially regulate RNA expression post-transcriptionally. Pseudouridine (Ψ) was recently established to be widespread and dynamically regulated on yeast mRNA, but less is known about Ψ presence, regulation, and biogenesis in mammalian mRNA. Here, we sought to characterize the Ψ landscape on mammalian mRNA, to identify the main Ψ-synthases (PUSs) catalyzing Ψ formation, and to understand the factors governing their specificity toward selected targets. We first developed a framework allowing analysis, evaluation, and integration of Ψ mappings, which we applied to >2.5 billion reads from 30 human samples. These maps, complemented with genetic perturbations, allowed us to uncover TRUB1 and PUS7 as the two key PUSs acting on mammalian mRNA and to computationally model the sequence and structural elements governing the specificity of TRUB1, achieving near-perfect prediction of its substrates (AUC = 0.974). We then validated and extended these maps and the inferred specificity of TRUB1 using massively parallel reporter assays in which we monitored Ψ levels at thousands of synthetically designed sequence variants comprising either the sequences surrounding pseudouridylation targets or systematically designed mutants perturbing RNA sequence and structure. Our findings provide an extensive and high-quality characterization of the transcriptome-wide distribution of pseudouridine in human and the factors governing it and provide an important resource for the community, paving the path toward functional and mechanistic dissection of this emerging layer of post-transcriptional regulation.
Despite its pivotal role in regulating transcription, our understanding of core promoter function, architecture, and cis-regulatory elements is lacking. Here, we devised a highthroughput assay to quantify the activity of ∼15,000 fully designed core promoters that we integrated and expressed from a fixed location within the human genome. We find that core promoters drive transcription unidirectionally, and that sequences originating from promoters exhibit stronger activity than sequences originating from enhancers. Testing multiple combinations and distances of core promoter elements, we observe a positive effect of TATA and Initiator, a negative effect of BREu and BREd, and a 10bp periodicity in the optimal distance between the TATA and the Initiator. By comprehensively screening TF binding-sites, we show that site orientation has little effect, that the effect of binding site number on expression is factor-specific, and that there is a striking agreement between the effect of binding site multiplicity in our assay and the tendency of the TF to appear in homotypic clusters throughout the genome. Overall, our results systematically assay the elements that drive expression in core- and proximal-promoter regions and shed light on organization principles of regulatory regions in the human genome.
Translation of mRNAs through Internal Ribosome Entry Sites (IRESs) has emerged as a prominent mechanism of cellular and viral initiation. It supports cap-independent translation of select cellular genes under normal conditions, and in conditions when cap-dependent translation is inhibited. IRES structure and sequence are believed to be involved in this process. However due to the small number of IRESs known, there have been no systematic investigations of the determinants of IRES activity. With the recent discovery of thousands of novel IRESs in human and viruses, the next challenge is to decipher the sequence determinants of IRES activity. We present the first in-depth computational analysis of a large body of IRESs, exploring RNA sequence features predictive of IRES activity. We identified predictive k-mer features resembling IRES trans-acting factor (ITAF) binding motifs across human and viral IRESs, and found that their effect on expression depends on their sequence, number and position. Our results also suggest that the architecture of retroviral IRESs differs from that of other viruses, presumably due to their exposure to the nuclear environment. Finally, we measured IRES activity of synthetically designed sequences to confirm our prediction of increasing activity as a function of the number of short IRES elements.
Transcriptome-wide mapping of N1-methyladenosine (m 1 A) at single-nucleotide resolution reveals m 1 A to be scarce in cytoplasmic mRNA, to inhibit translation, and to be highly dynamic at a single site in a mitochondrial mRNA.
To investigate gene specificity at the level of translation in both the human genome and viruses, we devised a high-throughput bicistronic assay to quantify cap-independent translation. We uncovered thousands of novel cap-independent translation sequences, and we provide insights on the landscape of translational regulation in both humans and viruses. We find extensive translational elements in the 3' untranslated region of human transcripts and the polyprotein region of uncapped RNA viruses. Through the characterization of regulatory elements underlying cap-independent translation activity, we identify potential mechanisms of secondary structure, short sequence motif, and base pairing with the 18S ribosomal RNA (rRNA). Furthermore, we systematically map the 18S rRNA regions for which reverse complementarity enhances translation. Thus, we make available insights into the mechanisms of translational control in humans and viruses.