Ribosome structure and activity are challenged at high temperatures, often demanding modifications to ribosomal RNAs (rRNAs) to retain translation fidelity. LC- MS/MS, bisulfite- sequencing, and high- resolution cryo-EM structures of the archaeal ribosome identified an RNA modification, N 4, N 4- dimethylcytidine (m42C), at the universally conserved C918 in the 16S rRNA helix 31 loop. Here, we characterize and structurally resolve a class of RNA methyltransferase that generates m42C whose function is critical for hyperthermophilic growth. m42C is synthesized by the activity of a unique family of RNA methyltransferase containing a Rossman-fold that targets only intact ribosomes. The phylogenetic distribution of the newly identified m42C synthase family implies that m42C is biologically relevant in each domain. Resistance of m42C to bisulfite- driven deamination suggests that efforts to capture m5C profiles via bisulfite sequencing are also capturing m4 2 C.
Ribosome structure and activity are challenged at high temperatures, often demanding modifications to ribosomal RNAs (rRNAs) to retain translation fidelity. LC-MS/MS, bisulfite-sequencing, and high-resolution cryo-EM structures of the archaeal ribosome identified an RNA modification, N 4, N 4-dimethylcytidine (m 4 2 C), at the universally conserved C918 in the 16S rRNA helix 31 loop. Here, we characterize and structurally resolve a class of RNA methyltransferase that generates m 4 2 C whose function is critical for hyperthermophilic growth. m 4 2 C is synthesized by the activity of a unique family of RNA methyltransferase containing a Rossman-fold that targets only intact ribosomes. The phylogenetic distribution of the newly identified m 4 2 C synthase family implies that m 4 2 C is biologically relevant in each domain. Resistance of m 4 2 C to bisulfite-driven deamination suggests that efforts to capture m 5 C profiles via bisulfite sequencing are also capturing m 4 2 C.
RNAs are often modified to invoke new activities. While many modifications are limited in frequency, restricted to non-coding RNAs, or present only in select organisms, 5-methylcytidine (m5C) is abundant across diverse RNAs and fitness-relevant across Domains of life, but the synthesis and impacts of m5C have yet to be fully investigated. Here, we map m5C in the model hyperthermophile, Thermococcus kodakarensis. We demonstrate that m5C is ~25x more abundant in T. kodakarensis than human cells, and the m5C epitranscriptome includes ~10% of unique transcripts. T. kodakarensis rRNAs harbor tenfold more m5C compared to Eukarya or Bacteria. We identify at least five RNA m5C methyltransferases (R5CMTs), and strains deleted for individual R5CMTs lack site-specific m5C modifications that limit hyperthermophilic growth. We show that m5C is likely generated through partial redundancy in target sites among R5CMTs. The complexity of the m5C epitranscriptome in T. kodakarensis argues that m5C supports life in the extremes. The epitranscriptome is fitness-relevant across Domains. Here, the authors map m5C in the model hyperthermophile, Thermococcus kodakarensis. The abundance and complexity of the m5C epitranscriptome in T. kodakarensis argues that m5C supports life in the extremes.
The N7-methyl guanosine cap structure is an essential 5' end modification of eukaryotic mRNA. It plays a critical role in many aspects of the life cycle of mRNA, including nuclear export, stability, and translation. Equipping synthetic transcripts with a 5' cap is paramount to the development of effective mRNA vaccines and therapeutics. Here, we report a simple and flexible workflow to selectively isolate and analyze structural features of the 5' end of an mRNA by means of DNA probe-directed enrichment with site-specific single-strand endoribonucleases. Specifically, we showed that the RNA cleavage by site-specific RNases can be effectively steered by a complementary DNA probe to recognition sites downstream of the probe-hybridized region, utilizing a flexible range of DNA probe designs. We applied this approach using human RNase 4 to isolate well-defined cleavage products from the 5' end of diverse uridylated and N1-methylpseudouridylated mRNA 5' end transcript sequences. hRNase 4 increases the precision of the RNA cleavage, reducing product heterogeneity while providing comparable estimates of capped products and their intermediaries relative to the widely used RNase H. Collectively, we demonstrated that this workflow ensures well-defined and predictable 5' end cleavage products suitable for analysis and relative quantitation of synthetic mRNA 5' cap structures by UHPLC-MS/MS.
The global deployment of mRNA vaccines against SARS-CoV-2 and the projected expansion of therapeutic applications of synthetic mRNA call for robust and high-precision analytical methods to evaluate attributes that are crucial to the safety and efficacy of the mRNA drug substances. Liquid chromatography–mass spectrometry (LC–MS) is one of the few techniques that can provide a direct and high-confidence readout of the identity and incorporation efficiency of the 5′ cap, length of the poly(A) tail, nucleotide sequence, and modification profile of synthetic mRNA molecules. Prior to LC–MS analysis, the RNA molecules are partially digested by specific endoribonucleases into oligonucleotides that are suitable for charge state-dependent fragmentation and mass deconvolution. The most commonly used endoribonuclease for RNA sequence mapping is the guanosine-specific RNase T1. RNase T1 has been employed for analysis of mRNA, rRNA, and tRNA as well as for mRNA poly(A) tail length verification. For mRNA 5′ cap analysis, selective excisions using probe-restrained RNase H or (deoxy)ribozymes are typically required. In this chapter, we will review the application of endoribonucleases for mRNA analysis, with emphasis on a recently characterized endoribonuclease derived from human RNase 4. We will also discuss the latest methods to assess 5′ cap and poly(A) tail incorporation in synthetic mRNA. Finally, we will highlight why more enzymatic tools are needed and how they can contribute to improving the quality of synthetic RNA analysis, and to help understand the biology of RNA modifications in the cell.
The chemical modification of RNA bases represents a ubiquitous activity that spans all domains of life. Pseudouridylation is the most common RNA modification and is observed within tRNA, rRNA, ncRNA and mRNAs. Pseudouridine synthase or 'PUS' enzymes include those that rely on guide RNA molecules and others that function as 'stand-alone' enzymes. Among the latter, several have been shown to modify mRNA transcripts. Although recent studies have defined the structural requirements for RNA to act as a PUS target, the mechanisms by which PUS1 recognizes these target sequences in mRNA are not well understood. Here we describe the crystal structure of yeast PUS1 bound to an RNA target that we identified as being a hot spot for PUS1-interaction within a model mRNA at 2.4 Å resolution. The enzyme recognizes and binds both strands in a helical RNA duplex, and thus guides the RNA containing the target uridine to the active site for subsequent modification of the transcript. The study also allows us to show the divergence of related PUS1 enzymes and their corresponding RNA target specificities, and to speculate on the basis by which PUS1 binds and modifies mRNA or tRNA substrates.
Abstract With the rapid growth of synthetic messenger RNA (mRNA)-based therapeutics and vaccines, the development of analytical tools for characterization of long, complex RNAs has become essential. Tandem liquid chromatography–mass spectrometry (LC–MS/MS) permits direct assessment of the mRNA primary sequence and modifications thereof without conversion to cDNA or amplification. It relies upon digestion of mRNA with site-specific endoribonucleases to generate pools of short oligonucleotides that are then amenable to MS-based sequence analysis. Here, we showed that the uridine-specific human endoribonuclease hRNase 4 improves mRNA sequence coverage, in comparison with the benchmark enzyme RNase T1, by producing a larger population of uniquely mappable cleavage products. We deployed hRNase 4 to characterize mRNAs fully substituted with 1-methylpseudouridine (m1Ψ) or 5-methoxyuridine (mo5U), as well as mRNAs selectively depleted of uridine–two key strategies to reduce synthetic mRNA immunogenicity. Lastly, we demonstrated that hRNase 4 enables direct assessment of the 5′ cap incorporation into in vitro transcribed mRNA. Collectively, this study highlights the power of hRNase 4 to interrogate mRNA sequence, identity, and modifications by LC–MS/MS.
The phosphorylated RNA polymerase II CTD interacting factor 1 (PCIF1) is a methyltransferase that adds a methyl group to the N6-position of 20O-methyladenosine (Am), generating N6, 20O-dimethyladenosine (m6Am) when Am is the cap-proximal nucleotide. In addition, PCIF1 has ancillary methylation activities on internal adenosines (both A and Am), although with much lower catalytic efficiency relative to that of its preferred cap substrate. The PCIF1 preference for 20Omethylated Am over unmodified A nucleosides is due mainly to increased binding affinity for Am. Importantly, it was recently reported that PCIF1 can methylate viral RNA. Although some viral RNA can be translated in the absence of a cap, it is unclear what roles PCIF1 modifications may play in the functionality of viral RNAs. Here we show, using in vitro assays of binding and methyltransfer, that PCIF1 binds an uncapped 50-Am oligonucleotide with approximately the same affinity as that of a cap analog (KM = 0.4 versus 0.3 mu M). In addition, PCIF1 methylates the uncapped 50-Am with activity decreased by only fivefold to sixfold compared with its preferred capped substrate. We finally discuss the relationship between PCIF1-catalyzed RNA methylation, shown here to have broader substrate specificity than previously appreciated, and that of the RNA demethylase fat mass and obesity-associated protein (FTO), which demonstrates PCIF1-opposing activities on capped RNAs.
Combinations of ribonucleases (RNases) are commonly used to digest RNA into oligoribonucleotide fragments prior to liquid chromatography-mass spectrometry (LC-MS) analysis. The distribution of the RNase target sequences or nucleobase sites within an RNA molecule is critical for achieving a high mapping coverage. Cusativin and MC1 are nucleotide-specific endoribonucleases encoded in the cucumber and bitter melon genomes, respectively. Their high specificity for cytidine (Cusativin) and uridine (MC1) make them ideal molecular biology tools for RNA modification mapping. However, heterogenous recombinant expression of either enzyme has been challenging because of their high toxicity to expression hosts and the requirement of posttranslational modifications. Here, we present two highly efficient and time-saving protocols that overcome these hurdles and enhance the expression and purification of these RNases. We first purified MC1 and Cusativin from bacteria by expressing and shuttling both enzymes to the periplasm as MBP-fusion proteins in T7 Express lysY/IqE. coli strain at low temperature. The RNases were enriched using amylose affinity chromatography, followed by a subsequent purification via a C-terminal 6xHIS tag. This fast, two-step purification allows for the purification of highly active recombinant RNases significantly surpassing yields reported in previous studies. In addition, we expressed and purified a Cusativin-CBD fusion enzyme in P. pastoris using chitin magnetic beads. Both Cusativin variants exhibited a similar sequence preference, suggesting that neither posttranslational modifications nor the epitope-tags have a substantial effect on the sequence specificity of the enzyme.
Quality control of mRNA represents an important regulatory mechanism for gene expression in eukaryotes. One component of this quality control is the nuclear retention and decay of misprocessed RNAs. Previously, we demonstrated that mature mRNAs containing a 5 ' splice site (5 ' SS) motif, which is typically found in misprocessed RNAs such as intronic polyadenylated (IPA) transcripts, are nuclear retained and degraded. Using high-throughput sequencing of cellular fractions, we now demonstrate that IPA transcripts require the zinc finger protein ZFC3H1 for their nuclear retention and degradation. Using reporter mRNAs, we demonstrate that ZFC3H1 promotes the nuclear retention of mRNAs with intact 5 ' SS motifs by sequestering them into nuclear speckles. Furthermore, we find that U1-70K, a component of the spliceosomal U1 snRNP, is also required for the nuclear retention of these reporter mRNAs and likely functions in the same pathway as ZFC3H1. Finally, we show that the disassembly of nuclear speckles impairs the nuclear retention of reporter mRNAs with 5 ' SS motifs. Our results highlight a splicing independent role of U1 snRNP and indicate that it works in conjunction with ZFC3H1 in preventing the nuclear export of misprocessed mRNAs by sequestering them into nuclear speckles.
The mammalian mRNA nuclear export process is thought to terminate at the cytoplasmic face of the nuclear pore complex through ribonucleoprotein remodeling. We conduct a stringent affinity-purification massspectrometry-based screen of the physical interactions of human RNA-binding E3 ubiquitin ligases. The resulting protein-interaction network reveals interactions between the RNA-binding E3 ubiquitin ligase MKRN2 and GLE1, a DEAD-box helicase activator implicated in mRNA export termination. We assess MKRN2 epistasis with GLE1 in a zebrafish model. Morpholino-mediated knockdown or CRISPR/Cas9-based knockout of MKRN2 partially rescue retinal developmental defects seen upon GLE1 depletion, consistent with a functional association between GLE1 and MKRN2. Using ribonomic approaches, we show that MKRN2 binds selectively to the 3' UTR of a diverse subset of mRNAs and that nuclear export of MKRN2-associated mRNAs is enhanced upon knockdown of MKRN2. Taken together, we suggest that MKRN2 interacts with GLE1 to selectively regulate mRNA nuclear export and retinal development.
Abstract While splicing has been shown to enhance nuclear export, it has remained unclear whether mRNAs generated from intronless genes use specific machinery to promote their export. Here, we investigate the role of the major nuclear pore basket protein, TPR, in regulating mRNA and lncRNA nuclear export in human cells. By sequencing mRNA from the nucleus and cytosol of control and TPR-depleted cells, we provide evidence that TPR is required for the efficient nuclear export of mRNAs and lncRNAs that are generated from short transcripts that tend to have few introns, and we validate this with reporter constructs. Moreover, in TPR-depleted cells reporter mRNAs generated from short transcripts accumulate in nuclear speckles and are bound to Nxf1. These observations suggest that TPR acts downstream of Nxf1 recruitment and may allow mRNAs to leave nuclear speckles and properly dock with the nuclear pore. In summary, our study provides one of the first examples of a factor that is specifically required for the nuclear export of intronless and intron-poor mRNAs and lncRNAs.
Protein complexes are key macromolecular machines of the cell, but their description remains incomplete. We and others previously reported an experimental strategy for global characterization of native protein assemblies based on chromatographic fractionation of biological extracts coupled to precision mass spectrometry analysis (chromatographic fractionation-mass spectrometry, CF-MS), but the resulting data are challenging to process and interpret. Here, we describe EPIC (elution profile-based inference of complexes), a software toolkit for automated scoring of large-scale CF-MS data to define high-confidence multi-component macromolecules from diverse biological specimens. As a case study, we used EPIC to map the global inter-actome of Caenorhabditis elegans, defining 612 putative worm protein complexes linked to diverse biological processes. These included novel subunits and assemblies unique to nematodes that we validated using orthogonal methods. The open source EPIC software is freely available as a Jupyter notebook packaged in a Docker container (https://hub.docker.com/r/baderlab/bio-epic/).