Cells must limit RNA-RNA interactions to avoid irreversible RNA entanglement. Cells may prevent deleterious RNA-RNA interactions by genome organization to avoid complementarity however, RNA viruses generate long, perfectly complementary antisense RNA during replication. How do viral RNAs avoid irreversible entanglement? One possibility is RNA sequestration into biomolecular condensates. To test this, we reconstituted critical SARS-CoV-2 RNA-RNA interactions in Nucleocapsid condensates. We observed that RNAs with low propensity RNA-RNA interactions resulted in more round, liquid-like condensates while those with high sequence complementarity resulted in more heterogeneous networked morphology independent of RNA structure stability. Residue-resolution molecular simulations and direct sequencing-based detection of RNA-RNA interactions support that these properties arise from degree of trans RNA contacts. We propose that extensive RNA-RNA interactions in cell and viral replication are controlled via a combination of genome organization, timing, RNA sequence content, RNA production ratios, and emergent biomolecular condensate material properties.
In this issue of Molecular Cell, Trussina et al. and Parker et al. present complementary lines of evidence that suggest that stress granules regulate trans RNA-RNA interactions through the action of helicase proteins, which destabilize persistent RNA-RNA interactions that are sufficient to maintain condensate integrity.1,2.
Biomolecular condensates are nonmembrane-bound assemblies of biological polymers such as protein and nucleic acids. An increasingly accepted paradigm across the viral tree of life is (a) that viruses form biomolecular condensates and (b) that the formation is required for the virus. Condensates can promote viral replication by promoting packaging, genome compaction, membrane bending, and co-opting of host translation. This review is primarily concerned with exploring methodologies for assessing virally encoded biomolecular condensates. The goal of this review is to provide an experimental framework for virologists to consider when designing experiments to (a) identify viral condensates and their components, (b) reconstitute condensation cell free from minimal components, (c) ask questions about what conditions lead to condensation, (d) map these questions back to the viral life cycle, and (e) design and test inhibitors/modulators of condensation as potential therapeutics. This experimental framework attempts to integrate virology, cell biology, and biochemistry approaches.
Nucleocapsid protein (N-protein) is required for multiple steps in betacoronaviruses replication. SARS-CoV-2-N-protein condenses with specific viral RNAs at particular temperatures making it a powerful model for deciphering RNA sequence specificity in condensates. We identify two separate and distinct double-stranded, RNA motifs (dsRNA stickers) that promote N-protein condensation. These dsRNA stickers are separately recognized by N-protein's two RNA binding domains (RBDs). RBD1 prefers structured RNA with sequences like the transcription-regulatory sequence (TRS). RBD2 prefers long stretches of dsRNA, independent of sequence. Thus, the two N-protein RBDs interact with distinct dsRNA stickers, and these interactions impart specific droplet physical properties that could support varied viral functions. Specifically, we find that addition of dsRNA lowers the condensation temperature dependent on RBD2 interactions and tunes translational repression. In contrast RBD1 sites are sequences critical for sub-genomic (sg) RNA generation and promote gRNA compression. The density of RBD1 binding motifs in proximity to TRS-L/B sequences is associated with levels of sub-genomic RNA generation. The switch to packaging is likely mediated by RBD1 interactions which generate particles that recapitulate the packaging unit of the virion. Thus, SARS-CoV-2 can achieve biochemical complexity, performing multiple functions in the same cytoplasm, with minimal protein components based on utilizing multiple distinct RNA motifs that control N-protein interactions.
One proposed role for biomolecular condensates that contain RNA is translation regulation. In several specific contexts, translation has been shown to be modulated by the presence of a phase-separating protein and under conditions which promote phase separation, and likely many more await discovery. A powerful tool for determining the rules for condensate-dependent translation is the use of engineered RNA sequences, which can serve as reporters for translation efficiency. This Perspective will discuss design features to consider in engineering RNA reporters to determine the role of phase separation in translational regulation. Specifically, we will cover (i) how to engineer RNA sequence to recapitulate native protein/RNA interactions, (ii) the advantages and disadvantages for commonly used reporter RNA sequences, and (iii) important control experiments to distinguish between binding- and condensation-dependent translational repression. The goal of this review is to promote the design and application of faithful translation reporters to demonstrate a physiological role of biomolecular condensates in translation.
Betacoronavirus SARS-CoV-2 infections caused the global Covid-19 pandemic. The nucleocapsid protein (N-protein) is required for multiple steps in the betacoronavirus replication cycle. SARS-CoV-2-N-protein is known to undergo liquid-liquid phase separation (LLPS) with specific RNAs at particular temperatures to form condensates. We show that N-protein recognizes at least two separate and distinct RNA motifs, both of which require double-stranded RNA (dsRNA) for LLPS. These motifs are separately recognized by N-protein's two RNA binding domains (RBDs). Addition of dsRNA accelerates and modifies N-protein LLPS in vitro and in cells and controls the temperature condensates form. The abundance of dsRNA tunes N-protein-mediated translational repression and may confer a switch from translation to genome packaging. Thus, N-protein's two RBDs interact with separate dsRNA motifs, and these interactions impart distinct droplet properties that can support multiple viral functions. These experiments demonstrate a paradigm of how RNA structure can control the properties of biomolecular condensates.
Viruses must efficiently and specifically package their genomes while excluding cellular nucleic acids and viral sub-genomic fragments. Some viruses use specific packaging signals, which are conserved sequence/structure motifs present only in the full-length genome. Recent work has shown that viral proteins important for packaging can undergo liquid-liquid phase separation (LLPS), where one or two viral nucleic acid binding proteins condense with the genome. The compositional simplicity of viral components lends itself well to theoretical modeling compared to more complex cellular organelles. Viral LLPS can be limited to one or two viral proteins and a single genome that is enriched in LLPS-promoting features. In our previous study, we observed that LLPS-promoting sequences of SARS-CoV-2 are located at the 5ʹ and 3ʹ ends of the genome, whereas the middle of the genome is predicted to consist mostly of solubilizing elements. Is this arrangement sufficient to drive single genome packaging, genome compaction, and genome cyclization? We addressed these questions using a coarse-grained polymer model, LASSI, to study the LLPS of nucleocapsid protein with RNA sequences that either promote LLPS or solubilization. With respect to genome cyclization, we find the most optimal arrangement restricts LLPS-promoting elements to the 5ʹ and 3ʹ ends of the genome, consistent with the native spatial patterning. Genome compaction is enhanced by clustered LLPS-promoting binding sites, while single genome packaging is most efficient when binding sites are distributed throughout the genome. These results suggest that many and variably positioned LLPS-promoting signals can support packaging in the absence of a singular packaging signal which argues against necessity of such a feature. We hypothesize that this model should be generalizable to multiple viruses as well as cellular organelles like paraspeckles, which enrich specific, long RNA sequences in a defined arrangement. Statement of significance The COVID-19 pandemic has motivated research of the basic mechanisms of coronavirus replication. A major challenge faced by viruses such as SARS-CoV-2 is the selective packaging of a large genome in a relatively small capsid while excluding host and sub-genomic nucleic acids. Genomic RNA of SARS-CoV-2 can condense with the Nucleocapsid (N-protein), a structural protein component critical for packaging of many viruses. Notably, certain regions of the genomic RNA drive condensation of N-protein while other regions solubilize it. Here, we explore how the spatial patterning of these opposing elements promotes single genome compaction, packaging, and cyclization. This model informs future in silico experiments addressing spatial patterning of genomic features that are experimentally intractable because of the length of the genome.
A mechanistic understanding of the SARS-CoV-2 viral replication cycle is essential to develop new therapies for the COVID-19 global health crisis. In this study, we show that the SARS-CoV-2 nucleocapsid protein (N-protein) undergoes liquid-liquid phase separation (LLPS) with the viral genome, and propose a model of viral packaging through LLPS. N-protein condenses with specific RNA sequences in the first 1000 nts (5’-End) under physiological conditions and is enhanced at human upper airway temperatures. N-protein condensates exclude non-packaged RNA sequences. We comprehensively map sites bound by N-protein in the 5’-End and find preferences for single-stranded RNA flanked by stable structured elements. Liquid-like N-protein condensates form in mammalian cells in a concentration-dependent manner and can be altered by small molecules. Condensation of N-protein is sequence and structure specific, sensitive to human body temperature, and manipulatable with small molecules thus presenting screenable processes for identifying antiviral compounds effective against SARS-CoV-2.
Fragile-X mental retardation autosomal homologue-1 (FXR1) is a muscle-enriched RNA-binding protein. FXR1 depletion is perinatally lethal in mice, Xenopus, and zebrafish; however, the mechanisms driving these phenotypes remain unclear. The FXR1 gene undergoes alternative splicing, producing multiple protein isoforms and mis-splicing has been implicated in disease. Furthermore, mutations that cause frameshifts in muscle-specific isoforms result in congenital multi-minicore myopathy. We observed that FXR1 alternative splicing is pronounced in the serine- and arginine-rich intrinsically disordered domain; these domains are known to promote biomolecular condensation. Here, we show that tissue-specific splicing of fxr1 is required for Xenopus development and alters the disordered domain of FXR1. FXR1 isoforms vary in the formation of RNA-dependent biomolecular condensates in cells and in vitro. This work shows that regulation of tissue-specific splicing can influence FXR1 condensates in muscle development and how mis-splicing promotes disease.
MicroRNA (miRNA) expression patterns are highly variable across human tissues and across cancer specimens. The intuitive assumption is that transcription is the main contributor to mature miRNA expression patterns, with post-transcriptional processes further modifying miRNA expression levels. Here we report the surprising model that, on the global level, post-transcriptional regulation dominates over transcriptional regulation in determining mature miRNA expression patterns in both normal tissues and cancer. Taking advantage of large genomic datasets in which the expression of both mature miRNAs and their host genes have been quantified, we establish and validate transcriptional and post-transcriptional metrics, with miRNA host gene expression estimating transcriptional regulation and mature miRNA to host gene ratio estimating post-transcriptional regulation. On average, the post-transcriptional metric contributes 2.8-fold more than the transcriptional metric to the variance of mature miRNA expression. The variation of the balance between the two mature miRNAs (5p and 3p miRNAs) produced from the same precursor hairpin is a non-negligible contributor to miRNA expression, explaining ∼27% of the variance of miRNAs’ post-transcriptional metric. Data of normal tissues yield similar results as cancer specimens. Additionally, the post-transcriptional metric is superior to the transcriptional metric in classifying cancer types. We further demonstrate that the post-transcriptional metric separates miRNAs into distinct groups, suggesting that there are groups of miRNAs that are co-regulated on the post-transcriptional level. Our data support a model in which the post-transcriptional regulation is the major driver of miRNA expression variation, and paves a way toward better mechanistic understanding of post-transcriptional regulation of mature miRNA expression.
We report that the SARS-CoV-2 nucleocapsid protein (N-protein) undergoes liquid-liquid phase separation (LLPS) with viral RNA. N-protein condenses with specific RNA genomic elements under physiological buffer conditions and condensation is enhanced at human body temperatures (33°C and 37°C) and reduced at room temperature (22°C). RNA sequence and structure in specific genomic regions regulate N-protein condensation while other genomic regions promote condensate dissolution, potentially preventing aggregation of the large genome. At low concentrations, N-protein preferentially crosslinks to specific regions characterized by single-stranded RNA flanked by structured elements and these features specify the location, number, and strength of N-protein binding sites (valency). Liquid-like N-protein condensates form in mammalian cells in a concentration-dependent manner and can be altered by small molecules. Condensation of N-protein is RNA sequence and structure specific, sensitive to human body temperature, and manipulatable with small molecules, and therefore presents a screenable process for identifying antiviral compounds effective against SARS-CoV-2.
Biomolecular condensation partitions cellular contents and has important roles in stress responses, maintaining homeostasis, development and disease. Many nuclear and cytoplasmic condensates are rich in RNA and RNA-binding proteins (RBPs), which undergo liquid–liquid phase separation (LLPS). Whereas the role of RBPs in condensates has been well studied, less attention has been paid to the contribution of RNA to LLPS. In this Review, we discuss the role of RNA in biomolecular condensation and highlight considerations for designing condensate reconstitution experiments. We focus on RNA properties such as composition, length, structure, modifications and expression level. These properties can modulate the biophysical features of native condensates, including their size, shape, viscosity, liquidity, surface tension and composition. We also discuss the role of RNA–protein condensates in development, disease and homeostasis, emphasizing how their properties and function can be determined by RNA. Finally, we discuss the multifaceted cellular functions of biomolecular condensates, including cell compartmentalization through RNA transport and localization, supporting catalytic processes, storage and inheritance of specific molecules, and buffering noise and responding to stress. Recent studies have highlighted the contribution of RNA to cellular liquid–liquid phase separation and condensate formation. RNA features modulate the composition and biophysical properties of RNA–protein condensates, which have various cellular functions, including RNA transport and localization, supporting catalytic processes and responding to stress.
Measuring multiple omics profiles from the same single cell opens up the opportunity to decode molecular regulation that underlies intercellular heterogeneity in development and disease. Here, we present co-sequencing of microRNAs and mRNAs in the same single cell using a half-cell genomics approach. This method demonstrates good robustness (~95% success rate) and reproducibility ( R 2 = 0.93 for both microRNAs and mRNAs), yielding paired half-cell microRNA and mRNA profiles, which we can independently validate. By linking the level of microRNAs to the expression of predicted target mRNAs across 19 single cells that are phenotypically identical, we observe that the predicted targets are significantly anti-correlated with the variation of abundantly expressed microRNAs. This suggests that microRNA expression variability alone may lead to non-genetic cell-to-cell heterogeneity. Genome-scale analysis of paired microRNA-mRNA co-profiles further allows us to derive and validate regulatory relationships of cellular pathways controlling microRNA expression and intercellular variability.
SUMMARYFragile-X mental retardation autosomal homolog-1 (FXR1) is a muscle-enriched RNA-binding protein. FXR1 depletion is perinatally lethal in mice, Xenopus, and zebrafish; however, the mechanisms driving these phenotypes remain unclear. The FXR1 gene undergoes alternative splicing, producing multiple protein isoforms and mis-splicing has been implicated in disease. Furthermore, mutations that cause frameshifts in muscle-specific isoforms result in congenital multi-minicore myopathy. We observed that FXR1 alternative splicing is pronounced in the serine and arginine-rich intrinsically-disordered domain; these domains are known to promote biomolecular condensation. Here, we show that tissue-specific splicing of fxr1 is required for Xenopus development and alters the disordered domain of FXR1. FXR1 isoforms vary in the formation of RNA-dependent biomolecular condensates in cells and in vitro. This work shows that regulation of tissue-specific splicing can influence FXR1 condensates in muscle development and how mis-splicing promotes disease.HIGHLIGHTSThe muscle-specific exon 15 impacts FXR1 functionsAlternative splicing of FXR1 is tissue- and developmental stage specificFXR1 forms RNA-dependent condensatesSplicing regulation changes FXR1 condensate properties
Department of Genetics, Yale University School of Medicine, New Haven, CT; Yale Stem Cell Center, Yale Cancer Center, New Haven, CT; Chinese People’s Liberation Army General Hospital, Beijing, China; Department of Biomedical Engineering, Yale University, New Haven, CT; Department of Cell Biology, and Department of Surgery, Yale University School of Medicine, New Haven, CT; Jiangsu Institute of Hematology, The First Affiliated Hospital of Soochow University, Suzhou, China; Lillehei Heart Institute and Department of Pediatrics, University of Minnesota, Minneapolis, MN; and Yale Center for RNA Science and Medicine, New Haven, CT
The hematopoietic stem cell-enriched miR-125 family microRNAs (miRNAs) are critical regulators of hematopoiesis. Overexpression of miR-125a or miR-125b is frequent in human acute myeloid leukemia (AML), and the overexpression of these miRNAs in mice leads to expansion of hematopoietic stem cells accompanied by perturbed hematopoiesis with mostly myeloproliferative phenotypes. However, whether and how miR-125 family miRNAs cooperate with known AML oncogenes in vivo, and how the resultant leukemia is dependent on miR-125 overexpression, are not well understood. We modeled the frequent co-occurrence of miR-125b overexpression and MLL translocations by examining functional cooperation between miR-125b and MLL-AF9 By generating a knock-in mouse model in which miR-125b overexpression is controlled by doxycycline induction, we demonstrated that miR-125b significantly enhances MLL-AF9-driven AML in vivo, and the resultant leukemia is partially dependent on continued overexpression of miR-125b Surprisingly, miR-125b promotes AML cell expansion and suppresses apoptosis involving a non-cell-intrinsic mechanism. MiR-125b expression enhances VEGFA expression and production from leukemia cells, in part by suppressing TET2 Recombinant VEGFA recapitulates the leukemia-promoting effects of miR-125b, whereas knockdown of VEGFA or inhibition of VEGF receptor 2 abolishes the effects of miR-125b In addition, significant correlation between miR-125b and VEGFA expression is observed in human AMLs. Our data reveal cooperative and dependent relationships between miR-125b and the MLL oncogene in AML leukemogenesis, and demonstrate a miR-125b-TET2-VEGFA pathway in mediating non-cell-intrinsic leukemia-promoting effects by an oncogenic miRNA.
Mature microRNAs (miRNAs) are processed from hairpin-containing primary miRNAs (pri-miRNAs). However, rules that distinguish pri-miRNAs from other hairpin-containing transcripts in the genome are incompletely understood. By developing a computational pipeline to systematically evaluate 30 structural and sequence features of mammalian RNA hairpins, we report several new rules that are preferentially utilized in miRNA hairpins and govern efficient pri-miRNA processing. We propose that a hairpin stem length of 36 ± 3 nt is optimal for pri-miRNA processing. We identify two bulge-depleted regions on the miRNA stem, located ∼16-21 nt and ∼28-32 nt from the base of the stem, that are less tolerant of unpaired bases. We further show that the CNNC primary sequence motif selectively enhances the processing of optimal-length hairpins. We predict that a small but significant fraction of human single-nucleotide polymorphisms (SNPs) alter pri-miRNA processing, and confirm several predictions experimentally including a disease-causing mutation. Our study enhances the rules governing mammalian pri-miRNA processing and suggests a diverse impact of human genetic variation on miRNA biogenesis.
Ten-Eleven-Translocation-2 (Tet2) is a DNA methylcytosine dioxygenase that functions as a tumor suppressor in hematopoietic malignancies. We examined the role of Tet2 in tumor-tissue myeloid cells and found that Tet2 sustains the immunosuppressive function of these cells. We found that Tet2 expression is increased in intratumoral myeloid cells both in mouse models of melanoma and in melanoma patients and that this increased expression is dependent on an IL-1R-MyD88 pathway. Ablation of Tet2 in myeloid cells suppressed melanoma growth in vivo and shifted the immunosuppressive gene expression program in tumor-associated macrophages to a proinflammatory one, with a concomitant reduction of the immunosuppressive function. This resulted in increased numbers of effector T cells in the tumor, and T cell depletion abolished the reduced tumor growth observed upon myeloid-specific deletion of Tet2. Our findings reveal a non-cell-intrinsic, tumor-promoting function for Tet2 and suggest that Tet2 may present a therapeutic target for the treatment of non-hematologic malignancies.
A large number of microRNAs (miRNAs) are grouped into families derived from the same phylogenetic ancestors. miRNAs within a family often share the same physiological functions despite differences in their primary sequences, secondary structures, or chromosomal locations. Consequently, the generation of animal models to analyze the activity of miRNA families is extremely challenging. Using zebrafish as a model system, we successfully provide experimental evidence that a large number of miRNAs can be simultaneously mutated to abrogate the activity of an entire miRNA family. We show that injection of the Cas9 nuclease and two, four, ten, and up to twenty-four multiplexed single guide RNAs (sgRNAs) can induce mutations in 90% of the miRNA genomic sequences analyzed. We performed a survey of these 45 mutations in 10 miRNA genes, analyzing the impact of our mutagenesis strategy on the processing of each miRNA both computationally and in vivo. Our results offer an effective approach to mutate and study the activity of miRNA families and pave the way for further analysis on the function of complex miRNA families in higher multicellular organisms.
Clustered regularly-interspaced palindromic repeats (CRISPR)-based genetic screens using single-guide-RNA (sgRNA) libraries have proven powerful to identify genetic regulators. Applying CRISPR screens to interrogate functional elements in noncoding regions requires generating sgRNA libraries that are densely covering, and ideally inexpensive, easy to implement and flexible for customization. Here we present a Molecular Chipper technology for generating dense sgRNA libraries for genomic regions of interest, and a proof-of-principle screen that identifies novel cis -regulatory domains for miR-142 biogenesis. The Molecular Chipper approach utilizes a combination of random fragmentation and a type III restriction enzyme to derive a densely covering sgRNA library from input DNA. Applying this approach to 17 microRNAs and their flanking regions and with a reporter for miR-142 activity, we identify both the pre-miR-142 region and two previously unrecognized cis -domains important for miR-142 biogenesis, with the latter regulating miR-142 processing. This strategy will be useful for identifying functional noncoding elements in mammalian genomes.