A-to-I RNA editing is a cellular mechanism that generates transcriptomic and proteomic diversity, which is essential for neuronal and immune functions. It involves the conversion of specific adenosines in RNA molecules to inosines, which are recognized as guanosines by cellular machinery. Despite the vast number of editing sites observed across the animal kingdom, pinpointing critical sites and understanding their in vivo functions remains challenging. Here, we study the function of an evolutionary conserved editing site in Drosophila , located in glutamate-gated chloride channel ( GluCl α). Our findings reveal that flies lacking editing at this site exhibit reduced olfactory responses to odors and impaired pheromone-dependent social interactions. Moreover, we demonstrate that editing of this site is crucial for the proper processing of olfactory information in projection neurons. Our results highlight the value of using evolutionary conservation as a criterion for identifying editing events with potential functional significance and paves the way for elucidating the intricate link between RNA modification, neuronal physiology, and behavior.
ADAR RNA editing enzymes are high-affinity dsRNA-binding proteins that deaminate adenosines to inosines in pre-mRNA hairpins and also exert editing-independent effects. We generated a Drosophila Adar(E374A) mutant strain encoding a catalytically inactive Adar with CRISPR/Cas9. We demonstrate that Adar adenosine deamination activity is necessary for normal locomotion and prevents age-dependent neurodegeneration. The catalytically inactive protein, when expressed at a higher than physiological level, can rescue neurodegeneration in Adar mutants, suggesting also editing-independent effects. Furthermore, loss of Adar RNA editing activity leads to innate immune induction, indicating that Drosophila Adar, despite being the homolog of mammalian ADAR2, also has functions similar to mammalian ADAR1. The innate immune induction in fly Adar mutants is suppressed by silencing of Dicer-2, which has a RNA helicase domain similar to MDA5 that senses unedited dsRNAs in mammalian Adar1 mutants. Our work demonstrates that the single Adar enzyme in Drosophila unexpectedly has dual functions. Human RNA editing enzymes ADAR1 and ADAR2 are required for innate immune functions and neurological functions, respectively. Here, the authors show that Drosophila Adar has both innate immune and brain functions, despite being the homolog of mammalian ADAR2.
Adenosine-to-inosine (A-to-I) RNA editing is an important post-transcriptional modification that affects the information encoded from DNA to RNA to protein. RNA editing can generate a multitude of transcript isoforms and can potentially be used to optimize protein function in response to varying conditions. In light of this and the fact that millions of editing sites have been identified in many different species, it is interesting to examine the extent to which these sites have evolved to be functionally important. In this review, we discuss results pertaining to the evolution of RNA editing, specifically in humans, cephalopods, and Drosophila. We focus on how comparative genomics approaches have aided in the identification of sites that are likely to be advantageous. The use of RNA editing as a mechanism to adapt to varying environmental conditions will also be reviewed.
Adenosine-to-inosine RNA editing diversifies the transcriptome and promotes functional diversity, particularly in the brain. A plethora of editing sites has been recently identified; however, how they are selected and regulated and which are functionally important are largely unknown. Here we show the cis-regulation and stepwise selection of RNA editing during Drosophila evolution and pinpoint a large number of functional editing sites. We found that the establishment of editing and variation in editing levels across Drosophila species are largely explained and predicted by cis-regulatory elements. Furthermore, editing events that arose early in the species tree tend to be more highly edited in clusters and enriched in slowly-evolved neuronal genes, thus suggesting that the main role of RNA editing is for fine-tuning neurological functions. While nonsynonymous editing events have been long recognized as playing a functional role, in addition to nonsynonymous editing sites, a large fraction of 3'UTR editing sites is evolutionarily constrained, highly edited, and thus likely functional. We find that these 3'UTR editing events can alter mRNA stability and affect miRNA binding and thus highlight the functional roles of noncoding RNA editing. Our work, through evolutionary analyses of RNA editing in Drosophila, uncovers novel insights of RNA editing regulation as well as its functions in both coding and non-coding regions.
Adenosine-to-inosine RNA editing modifies maturing mRNAs through the binding of adenosine deaminase acting on RNA (Adar) proteins to double-stranded RNA structures in a process critical for neuronal function. Editing levels at individual editing sites span a broad range and are mediated by both cis-acting elements (surrounding RNA sequence and secondary structure) and trans-acting factors. Here, we aim to determine the roles that cis-acting elements and trans-acting factors play in regulating editing levels. Using two closely related Drosophila species, D. melanogaster and D. sechellia, and their F1 hybrids, we dissect the effects of cis sequences from trans regulators on editing levels by comparing species-specific editing in parents and their hybrids. We report that cis sequence differences are largely responsible for editing level differences between these two Drosophila species. This study presents evidence for cis sequence and structure changes as the dominant evolutionary force that modulates RNA editing levels between these Drosophila species.
Adenosine-to-inosine (A-to-I) RNA editing, catalysed by ADAR enzymes conserved in metazoans, plays an important role in neurological functions. Although the fine-tuning mechanism provided by A-to-I RNA editing is important, the underlying rules governing ADAR substrate recognition are not well understood. We apply a quantitative trait loci (QTL) mapping approach to identify genetic variants associated with variability in RNA editing. With very accurate measurement of RNA editing levels at 789 sites in 131 Drosophila melanogaster strains, here we identify 545 editing QTLs (edQTLs) associated with differences in RNA editing. We demonstrate that many edQTLs can act through changes in the local secondary structure for edited dsRNAs. Furthermore, we find that edQTLs located outside of the edited dsRNA duplex are enriched in secondary structure, suggesting that distal dsRNA structure beyond the editing site duplex affects RNA editing efficiency. Our work will facilitate the understanding of the cis-regulatory code of RNA editing.
The CRISPR/Cas9 system has recently emerged as a powerful tool for functional genomic studies in Drosophila melanogaster. However, single-guide RNA (sgRNA) parameters affecting the specificity and efficiency of the system in flies are still not clear. Here, we found that off-target effects did not occur in regions of genomic DNA with three or more nucleotide mismatches to sgRNAs. Importantly, we document for a strong positive correlation between mutagenesis efficiency and sgRNA GC content of the six protospacer-adjacent motif-proximal nucleotides (PAMPNs). Furthermore, by injecting well-designed sgRNA plasmids at the optimal concentration we determined, we could efficiently generate mutations in four genes in one step. Finally, we generated null alleles of HP1a using optimized parameters through homology-directed repair and achieved an overall mutagenesis rate significantly higher than previously reported. Our work demonstrates a comprehensive optimization of sgRNA and promises to vastly simplify CRISPR/Cas9 experiments in Drosophila.
The CRISPR/Cas9 system has recently emerged as a powerful tool for functional genomic studies in Drosophila melanogaster. However, single-guide RNA (sgRNA) parameters affecting the specificity and efficiency of the system in flies are still not clear. Here, we found that off-target effects did not occur in regions of genomic DNA with three or more nucleotide mismatches to sgRNAs. Importantly, we document for a strong positive correlation between mutagenesis efficiency and sgRNA GC content of the six protospacer-adjacent motif-proximal nucleotides (PAMPNs). Furthermore, by injecting well-designed sgRNA plasmids at the optimal concentration we determined, we could efficiently generate mutations in four genes in one step. Finally, we generated null alleles of HP1a using optimized parameters through homology-directed repair and achieved an overall mutagenesis rate significantly higher than previously reported. Our work demonstrates a comprehensive optimization of sgRNA and promises to vastly simplify CRISPR/Cas9 experiments in Drosophila.
P<P Published online December 17, 2013 in advance of the print journal. Preprint Accepted likely to differ from the final, published version. Peer-reviewed and accepted for publication but not copyedited or typeset; preprint is Open Access Open Access option. Genome Research Freely available online through the License Commons Creative . http://creativecommons.org/licenses/by-nc/3.0/ Unported), as described at available under a Creative Commons License (Attribution-NonCommercial 3.0 , is Genome Research This manuscript is Open Access.This article, published in
We show that RNA editing sites can be called with high confidence using RNA sequencing data from multiple samples across either individuals or species, without the need for matched genomic DNA sequence. We identified many previously unidentified editing sites in both humans and Drosophila; our results nearly double the known number of human protein recoding events. We also found that human genes harboring conserved editing sites within Alu repeats are enriched for neuronal functions.
RNA molecules transmit the information encoded in the genome and generally reflect its content. Adenosine-to-inosine (A-to-I) RNA editing by ADAR proteins converts a genomically encoded adenosine into inosine. It is known that most RNA editing in human takes place in the primate-specific Alu sequences, but the extent of this phenomenon and its effect on transcriptome diversity are not yet clear. Here, we analyzed large-scale RNA-seq data and detected ∼1.6 million editing sites. As detection sensitivity increases with sequencing coverage, we performed ultradeep sequencing of selected Alu sequences and showed that the scope of editing is much larger than anticipated. We found that virtually all adenosines within Alu repeats that form double-stranded RNA undergo A-to-I editing, although most sites exhibit editing at only low levels (<1%). Moreover, using high coverage sequencing, we observed editing of transcripts resulting from residual antisense expression, doubling the number of edited sites in the human genome. Based on bioinformatic analyses and deep targeted sequencing, we estimate that there are over 100 million human Alu RNA editing sites, located in the majority of human genes. These findings set the stage for exploring how this primate-specific massive diversification of the transcriptome is utilized.
We show that RNA editing sites can be called with high confidence using RNA sequencing data from multiple samples across either individuals or species, without the need for matched genomic DNA sequence. We identified many previously unidentified editing sites in both humans and Drosophila; our results nearly double the known number of human protein recoding events. We also found that human genes harboring conserved editing sites within Alu repeats are enriched for neuronal functions. RNA editing is the postor co-transcriptional modification of RNA nucleotides from their genome-encoded sequence. In humans, the most prevalent type is adenosine-to-inosine (Ato-I) editing, catalyzed by the adenosine deaminase acting on RNA (ADAR) family of enzymes1. The ADAR enzymes bind double-stranded RNAs and deaminate adenosine to inosine, which is recognized as guanosine by the cellular machinery. A-to-I editing is pervasive in Alu repeats because of the double-stranded RNA structures formed by inverted Alu repeats in many genes2,3. However, only a few dozen human RNA editing targets that change amino acids in nonrepetitive regions have been identified4, and most of them were identified in nervous system tissues5. High-throughput RNA sequencing (RNA-seq) has enabled transcriptome-wide identification of A-to-I editing sites. The major challenge in identifying RNA editing sites using RNA-seq data is the discrimination of RNA editing sites from genome-encoded single-nucleotide polymorphisms (SNPs) and technical artifacts caused by sequencing or read-mapping errors. Recently, we and others have developed computational frameworks to identify RNA editing sites by comparing the sequence differences between RNA-seq and matched genomic DNA sequencing from a single individual6–8. This approach is robust in minimizing erroneous variant calls caused by sequencing or read-mapping errors, but it requires deep sequencing of both the transcriptome and the genome from the same sample. Samples with such data are Correspondence should be addressed to J.B.L. (jin.billy.li@stanford.edu). 3These authors contributed equally to this work. Accession codes. Gene Expression Omnibus: GSE42815 (sequencing data for wild-type and Adar5G1 Drosophila strains). Note: Supplementary information is available in the online version of the paper. AUTHOR CONTRIBUTIONS G.R. and R.Z. performed computational analyses with help from R.P., P.D. and J.B.L.; R.Z. and G.R. carried out the validation experiments; L.P.K. and M.A.O. generated RNA-seq data for wild-type and Adar−/− flies; and G.R., R.Z. and J.B.L. wrote the paper with input from other authors. COMPETING FINANCIAL INTERESTS The authors declare no competing financial interests. NIH Public Access Author Manuscript Nat Methods. Author manuscript; available in PMC 2013 August 01. Published in final edited form as: Nat Methods. 2013 February ; 10(2): 128–132. doi:10.1038/nmeth.2330. N IH PA Athor M anscript N IH PA Athor M anscript N IH PA Athor M anscript relatively uncommon, and are currently biased toward lymphocyte cell lines, which may not be biologically relevant for RNA editing studies. To use the multitude of publicly available RNA-seq data sets for discovery of RNA editing sites, we developed two related and complementary methods to accurately identify RNA editing sites using RNA-seq data from multiple individuals in a single species. In the first method (‘separate samples method’; Fig. 1a), RNA variants are called separately in each RNA-seq sample after mapping sequencing reads to a (nonmatched) genomic reference sequence, and known common genomic SNPs are removed. To distinguish RNA editing sites from rare SNPs in the remaining pool of RNA variants, we took advantage of the fact that the same editing sites are often present in different individuals whereas rare SNPs are most likely not. In the second method (‘pooled samples method’; Fig. 1b), RNA-seq alignments from different individuals are pooled together to achieve higher read coverage, enhancing the sensitivity for calling RNA variants. RNA variants are called, and common SNPs are removed, similarly to the separate samples method. As rare SNPs are unlikely to be present in multiple individuals, they exist at a very low frequency in the pooled alignment file. The method for mapping RNA-seq reads and calling variants is based on our previously published computational pipeline8 (Online Methods). The hallmark of our pipeline is separate filtering criteria for variants occurring in Alu repeats and variants occurring in nonAlu regions of the genome, resulting in much greater sensitivity in detecting editing sites in Alu repeats (where A-to-I editing is prevalent) and drastically improved specificity for detecting editing sites in non-Alu regions as compared to other methods8. The major modification from our previous pipeline is the use of the Genome Analysis ToolKit (GATK)9 instead of empirically determined parameters for variant calling to provide a uniform statistical framework for variant calling that can be applied to diverse RNA-seq data sets. We noticed that variant calling using empirical parameters instead of using GATK resulted in an abundance of false positive mismatches, especially when the proportion of transcripts being edited, or editing level, is very low (see below). As a proof of concept, we applied our two methods to identify RNA editing sites using RNA-seq data obtained from 40 human lymphoblastoid cell lines (Supplementary Note 1 and Supplementary Table 1). We found that the majority of mismatches identified using both methods were A-to-G mismatches, indicative of A-to-I editing (Supplementary Fig. 1). We observed a slight enrichment in T-to-C mismatches, the majority of which are incorrectly annotated A-to-G mismatches (Supplementary Note 1 and Supplementary Fig. 2). These same 40 RNA-seq data sets have been used in a previous study10 that provided evidence to support the possibility of noncanonical editing mechanisms. However, more recent studies have shown that these noncanonical mismatches are false positives8,11–16. Our results support the observation that all non–A-to-G mismatches are false positives. Furthermore, we analyzed the same lymphoblastoid RNA-seq data that had been used in the above-mentioned study10, and we only found evidence to support A-to-I editing in these samples. Overall, we identified 303,624 A-to-G variants in Alu repeats, 2,796 A-to-G variants in non-Alu repeats and 2,815 A-to-G variants in nonrepetitive regions using RNAseq data from lymphoblastoid cell lines (Supplementary Tables 2,3 and Supplementary Data 1,2). We found that more RNA editing sites were called using RNA-seq data only than by comparing sequence differences between RNA and DNA sequencing data using our previous method8 (Supplementary Note 1 and Supplementary Fig. 3). We greatly enhanced sensitivity to detect editing sites by using multiple RNA-seq samples, which allowed us to accurately identify RNA editing sites supported by only one mismatched read in a particular sample (Supplementary Fig. 4). Ramaswami et al. Page 2 Nat Methods. Author manuscript; available in PMC 2013 August 01. N IH PA Athor M anscript N IH PA Athor M anscript N IH PA Athor M anscript Next, we applied our approaches to identify RNA editing sites using RNA-seq data obtained from brain tissues of 50 human individuals (Fig. 1 and Supplementary Table 4). Using the separate samples method, we found that RNA variants present in one or more samples in Alu repeats and RNA variants present in two or more samples in non-Alu regions were highly enriched for potential A-to-I editing sites (Fig. 1c,d). Using the pooled samples method, we found that RNA variants with one or more variant reads in Alu repeats and RNA variants with two or more variant reads in non-Alu regions were highly enriched for potential A-to-I editing sites (Fig. 1e,f). We identified 612,573, 13,724, and 12,160 A-to-G variants in Alu repeats, non-Alu repeats, and nonrepetitive regions, repectively using RNAseq data from human brain tissues (Supplementary Fig. 5, Supplementary Table 2 and Supplementary Data 3,4). As expected for A-to-I editing sites17, these A-to-G variants spanned a wide spectrum of editing levels (Supplementary Fig. 6) and were associated with an underand overrepresentation of guanosines immediately 5′ and 3′ of the edited adenosine, respectively, although the sequence preferences at these two positions were not completely independent (Supplementary Fig. 7). We also identified RNA editing sites from other human tissues (Supplementary Note 2, Supplementary Fig. 8, Supplementary Tables 2,5 and Supplementary Data 5,6). Altogether, from human RNA-seq data alone we identified 996,012, 16,622 and 15,020 A-to-I RNA editing sites in Alu, repetitive non-Alu and non-repetitive regions, respectively, most of which we identified in the brain samples only (Supplementary Fig. 9). As large numbers of RNA editing sites are identified, it is difficult to pinpoint the functionally important ones. Additionally, the accuracy (proportion of total variants that are A-to-G type, or A-to-G fraction) of the two methods described above in functionally important regions, such as in nonrepetitive coding regions, is not as good as in intronic or untranslated regions (Supplementary Table 3), most likely because of challenges in mapping reads to spliced exons. To address these challenges, we developed a cross-species transcriptome comparison method based on the fact that functionally relevant RNA editing events tend to be conserved between related species, whereas SNPs or false positives, mainly from errors in DNA sequencing and computational mapping, are unlikely to be common to unrelated species (Supplementary Fig. 10). To enrich for functionally relevant editing sites, we focused on identifying conserved RNA variants in exonic regions. We first applied this method to the primate lineage to identify human RNA editing sites c