Long-read RNA sequencing has arisen as a counterpart to short-read sequencing, with the potential to capture full-length isoforms, albeit at the cost of lower depth. Yet this potential is not fully realized due to inherent limitations of current long-read assembly methods and underdeveloped approaches to integrate short-read data. Here, we critically compare the existing methods and develop a new integrative approach to characterize a particularly challenging pool of low-abundance long noncoding RNA (lncRNA) transcripts from short- and long-read sequencing in two distinct cell lines. Our analysis reveals severe limitations in each of the sequencing platforms. For short-read assemblies, coverage declines at transcript termini resulting in ambiguous ends, and uneven low coverage results in segmentation of a single transcript into multiple transcripts. Conversely, long-read sequencing libraries lack depth and strand-of-origin information in cDNA-based methods, culminating in erroneous assembly and quantitation of transcripts. We also discover a cDNA synthesis artifact in long-read datasets that markedly impacts the identity and quantitation of assembled transcripts. Towards remediating these problems, we develop a computational pipeline to "strand" long-read cDNA libraries that rectifies inaccurate mapping and assembly of long-read transcripts. Leveraging the strengths of each platform and our computational stranding, we also present and benchmark a hybrid assembly approach that drastically increases the sensitivity and accuracy of full-length transcript assembly on the correct strand and improves detection of biological features of the transcriptome. When applied to a challenging set of under-annotated and cell-type variable lncRNA, our method resolves the segmentation problem of short-read sequencing and the depth problem of long-read sequencing, resulting in the assembly of coherent transcripts with precise 5' and 3' ends. Our workflow can be applied to existing datasets for superior demarcation of transcript ends and refined isoform structure, which can enable better differential gene expression analyses and molecular manipulations of transcripts.
Long introns with short exons in vertebrate genes are thought to require spliceosome assembly across exons (exon definition), rather than introns, thereby requiring transcription of an exon to splice an upstream intron. Here, we developed CoLa-seq (co-transcriptional lariat sequencing) to investigate the timing and determinants of co-transcriptional splicing genome wide. Unexpectedly, 90% of all introns, including long introns, can splice before transcription of a downstream exon, indicating that exon definition is not obligatory for most human introns. Still, splicing timing varies dramatically across introns, and various genetic elements determine this variation. Strong U2AF2 binding to the polypyrimidine tract predicts early splicing, explaining exon definition-independent splicing. Together, our findings question the essentiality of exon definition and reveal features beyond intron and exon length that are determinative for splicing timing.
Withdrawal StatementThe corresponding author has withdrawn this preprint owing to inability to reproduce some of the data, instances of inappropriate data exclusion, and loss of much of the primary experimental records/data. Specifically, Figures 2B,E,H; 3; 4; 5A,B,D; 6; S2A,C; S4A; S5; and S6A,B and attendant text contain analyses for which the primary record and/or raw data no longer exist; the analyses, where still available, suffer from inappropriate data exclusion and thus should not be construed to be an accurate reflection of the experiments. Attempts by others in the lab to repeat several of the experiments in these indicated panels have failed reproduce the presented effects, despite showing much greater precision. Therefore the authors do not wish this work to be cited as reference for the project. If you have any questions, please contact the corresponding author.
As splicing is intimately coupled with transcription, understanding splicing mechanisms requires an understanding of splicing timing, which is currently limited. Here, we developed CoLa-seq ( co -transcriptional la riat seq uencing), a genomic assay that reports splicing timing relative to transcription through analysis of nascent lariat intermediates. In human cells, we mapped 165,282 branch points and characterized splicing timing for over 70,000 introns. Splicing timing varies dramatically across introns, with regulated introns splicing later than constitutive introns. Machine learning-based modeling revealed genetic elements predictive of splicing timing, notably the polypyrimidine tract, intron length, and regional GC content, which illustrate the significance of the broader genomic context of an intron and the impact of co-transcriptional splicing. The importance of the splicing factor U2AF in early splicing rationalizes surprising observations that most introns can splice independent of exon definition. Together, these findings establish a critical framework for investigating the mechanisms and regulation of co-transcriptional splicing. Highlights CoLa-seq enables cell-type specific, genome-wide branch point annotation with unprecedented efficiency. CoLa-seq captures co-transcriptional splicing for tens of thousands of introns and reveals splicing timing varies dramatically across introns. Modeling uncovers key genetic determinants of splicing timing, most notably regional GC content, intron length, and the polypyrimidine tract, the binding site for U2AF2. Early splicing precedes transcription of a downstream 5’ SS and in some cases accessibility of the upstream 3’ SS, precluding exon definition.
Long noncoding RNAs (lncRNAs) localize in the cell nucleus and influence gene expression through a variety of molecular mechanisms. Chromatin-enriched RNAs (cheRNAs) are a unique class of lncRNAs that are tightly bound to chromatin and putatively function to locally cis-activate gene transcription. CheRNAs can be identified by biochemical fractionation of nuclear RNA followed by RNA sequencing, but until now, a rigorous analytic pipeline for nuclear RNA-seq has been lacking. In this study, we survey four computational strategies for nuclear RNA-seq data analysis and develop a new pipeline, Tuxedo-ch, which outperforms other approaches. Tuxedo-ch assembles a more complete transcriptome and identifies cheRNA with higher accuracy than other approaches. We used Tuxedo-ch to analyze benchmark datasets of K562 cells and further characterize the genomic features of intergenic cheRNA (icheRNA) and their similarity to enhancer RNAs (eRNAs). We quantify the transcriptional correlation of icheRNA and adjacent genes and show that icheRNA is more positively associated with neighboring gene expression than eRNA or cap analysis of gene expression (CAGE) signals. We also explore two novel genomic associations of cheRNA, which indicate that cheRNAs may function to promote or repress gene expression in a context-dependent manner. IcheRNA loci with significant levels of H3K9me3 modifications are associated with active enhancers, consistent with the hypothesis that enhancers are derived from ancient mobile elements. In contrast, antisense cheRNA (as-cheRNA) may play a role in local gene repression, possibly through local RNA:DNA:DNA triple-helix formation.
The noncoding genome is pervasively transcribed. Noncoding RNAs (ncRNAs) generated from enhancers have been proposed as a general facet of enhancer function and some have been shown to be required for enhancer activity. Here we examine the transcription-factor-(TF)-dependence of ncRNA expression to define enhancers and enhancer-associated ncRNAs that are involved in a TF-dependent regulatory network. TBX5, a cardiac TF, regulates a network of cardiac channel genes to maintain cardiac rhythm. We deep sequenced wildtype and Tbx5-mutant mouse atria, identifying ~2600 novel Tbx5-dependent ncRNAs. Tbx5-dependent ncRNAs were enriched for tissue-specific marks of active enhancers genome-wide. Tbx5-dependent ncRNAs emanated from regions that are enriched for TBX5-binding and that demonstrated Tbx5-dependent enhancer activity. Tbx5-dependent ncRNA transcription provided a quantitative metric of Tbx5-dependent enhancer activity, correlating with target gene expression. We identified RACER, a novel Tbx5-dependent long noncoding RNA (lncRNA) required for the expression of the calcium-handling gene Ryr2. We illustrate that TF-dependent enhancer transcription can illuminate components of TF-dependent gene regulatory networks.
Why do dead batteries bounce considerably higher than fresh batteries? This phenomenon has a chemical explanation that can be used to teach students about the chemistry of alkaline Zn/MnO2 cells. Batteries discharged to various extents can be tested for bounciness and conversion of Zn to ZnO. These measurements allow students to connect the chemistry that powers these batteries with the increased bouncing effect. The experiments can be presented as a teacher-led demonstration or hands-on laboratory for students.