Most approaches to transcript quantification rely on fixed reference annotations; however, the transcriptome is dynamic and depending on the context, such static annotations contain inactive isoforms for some genes, whereas they are incomplete for others. Here we present Bambu, a method that performs machine-learning-based transcript discovery to enable quantification specific to the context of interest using long-read RNA-sequencing. To identify novel transcripts, Bambu estimates the novel discovery rate, which replaces arbitrary per-sample thresholds with a single, interpretable, precision-calibrated parameter. Bambu retains the full-length and unique read counts, enabling accurate quantification in presence of inactive isoforms. Compared to existing methods for transcript discovery, Bambu achieves greater precision without sacrificing sensitivity. We show that context-aware annotations improve quantification for both novel and known transcripts. We apply Bambu to quantify isoforms from repetitive HERVH-LTR7 retrotransposons in human embryonic stem cells, demonstrating the ability for context-specific transcript expression analysis. Leveraging long-read RNA-seq data and machine learning, Bambu facilitates accurate transcript discovery and quantification.
List of ncNATs expressed in TCGA tumors, for which their associated sense genes have two active alternative promoters
List of 353 unique ncNATs expressed in tumors of one or multiple tumor types from TCGA.
List of ncNATs expressed in TCGA and GTEx non-tumor samples, for which their associated sense genes have two active alternative promoters
List of 521 unique ncNATs expressed in non-tumor samples of one or multiple cancer/tissue types from TCGA and GTEx.
The human genome contains more than 200,000 gene isoforms. However, different isoforms can be highly similar, and with an average length of 1.5kb remain difficult to study with short read sequencing. To systematically evaluate the ability to study the transcriptome at a resolution of individual isoforms we profiled 5 human cell lines with short read cDNA sequencing and Nanopore long read direct RNA, amplification-free direct cDNA, PCR-cDNA sequencing. The long read protocols showed a high level of consistency, with amplification-free RNA and cDNA sequencing being most similar. While short and long reads generated comparable gene expression estimates, they differed substantially for individual isoforms. We find that increased read length improves read-to-transcript assignment, identifies interactions between alternative promoters and splicing, enables the discovery of novel transcripts from repetitive regions, facilitates the quantification of full-length fusion isoforms and enables the simultaneous profiling of m6A RNA modifications when RNA is sequenced directly. Our study demonstrates the advantage of long read RNA sequencing and provides a comprehensive resource that will enable the development and benchmarking of computational methods for profiling complex transcriptional events at isoform-level resolution.
Abstract Multiple noncoding natural antisense transcripts (ncNAT) are known to modulate key biological events such as cell growth or differentiation. However, the actual impact of ncNATs on cancer progression remains largely unknown. In this study, we identified a complete list of differentially expressed ncNATs in hepatocellular carcinoma. Among them, a previously undescribed ncNAT HNF4A-AS1L suppressed cancer cell growth by regulating its sense gene HNF4A, a well-known cancer driver, through a promoter-specific mechanism. HNF4A-AS1L selectively activated the HNF4A P1 promoter via HNF1A, which upregulated expression of tumor suppressor P1-driven isoforms, while having no effect on the oncogenic P2 promoter. RNA-seq data from 23 tissue and cancer types identified approximately 100 ncNATs whose expression correlated specifically with the activity of one promoter of their associated sense gene. Silencing of two of these ncNATs ENSG00000259357 and ENSG00000255031 (antisense to CERS2 and CHKA, respectively) altered the promoter usage of CERS2 and CHKA. Altogether, these results demonstrate that promoter-specific regulation is a mechanism used by ncNATs for context-specific control of alternative isoform expression of their counterpart sense genes. Significance: This study characterizes a previously unexplored role of ncNATs in regulation of isoform expression of associated sense genes, highlighting a mechanism of alternative promoter usage in cancer.
SummaryDothistroma needle blight is one of the most devastating pine tree diseases worldwide. New and emerging epidemics have been frequent over the last 25 years, particularly in the Northern Hemisphere, where they are in part associated with changing weather patterns. One of the main Dothistroma needle blight pathogens, Dothistroma septosporum, has a global distribution but most molecular plant pathology research has been confined to Southern Hemisphere populations that have limited genetic diversity. Extensive genomic and transcriptomic data are available for a D. septosporum reference strain from New Zealand, where an introduced clonal population of the pathogen predominates. Due to the global importance of this pathogen, we determined whether the genome of this reference strain is representative of the species worldwide by sequencing the genomes of 18 strains sampled globally from different pine hosts. Genomic polymorphism shows substantial variation within the species, clustered into two distinct groups of strains with centres of diversity in Central and South America. A reciprocal chromosome translocation uniquely identifies the New Zealand strains. Globally, strains differ in their production of the virulence factor dothistromin, with extremely high production levels in strain ALP3 from Germany. Comparisons with the New Zealand reference revealed that several strains are aneuploids; for example, ALP3 has duplications of three chromosomes. Increased gene copy numbers therefore appear to contribute to increased production of dothistromin, emphasizing that studies of population structure are a necessary adjunct to functional analyses of genetic polymorphisms to identify the molecular basis of virulence in this important forest pathogen.
Fungal secondary metabolites have many important biological roles and some, like the toxic polyketide aflatoxin, have been intensively studied at the genetic level. Complete sets of polyketide synthase (PKS) genes can now be identified in fungal pathogens by whole genome sequencing and studied in order to predict the biosynthetic potential of those fungi. The pine needle pathogen Dothistroma septosporum is predicted to have only three functional PKS genes, a small number for a hemibiotrophic fungus. One of these genes is required for production of dothistromin, a polyketide virulence factor related to aflatoxin, whose biosynthetic genes are dispersed across one chromosome rather than being clustered. Here we evaluated the evolution of the other two genes, and their predicted gene clusters, using phylogenetic and population analyses. DsPks1 and its gene cluster are quite conserved amongst related fungi, whilst DsPks2 appears to be novel. The DsPks1 protein was predicted to be required for dihydroxynaphthalene (DHN) melanin biosynthesis but functional analysis of DsPks1 mutants showed that D. septosporum produced mainly dihydroxyphenylalanine (DOPA) melanin, which is produced by a PKS-independent pathway. Although the secondary metabolites made by these two PKS genes are not known, comparisons between strains of D. septosporum from different regions of the world revealed that both PKS core genes are under negative selection and we suggest they may have important cryptic roles in planta.
At least since the Neolithic, humans have largely lived in networks of small, traditional communities. Often socially isolated, these groups evolved distinct languages and cultures over microgeographic scales of just tens of kilometers. Population genetic theory tells us that genetic drift should act quickly in such isolated groups, thus raising the question: do networks of small human communitiesmaintain levels of genetic diversity over microgeographic scales? This question can no longer be asked in most parts of the world, which have been heavily impacted by historical events that make traditional society structures the exception. However, such studies remain possible in parts of Island Southeast Asia and Oceania, where traditional ways of life are still practiced. We captured genome-wide genetic data, together with linguistic records, for a case-study system-eight villages distributed across Sumba, a small, remote island in eastern Indonesia. More than 4,000 years after these communities were established during the Neolithic period, most speak different languages and can be distinguished genetically. Yet their nuclear diversity is not reduced, instead being comparable to other, evenmuch larger, regional groups. Modeling reveals a separation of time scales: while languages and culture can evolve quickly, creating social barriers, sporadic migration averaged over many generations is sufficient to keep villages linked genetically. This loosely-connected network structure, once the global norm and still extant on Sumba today, provides a living proxy to explore fine-scale genome dynamics in the sort of small traditional communities within which the most recent episodes of human evolution occurred.
Background Prior to egg laying the parasitoid wasp Nasonia vitripennis envenomates its pupal host with a complex mixture of venom peptides. This venom induces several dramatic changes in the host, including developmental arrest, immunosuppression, and altered metabolism. The diverse and potent bioactivity of N. vitripennis venom provides opportunities for the development of novel acting pharmaceuticals based on these molecules. However, currently very little is known about the specific functions of individual venom peptides or what mechanisms underlie the hosts response to envenomation. Many of the venom peptides also lack bioinformatically derived annotations because no homologs can be identified in the sequences databases. The RNA interference system of N. vitripennis provides a method for functional characterisation of venom protein encoding genes, however working with the current list of 79 candidates represents a daunting task. For this reason we were interested in determining the expression levels of venom encoding genes in the venom gland, as this information could be used to rank candidates for further study. To do this we carried out deep transcriptome sequencing of the venom gland and ovary tissue and used RNA-seq to rank the venom protein encoding genes by expression level. The generation of a specific venom gland transcriptome dataset also provides further opportunities to investigate novel features of this specialised organ. Results RNA-seq revealed that the highest expressed venom encoding gene in the venom gland was ‘Venom protein Y’. The highest expressed annotated gene in this tissue was serine protease Nasvi2EG007167 , which has previously been implicated in the apoptotic activity of N. vitripennis venom. As expected the RNA-seq confirmed that venom encoding genes are almost exclusively expressed in the venom gland relative to the neighbouring ovary tissue. Novel genes appear to perform key roles in N. vitripennis venom function, with over half of the 15 highest expressed venom encoding loci lacking bioinformatic annotations. The high throughput sequencing data also provided evidence for the existence of an additional 472 previously undescribed transcribed regions in the N. vitripennis genome. Finally, metatranscriptomic analysis of the venom gland transcriptome finds little evidence for the role of Wolbachia in the venom system. Conclusions The expression level information provided here for the N. vitripennis venom protein encoding genes represents a valuable dataset that can be used by the research community to rank candidates for further functional characterisation. These candidates represent bioactive peptides valuable in the development of new pharmaceuticals.
SummaryWe present genome‐wide gene expression patterns as a time series through the infection cycle of the fungal pine needle blight pathogen, Dothistroma septosporum, as it invades its gymnosperm host, Pinus radiata. We determined the molecular changes at three stages of the disease cycle: epiphytic/biotrophic (early), initial necrosis (mid) and mature sporulating lesion (late). Over 1.7 billion combined plant and fungal reads were sequenced to obtain 3.2 million fungal‐specific reads, which comprised as little as 0.1% of the sample reads early in infection. This enriched dataset shows that the initial biotrophic stage is characterized by the up‐regulation of genes encoding fungal cell wall‐modifying enzymes and signalling proteins. Later necrotrophic stages show the up‐regulation of genes for secondary metabolism, putative effectors, oxidoreductases, transporters and starch degradation. This in‐depth through‐time transcriptomic study provides our first snapshot of the gene expression dynamics that characterize infection by this fungal pathogen in its gymnosperm host.