The advent of single-molecule, long-read sequencing (LRS) technologies by Oxford Nanopore Technologies and Pacific Biosciences has revolutionized genomics, transcriptomics and, more recently, epigenomics research. These technologies offer distinct advantages, including the direct detection of methylated DNA and simultaneous assessment of DNA sequences spanning multiple kilobases along with their modifications at the single-molecule level. This has enabled the development of new assays for analyzing chromatin states and made it possible to integrate data for DNA methylation, chromatin accessibility, transcription factor binding and histone modifications, thereby facilitating comprehensive epigenomic profiling. Owing to recent advancements, alternative, nascent and translating transcripts can be detected using LRS approaches. This Review discusses LRS-based experimental and computational strategies for characterizing chromatin states and highlights their advantages over short-read sequencing methods. Furthermore, we demonstrate how various long-read methods can be integrated to design multi-omics studies to investigate the relationship between chromatin states and transcriptional dynamics. Long-read sequencing technologies have revolutionized genomics and transcriptomics, and more recently enabled comprehensive epigenomic profiling. These advances now also allow investigation of the relationship between chromatin states and transcriptional dynamics.
SQANTI3 is a tool designed for the quality control, curation and annotation of long-read transcript models obtained with third-generation sequencing technologies. Leveraging its annotation framework, SQANTI3 calculates quality descriptors of transcript models, junctions and transcript ends. With this information, potential artifacts can be identified and replaced with reliable sequences. Furthermore, the integrated functional annotation feature enables subsequent functional iso-transcriptomics analyses.
The Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.
Long-read RNA sequencing has emerged as a powerful tool for transcript discovery, even in well-annotated organisms. However, assessing the accuracy of different methods in identifying annotated and novel transcripts remains a challenge. Here, we present SQANTI-SIM, a versatile tool that wraps around popular long-read simulators to allow precise management of transcript novelty based on the structural categories defined by SQANTI3. By selectively excluding specific transcripts from the reference dataset, SQANTI-SIM effectively emulates scenarios involving unannotated transcripts. Furthermore, the tool provides customizable features and supports the simulation of additional types of data, representing the first multi-omics simulation tool for the lrRNA-seq field.
The emergence of long-read RNA sequencing (lrRNA-seq) has provided an unprecedented opportunity to analyze transcriptomes at isoform resolution. However, the technology is not free from biases, and transcript models inferred from these data require quality control and curation. In this study, we introduce SQANTI3, a tool specifically designed to perform quality analysis on transcriptomes constructed using lrRNA-seq data. SQANTI3 provides an extensive naming framework to describe transcript model diversity in comparison to the reference transcriptome. Additionally, the tool incorporates a wide range of metrics to characterize various structural properties of transcript models, such as transcription start and end sites, splice junctions, and other structural features. These metrics can be utilized to filter out potential artifacts. Moreover, SQANTI3 includes a Rescue module that prevents the loss of known genes and transcripts exhibiting evidence of expression but displaying low-quality features. Lastly, SQANTI3 incorporates IsoAnnotLite, which enables functional annotation at the isoform level and facilitates functional iso-transcriptomics analyses. We demonstrate the versatility of SQANTI3 in analyzing different data types, isoform reconstruction pipelines, and sequencing platforms, and how it provides novel biological insights into isoform biology. The SQANTI3 software is available at https://github.com/ConesaLab/SQANTI3 .
PaintOmics is a web server for the integrative analysis and visualisation of multi-omics datasets using biological pathway maps. PaintOmics 4 has several notable updates that improve and extend analyses. Three pathway databases are now supported: KEGG, Reactome and MapMan, providing more comprehensive pathway knowledge for animals and plants. New metabolite analysis methods fill gaps in traditional pathway-based enrichment methods. The metabolite hub analysis selects compounds with a high number of significant genes in their neighbouring network, suggesting regulation by gene expression changes. The metabolite class activity analysis tests the hypothesis that a metabolic class has a higher-than-expected proportion of significant elements, indicating that these compounds are regulated in the experiment. Finally, PaintOmics 4 includes a regulatory omics module to analyse the contribution of trans-regulatory layers (microRNA and transcription factors, RNA-binding proteins) to regulate pathways. We show the performance of PaintOmics 4 on both mouse and plant data to highlight how these new analysis features provide novel insights into regulatory biology. PaintOmics 4 is available at https://paintomics.org/.
Worldwide COVID-19 epidemiology data indicate differences in disease incidence amongst sex and gender demographic groups. Specifically, male patients are at a higher death risk than female patients, and the older population is significantly more affected than young individuals. Whether this difference is a consequence of a pre-existing differential response to the virus, has not been studied in detail. We created DeCovid, an R shiny app that combines gene expression (GE) data of different human tissue from the Genotype-Tissue Expression (GTEx) project along with the COVID-19 Disease Map and COVID-19 related pathways gene collections to explore basal GE differences across healthy demographic groups. We used this app to study differential gene expression of COVID-19 associated genes in different age and sex groups. We identified that healthy women show higher expression-levels of interferon genes. Conversely, healthy men exhibit higher levels of proinflammatory cytokines. Additionally, young people present a stronger complement system and maintain a high level of matrix metalloproteases than older adults. Our data suggest the existence of different basal immunophenotypes amongst different demographic groups, which are relevant to COVID-19 progression and may contribute to explaining sex and age biases in disease severity. The DeCovid app is an effective and easy to use tool for exploring the GE levels relevant to COVID-19 across demographic groups and tissues.
MOTIVATION microRNAs (miRNAs) are essential components of gene expression regulation at the post-transcriptional level. miRNAs have a well-defined molecular structure and this has facilitated the development of computational and high-throughput approaches to predict miRNAs genes. However, due to their short size, miRNAs have often been incorrectly annotated in both plants and animals. Consequently, published miRNA annotations and miRNA databases are enriched for false miRNAs, jeopardizing their utility as molecular information resources. To address this problem, we developed MirCure, a new software for quality control, filtering and curation of miRNA candidates. MirCure is an easy-to-use tool with a graphical interface that allows both scoring of miRNA reliability and browsing of supporting evidence by manual curators. RESULTS Given a list of miRNA candidates, MirCure evaluates a number of miRNA-specific features based on gene expression, biogenesis and conservation data, and generates a score that can be used to discard poorly supported miRNA annotations. MirCure can also curate and adjust the annotation of the 5p and 3p arms based on user-provided small RNA-seq data. We evaluated MirCure on a set of manually curated animal and plant miRNAs and demonstrated great accuracy. Moreover, we show that MirCure can be used to revisit previous bona fide miRNAs annotations to improve miRNA databases. AVAILABILITY AND IMPLEMENTATION The MirCure software and all the additional scripts used in this project are publicly available at https://github.com/ConesaLab/MirCure. A Docker image of MirCure is available at https://hub.docker.com/r/conesalab/mircure. SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.