The Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.
Background Spatial transcriptomics allows gene expression to be measured within complex tissue contexts. Among the array of spatial capture technologies available is 10x Genomics’ Visium platform, a popular method which enables transcriptomewide profiling of tissue sections. Visium offers a range of sample handling and library construction methods which introduces a need for benchmarking to compare data quality and assess how well the technology can recover expected tissue features and biological signatures. Results Here we present SpatialBench , a unique reference dataset generated from spleen tissue of mice responding to malaria infection spanning several tissue preparation protocols (both fresh frozen and FFPE samples, with and without CytAssist tissue placement). We noted better quality control metrics in reference samples prepared using probe-based capture methods, particularly those processed with CytAssist, validating the improvement in data quality produced with the platform. Our analysis of replicate samples extends to explore spatially variable gene detection, the outcomes of clustering and cell deconvolution using matched single-cell RNA-sequencing data and publicly available reference data to identify cell types and tissue regions expected in the spleen. Multi-sample differential expression analysis recovered known gene signatures related to biological sex or gene knockout. Conclusions We framed a comprehensive multi-sample analysis workflow that allowed us to generate consistent results both within and between different subsets of replicate samples, enabling broader comparisons and interpretations to be made at the group-level. Our SpatialBench dataset, analysis, and workflow can serve as a practical guide for Visium users and may prove valuable in other benchmarking studies.
MOTIVATION:The process of analyzing high throughput sequencing data often requires the identification and extraction of specific target sequences. This could include tasks, such as identifying cellular barcodes and UMIs in single-cell data, and specific genetic variants for genotyping. However, existing tools, which perform these functions are often task-specific, such as only demultiplexing barcodes for a dedicated type of experiment, or are not tolerant to noise in the sequencing data. RESULTS:To overcome these limitations, we developed Flexiplex, a versatile and fast sequence searching and demultiplexing tool for omics data, which is based on the Levenshtein distance and thus allows imperfect matches. We demonstrate Flexiplex's application on three use cases, identifying cell-line-specific sequences in Illumina short-read single-cell data, and discovering and demultiplexing cellular barcodes from noisy long-read single-cell RNA-seq data. We show that Flexiplex achieves an excellent balance of accuracy and computational efficiency compared to leading task-specific tools. AVAILABILITY AND IMPLEMENTATION:Flexiplex is available at https://davidsongroup.github.io/flexiplex/.
A modified Chromium 10x droplet-based protocol that subsamples cells for both short-read and long-read (nanopore) sequencing together with a new computational pipeline (FLAMES) is developed to enable isoform discovery, splicing analysis, and mutation detection in single cells. We identify thousands of unannotated isoforms and find conserved functional modules that are enriched for alternative transcript usage in different cell types and species, including ribosome biogenesis and mRNA splicing. Analysis at the transcript level allows data integration with scATAC-seq on individual promoters, improved correlation with protein expression data, and linked mutations known to confer drug resistance to transcriptome heterogeneity.