Virus-derived circular RNA molecules (VcircRNAs) are expressed by many RNA viruses during infection. Putative functions include modulating viral replication and interacting with the host immune response. Some function as non-coding RNA fragments that regulate gene expression through binding to complementary RNA sequences, whereas others contain internal ribosomal entry site (IRES) sequences or non-canonical modifications that allow them to be translated. Here, we confirm the expression of a distinct SARS-CoV-2 VcircRNA molecule, circ7b8N, that has not been previously identified. We found that circ7b8N is expressed and detectable in cell culture infections and in acute infections across SARS-CoV-2 variants and shows promise for detection in post-acute clinical samples. Conservation of circ7b8N junctions is limited to the nearest phylogenetic relatives within the betacoronavirus genus but are present in other human and bat-infecting coronaviruses. Host cell gene expression is modulated by the treatment with circ7b8N agnostic of viral infection. The discovery and subsequent confirmation of circ7b8N expressed by SARS-CoV-2 provides a new biomarker for infection, and its conservation across variants suggests functional importance.
Single-cell RNA sequencing (scRNA-seq) analyses typically focus on alignment and differential gene expression, being blind to other transcriptome variation. Here we present sc-SPLASH for statistics-first, reference-free discovery on barcoded scRNA-seq and spatial transcriptomics; its independent BKC submodule optimizes barcoded data preprocessing and is approximately 50-fold faster than UMI-tools. sc-SPLASH discovers secreted repeat proteins in immune-like cells of sponge (Spongilla; missing from the reference) and tunicate (Ciona). sc-SPLASH extends reference-free analysis to barcoded single-cell and spatial transcriptomics data.
Most plant genomes and their (post-)transcriptional regulation remain unknown. We used SPLASH-a new, reference genome-free sequence variation detection algorithm-to analyze transcriptional and post-transcriptional regulation from RNA-seq data. We discovered allelic variation in expression during maize pollen development and imbibition-dependent cryptic splicing in Arabidopsis seeds. SPLASH enables discovery of novel regulatory mechanisms, including differential regulation of genes from parental haplotypes of hybrids, without the use of alignment to a reference genome.
Genome-wide association studies (GWAS) map genetic variation to a reference genome and correlate variants to phenotypes. Yet, GWAS and similar procedures have limitations, including an inability to predict phenotype on variants never seen during the discovery phase and difficulty integrating structural variants. Deep and machine learning alternatives have not been successful at consistent prediction of resistance phenotypes (Hu et al. 2024). Here, we introduce FLASH: a new interpretable, statistically-based deep learning framework that operates directly on raw sequencing reads. In over 35,000 isolates of bacteria, fungi and viruses, FLASH achieves uniformly high accuracy on independent test data, including on variation never seen in training, meeting or exceeding bespoke state of the art methods. FLASH identifies canonical drug targets ab initio and new pan-species predictors of virulence, including those lacking annotation and those only partially aligned to NCBI reference databases. Further, FLASH can predict phenotypes beyond the possibility of GWAS, such as bacterial host range of phage, a task that to our knowledge is impossible today. FLASH is simple to run, highly efficient and constitutes a new approach for predicting gene function and phenotype across the tree of life. It is especially valuable when bioethical concerns and the vast genetic complexity of pathogenic microbes limit the feasibility of experimental validation.
Realizing the promise of precision medicine will require the highest standards of accuracy in genome sequencing and analysis. Here we describe challenges and opportunities for the field through the lens of genome data quality. We present recommendations in the context of specific areas of application for genomic sequencing in which isolated standards have arisen: germline sequencing, tumour sequencing, cell-free DNA testing, and sequencing for quality control in genetic therapy. Despite these distinct clinical contexts, technical challenges are often similar; for example, accurately detecting low-frequency genetic variants in tumour sequencing or gene-edited cells. We call for increased synchronization among these communities to establish new medical genome standards that promote confidence in genomic diagnostics and genetic therapies in a time of rapid technology-driven change. We suggest practical approaches for implementing these genome standards across contexts, and identify key areas that require further development.
Mouse lemurs (Microcebus spp.) are an emerging primate model organism, but their genetics, cellular and molecular biology remain largely unexplored. In an accompanying paper1, we performed large-scale single-cell RNA sequencing of 27 organs from mouse lemurs. We identified more than 750 molecular cell types, characterized their transcriptomic profiles and provided insight into primate evolution of cell types. Here we use the generated atlas to characterize mouse lemur genes, physiology, disease and mutations. We uncover thousands of previously unidentified lemur genes and hundreds of thousands of new splice junctions including over 85,000 primate splice junctions missing in mice. We systematically explore the lemur immune system by comparing global expression profiles of key immune genes in health and disease, and by mapping immune cell development, trafficking and activation. We characterize primate-specific and lemur-specific physiology and disease, including molecular features of the immune program, lemur adipocytes and metastatic endometrial cancer that resembles the human malignancy. We present expression patterns of more than 400 primate genes missing in mice, many with similar expression patterns to humans and some implicated in human disease. Finally, we provide an experimental framework for reverse genetic analysis by identifying naturally occurring nonsense mutations in three primate immune genes missing in mice and by analysing their transcriptional phenotypes. This work establishes a foundation for molecular and genetic analyses of mouse lemurs and prioritizes primate genes, isoforms, physiology and disease for future study.
AbstractMyriad mechanisms diversify the sequence content of eukaryotic transcripts at both the DNA and RNA levels, leading to profound functional consequences. Examples of this diversity include RNA splicing and V(D)J recombination. Currently, these mechanisms are detected using fragmented bioinformatic tools that require predefining a form of transcript diversification and rely on alignment to an incomplete reference genome, filtering out unaligned sequences, potentially crucial for novel discoveries. Here, we present SPLASH+, significantly advancing biological discovery possible with SPLASH, our recently introduced efficient, reference-free statistical approach. Integrating a micro-assembly and biological interpretation framework, SPLASH+ enables new discoveries including broad and novel examples of transcript diversification in single cellsde novo, without the need for cell type metadata, which is impossible with current algorithms. Applied to 10,326 primary human single cells across 19 tissues profiled with SmartSeq2, SPLASH+ discovers a set of splicing and histone regulators with highly conserved intronic regions that are themselves subject to complex splicing regulation. Additionally, it reveals unreported transcript diversity in the heat shock proteinHSP90AA1, as well as diversification in centromeric RNA expression, V(D)J recombination, RNA editing, and repeat expansion, all missed by existing methods. SPLASH+ is highly efficient, enabling the discovery of an unprecedented breadth of RNA regulation and diversification in single cells through a new automated paradigm of unbiased transcriptomic analysis.
The detection of circular RNA molecules (circRNAs) is typically based on short-read RNA sequencing data processed using computational tools. Numerous such tools have been developed, but a systematic comparison with orthogonal validation is missing. Here, we set up a circRNA detection tool benchmarking study, in which 16 tools detected more than 315,000 unique circRNAs in three deeply sequenced human cell types. Next, 1,516 predicted circRNAs were validated using three orthogonal methods. Generally, tool-specific precision is high and similar (median of 98.8%, 96.3% and 95.5% for qPCR, RNase R and amplicon sequencing, respectively) whereas the sensitivity and number of predicted circRNAs (ranging from 1,372 to 58,032) are the most significant differentiators. Of note, precision values are lower when evaluating low-abundance circRNAs. We also show that the tools can be used complementarily to increase detection sensitivity. Finally, we offer recommendations for future circRNA detection and validation. This study describes benchmarking and validation of computational tools for detecting circRNAs, finding most to be highly precise with variations in sensitivity and total detection. The study also finds over 315,000 putative human circRNAs.
Pre-training is a powerful paradigm in machine learning to pass information across models. For example, suppose one has a modest-sized dataset of images of cats and dogs and plans to fit a deep neural network to classify them. With pre-training, we start with a neural network trained on a large corpus of images of not just cats and dogs but hundreds of classes. We fix all network weights except the top layer(s) and fine tune on our dataset. This often results in dramatically better performance than training solely on our dataset. Here, we ask: ‘Can pre-training help the lasso?’. We propose a framework where the lasso is fit on a large dataset and then fine-tuned on a smaller dataset. The latter can be a subset of the original, or have a different but related outcome. This framework has a wide variety of applications, including stratified and multi-response models. In the stratified model setting, lasso pre-training first estimates coefficients common to all groups, then estimates group-specific coefficients during fine-tuning. Under appropriate assumptions, support recovery of the common coefficients is superior to the usual lasso trained on individual groups. This separate identification of common and individual coefficients also aids scientific understanding.
Bacteria comprise > 12% of Earth’s biomass and profoundly impact human and planetary health.[1][1] Many key biological functions of microbes, and functions differentiating strains, are conferred or modified by genome plasticity including mobilization of genetic elements, phage integration, and CRISPR arrays. Characterizing each of these processes is time-consuming and requires custom bioinformatic workflows ill-suited to enable discovery of new sources of genetic diversity or to uncover which elements are active. Further, strain typing of bacterial species and approaches to discriminate sub-populations remain time-consuming and resource intensive. Here, we show that SPLASH, our published approach for reference-free discovery and analysis directly from raw reads, and an improved statistical assembly algorithm, compactors, unify diverse tasks in microbial sequence analysis: discovering new mobile elements and CRISPR arrays missing from any reference, and generating rapid, metadata-free strain typing of diverse bacteria. SPLASH and compactors together constitute a new general discovery tool for biological discovery in the microbial world.### Competing Interest StatementThe authors have declared no competing interest. [1]: #ref-1
SPLASH is an unsupervised, reference-free, and unifying algorithm that discovers regulated sequence variation through statistical analysis of k -mer composition, subsuming many application-specific methods. Here, we introduce SPLASH2, a fast, scalable implementation of SPLASH based on an efficient k -mer counting approach. SPLASH2 enables rapid analysis of massive datasets from a wide range of sequencing technologies and biological contexts, delivering unparalleled scale and speed. The SPLASH2 algorithm unveils new biology (without tuning) in single-cell RNA-sequencing data from human muscle cells, as well as bulk RNA-seq from the entire Cancer Cell Line Encyclopedia (CCLE), including substantial unannotated alternative splicing in cancer transcriptome. The same untuned SPLASH2 algorithm recovers the BCR-ABL gene fusion, and detects circRNA sensitively and specifically, underscoring SPLASH2’s unmatched precision and scalability across diverse RNA-seq detection tasks.
RNA secondary and tertiary structure is critically involved in ribozyme and ribosomal rRNA function, as well as viral and cellular regulation. Traditional experimental methods for RNA structure determination such as X-ray crystallography or chemical mapping are incisive; however, these approaches suffer from low-throughput and low-dimensionality, respectively. Computational approaches, leveraging evolutionary signals from correlated positions' mutations, provide an alternative means to infer RNA structures. However, these methods require assembly, and face challenges due to statistical biases inherent in multiple sequence alignment (MSA). Furthermore, these methods cannot make use of the full spectrum of natural variations seen for a given RNA element. Here, we introduce SPLASH-structure, a direct assembly-free, MSA-free, and metadata-free statistical method for identifying conserved RNA structures by analyzing raw sequencing data, quantifying compensatory mutations or stem variation exclusion in the putative RNA structures. We show SPLASH-structure rediscovers known HIV structural elements and identifies conserved rRNA structures in metatranscriptomics samples. Moreover, SPLASH-structure finds Culex narnavirus 1, Gordis virus, and Culex mosquito virus 4, as well as previously unannotated viral genomes in mosquito metatranscriptomics samples de novo, highlighting the method's potential for viral discovery. SPLASH-structure is an ultra-fast, easy to use, and robust tool that excels in high-throughput RNA structure prediction and hypothesis generation, presenting a novel approach for discovering structural RNA elements. ### Competing Interest Statement The authors have declared no competing interest.
Early stages of deadly respiratory diseases including COVID-19 are challenging to elucidate in humans. Here, we define cellular tropism and transcriptomic effects of SARS-CoV-2 virus by productively infecting healthy human lung tissue and using scRNA-seq to reconstruct the transcriptional program in "infection pseudotime" for individual lung cell types. SARS-CoV-2 predominantly infected activated interstitial macrophages (IMs), which can accumulate thousands of viral RNA molecules, taking over 60% of the cell transcriptome and forming dense viral RNA bodies while inducing host profibrotic (TGFB1, SPP1) and inflammatory (early interferon response, CCL2/7/8/13, CXCL10, and IL6/10) programs and destroying host cell architecture. Infected alveolar macrophages (AMs) showed none of these extreme responses. Spike-dependent viral entry into AMs used ACE2 and Sialoadhesin/CD169, whereas IM entry used DC-SIGN/CD209. These results identify activated IMs as a prominent site of viral takeover, the focus of inflammation and fibrosis, and suggest targeting CD209 to prevent early pathology in COVID-19 pneumonia. This approach can be generalized to any human lung infection and to evaluate therapeutics.
Typical high-throughput single-cell RNA-sequencing (scRNA-seq) analyses are primarily conducted by (pseudo)alignment, through the lens of annotated gene models, and aimed at detecting differential gene expression. This misses diversity generated by other mechanisms that diversify the transcriptome such as splicing and V(D)J recombination, and is blind to sequences missing from imperfect reference genomes. Here, we present sc-SPLASH, a highly efficient pipeline that extends our SPLASH framework for statistics-first, reference-free discovery to barcoded scRNA-seq (10x Chromium) and spatial transcriptomics (10x Visium); we also provide its optimized module for preprocessing and k-mer counting in barcoded data, BKC, as a standalone tool. sc-SPLASH rediscovers known biology including V(D)J recombination and cell-type-specific alternative splicing in human and trans-splicing in tunicate (Ciona) and when applied to spatial datasets, detects sequence variation including tumor-specific somatic mutation. In sponge (Spongilla) and tunicate (Ciona), we uncover secreted repeat proteins expressed in immune-type cells and regulated during development; the sponge genes were absent from the reference assembly. sc-SPLASH provides a powerful alternative tool for exploring transcriptomes that is applicable to the breadth of life's diversity.
Targeted low-throughput studies have previously identified subcellular RNA localization as necessary for cellular functions including polarization, and translocation. Furthermore, these studies link localization to RNA isoform expression, especially 3’ Untranslated Region (UTR) regulation. The recent introduction of genome-wide spatial transcriptomics techniques enables the potential to test if subcellular localization is regulated in situ pervasively. In order to do this, robust statistical measures of subcellular localization and alternative poly-adenylation (APA) at single-cell resolution are needed. Developing a new statistical framework called SPRAWL, we detect extensive cell-type specific subcellular RNA localization regulation in the mouse brain and to a lesser extent mouse liver. We integrated SPRAWL with a new approach to measure cell-type specific regulation of alternative 3’ UTR processing and detected examples of significant correlations between 3’ UTR length and subcellular localization. Included examples, Timp3 , Slc32a1 , Cxcl14 , and Nxph1 have subcellular localization in the mouse brain highly correlated with regulated 3’ UTR processing that includes the use of unannotated, but highly conserved, 3’ ends. Together, SPRAWL provides a statistical framework to integrate multi-omic single-cell resolved measurements of gene-isoform pairs to prioritize an otherwise impossibly large list of candidate functional 3’ UTRs for functional prediction and study. In these studies of data from mice, SPRAWL predicts that 3’ UTR regulation of subcellular localization may be more pervasive than currently known.
We introduce SPLASH2, a fast, scalable implementation of SPLASH based on an efficient k-mer counting approach for regulated sequence variation detection in massive datasets from a wide range of sequencing technologies and biological contexts. We demonstrate biological discovery by SPLASH2 in single-cell RNA sequencing (RNA-seq) data and in bulk RNA-seq data from the Cancer Cell Line Encyclopedia, including unannotated alternative splicing in cancer transcriptomes and sensitive detection of circular RNA. SPLASH2 speeds up analysis of sequence variation in massive datasets.
Contingency tables, data represented as counts matrices, are ubiquitous across quantitative research and data-science applications. Existing statistical tests are insufficient however, as none are simultaneously computationally efficient and statistically valid for a finite number of observations. In this work, motivated by a recent application in reference-free genomic inference [K. Chaunget al.,Cell186, 5440–5456 (2023)], we develop Optimized Adaptive Statistic for Inferring Structure (OASIS), a family of statistical tests for contingency tables. OASIS constructs a test statistic which is linear in the normalized data matrix, providing closed-formP-value bounds through classical concentration inequalities. In the process, OASIS provides a decomposition of the table, lending interpretability to its rejection of the null. We derive the asymptotic distribution of the OASIS test statistic, showing that these finite-sample bounds correctly characterize the test statistic’sP-value up to a variance term. Experiments on genomic sequencing data highlight the power and interpretability of OASIS. Using OASIS, we develop a method that can detect SARS-CoV-2 andMycobacterium tuberculosisstrains de novo, which existing approaches cannot achieve. We demonstrate in simulations that OASIS is robust to overdispersion, a common feature in genomic data like single-cell RNA sequencing, where under accepted noise models OASIS provides good control of the false discovery rate, while Pearson’sX2consistently rejects the null. Additionally, we show in simulations that OASIS is more powerful than Pearson’sX2in certain regimes, including for some important two group alternatives, which we corroborate with approximate power calculations.
Today's genomics workflows typically require alignment to a reference sequence, which limits discovery. We introduce a unifying paradigm, SPLASH (Statistically Primary aLignment Agnostic Sequence Homing), which directly analyzes raw sequencing data, using a statistical test to detect a signature of regulation: sample-specific sequence variation. SPLASH detects many types of variation and can be efficiently run at scale. We show that SPLASH identifies complex mutation patterns in SARS-CoV-2, discovers regulated RNA isoforms at the single-cell level, detects the vast sequence diversity of adaptive immune receptors, and uncovers biology in non-model organisms undocumented in their reference genomes: geographic and seasonal variation and diatom association in eelgrass, an oceanic plant impacted by climate change, and tissue-specific transcripts in octopus. SPLASH is a unifying approach to genomic analysis that enables expansive discovery without metadata or references.
The authors have withdrawn this manuscript due to a duplicate posting of manuscript number BIORXIV/2022/497555. Therefore, the authors do not wish this work to be cited as reference for the project. If you have any questions, please contact the corresponding author. The correct preprint can be found at doi:https://doi.org/10.1101/2022.06.24.497555