Spatial transcriptomics has been widely applied to study the spatial distribution of cell types, cell states, and specific gene expression in tissue samples. However, we show that there is a prevalent transcript leakage problem in spatial transcriptomics data, where transcripts expressed by a cell diffuse to its neighborhood and are recurrently detected in the nearby cells. By analyzing published data sets, we show that this problem is general across data produced from different tissues and different species using different imaging-based and sequencing-based spatial transcriptomics platforms. It affects both upstream tasks such as expression quantification as well as downstream tasks such as cell-type annotation and detection of spatially-dependent gene expression. To tackle the transcript leakage problem, we propose a reference-free Bayesian model-based method, DeLeakage, which cleans up the data much more effectively than existing denoising methods. DeLeakage also improves cell-type annotation and avoids false detection of spatially dependent expression.
The sequence of the human genome provides a foundation for understanding cellular processes in health and disease1. The organisation of this primary genetic information into cell-specific structure and function is critical to understanding the cell type-specific interpretation and execution of the genome. Epigenetic processes are essential for packaging and higher-level functional organisation of the genome, and changes therein are increasingly recognised as contributors to human disease. Building on primary data generated by multinational consortia, the International Human Epigenome Consortium2 (IHEC) has uniformly processed a collection of more than 2000 comprehensive human reference epigenomes, collectively referred to as EpiATLAS. This effort involved the development of standardised molecular and bioinformatics protocols, metadata models, and analytical tools to manage, integrate, display, and share vast amounts of epigenomic data. This includes the creation of a publicly available Epigenome Reference Registry, which provides a system for accessing protected human subject datasets and facilitates open searching of de-identified samples and experimental data. The integrated EpiATLAS ecosystem and its comprehensive human reference epigenome maps provide an unprecedented resource for the biosciences, expanding the annotated epigenomic landscape while uncovering previously unappreciated relationships among regulatory layers and revealing how epigenetic inputs underpin fundamental cellular functions and disease associations.
A role for the trafficking receptor SORLA (Sortilin-related receptor containing LDLR class A repeats) in reducing Aβ levels has been well established; however, relatively little is known with respect to whether and how SORLA can potentially affect tau pathology in vivo. Here, we show that transgenic SORLA up-regulation (SORLA TG) can attenuate pathological effects in aged PS19 (P301S tau) mouse brain, including tau phosphorylation and seeding, ventricle dilation, synapse loss, long-term potentiation (LTP) impairment, and glial hyperactivation. Proteomic analysis indicates attenuation of PS19 profiles in PS19/SORLA TG hippocampus, including pathological changes in synapse-related proteins and key drivers of synaptic dysfunction such as ApoE and C1q. Single-nucleus RNA sequencing analysis reveals suppression of PS19 signatures with SORLA up-regulation and identifies a previously unrecognized involvement of Sema4D-PlexinB1/B2 signaling in glial pathology. In contrast, SORLA deletion exacerbates tau seeding and aggregation, glial hyperactivation, and PlxnB1/B2 induction in PS19 hippocampus. These results indicate that SORLA confers neuroprotection against tau toxicity in the PS19 mouse brain.
Alzheimer’s disease (AD) is a complex neurodegenerative disorder that is associated with cognitive decline in the elderly. While β-amyloid (Aβ) plaques and neurofibrillary tau tangles have been used to define and stage AD onset in human brain, how these pathologies can affect various cell types nearby remains a subject of intense interest in the field. Recent developments in spatial transcriptomic technology have seen accelerated growth, and spatial transcriptomic platforms have been used independently or together with single cell transcriptomic methods to characterize cellular changes in the AD brain from mouse models and human. Here, we review current-era spatial transcriptomic technologies, analytical pipelines and their implementation in AD research. We summarize findings from spatial transcriptomics in AD, and discuss limitations and challenges associated with various spatial transcriptomic platforms. The pathological hallmarks for AD were described by Alois Alzheimer over a century ago; from the convergence of a century of AD research and technological advances in imaging and transcriptomics, a new era in AD has emerged. Although current-era spatial platforms feature limitations and challenges, evolution of spatial transcriptomics and its combined implementation with other data modalities promises significant strides in AD and related neurodegenerative disorders.
A holy grail in computational biology is accurate modeling of transcript expression levels using epigenetic features, which would provide a quantitative way to study gene regulation in normal and disease states. Previous studies relied heavily on immortalized cell lines that exhibit properties different from cells in natural tissue environments. Most studies also quantified the expression of each gene by a single expression level, which fails to capture separate expression levels of different transcript isoforms of the same gene. In this study, making use of the latest large-scale dataset of paired transcriptomic and epigenomic data of human samples produced by the International Human Epigenome Consortium (IHEC), we computationally modeled the expression levels of individual transcript isoforms in 324 samples from 29 tissue types. We constructed the models using graph-based methods that integrate both location-specific epigenomic features and multiple types of gene-gene relationships. We found that to infer transcript isoform expression levels in a sample, a model that integrates information from many samples of other tissue types consistently outperforms a model trained on data from this sample itself, providing strong support that it is possible to construct a "universal" model that can accurately infer transcript isoform expression levels across tissue types.
A role for the trafficking receptor SORLA in reducing Aβ levels has been well-established, however, relatively little is known with respect to whether and how SORLA can potentially affect tau pathology in vivo. Here, we show that transgenic SORLA upregulation (SORLA TG) can reverse pathological effects in aged PS19 (P301S tau) mouse brain, including tau phosphorylation and seeding, ventricle dilation, synapse loss, LTP impairment and glial hyperactivation. Proteomic analysis indicates reversion of PS19 profiles in PS19/SORLA TG hippocampus, including pathological changes in synapse-related proteins as well as key drivers of synaptic dysfunction such as Apoe and C1q. snRNA-seq analysis reveals suppression of PS19- signatures with SORLA upregulation, including proinflammatory induction of Plxnb1/Plxnb2 in glia. Tau seeding and aggregation, neuroinflammation, as well as PlxnB1/B2 induction are exacerbated in PS19 hippocampus with SORLA deletion. These results implicate a global role for SORLA in neuroprotection from tau toxicity in PS19 mouse brain.
Evidence indicates that the integrity of in utero development influences late life healthy or unhealthy aging; however, specific links between them are unclear. Histone chaperone HIRA is thought to play a role in both life stages, and here, we explore this role using the murine pigmentary system by investigating and comparing the effects of its lineage-specific knockout, either conditionally during embryogenesis or postnatally. Embryonic knockout of Hira in tyrosinase+ neural crest-derived lineages, including melanoblasts, led to reduced melanoblast numbers during embryogenesis, with single-cell RNA sequencing analysis indicating evidence of lineage-specificity defects. This was supported in an in vitro model using melb-a melanoblasts in which Hira knockdown affected lineage identity and melanoblast differentiation potential, with ATAC-seq data indicating a role of HIRA in orchestrating chromatin accessibility. Interestingly, however, newborn Hira knockout mice had wild type numbers of differentiated melanocytes, albeit functionally defective, as demonstrated by very mild hypopigmentation of the first hair coat, increased melanocyte telomere-associated DNA damage foci, and impaired response to proliferative challenge. Moreover, as they aged, mice with embryonic melanoblast Hira knockout displayed marked defects in melanocyte stem cell maintenance and premature hair graying. Importantly, this phenotype was not observed after postnatal inducible knockout, indicating an essential role for HIRA at embryonic stages that is transmitted to adulthood, rather than a direct postnatal requirement within the pigmentary system. This genetic model shows that HIRA function during early development lays a foundation for maintaining lineage identity and subsequent maintenance of adult tissue-specific stem cells during aging.
Cellular senescence, a stress-induced stable proliferation arrest associated with an inflammatory senescence-associated secretory phenotype (SASP), is a cause of aging. In senescent cells, cytoplasmic chromatin fragments (CCFs) activate SASP via the anti-viral cGAS/STING pathway. Promyelocytic leukemia (PML) protein organizes PML nuclear bodies (NBs), which are also involved in senescence and anti-viral immunity. The HIRA histone H3.3 chaperone localizes to PML NBs in senescent cells. Here, we show that HIRA and PML are essential for SASP expression, tightly linked to HIRA’s localization to PML NBs. Inactivation of HIRA does not directly block expression of nuclear factor κB (NF-κB) target genes. Instead, an H3.3-independent HIRA function activates SASP through a CCF-cGAS-STING-TBK1-NF-κB pathway. HIRA physically interacts with p62/SQSTM1, an autophagy regulator and negative SASP regulator. HIRA and p62 co-localize in PML NBs, linked to their antagonistic regulation of SASP, with PML NBs controlling their spatial configuration. These results outline a role for HIRA and PML in the regulation of SASP.
In this report, we present OLAF-Seq, a novel strategy to construct a long-read sequencing library such that adjacent fragments are linked with end-terminal duplications. We use the CRISPR-Cas9 nickase enzyme and a pool of multiple sgRNAs to perform non-random fragmentation of targeted long DNA molecules (> 300kb) into smaller library-sized fragments (about 20 kbp) in a manner so as to retain physical linkage information (up to 1000 bp) between adjacent fragments. DNA molecules targeted for fragmentation are preferentially ligated with adaptors for sequencing, so this method can enrich targeted regions while taking advantage of the long-read sequencing platforms. This enables the sequencing of target regions with significantly lower total coverage, and the genome sequence within linker regions provides information for assembly and phasing. We demonstrated the validity and efficacy of the method first using phage and then by sequencing a panel of 100 full-length cancer-related genes (including both exons and introns) in the human genome. When the designed linkers contained heterozygous genetic variants, long haplotypes could be established. This sequencing strategy can be readily applied in both PacBio and Oxford Nanopore platforms for both long and short genes with an easy protocol. This economically viable approach is useful for targeted enrichment of hundreds of target genomic regions and where long no-gap contigs need deep sequencing.
ABSTRACT In this report, we present linked-pair sequencing, a novel strategy to construct a long-read sequencing library such that adjacent fragments are linked with end-terminal duplications. We use the CRISPR-Cas9 nickase enzyme and a pool of multiple sgRNAs to perform non-random fragmentation of targeted long DNA molecules (>300kb) into smaller library-sized fragments (about 20 kbp) in a manner so as to retain physical linkage information (up to 1000 bp) between adjacent fragments. DNA molecules targeted for fragmentation are preferentially ligated with adaptors for sequencing, so this method can enrich targeted regions while taking advantage of the long-read sequencing platforms. This enables the sequencing of target regions with significantly lower total coverage, and the genome sequence within linker regions provides information for assembly and phasing. We demonstrated the validity and efficacy of the method first using phage and then by sequencing a panel of 100 full-length cancer-related genes (including both exons and introns) in the human genome. When the designed linkers contained heterozygous genetic variants, long haplotypes could be established. This sequencing strategy can be readily applied in both PacBio and Oxford Nanopore platforms. This economically viable approach is useful for targeted enrichment of hundreds of target genomic regions and where long no-gap contigs need deep sequencing.
Circular RNAs (circRNAs) are abundantly expressed in cancer. Their resistance to exonucleases enables them to have potentially stable interactions with different types of biomolecules. Alternative splicing can create different circRNA isoforms that have different sequences and unequal interaction potentials. The study of circRNA function thus requires knowledge of complete circRNA sequences. Here we describe psirc, a method that can identify full-length circRNA isoforms and quantify their expression levels from RNA sequencing data. We confirm the effectiveness and computational efficiency of psirc using both simulated and actual experimental data. Applying psirc on transcriptome profiles from nasopharyngeal carcinoma and normal nasopharynx samples, we discover and validate circRNA isoforms differentially expressed between the two groups. Compared with the assumed circular isoforms derived from linear transcript annotations, some of the alternatively spliced circular isoforms have 100 times higher expression and contain substantially fewer microRNA response elements, showing the importance of quantifying full-length circRNA isoforms.
MOTIVATION:In de novo sequence assembly, a standard pre-processing step is k-mer counting, which computes the number of occurrences of every length-k sub-sequence in the sequencing reads. Sequencing errors can produce many k-mers that do not appear in the genome, leading to the need for an excessive amount of memory during counting. This issue is particularly serious when the genome to be assembled is large, the sequencing depth is high, or when the memory available is limited.RESULTS:Here, we propose a fast near-exact k-mer counting method, CQF-deNoise, which has a module for dynamically removing noisy false k-mers. It automatically determines the suitable time and number of rounds of noise removal according to a user-specified wrong removal rate. We tested CQF-deNoise comprehensively using data generated from a diverse set of genomes with various data properties, and found that the memory consumed was almost constant regardless of the sequencing errors while the noise removal procedure had minimal effects on counting accuracy. Compared with four state-of-the-art k-mer counting methods, CQF-deNoise consistently performed the best in terms of memory usage, consuming 49-76% less memory than the second best method. When counting the k-mers from a human dataset with around 60× coverage, the peak memory usage of CQF-deNoise was only 10.9 GB (gigabytes) for k = 28 and 21.5 GB for k = 55. De novo assembly of 106× human sequencing data using CQF-deNoise for k-mer counting required only 2.7 h and 90 GB peak memory.AVAILABILITY AND IMPLEMENTATION:The source codes of CQF-deNoise and SH-assembly are available at https://github.com/Christina-hshi/CQF-deNoise.git and https://github.com/Christina-hshi/SH-assembly.git, respectively, both under the BSD 3-Clause license.
K-mer counting has many applications in sequencing data processing and analysis. However, sequencing errors can produce many false k-mers that substantially increase the memory requirement during counting. We propose a fast k-mer counting method, CQF-deNoise, which has a novel component for dynamically identifying and removing false k-mers while preserving counting accuracy. Compared with four state-of-the-art k-mer counting methods, CQF-deNoise consumed 49-76% less memory than the second best method, but still ran competitively fast. The k-mer counts from CQF-deNoise produced cell clusters from single-cell RNA-seq data highly consistent with CellRanger but required only 5% of the running time at the same memory consumption, suggesting that CQF-deNoise can be used for a preview of cell clusters for an early detection of potential data problems, before running a much more time-consuming full analysis pipeline.
Motivation: The three-dimensional structure of genomes makes it possible for genomic regions not adjacent in the primary sequence to be spatially proximal. These DNA contacts have been found to be related to various molecular activities. Previous methods for analyzing DNA contact maps obtained from Hi-C experiments have largely focused on studying individual interactions, forming spatial clusters composed of contiguous blocks of genomic locations, or classifying these clusters into general categories based on some global properties of the contact maps. Results: Here, we describe a novel computational method that can flexibly identify small clusters of spatially proximal genomic regions based on their local contact patterns. Using simulated data that highly resemble Hi-C data obtained from real genome structures, we demonstrate that our method identifies spatial clusters that are more compact than methods previously used for clustering genomic regions based on DNA contact maps. The clusters identified by our method enable us to confirm functionally related genomic regions previously reported to be spatially proximal in different species. We further show that each genomic region can be assigned a numeric affinity value that indicates its degree of participation in each local cluster, and these affinity values correlate quantitatively with DNase I hypersensitivity, gene expression, super enhancer activities and replication timing in a cell type specific manner. We also show that these cluster affinity values can precisely define boundaries of reported topologically associating domains, and further define local sub-domains within each domain. Availability and implementation: The source code of BNMF and tutorials on how to use the software to extract local clusters from contact maps are available at http://yiplab.cse.cuhk.edu.hk/bnmf/. Contact: kevinyip@cse.cuhk.edu.hk Supplementary information: Supplementary data are available at Bioinformatics online.