The three-dimensional (3D) organization of cis-regulatory elements (CREs) is critical in transcription control. However, capturing transcriptome, epigenome and 3D genome from the same single cells remains challenging. Here we present scHiCAR (single-cell Hi-C with assay for transposase-accessible chromatin and RNA sequencing), a plate-based combinatorial barcoding method that simultaneously profiles mRNA, open chromatin and chromosome conformation capture from the same cells. Compared to existing single-cell 3D genome methods, scHiCAR more efficiently enriches long-range cis-interactions anchored at candidate CREs (cCREs). Applied to 1.62 million mouse brain cells and complemented with a deep-learning-based loop caller, scHiCAR accurately defines cell-type-specific transcriptomes, accessible cCREs and 5-kb-resolution enhancer-promoter pairs across 22 brain cell types. scHiCAR also performs robustly in challenging tissues such as skeletal muscle, enabling trimodal single-cell-level analysis of gene regulation dynamics during muscle stem cell regeneration. By providing a scalable and cost-effective system for single-cell trimodal analysis of gene-regulatory landscapes in complex tissues, scHiCAR reveals gene-locus-specific regulatory roles of 3D genome reorganization in transcriptional control.
Divergence of cis- regulatory elements drives species-specific traits 1 , but how this manifests in the evolution of the neocortex at the molecular and cellular level remains unclear. Here we investigated the gene regulatory programs in the primary motor cortex of human, macaque, marmoset and mouse using single-cell multiomics assays, generating gene expression, chromatin accessibility, DNA methylome and chromosomal conformation profiles from a total of over 200,000 cells. From these data, we show evidence that divergence of transcription factor expression corresponds to species-specific epigenome landscapes. We find that conserved and divergent gene regulatory features are reflected in the evolution of the three-dimensional genome. Transposable elements contribute to nearly 80% of the human-specific candidate cis- regulatory elements in cortical cells. Through machine learning, we develop sequence-based predictors of candidate cis- regulatory elements in different species and demonstrate that the genomic regulatory syntax is highly preserved from rodents to primates. Finally, we show that epigenetic conservation combined with sequence similarity helps to uncover functional cis- regulatory elements and enhances our ability to interpret genetic variants contributing to neurological disease and traits.
Single-cell sequencing could help to solve the fundamental challenge of linking millions of cell-type-specific enhancers with their target genes. However, this task is confounded by patterns of gene co-expression in much the same way that genetic correlation due to linkage disequilibrium confounds fine-mapping in genome-wide association studies (GWAS). We developed a non-parametric permutation-based procedure to establish stringent statistical criteria to control the risk of false-positive associations in enhancer-gene association studies (EGAS). We applied our procedure to large-scale transcriptome and epigenome data from multiple tissues and species, including the mouse and human brain, to predict enhancer-gene associations genome wide. We tested the functional validity of our predictions by comparing them with chromatin conformation data and causal enhancer perturbation experiments. Our study shows how controlling for gene co-expression enables robust enhancer-gene linkage using single-cell sequencing data.
Sequence divergence of cis- regulatory elements drives species-specific traits, but how this manifests in the evolution of the neocortex at the molecular and cellular level remains to be elucidated. We investigated the gene regulatory programs in the primary motor cortex of human, macaque, marmoset, and mouse with single-cell multiomics assays, generating gene expression, chromatin accessibility, DNA methylome, and chromosomal conformation profiles from a total of over 180,000 cells. For each modality, we determined species-specific, divergent, and conserved gene expression and epigenetic features at multiple levels. We find that cell type-specific gene expression evolves more rapidly than broadly expressed genes and that epigenetic status at distal candidate cis -regulatory elements (cCREs) evolves faster than promoters. Strikingly, transposable elements (TEs) contribute to nearly 80% of the human-specific cCREs in cortical cells. Through machine learning, we develop sequence-based predictors of cCREs in different species and demonstrate that the genomic regulatory syntax is highly preserved from rodents to primates. Lastly, we show that epigenetic conservation combined with sequence similarity helps uncover functional cis -regulatory elements and enhances our ability to interpret genetic variants contributing to neurological disease and traits.
Delineating the gene-regulatory programs underlying complex cell types is fundamental for understanding brain function in health and disease. Here, we comprehensively examined human brain cell epigenomes by probing DNA methylation and chromatin conformation at single-cell resolution in 517 thousand cells (399 thousand neurons and 118 thousand non-neurons) from 46 regions of three adult male brains. We identified 188 cell types and characterized their molecular signatures. Integrative analyses revealed concordant changes in DNA methylation, chromatin accessibility, chromatin organization, and gene expression across cell types, cortical areas, and basal ganglia structures. We further developed single-cell methylation barcodes that reliably predict brain cell types using the methylation status of select genomic sites. This multimodal epigenomic brain cell atlas provides new insights into the complexity of cell-type–specific gene regulation in adult human brains.
Recent advances in single-cell technologies have led to the discovery of thousands of brain cell types; however, our understanding of the gene regulatory programs in these cell types is far from complete1-4. Here we report a comprehensive atlas of candidate cis-regulatory DNA elements (cCREs) in the adult mouse brain, generated by analysing chromatin accessibility in 2.3 million individual brain cells from 117 anatomical dissections. The atlas includes approximately 1 million cCREs and their chromatin accessibility across 1,482 distinct brain cell populations, adding over 446,000 cCREs to the most recent such annotation in the mouse genome. The mouse brain cCREs are moderately conserved in the human brain. The mouse-specific cCREs-specifically, those identified from a subset of cortical excitatory neurons-are strongly enriched for transposable elements, suggesting a potential role for transposable elements in the emergence of new regulatory programs and neuronal diversity. Finally, we infer the gene regulatory networks in over 260 subclasses of mouse brain cells and develop deep-learning models to predict the activities of gene regulatory elements in different brain cell types from the DNA sequence alone. Our results provide a resource for the analysis of cell-type-specific gene regulation programs in both mouse and human brains. An atlas of candidate cis-regulatory DNA elements (cCREs) in the adult mouse brain unravels the transcriptional regulatory programs that drive the heterogeneity and complexity of brain structure and function.
Genome-wide association studies (GWAS) have identified thousands of non-coding variants that contribute to psychiatric disease risks, likely by perturbing cis -regulatory elements (CREs). However, our ability to interpret and explore their mechanisms of action is hampered by a lack of annotation of functional CREs (fCREs) in neural cell types. Here, through genome-scale CRISPR screens of 22,000 candidate CREs (cCREs) in human induced pluripotent stem cells (iPSCs) undergoing differentiation to excitatory neurons, we identify 2,847 and 5,540 fCREs essential for iPSC fitness and neuronal differentiation, respectively. These fCREs display dynamic epigenomic features and exhibit increased numbers and genomic spans of chromatin interactions following terminal neuronal differentiation. Furthermore, fCREs essential for neuronal differentiation show significantly greater enrichment of genetic heritability for neurodevelopmental diseases including schizophrenia (SCZ), attention deficit hyperactivity disorder (ADHD), and autism spectrum disorders (ASD) than cCREs. Using high-throughput prime editing screens we experimentally confirm 45 SCZ risk variants that act by affecting the function of fCREs. The extensive and in-depth functional annotation of cCREs in neuronal types therefore provides a crucial resource for interpreting non-coding risk variants of neuropsychiatric disorders.
Recent advances in single-cell transcriptomics have illuminated the diverse neuronal and glial cell types within the human brain. However, the regulatory programs governing cell identity and function remain unclear. Using a single-nucleus assay for transposase-accessible chromatin using sequencing (snATAC-seq), we explored open chromatin landscapes across 1.1 million cells in 42 brain regions from three adults. Integrating this data unveiled 107 distinct cell types and their specific utilization of 544,735 candidate cis-regulatory DNA elements (cCREs) in the human genome. Nearly a third of the cCREs demonstrated conservation and chromatin accessibility in the mouse brain cells. We reveal strong links between specific brain cell types and neuropsychiatric disorders including schizophrenia, bipolar disorder, Alzheimer’s disease (AD), and major depression, and have developed deep learning models to predict the regulatory roles of noncoding risk variants in these disorders.
We previously reported Paired-Tag, a combinatorial indexing-based method that can simultaneously map histone modifications and gene expression at single-cell resolution at scale. However, the lengthy procedure of Paired-Tag has hindered its general adoption in the community. To address this bottleneck, we developed a droplet-based Paired-Tag protocol that is faster and more accessible than the previous method. Using cultured mammalian cells and primary brain tissues, we demonstrate its superior performance at identifying candidate cis -regulatory elements and associating their dynamic chromatin state to target gene expression in each constituent cell type in a complex tissue.
In 2021, the World Health Organization reclassified glioblastoma, the most common form of adult brain cancer, into isocitrate dehydrogenase (IDH)-wild-type glioblastomas and grade IV IDH mutant (G4 IDHm) astrocytomas. For both tumor types, intratumoral heterogeneity is a key contributor to therapeutic failure. To better define this heterogeneity, genome-wide chromatin accessibility and transcription profiles of clinical samples of glioblastomas and G4 IDHm astrocytomas were analyzed at single-cell resolution. These profiles afforded resolution of intratumoral genetic heterogeneity , including delineation of cell-to-cell variations in distinct cell states, focal gene amplifications, as well as extrachromosomal circular DNAs. Despite differences in IDH mutation status and significant intratumoral heterogeneity, the profiled tumor cells shared a common chromatin structure defined by open regions enriched for nuclear factor 1 transcription factors (NFIA and NFIB). Silencing of NFIA or NFIB suppressed in vitro and in vivo growths of patient-derived glioblastomas and G4 IDHm astrocytoma models. These findings suggest that despite distinct genotypes and cell states, glioblastoma/ G4 astrocytoma cells share dependency on core transcriptional programs, yielding an attractive platform for addressing therapeutic challenges associated with intratumoral heterogeneity.
Cytosine DNA methylation is essential in brain development and has been implicated in various neurological disorders. A comprehensive understanding of DNA methylation diversity across the entire brain in the context of the brain's 3D spatial organization is essential for building a complete molecular atlas of brain cell types and understanding their gene regulatory landscapes. To this end, we employed optimized single-nucleus methylome (snmC-seq3) and multi-omic (snm3C-seq1) sequencing technologies to generate 301,626 methylomes and 176,003 chromatin conformation/methylome joint profiles from 117 dissected regions throughout the adult mouse brain. Using iterative clustering and integrating with companion whole-brain transcriptome and chromatin accessibility datasets, we constructed a methylation-based cell type taxonomy that contains 4,673 cell groups and 261 cross-modality-annotated subclasses. We identified millions of differentially methylated regions (DMRs) across the genome, representing potential gene regulation elements. Notably, we observed spatial cytosine methylation patterns on both genes and regulatory elements in cell types within and across brain regions. Brain-wide multiplexed error-robust fluorescence in situ hybridization (MERFISH2) data validated the association of this spatial epigenetic diversity with transcription and allowed the mapping of the DNA methylation and topology information into anatomical structures more precisely than our dissections. Furthermore, multi-scale chromatin conformation diversities occur in important neuronal genes, highly associated with DNA methylation and transcription changes. Brain-wide cell type comparison allowed us to build a regulatory model for each gene, linking transcription factors, DMRs, chromatin contacts, and downstream genes to establish regulatory networks. Finally, intragenic DNA methylation and chromatin conformation patterns predicted alternative gene isoform expression observed in a companion whole-brain SMART-seq3 dataset. Our study establishes the first brain-wide, single-cell resolution DNA methylome and 3D multi-omic atlas, providing an unparalleled resource for comprehending the mouse brain's cellular-spatial and regulatory genome diversity.
The primary motor cortex (M1) is essential for voluntary fine-motor control and is functionally conserved across mammals1. Here, using high-throughput transcriptomic and epigenomic profiling of more than 450,000 single nuclei in humans, marmoset monkeys and mice, we demonstrate a broadly conserved cellular makeup of this region, with similarities that mirror evolutionary distance and are consistent between the transcriptome and epigenome. The core conserved molecular identities of neuronal and non-neuronal cell types allow us to generate a cross-species consensus classification of cell types, and to infer conserved properties of cell types across species. Despite the overall conservation, however, many species-dependent specializations are apparent, including differences in cell-type proportions, gene expression, DNA methylation and chromatin state. Few cell-type marker genes are conserved across species, revealing a short list of candidate genes and regulatory mechanisms that are responsible for conserved features of homologous cell types, such as the GABAergic chandelier cells. This consensus transcriptomic classification allows us to use patch-seq (a combination of whole-cell patch-clamp recordings, RNA sequencing and morphological characterization) to identify corticospinal Betz cells from layer 5 in non-human primates and humans, and to characterize their highly specialized physiology and anatomy. These findings highlight the robust molecular underpinnings of cell-type diversity in M1 across mammals, and point to the genes and regulatory pathways responsible for the functional identity of cell types and their species-specific adaptations.
Non-coding variation in complex human disease has been well established by genome-wide association studies, and is thought to involve regulatory elements, such as enhancers, whose variation affects the expression of the gene responsible for the disease. The regulatory elements often lie far from the gene they regulate, or within introns of genes differing from the regulated gene, making it difficult to identify the gene whose function is affected by a given enhancer variation. Enhancers are connected to their target gene promoters via long-range physical interactions (loops). In our study, we re-mapped, onto the human genome, more than 10,000 enhancers connected to promoters via long-range interactions, that we had previously identified in mouse brain-derived neural stem cells by RNApolII-ChIA-PET analysis, coupled to ChIP-seq mapping of DNA/chromatin regions carrying epigenetic enhancer marks. These interactions are thought to be functionally relevant. We discovered, in the human genome, thousands of DNA regions syntenic with the interacting mouse DNA regions (enhancers and connected promoters). We further annotated these human regions regarding their overlap with sequence variants (single nucleotide polymorphisms, SNPs; copy number variants, CNVs), that were previously associated with neurodevelopmental disease in humans. We document various cases in which the genetic variant, associated in humans to neurodevelopmental disease, affects an enhancer involved in long-range interactions: SNPs, previously identified by genome-wide association studies to be associated with schizophrenia, bipolar disorder, and intelligence, are located within our human syntenic enhancers, and alter transcription factor recognition sites. Similarly, CNVs associated to autism spectrum disease and other neurodevelopmental disorders overlap with our human syntenic enhancers. Some of these enhancers are connected (in mice) to homologs of genes already associated to the human disease, strengthening the hypothesis that the gene is indeed involved in the disease. Other enhancers are connected to genes not previously associated with the disease, pointing to their possible pathogenetic involvement. Our observations provide a resource for further exploration of neural disease, in parallel with the now widespread genome-wide identification of DNA variants in patients with neural disease.
Abstract INTRODUCTION In 2021, the World Health Organization (WHO) reclassified glioblastoma, the most common form of adult brain cancer, into isocitrate dehydrogenase (IDH) wild-type glioblastomas and grade IV IDH mutant (G4 IDHm) astrocytomas. For both tumor types, intra-tumoral heterogeneity is a key contributor to therapeutic failure. METHODS we applied integrated genome-wide chromatin accessibility (snATACseq) and transcription (snRNAseq) profiles to clinical specimens derived IDHwt glioblastomas and G4 IDHm) astrocytomas, with goal of therapeutic target discovery. RESULTS The integrated analysis achieved resolution of intra-tumoral heterogeneity not previously possible, providing a molecular landscape of extensive regional and cellular variability. snATACseq delineated focal amplification down to an ~40 KB resolution. The snRNA analysis elucidated distinct cell types and cell states (neural progenitor/oligodendrocyte cell-like or astrocyte/mesenchymal cell-like) that were superimposable onto the snATACseq landscape. Paired-seq (parallel snATACseq and snRNAseq using the same clinical sample) provided high resolution delineation of extrachromosomal circular DNA (ecDNA), harboring oncogenes including CCND1 and EGFR. Importantly, the copy number of ecDNA genes correlated closely with the level of RNA expression. Integrated analysis across all specimens profiled suggests that IDHm grade 4 astrocytoma and IDHwt glioblastoma cells shared a common chromatin structure defined by open regions enriched for Nuclear Factor 1 transcription factors (NFIA and NFIB). Silencing of NF1A or NF1B suppressed in vitro and in vivo growth of patient-derived IDHwt glioblastomas and G4 IDHm astrocytoma models that mimic distinct glioblastoma cell states. CONCLUSION Our findings suggest despite distinct genotypes and cell states, glablastoma/G4 astrocytoma cells share dependency on core transcriptional programs, yielding an attractive platform for addressing therapeutic challenges associated with intra-tumoral heterogeneity.
We describe here Paired-Tag, a high-throughput multi-omics method for joint profiling of histone modifications and gene expressions in single cells. The assay is based on a combinatorial barcoding indexing strategy that does not require special instruments. It can be performed with nuclei extracted from cultured cells or frozen tissues, in standard molecular biology laboratories.
Dysregulation of cardiac transcription programs has been identified in patients and families with heart failure, as well as those with morphological and functional forms of congenital heart defects. Mediator is a multi-subunit complex that plays a central role in transcription initiation by integrating regulatory signals from gene-specific transcriptional activators to RNA polymerase II (Pol II). Recently, Mediator subunit 30 (MED30), a metazoan specific Mediator subunit, has been associated with Langer-Giedion syndrome (LGS) Type II and Cornelia de Lange syndrome-4 (CDLS4), characterized by several abnormalities including congenital heart defects. A point mutation in MED30 has been identified in mouse and is associated with mitochondrial cardiomyopathy. Very recent structural analyses of Mediator revealed that MED30 localizes to the proximal Tail, anchoring Head and Tail modules, thus potentially influencing stability of the Mediator core. However, in vivo cellular and physiological roles of MED30 in maintaining Mediator core integrity remain to be tested. Here, we report that deletion of MED30 in embryonic or adult cardiomyocytes caused rapid development of cardiac defects and lethality. Importantly, cardiomyocyte specific ablation of MED30 destabilized Mediator core subunits, while the kinase module was preserved, demonstrating an essential role of MED30 in stability of the overall Mediator complex. RNAseq analyses of constitutive cardiomyocyte specific Med30 knockout (cKO) embryonic hearts and inducible cardiomyocyte specific Med30 knockout (icKO) adult cardiomyocytes further revealed critical transcription networks in cardiomyocytes controlled by Mediator. Taken together, our results demonstrated that MED30 is essential for Mediator stability and transcriptional networks in both developing and adult cardiomyocytes. Our results affirm the key role of proximal Tail modular subunits in maintaining core Mediator stability in vivo.
Genome-wide profiling of histone modifications can reveal not only the location and activity state of regulatory elements, but also the regulatory mechanisms involved in cell-type-specific gene expression during development and disease pathology. Conventional assays to profile histone modifications in bulk tissues lack single-cell resolution. Here we describe an ultra-high-throughput method, Paired-Tag, for joint profiling of histone modifications and transcriptome in single cells to produce cell-type-resolved maps of chromatin state and transcriptome in complex tissues. We used this method to profile five histone modifications jointly with transcriptome in the adult mouse frontal cortex and hippocampus. Integrative analysis of the resulting maps identified distinct groups of genes subject to divergent epigenetic regulatory mechanisms. Our single-cell multiomics approach enables comprehensive analysis of chromatin state and gene regulation in complex tissues and characterization of gene regulatory programs in the constituent cell types.
Single cell transcriptomics has transformed the characterization of brain cell identity by providing quantitative molecular signatures for large, unbiased samples of brain cell populations. With the proliferation of taxonomies based on individual datasets, a major challenge is to integrate and validate results toward defining biologically meaningful cell types. We used a battery of single-cell transcriptome and epigenome measurements generated by the BRAIN Initiative Cell Census Network (BICCN) to comprehensively assess the molecular signatures of cell types in the mouse primary motor cortex (MOp). We further developed computational and statistical methods to integrate these multimodal data and quantitatively validate the reproducibility of the cell types. The reference atlas, based on more than 600,000 high quality single-cell or -nucleus samples assayed by six molecular modalities, is a comprehensive molecular account of the diverse neuronal and non-neuronal cell types in MOp. Collectively, our study indicates that the mouse primary motor cortex contains over 55 neuronal cell types that are highly replicable across analysis methods, sequencing technologies, and modalities. We find many concordant multimodal markers for each cell type, as well as thousands of genes and gene regulatory elements with discrepant transcriptomic and epigenomic signatures. These data highlight the complex molecular regulation of brain cell types and will directly enable design of reagents to target specific MOp cell types for functional analysis.
Mammalian brain cells show remarkable diversity in gene expression, anatomy and function, yet the regulatory DNA landscape underlying this extensive heterogeneity is poorly understood. Here we carry out a comprehensive assessment of the epigenomes of mouse brain cell types by applying single-nucleus DNA methylation sequencing 1 , 2 to profile 103,982 nuclei (including 95,815 neurons and 8,167 non-neuronal cells) from 45 regions of the mouse cortex, hippocampus, striatum, pallidum and olfactory areas. We identified 161 cell clusters with distinct spatial locations and projection targets. We constructed taxonomies of these epigenetic types, annotated with signature genes, regulatory elements and transcription factors. These features indicate the potential regulatory landscape supporting the assignment of putative cell types and reveal repetitive usage of regulators in excitatory and inhibitory cells for determining subtypes. The DNA methylation landscape of excitatory neurons in the cortex and hippocampus varied continuously along spatial gradients. Using this deep dataset, we constructed an artificial neural network model that precisely predicts single neuron cell-type identity and brain area spatial location. Integration of high-resolution DNA methylomes with single-nucleus chromatin accessibility data 3 enabled prediction of high-confidence enhancer–gene interactions for all identified cell types, which were subsequently validated by cell-type-specific chromatin conformation capture experiments 4 . By combining multi-omic datasets (DNA methylation, chromatin contacts, and open chromatin) from single nuclei and annotating the regulatory genome of hundreds of cell types in the mouse brain, our DNA methylation atlas establishes the epigenetic basis for neuronal diversity and spatial organization throughout the mouse cerebrum.
The mammalian cerebrum performs high-level sensory perception, motor control and cognitive functions through highly specialized cortical and subcortical structures 1 . Recent surveys of mouse and human brains with single-cell transcriptomics 2 – 6 and high-throughput imaging technologies 7 , 8 have uncovered hundreds of neural cell types distributed in different brain regions, but the transcriptional regulatory programs that are responsible for the unique identity and function of each cell type remain unknown. Here we probe the accessible chromatin in more than 800,000 individual nuclei from 45 regions that span the adult mouse isocortex, olfactory bulb, hippocampus and cerebral nuclei, and use the resulting data to map the state of 491,818 candidate cis -regulatory DNA elements in 160 distinct cell types. We find high specificity of spatial distribution for not only excitatory neurons, but also most classes of inhibitory neurons and a subset of glial cell types. We characterize the gene regulatory sequences associated with the regional specificity within these cell types. We further link a considerable fraction of the cis -regulatory elements to putative target genes expressed in diverse cerebral cell types and predict transcriptional regulators that are involved in a broad spectrum of molecular and cellular pathways in different neuronal and glial cell populations. Our results provide a foundation for comprehensive analysis of gene regulatory programs of the mammalian brain and assist in the interpretation of noncoding risk variants associated with various neurological diseases and traits in humans.