Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression1. Previously, the ENCODE consortium mapped biochemical signals across hundreds of cell types and tissues and integrated these data to develop a registry containing 0.9 million human and 300,000 mouse candidate cis-regulatory elements (cCREs) annotated with potential functions2. Here we have expanded the registry to include 2.37 million human and 967,000 mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays such as STARR-seq3, massively parallel reporter assay4, CRISPR perturbation5,6 and transgenic mouse assays7 have profiled more than 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer and silencer roles in different cellular contexts. Integrating the registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by the identification of KLF1 as a novel causal gene for red blood cell traits. This expanded registry is a valuable resource for studying the regulatory genome and its impact on health and disease.
Heterozygous truncating variants in the sarcomere protein titin (TTN) are the most common genetic cause of heart failure. To understand mechanisms that regulate abundant cardiomyocyte (CM) TTN expression, we characterized highly conserved intron 1 sequences that exhibited dynamic changes in chromatin accessibility during differentiation of human CMs from induced pluripotent stem cells (hiPSC-CMs). Homozygous deletion of these sequences in mice caused embryonic lethality, whereas heterozygous mice showed an allele-specific reduction in Ttn expression. A 296 bp fragment of this element, denoted E1, was sufficient to drive expression of a reporter gene in hiPSC-CMs. Deletion of E1 downregulated TTN expression, impaired sarcomerogenesis, and decreased contractility in hiPSC-CMs. Site-directed mutagenesis of predicted binding sites of NK2 homeobox 5 (NKX2-5) and myocyte enhancer factor 2 (MEF2) within E1 abolished its transcriptional activity. In embryonic mice expressing E1 reporter gene constructs, we validated in vivo cardiac-specific activity of E1 and the requirement for NKX2-5- and MEF2-binding sequences. Moreover, isogenic hiPSC-CMs containing a rare E1 variant in the predicted MEF2-binding motif that was identified in a patient with unexplained dilated cardiomyopathy (DCM) showed reduced TTN expression. Together, these discoveries define an essential, functional enhancer that regulates TTN expression. Manipulation of this element may advance therapeutic strategies to treat DCM caused by TTN haploinsufficiency.
Transcription factors (TFs) bind combinatorially to cis-regulatory elements, orchestrating transcriptional programs. Although studies of chromatin state and chromosomal interactions have demonstrated dynamic neurodevelopmental cis-regulatory landscapes, parallel understanding of TF interactions lags. To elucidate combinatorial TF binding driving mouse basal ganglia development, we integrated chromatin immunoprecipitation sequencing (ChIP-seq) for twelve TFs, H3K4me3-associated enhancer-promoter interactions, chromatin and gene expression data, and functional enhancer assays. We identified sets of putative regulatory elements with shared TF binding (TF-pRE modules) that orchestrate distinct processes of GABAergic neurogenesis and suppress other cell fates. The majority of pREs were bound by one or two TFs; however, a small proportion were extensively bound. These sequences had exceptional evolutionary conservation and motif density, complex chromosomal interactions, and activity as in vivo enhancers. Our results provide insights into the combinatorial TF-pRE interactions that activate and repress expression programs during telencephalon neurogenesis and demonstrate the value of TF binding toward modeling developmental transcriptional wiring.
Although most mammalian transcriptional enhancers regulate their cognate promoters over distances of tens of kilobases, some enhancers act over distances in the megabase range1. The sequence features that enable such long-distance enhancer-promoter interactions remain unclear. Here we used in vivo enhancer-replacement experiments at the mouse Shh locus to show that short- and medium-range limb enhancers cannot initiate gene expression at long-distance range. We identify a cis-acting element, range extender (REX), that confers long-distance regulatory activity and is located next to a long-range limb enhancer of Sall1. The REX element has no endogenous enhancer activity. However, addition of the REX to other short- and mid-range limb enhancers substantially increases their genomic interaction range. In the most extreme example observed, addition of REX increased the range of an enhancer by an order of magnitude from its native 73 kb to 848 kb. The REX element contains highly conserved [C/T]AATTA homeodomain motifs that are critical for its activity. These motifs are enriched in long-range limb enhancers genome-wide, including the ZRS (zone of polarizing activity (ZPA) regulatory sequence), a benchmark long-range limb enhancer of Shh2. The ZRS enhancer with mutated [C/T]AATTA motifs maintains limb activity at short range, but loses its long-range activity, resulting in severe limb reduction in knock-in mice. In summary, we identify a sequence signature associated with long-range enhancer-promoter interactions and describe a prototypical REX element that is necessary and sufficient to confer long-distance activation by remote enhancers.
Chromatin signatures are widely used to identify tissue-specific in vivo enhancers, but their sensitivity and specificity remains unclear. Here we show that many developmental enhancers remain undetectable using currently available chromatin data. In an initial comparison of over 1200 developmental enhancers with tissue-matched chromatin data, 14% ( n = 285) lacked canonical enhancer-associated chromatin signatures. To further assess the prevalence of enhancers missed by chromatin profiling approaches, we used a high-throughput transgenic enhancer assay to screen the regulatory landscapes of two key developmental genes at 5 kb resolution, spanning 1.3 Mb of mouse sequence in total. We observed that 23 of 88 (26%) in vivo enhancers discovered by this approach lacked enhancer-associated chromatin signatures in the respective tissue. Our findings suggest the existence of tens of thousands of enhancers that remain undiscovered by currently available chromatin data, underscoring the continued need for expanding resources for enhancer discovery.
Distant-acting enhancers are central to human development1. However, our limited understanding of their functional sequence features prevents the interpretation of enhancer mutations in disease2. Here we determined the functional sensitivity to mutagenesis of human developmental enhancers in vivo. Focusing on seven enhancers that are active in the developing brain, heart, limb and face, we created over 1,700 transgenic mice for over 260 mutagenized enhancer alleles. Systematic mutation of 12-base-pair blocks collectively altered each sequence feature in each enhancer at least once. We show that 69% of all blocks are required for normal in vivo activity, with mutations more commonly resulting in loss (60%) than in gain (9%) of function. Using predictive modelling, we annotated critical nucleotides at the base-pair resolution. The vast majority of motifs predicted by these machine learning models (88%) coincided with changes in in vivo function, and the models showed considerable sensitivity, identifying 59% of all functional blocks. Taken together, our results reveal that human enhancers contain a high density of sequence features that are required for their normal in vivo function and provide a rich resource for further exploration of human enhancer logic.
While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.
The genetic basis of human facial variation and craniofacial birth defects remains poorly understood. Distant-acting transcriptional enhancers control the fine-tuned spatiotemporal expression of genes during critical stages of craniofacial development. However, a lack of accurate maps of the genomic locations and cell type-resolved activities of craniofacial enhancers prevents their systematic exploration in human genetics studies. Here, we combine histone modification, chromatin accessibility, and gene expression profiling of human craniofacial development with single-cell analyses of the developing mouse face to define the regulatory landscape of facial development at tissue- and single cell-resolution. We provide temporal activity profiles for 14,000 human developmental craniofacial enhancers. We find that 56% of human craniofacial enhancers share chromatin accessibility in the mouse and we provide cell population- and embryonic stage-resolved predictions of their in vivo activity. Taken together, our data provide an expansive resource for genetic and developmental studies of human craniofacial development.
Regulatory elements (enhancers) are major drivers of gene expression in mammals and harbor many genetic variants associated with human diseases. Here, we present an updated VISTA Enhancer Browser (https://enhancer.lbl.gov), a database of transgenic enhancer assays conducted in developing mouse embryos in vivo. Since the original publication in 20 07, the database grew nearly 20-fold from 250 to over 4500 experiments and currently harbors over 23 500 images. The updated database provides structured information on experiments conducted at different stages of embryonic development, including enhancer activities of human pathogenic and synthetic variants and sequences derived from a variety of species. In addition to manually curated results of thousands of individual experiments, the new database also features hundreds of manually curated comparisons between alleles. The VISTA Enhancer Browser provides a crucial resource for study of human genetic variation, gene regulation and developmental biology. [GRAPHICS]
While most mammalian enhancers regulate their cognate promoters over moderate distances of tens of kilobases (kb), some enhancers act over distances in the megabase range. The sequence features enabling such extreme-distance enhancer-promoter interactions remain elusive. Here, we used in vivo enhancer replacement experiments in mice to show that short- and medium-range enhancers cannot initiate gene expression at extreme-distance range. We uncover a novel conserved cis-acting element, Range EXtender (REX), that confers extreme-distance regulatory activity and is located next to a long-range enhancer of Sall1. The REX element itself has no endogenous enhancer activity. However, addition of the REX to other short- and mid-range enhancers substantially increases their genomic interaction range. In the most extreme example observed, addition of the REX increased the range of an enhancer by an order of magnitude, from its native 71kb to 840kb. The REX element contains highly conserved [C/T]AATTA homeodomain motifs. These motifs are enriched around long-range limb enhancers genome-wide, including the ZRS, a benchmark long-range limb enhancer of Shh. Mutating the [C/T]AATTA motifs within the ZRS does not affect its limb-specific enhancer activity at short range, but selectively abolishes its long-range activity, resulting in severe limb reduction in knock-in mice. In summary, we identify a sequence signature globally associated with long-range enhancer-promoter interactions and describe a prototypical REX element that is necessary and sufficient to confer extreme-distance gene activation by remote enhancers.
Remote enhancers are thought to interact with their target promoters via physical proximity, yet the importance of this proximity for enhancer function remains unclear. Here we investigate the three-dimensional (3D) conformation of enhancers during mammalian development by generating high-resolution tissue-resolved contact maps for nearly a thousand enhancers with characterized in vivo activities in ten murine embryonic tissues. Sixty-one percent of developmental enhancers bypass their neighboring genes, which are often marked by promoter CpG methylation. The majority of enhancers display tissue-specific 3D conformations, and both enhancer–promoter and enhancer–enhancer interactions are moderately but consistently increased upon enhancer activation in vivo. Less than 14% of enhancer–promoter interactions form stably across tissues; however, these invariant interactions form in the absence of the enhancer and are likely mediated by adjacent CTCF binding. Our results highlight the general importance of enhancer–promoter physical proximity for developmental gene activation in mammals.
There is evidence that transcription factor (TF) encoding genes, which temporally control development in multiple cell types, can have tens of enhancers that regulate their expression. The NR2F1 TF developmentally promotes caudal and ventral cortical regional fates. Here, we epigenomically compared the activity of Nr2f1’s enhancers during mouse cortical development with their activity in a transgenic assay. We identified at least six that are likely to be important in prenatal cortical development, with three harboring de novo mutants identified in ASD individuals. We chose to study the function of two of the most robust enhancers by deleting them singly or together. We found that they have distinct and overlapping functions in driving Nr2f1’s regional and laminar expression in the developing cortex. Thus, these two enhancers, probably in combination with the others that we defined epigenetically, precisely tune Nr2f1’s regional, cell type, and temporal expression during corticogenesis.
Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, The ENCODE consortium mapped biochemical signals across many cell types and tissues and integrated these data to develop a Registry of 0.9 million human and 300 thousand mouse candidate cis-Regulatory Elements (cCREs) annotated with potential functions1. We have expanded the Registry to include 2.35 million human and 927 thousand mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded Registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays like STARR-seq, MPRA, CRISPR perturbation, and transgenic mouse assays now cover over 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer/silencer roles in different cellular contexts. Integrating the Registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by discovering KLF1 as a novel causal gene for red blood cell traits. This expanded Registry is a valuable resource for studying the regulatory genome and its impact on health and disease.
Approximately a quarter of the human genome consists of gene deserts, large regions devoid of genes often located adjacent to developmental genes and thought to contribute to their regulation. However, defining the regulatory functions embedded within these deserts is challenging due to their large size. Here, we explore the cis-regulatory architecture of a gene desert flanking the Shox2 gene, which encodes a transcription factor indispensable for proximal limb, craniofacial, and cardiac pacemaker development. We identify the gene desert as a regulatory hub containing more than 15 distinct enhancers recapitulating anatomical subdomains of Shox2 expression. Ablation of the gene desert leads to embryonic lethality due to Shox2 depletion in the cardiac sinus venosus, caused in part by the loss of a specific distal enhancer. The gene desert is also required for stylopod morphogenesis, mediated via distributed proximal limb enhancers. In summary, our study establishes a multi-layered role of the Shox2 gene desert in orchestrating pleiotropic developmental expression through modular arrangement and coordinated dynamics of tissue-specific enhancers.
A lingering question in developmental biology has centered on how transcription factors with widespread distribution in vertebrate embryos can perform tissue-specific functions. Here, using the murine hindlimb as a model, we investigate the elusive mechanisms whereby PBX TALE homeoproteins, viewed primarily as HOX cofactors, attain context-specific developmental roles despite ubiquitous presence in the embryo. We first demonstrate that mesenchymal-specific loss of PBX1/2 or the transcriptional regulator HAND2 generates similar limb phenotypes. By combining tissue-specific and temporally controlled mutagenesis with multi-omics approaches, we reconstruct a gene regulatory network (GRN) at organismal-level resolution that is collaboratively directed by PBX1/2 and HAND2 interactions in subsets of posterior hindlimb mesenchymal cells. Genome-wide profiling of PBX1 binding across multiple embryonic tissues further reveals that HAND2 interacts with subsets of PBX-bound regions to regulate limb-specific GRNs. Our research elucidates fundamental principles by which promiscuous transcription factors cooperate with cofactors that display domain-restricted localization to instruct tissue-specific developmental programs.
The genetic basis of craniofacial birth defects and general variation in human facial shape remains poorly understood. Distant-acting transcriptional enhancers are a major category of non-coding genome function and have been shown to control the fine-tuned spatiotemporal expression of genes during critical stages of craniofacial development 1–3 . However, a lack of accurate maps of the genomic location and cell type-specific in vivo activities of all craniofacial enhancers prevents their systematic exploration in human genetics studies. Here, we combined histone modification and chromatin accessibility profiling from different stages of human craniofacial development with single-cell analyses of the developing mouse face to create a comprehensive catalogue of the regulatory landscape of facial development at tissue- and single cell-resolution. In total, we identified approximately 14,000 enhancers across seven developmental stages from weeks 4 through 8 of human embryonic face development. We used transgenic mouse reporter assays to determine the in vivo activity patterns of human face enhancers predicted from these data. Across 16 in vivo validated human enhancers, we observed a rich diversity of craniofacial subregions in which these enhancers are active in vivo . To annotate the cell type specificities of human-mouse conserved enhancers, we performed single-cell RNA-seq and single-nucleus ATAC-seq of mouse craniofacial tissues from embryonic days e11.5 to e15.5. By integrating these data across species, we find that the majority (56%) of human craniofacial enhancers are functionally conserved in mice, providing cell type- and embryonic stage-resolved predictions of their in vivo activity profiles. Using retrospective analysis of known craniofacial enhancers in combination with single cell-resolved transgenic reporter assays, we demonstrate the utility of these data for predicting the in vivo cell type specificity of enhancers. Taken together, our data provide an expansive resource for genetic and developmental studies of human craniofacial development. Graphical Abstract
Topologically associating domain (TAD) boundaries partition the genome into distinct regulatory territories. Anecdotal evidence suggests that their disruption may interfere with normal gene expression and cause disease phenotypes(1-3), but the overall extent to which this occurs remains unknown. Here we demonstrate that targeted deletions of TAD boundaries cause a range of disruptions to normal in vivo genome function and organismal development. We used CRISPR genome editing in mice to individually delete eight TAD boundaries (11-80 kb in size) from the genome. All deletions examined resulted in detectable molecular or organismal phenotypes, which included altered chromatin interactions or gene expression, reduced viability, and anatomical phenotypes. We observed changes in local 3D chromatin architecture in 7 of 8 (88%) cases, including the merging of TADs and altered contact frequencies within TADs adjacent to the deleted boundary. For 5 of 8 (63%) loci examined, boundary deletions were associated with increased embryonic lethality or other developmental phenotypes. For example, a TAD boundary deletion near Smad3/Smad6 caused complete embryonic lethality, while a deletion near Tbx5/Lhx5 resulted in a severe lung malformation. Our findings demonstrate the importance of TAD boundary sequences for in vivo genome function and reinforce the critical need to carefully consider the potential pathogenicity of noncoding deletions affecting TAD boundaries in clinical genetics screening.
Mouse models are a critical tool for studying human diseases, particularly developmental disorders1. However, conventional approaches for phenotyping may fail to detect subtle defects throughout the developing mouse2. Here we set out to establish single-cell RNA sequencing of the whole embryo as a scalable platform for the systematic phenotyping of mouse genetic models. We applied combinatorial indexing-based single-cell RNA sequencing3 to profile 101 embryos of 22 mutant and 4 wild-type genotypes at embryonic day 13.5, altogether profiling more than 1.6 million nuclei. The 22 mutants represent a range of anticipated phenotypic severities, from established multisystem disorders to deletions of individual regulatory regions4,5. We developed and applied several analytical frameworks for detecting differences in composition and/or gene expression across 52 cell types or trajectories. Some mutants exhibit changes in dozens of trajectories whereas others exhibit changes in only a few cell types. We also identify differences between widely used wild-type strains, compare phenotyping of gain- versus loss-of-function mutants and characterize deletions of topological associating domain boundaries. Notably, some changes are shared among mutants, suggesting that developmental pleiotropy might be ‘decomposable’ through further scaling of this approach. Overall, our findings show how single-cell profiling of whole embryos can enable the systematic molecular and cellular phenotypic characterization of mouse mutants with unprecedented breadth and resolution. A study reports single-cell RNA-sequencing profiles for more than 1.6 million cell nuclei from 101 whole mouse embryos including 22 mutant and 4 wild-type genotypes, from one experiment.
Transcriptional enhancers are a predominant class of noncoding regulatory elements that activate cell type-specific gene expression. Tissue-specific enhancer-associated chromatin signatures have proven useful to identify candidate enhancer elements at a genome-wide scale, but their sensitivity for the comprehensive detection of all enhancers active in a given tissue in vivo remains unclear. Here we show that a substantial proportion of in vivo enhancers are hidden from discovery by conventional chromatin profiling methods. In an initial comparison of over 1,200 in vivo validated tissue-specific enhancers with tissue-matched mouse developmental epigenome data, 14% (n=286) of active enhancers did not show canonical enhancer-associated chromatin signatures in the tissue in which they are active. To assess the prevalence of enhancers not detectable by conventional chromatin profiling approaches in more detail, we used a high throughput transgenic enhancer reporter assay to systematically screen over 1.3 Mb of mouse genomic sequence at two critical developmental loci, assessing a total of 281 consecutive 5kb regions for in vivo enhancer activity in mouse embryos. We observed reproducible enhancer-reporter activity in 88 tissue-specific elements, 26% of which did not show canonical enhancer-associated chromatin signatures in the corresponding tissues. Overall, we find these hidden enhancers are indistinguishable from marked enhancers based on levels of evolutionary conservation, enrichment of transcription factor families, and genomic positioning relative to putative target genes. In combination, our retrospective and prospective studies assessed only 0.1% of the mouse genome and identified 309 tissue-specific enhancers that are hidden from current chromatin-based enhancer identification approaches. Our findings suggest the existence of tens of thousands of active enhancers throughout the genome that remain undetected by current chromatin profiling approaches and are an unappreciated source of additional genome function of import in interpreting growing whole human genome sequencing data.
The Plant Cell Atlas (PCA) community hosted a virtual symposium on December 9 and 10, 2021 on single cell and spatial omics technologies. The conference gathered almost 500 academic, industry, and government leaders to identify the needs and directions of the PCA community and to explore how establishing a data synthesis center would address these needs and accelerate progress. This report details the presentations and discussions focused on the possibility of a data synthesis center for a PCA and the expected impacts of such a center on advancing science and technology globally. Community discussions focused on topics such as data analysis tools and annotation standards; computational expertise and cyber-infrastructure; modes of community organization and engagement; methods for ensuring a broad reach in the PCA community; recruitment, training, and nurturing of new talent; and the overall impact of the PCA initiative. These targeted discussions facilitated dialogue among the participants to gauge whether PCA might be a vehicle for formulating a data synthesis center. The conversations also explored how online tools can be leveraged to help broaden the reach of the PCA (i.e., online contests, virtual networking, and social media stakeholder engagement) and decrease costs of conducting research (e.g., virtual REU opportunities). Major recommendations for the future of the PCA included establishing standards, creating dashboards for easy and intuitive access to data, and engaging with a broad community of stakeholders. The discussions also identified the following as being essential to the PCA's success: identifying homologous cell-type markers and their biocuration, publishing datasets and computational pipelines, utilizing online tools for communication (such as Slack), and user-friendly data visualization and data sharing. In conclusion, the development of a data synthesis center will help the PCA community achieve these goals by providing a centralized repository for existing and new data, a platform for sharing tools, and new analytical approaches through collaborative, multidisciplinary efforts. A data synthesis center will help the PCA reach milestones, such as community-supported data evaluation metrics, accelerating plant research necessary for human and environmental health.