Functional genomics data, such as chromatin state maps, provide critical insights into biological processes, but are hard to navigate and interpret. We present Epilogos to address this challenge by offering a simple information-theoretic framework for large-scale visualization, navigation and interpretation of functional genomics annotations, and apply it to over 2,000 genome-wide chromatin state maps in human and mouse. We construct intuitive visualizations of multi-tissue chromatin state maps, prioritize salient genomic regions, identify group-wise differential regions, and enable rapid similarity search given a region of interest. To facilitate usability, we provide a purpose-built web-based browser interface (http://epilogos.net) alongside open-source software for community access and adoption.
The Encyclopedia of DNA elements (ENCODE) project is a collaborative effort to create a comprehensive catalog of functional elements in the human genome. The current database comprises more than 19000 functional genomics experiments across more than 1000 cell lines and tissues using a wide array of experimental techniques to study the chromatin structure, regulatory and transcriptional landscape of the Homo sapiens and Mus musculus genomes. All experimental data, metadata, and associated computational analyses created by the ENCODE consortium are submitted to the Data Coordination Center (DCC) for validation, tracking, storage, and distribution to community resources and the scientific community. The ENCODE project has engineered and distributed uniform processing pipelines in order to promote data provenance and reproducibility as well as allow interoperability between genomic resources and other consortia. All data files, reference genome versions, software versions, and parameters used by the pipelines are captured and available via the ENCODE Portal. The pipeline code, developed using Docker and Workflow Description Language (WDL; https://openwdl.org/) is publicly available in GitHub, with images available on Dockerhub (https://hub.docker.com), enabling access to a diverse range of biomedical researchers. ENCODE pipelines maintained and used by the DCC can be installed to run on personal computers, local HPC clusters, or in cloud computing environments via Cromwell. Access to the pipelines and data via the cloud allows small labs the ability to use the data or software without access to institutional compute clusters. Standardization of the computational methodologies for analysis and quality control leads to comparable results from different ENCODE collections - a prerequisite for successful integrative analyses.
Methodological advances in conformation capture techniques have fundamentally changed our understanding of chromatin architecture. However, the nanoscale organization of chromatin and its cell-to-cell variance are less studied. Analyzing genome-wide data from 733 human cell and tissue samples, we identified 2 prototypical regions that exhibit high or absent hypersensitivity to deoxyribonuclease I, respectively. These regulatory active or inactive regions were examined in the lymphoblast cell line K562 by using high-throughput super-resolution microscopy. In both regions, we systematically measured the physical distance of 2 fluorescence in situ hybridization spots spaced by only 5 kb of DNA. Unexpectedly, the re -sulting distance distributions range from very compact to almost elongated configurations of more than 200-nm length for both the active and inactive regions. Monte Carlo simulations of a coarse-grained model of these chromatin regions based on pub-lished data of nucleosome occupancy in K562 cells were performed to understand the underlying mechanisms. There was no parameter set for the simulation model that can explain the microscopically measured distance distributions. Obviously, the chromatin state given by the strength of internucleosomal interaction, nucleosome occupancy, or amount of histone H1 differs from cell to cell, which results in the observed broad distance distributions. This large variability was not expected, especially in inactive regions. The results for the mechanisms for different distance distributions on this scale are important for understanding the contacts that mediate gene regulation. Microscopic measurements show that the inactive region investigated here is ex-pected to be embedded in a more compact chromatin environment. The simulation results of this region require an increase in the strength of internucleosomal interactions. It may be speculated that the higher density of chromatin is caused by the increased internucleosomal interaction strength.
Methodological advances in conformation capture techniques have fundamentally changed our understanding of chromatin architecture. However, the nanoscale organization of chromatin and its cell-to-cell variance are less studied. By using a combination of high throughput super-resolution microscopy and coarse-grained modelling we investigated properties of active and inactive chromatin in interphase nuclei. Using DNase I hypersensitivity as a criterion, we have selected prototypic active and inactive regions from ENCODE data that are representative for K-562 and more than 150 other cell types. By using oligoFISH and automated STED microscopy we systematically measured physical distances of the endpoints of 5kb DNA segments in these regions. These measurements result in high-resolution distance distributions which are right-tailed and range from very compact to almost elongated configurations of more than 200 nm length for both the active and inactive regions. Coarse-grained modeling of the respective DNA segments suggests that in regions with high DNase I hypersensitivity cell-to-cell differences in nucleosome occupancy determine the histogram shape. Simulations of the inactive region cannot sufficiently describe the compaction measured by microscopy, although internucleosomal interactions were elevated and the linker histone H1 was included in the model. These findings hint at further organizational mechanisms while the microscopy-based distance distribution indicates high cell-to-cell differences also in inactive chromatin regions. The analysis of the distance distributions suggests that direct enhancer-promoter contacts, which most models of enhancer action assume, happen for proximal regulatory elements in a probabilistic manner due to chromatin flexibility.
Combinatorial binding of transcription factors to regulatory DNA underpins gene regulation in all organisms. Genetic variation in regulatory regions has been connected with diseases and diverse phenotypic traits 1 , but it remains challenging to distinguish variants that affect regulatory function 2 . Genomic DNase I footprinting enables the quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin 3 – 6 . However, only a small fraction of such sites have been precisely resolved on the human genome sequence 6 . Here, to enable comprehensive mapping of transcription factor footprints, we produced high-density DNase I cleavage maps from 243 human cell and tissue types and states and integrated these data to delineate about 4.5 million compact genomic elements that encode transcription factor occupancy at nucleotide resolution. We map the fine-scale structure within about 1.6 million DNase I-hypersensitive sites and show that the overwhelming majority are populated by well-spaced sites of single transcription factor–DNA interaction. Cell-context-dependent cis -regulation is chiefly executed by wholesale modulation of accessibility at regulatory DNA rather than by differential transcription factor occupancy within accessible elements. We also show that the enrichment of genetic variants associated with diseases or phenotypic traits in regulatory regions 1 , 7 is almost entirely attributable to variants within footprints, and that functional variants that affect transcription factor occupancy are nearly evenly partitioned between loss- and gain-of-function alleles. Unexpectedly, we find increased density of human genetic variation within transcription factor footprints, revealing an unappreciated driver of cis -regulatory evolution. Our results provide a framework for both global and nucleotide-precision analyses of gene regulatory mechanisms and functional genetic variation.
Early mammalian development is orchestrated by genome-encoded regulatory elements populated by a changing complement of regulatory factors, creating a dynamic chromatin landscape. To define the spatiotemporal organization of regulatory DNA landscapes during mouse development and maturation, we generated nucleotide-resolution DNA accessibility maps from 15 tissues sampled at 9 intervals spanning post-conception day 9.5 through early adult, and integrated these with 41 adult-stage DNase-seq profiles to create a global atlas of mouse regulatory DNA. Collectively, we delineated >1.8 million DNase I hypersensitive sites (DHSs), with the vast majority displaying temporal and tissue-selective patterning. Here we show that tissue regulatory DNA compartments show sharp embryonic-to-fetal transitions characterized by wholesale turnover of DHSs and progressive domination by a diminishing number of transcription factors. We show further that aligning mouse and human fetal development on a regulatory axis exposes disease-associated variation enriched in early intervals lacking human samples. Our results provide an expansive new resource for decoding mammalian developmental regulatory programs. ### Competing Interest Statement The authors have declared no competing interest.
Combinatorial binding of transcription factors to regulatory DNA underpins gene regulation in all organisms. Genetic variation in regulatory regions has been connected with diseases and diverse phenotypic traits, yet it remains challenging to distinguish variants that impact regulatory function. Genomic DNase I footprinting enables quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin. However, to date only a small fraction of such sites have been precisely resolved on the human genome sequence. To enable comprehensive mapping of transcription factor footprints, we produced high-density DNase I cleavage maps from 243 human cell and tissue types and states and integrated these data to delineate at nucleotide resolution ~4.5 million compact genomic elements encoding transcription factor occupancy. We map the fine-scale structure of ~1.6 million DHS and show that the overwhelming majority is populated by well-spaced sites of single transcription factor:DNA interaction. Cell context-dependent cis-regulation is chiefly executed by wholesale actuation of accessibility at regulatory DNA versus by differential transcription factor occupancy within accessible elements. We show further that the well-described enrichment of disease- and phenotypic trait-associated genetic variants in regulatory regions is almost entirely attributable to variants localizing within footprints, and that functional variants impacting transcription factor occupancy are nearly evenly partitioned between loss- and gain-of-function alleles. Unexpectedly, we find that the global density of human genetic variation is markedly increased within transcription factor footprints, revealing an unappreciated driver of cis-regulatory evolution. Our results provide a new framework for both global and nucleotide-precision analyses of gene regulatory mechanisms and functional genetic variation.
DNase I hypersensitive sites (DHSs) are generic markers of regulatory DNA 1 – 5 and contain genetic variations associated with diseases and phenotypic traits 6 – 8 . We created high-resolution maps of DHSs from 733 human biosamples encompassing 438 cell and tissue types and states, and integrated these to delineate and numerically index approximately 3.6 million DHSs within the human genome sequence, providing a common coordinate system for regulatory DNA. Here we show that these maps highly resolve the cis -regulatory compartment of the human genome, which encodes unexpectedly diverse cell- and tissue-selective regulatory programs at very high density. These programs can be captured comprehensively by a simple vocabulary that enables the assignment to each DHS of a regulatory barcode that encapsulates its tissue manifestations, and global annotation of protein-coding and non-coding RNA genes in a manner orthogonal to gene expression. Finally, we show that sharply resolved DHSs markedly enhance the genetic association and heritability signals of diseases and traits. Rather than being confined to a small number of distal elements or promoters, we find that genetic signals converge on congruently regulated sets of DHSs that decorate entire gene bodies. Together, our results create a universal, extensible coordinate system and vocabulary for human regulatory DNA marked by DHSs, and provide a new global perspective on the architecture of human gene regulation.
DNase I hypersensitive sites (DHSs) are generic markers of regulatory DNA and harbor disease- and phenotypic trait-associated genetic variation. We established high-precision maps of DNase I hypersensitive sites from 733 human biosamples encompassing 439 cell and tissue types and states, and integrated these to precisely delineate and numerically index ~3.6 million DHSs encoded within the human genome, providing a common coordinate system for regulatory DNA. Here we show that the expansive scale of cell and tissue states sampled exposes an unprecedented degree of stereotyped actuation of large sets of elements, signaling the operation of distinct genome-scale regulatory programs. We show further that the complex actuation patterns of individual elements can be captured comprehensively by a simple regulatory vocabulary reflecting their dominant cellular manifestation. This vocabulary, in turn, enables comprehensive and quantitative regulatory annotation of both protein-coding genes and the vast array of well-defined but poorly-characterized non-coding RNA genes. Finally, we show that the combination of high-precision DHSs and regulatory vocabularies markedly concentrate disease- and trait-associated non-coding genetic signals both along the genome and across cellular compartments. Taken together, our results provide a common and extensible coordinate system and vocabulary for human regulatory DNA, and a new global perspective on the architecture of human gene regulation.
BackgroundTranscriptional dysregulation drives cancer formation but the underlying mechanisms are still poorly understood. Renal cell carcinoma (RCC) is the most common malignant kidney tumor which canonically activates the hypoxia-inducible transcription factor (HIF) pathway. Despite intensive study, novel therapeutic strategies to target RCC have been difficult to develop. Since the RCC epigenome is relatively understudied, we sought to elucidate key mechanisms underpinning the tumor phenotype and its clinical behavior.MethodsWe performed genome-wide chromatin accessibility (DNase-seq) and transcriptome profiling (RNA-seq) on paired tumor/normal samples from 3 patients undergoing nephrectomy for removal of RCC. We incorporated publicly available data on HIF binding (ChIP-seq) in a RCC cell line. We performed integrated analyses of these high-resolution, genome-scale datasets together with larger transcriptomic data available through The Cancer Genome Atlas (TCGA).FindingsThough HIF transcription factors play a cardinal role in RCC oncogenesis, we found that numerous transcription factors with a RCC-selective expression pattern also demonstrated evidence of HIF binding near their gene body. Examination of chromatin accessibility profiles revealed that some of these transcription factors influenced the tumor's regulatory landscape, notably the stem cell transcription factor POU5F1 (OCT4). Elevated POU5F1 transcript levels were correlated with advanced tumor stage and poorer overall survival in RCC patients. Unexpectedly, we discovered a HIF-pathway-responsive promoter embedded within a endogenous retroviral long terminal repeat (LTR) element at the transcriptional start site of the PSOR1C3 long non-coding RNA gene upstream of POU5F1. RNA transcripts are induced from this promoter and read through PSOR1C3 into POU5F1 producing a novel POU5F1 transcript isoform. Rather than being unique to the POU5F1 locus, we found that HIF binds to several other transcriptionally active LTR elements genome-wide correlating with broad gene expression changes in RCC.InterpretationIntegrated transcriptomic and epigenomic analysis of matched tumor and normal tissues from even a small number of primary patient samples revealed remarkably convergent shared regulatory landscapes. Several transcription factors appear to act downstream of HIF including the potent stem cell transcription factor POU5F1. Dysregulated expression of POU5F1 is part of a larger pattern of gene expression changes in RCC that may be induced by HIF-dependent reactivation of dormant promoters embedded within endogenous retroviral LTRs.
Transcriptional dysregulation drives cancer formation but the underlying mechanisms are still poorly understood. As a model system, we used renal cell carcinoma (RCC), the most common malignant kidney tumor which canonically activates the hypoxia-inducible transcription factor (HIF) pathway. We performed genome-wide chromatin accessibility and transcriptome profiling on paired tumor/normal samples and found that numerous transcription factors with a RCC-selective expression pattern also demonstrated evidence of HIF binding in the vicinity of their gene body. Some of these transcription factors influenced the tumor’s regulatory landscape, notably the stem cell transcription factor POU5F1 ( OCT4 ). Unexpectedly, we discovered a HIF-pathway-responsive cryptic promoter embedded within a human-specific retroviral repeat element that drives POU5F1 expression in RCC via a novel transcript. Elevat POU5F1 expression levels were correlated with advanced tumor stage and poorer overall survival in RCC patients. Thus, integrated transcriptomic and epigenomic analysis of even a small number of primary patient samples revealed remarkably convergent shared regulatory landscapes and a novel mechanism for dysregulated expression of POU5F1 in RCC.
Cancer is a disease potentiated by mutations in somatic cells. Cancer mutations are not distributed uniformly along the human genome. Instead, different human genomic regions vary by up to fivefold in the local density of cancer somatic mutations, posing a fundamental problem for statistical methods used in cancer genomics. Epigenomic organization has been proposed as a major determinant of the cancer mutational landscape. However, both somatic mutagenesis and epigenomic features are highly cell-type-specific. We investigated the distribution of mutations in multiple independent samples of diverse cancer types and compared them to cell-type-specific epigenomic features. Here we show that chromatin accessibility and modification, together with replication timing, explain up to 86% of the variance in mutation rates along cancer genomes. The best predictors of local somatic mutation density are epigenomic features derived from the most likely cell type of origin of the corresponding malignancy. Moreover, we find that cell-of-origin chromatin features are much stronger determinants of cancer mutation profiles than chromatin features of matched cancer cell lines. Furthermore, we show that the cell type of origin of a cancer can be accurately determined based on the distribution of mutations along its genome. Thus, the DNA sequence of a cancer genome encompasses a wealth of information about the identity and epigenomic features of its cell of origin.
The laboratory mouse shares the majority of its protein-coding genes with humans, making it the premier model organism in biomedical research, yet the two mammals differ in significant ways. To gain greater insights into both shared and species-specific transcriptional and cellular regulatory programs in the mouse, the Mouse ENCODE Consortium has mapped transcription, DNase I hypersensitivity, transcription factor binding, chromatin modifications and replication domains throughout the mouse genome in diverse cell and tissue types. By comparing with the human genome, we not only confirm substantial conservation in the newly annotated potential functional sequences, but also find a large degree of divergence of sequences involved in transcriptional regulation, chromatin state and higher order chromatin organization. Our results illuminate the wide range of evolutionary forces acting on genes and their regulatory regions, and provide a general resource for research into mammalian biology and mechanisms of human diseases.
To study the evolutionary dynamics of regulatory DNA, we mapped >1.3 million deoxyribonuclease I–hypersensitive sites (DHSs) in 45 mouse cell and tissue types, and systematically compared these with human DHS maps from orthologous compartments. We found that the mouse and human genomes have undergone extensive cis-regulatory rewiring that combines branch-specific evolutionary innovation and loss with widespread repurposing of conserved DHSs to alternative cell fates, and that this process is mediated by turnover of transcription factor (TF) recognition elements. Despite pervasive evolutionary remodeling of the location and content of individual cis-regulatory regions, within orthologous mouse and human cell types the global fraction of regulatory DNA bases encoding recognition sites for each TF has been strictly conserved. Our findings provide new insights into the evolutionary forces shaping mammalian regulatory DNA landscapes.
Hematopoietic protein-1 (Hem-1) is a hematopoietic cell specific member of the WAVE (Wiskott-Aldrich syndrome verprolin-homologous protein) complex, which regulates filamentous actin (F-actin) polymerization in many cell types including immune cells. However, the roles of Hem-1 and the WAVE complex in erythrocyte biology are not known. In this study, we utilized mice lacking Hem-1 expression due to a non-coding point mutation in the Hem1 gene to show that absence of Hem-1 results in microcytic, hypochromic anemia characterized by abnormally shaped erythrocytes with aberrant F-actin foci and decreased lifespan. We find that Hem-1 and members of the associated WAVE complex are normally expressed in wildtype erythrocyte progenitors and mature erythrocytes. Using mass spectrometry and global proteomics, Coomassie staining, and immunoblotting, we find that the absence of Hem-1 results in decreased representation of essential erythrocyte membrane skeletal proteins including α- and β- spectrin, dematin, p55, adducin, ankyrin, tropomodulin 1, band 3, and band 4.1. Hem1⁻/⁻ erythrocytes exhibit increased protein kinase C-dependent phosphorylation of adducin at Ser724, which targets adducin family members for dissociation from spectrin and actin, and subsequent proteolysis. Increased adducin Ser724 phosphorylation in Hem1⁻/⁻ erythrocytes correlates with decreased protein expression of the regulatory subunit of protein phosphatase 2A (PP2A), which is required for PP2A-dependent dephosphorylation of PKC targets. These results reveal a novel, critical role for Hem-1 in the homeostasis of structural proteins required for formation and stability of the actin membrane skeleton in erythrocytes.
DNase I hypersensitive sites (DHSs) are markers of regulatory DNA and have underpinned the discovery of all classes of cis-regulatory elements including enhancers, promoters, insulators, silencers and locus control regions. Here we present the first extensive map of human DHSs identified through genome-wide profiling in 125 diverse cell and tissue types. We identify ∼2.9 million DHSs that encompass virtually all known experimentally validated cis-regulatory sequences and expose a vast trove of novel elements, most with highly cell-selective regulation. Annotating these elements using ENCODE data reveals novel relationships between chromatin accessibility, transcription, DNA methylation and regulatory factor occupancy patterns. We connect ∼580,000 distal DHSs with their target promoters, revealing systematic pairing of different classes of distal DHSs and specific promoter types. Patterning of chromatin accessibility at many regulatory regions is organized with dozens to hundreds of co-activated elements, and the transcellular DNase I sensitivity pattern at a given region can predict cell-type-specific functional behaviours. The DHS landscape shows signatures of recent functional evolutionary constraint. However, the DHS compartment in pluripotent and immortalized cells exhibits higher mutation rates than that in highly differentiated cells, exposing an unexpected link between chromatin accessibility, proliferative potential and patterns of human variation. An extensive map of human DNase I hypersensitive sites, markers of regulatory DNA, in 125 diverse cell and tissue types is described; integration of this information with other ENCODE-generated data sets identifies new relationships between chromatin accessibility, transcription, DNA methylation and regulatory factor occupancy patterns. This paper describes the first extensive map of human DNaseI hypersensitive sites — markers of regulatory DNA — in 125 diverse cell and tissue types. Integration of this information with other data sets generated by ENCODE (Encyclopedia of DNA Elements) identified new relationships between chromatin accessibility, transcription, DNA methylation and regulatory-factor occupancy patterns. Evolutionary-conservation analysis revealed signatures of recent functional constraint within DNaseI hypersensitive sites.
Genome-wide association studies have identified many noncoding variants associated with common diseases and traits. We show that these variants are concentrated in regulatory DNA marked by deoxyribonuclease I (DNase I) hypersensitive sites (DHSs). Eighty-eight percent of such DHSs are active during fetal development and are enriched in variants associated with gestational exposure-related phenotypes. We identified distant gene targets for hundreds of variant-containing DHSs that may explain phenotype associations. Disease-associated variants systematically perturb transcription factor recognition sequences, frequently alter allelic chromatin states, and form regulatory networks. We also demonstrated tissue-selective enrichment of more weakly disease-associated variants within DHSs and the de novo identification of pathogenic cell types for Crohn's disease, multiple sclerosis, and an electrocardiogram trait, without prior knowledge of physiological mechanisms. Our results suggest pervasive involvement of regulatory DNA variation in common human disease and provide pathogenic insights into diverse disorders.