Chromosome conformation capture (3C) methods measure DNA contact frequencies based on nuclear proximity ligation, to uncover in vivo genomic folding patterns. 4C-seq is a derivative 3C method, designed to search the genome for sequences contacting a selected genomic site of interest. 4C-seq employs inverse PCR and next generation sequencing to amplify, identify and quantify its proximity ligated DNA fragments. It generates high-resolution contact profiles for selected genomic sites based on limited amounts of sequencing reads. 4C-seq can be used to study multiple aspects of genome organization. It primarily serves to identify specific long-range DNA contacts between individual regulatory DNA modules, forming for example regulatory chromatin loops between enhancers and promoters, or architectural chromatin loops between cohesin- and CTCF- associated domain boundaries. Additionally, 4C-seq contact profiles can reveal the contours of contact domains and can identify the structural domains that co-occupy the same nuclear compartment. Here, we present an improved step-by-step protocol for sample preparation and the generation of 4C-seq sequencing libraries, including an optimized PCR and 4C template purification strategy. In addition, a data processing pipeline is provided which processes multiplexed 4C-seq reads directly from FASTQ files and generates files compatible with standard genome browsers for visualization and further statistical analysis of the data such as peak calling using peakC. The protocols and the pipeline presented should readily allow anyone to generate, visualize and interpret their own high resolution 4C contact datasets.
Most disease-associated variants identified by population based genetic studies are non-coding, which compromises finding causative genes and mechanisms. Presumably they interact through looping with nearby genes to modulate transcription. Hi-C provides the most complete and unbiased method for genome-wide identification of potential regulatory interactions, but finding chromatin loops in Hi-C data remains difficult and tissue specific data are limited. We have generated Hi-C data from primary cardiac tissue and developed a method, peakHiC, for sensitive and quantitative loop calling to uncover the human heart regulatory interactome. We identify complex CTCF-dependent and -independent contact networks, with loops between coding and non-coding gene promoters, shared enhancers and repressive sites. Across the genome, enhancer interaction strength correlates with gene transcriptional output and loop dynamics follows CTCF, cohesin and H3K27Ac occupancy levels. Finally, we demonstrate that intersection of the human heart regulatory interactome with cardiovascular disease variants facilitates prioritizing disease-causative genes.
Mutations and variations in and around SCN5A, encoding the major cardiac sodium channel, influence impulse conduction and are associated with a broad spectrum of arrhythmia disorders. Here, we identify an evolutionary conserved regulatory cluster with super enhancer characteristics downstream of SCN5A, which drives localized cardiac expression and contains conduction velocity-associated variants. We use genome editing to create a series of deletions in the mouse genome and show that the enhancer cluster controls the conformation of a >0.5 Mb genomic region harboring multiple interacting gene promoters and enhancers. We find that this cluster and its individual components are selectively required for cardiac Scn5a expression, normal cardiac conduction and normal embryonic development. Our studies reveal physiological roles of an enhancer cluster in the SCN5A-SCN10A locus, show that it controls the chromatin architecture of the locus and Scn5a expression, and suggest that genetic variants affecting its activity may influence cardiac function.
Genome-wide association studies (GWASs) implicate the PHACTR1 locus (6p24) in risk for five vascular diseases, including coronary artery disease, migraine headache, cervical artery dissection, fibromuscular dysplasia, and hypertension. Through genetic fine mapping, we prioritized rs9349379, a common SNP in the third intron of the PHACTR1 gene, as the putative causal variant. Epigenomic data from human tissue revealed an enhancer signature at rs9349379 exclusively in aorta, suggesting a regulatory function for this SNP in the vasculature. CRISPR-edited stem cell-derived endothelial cells demonstrate rs9349379 regulates expression of endothelin 1 (EDN1), a gene located 600 kb upstream of PHACTR1. The known physiologic effects of EDN1 on the vasculature may explain the pattern of risk for the five associated diseases. Overall, these data illustrate the integration of genetic, phenotypic, and epigenetic analysis to identify the biologic mechanism by which a common, non-coding variant can distally regulate a gene and contribute to the pathogenesis of multiple vascular diseases.
The noncoding genome is pervasively transcribed. Noncoding RNAs (ncRNAs) generated from enhancers have been proposed as a general facet of enhancer function and some have been shown to be required for enhancer activity. Here we examine the transcription-factor-(TF)-dependence of ncRNA expression to define enhancers and enhancer-associated ncRNAs that are involved in a TF-dependent regulatory network. TBX5, a cardiac TF, regulates a network of cardiac channel genes to maintain cardiac rhythm. We deep sequenced wildtype and Tbx5-mutant mouse atria, identifying ~2600 novel Tbx5-dependent ncRNAs. Tbx5-dependent ncRNAs were enriched for tissue-specific marks of active enhancers genome-wide. Tbx5-dependent ncRNAs emanated from regions that are enriched for TBX5-binding and that demonstrated Tbx5-dependent enhancer activity. Tbx5-dependent ncRNA transcription provided a quantitative metric of Tbx5-dependent enhancer activity, correlating with target gene expression. We identified RACER, a novel Tbx5-dependent long noncoding RNA (lncRNA) required for the expression of the calcium-handling gene Ryr2. We illustrate that TF-dependent enhancer transcription can illuminate components of TF-dependent gene regulatory networks.