In response to hypoxic stress at high altitudes, variation at the EPAS1 locus has experienced strong selection in Tibetans. Functional dissection of the selection signals at this locus identified ENH5, an enhancer within the adaptive haplotype that has a blunted response to hypoxic stress in Tibetans. ENH5 was shown to be pleiotropic in several tissues related to hypoxia response, suggesting that a possible mechanism behind the strong selection signatures could be adaptive pleiotropy. Tibetans not only experience hypoxic conditions, but also cold temperatures due to the altitude and climate of the Tibetan Plateau. However, it is unclear whether cold temperatures affect ENH5 activity possibly contributing to the selective pressure at this locus. Here, we further characterized the role of ENH5 in subcutaneous white adipose tissue, an important tissue type that regulates body temperature in response to cold temperatures by releasing stored fat as heat through a process called thermogenesis. In this work, we investigated the role of ENH5 in adipocytes using ENH5 knockout mice (ENH5 KO), which phenocopy the reduced activity of the Tibetan allele. We show that ENH5 KO mice at normoxia and room temperature do not have significant differences in organismal phenotypes related to adiposity and metabolism compared to WT mice on a high fat diet. However, we detected effects of ENH5 conditional on thermogenic stimulation and hypoxia exposure, independently, in adipocytes cultured in vitro . Under either of these conditions, ENH5 KO has stronger differential expression of key genes involved in thermogenesis activity and adipocyte differentiation compared to WT. This differential response to thermogenic stimulation expands on the pleiotropic effects of the Tibetan ENH5 allele(s), in addition to those previously shown in well-established hypoxia-responsive tissues. Our results raise the possibility that pleiotropic effects of ENH5 may implicate unforeseen mechanisms, such as cellular energetics and thermogenesis, possibly contributing to the phenotypic adaptation to high altitude in Tibetans.
Deciphering which genes are most important to disease etiology is a central challenge in human genetics. While genome-wide association studies have cataloged thousands of variants, it’s been proposed that most are indirect regulators of a limited, currently unidentified set of central disease-driving genes, defined here as disease-proximal genes (DPGs). Here, we introduce DANDELION, a mediation-inspired statistical framework that prioritizes DPGs by integrating trans-regulatory effects from disease-relevant tissues with gene-level burden from whole-exome sequencing. Applying DANDELION to asthma uncovers novel DPGs that escape detection by conventional methods. CRISPR screens in epithelial and T cells find that most DPGs regulate key asthma-related cellular phenotypes. We also demonstrate that loss of two DPGs, SLC27A3 and SCD, affects inflammation and airway remodeling in a mouse model of allergic asthma. Our study establishes DANDELION as a powerful framework for prioritizing novel, therapeutically actionable genes and pathways underlying disease pathogenesis.
Background:Genome-wide association studies (GWAS) have identified hundreds of loci underlying adult-onset asthma (AOA) and childhood-onset asthma (COA). However, the causal variants, regulatory elements, and effector genes at these loci are largely unknown. Methods:We performed heritability enrichment analysis to determine relevant cell types for AOA and COA, respectively. Next, we fine-mapped putative causal variants at AOA and COA loci. To improve the resolution of fine-mapping, we integrated ATAC-seq data in blood and lung cell types to annotate variants in candidate cis-regulatory elements (CREs). We then computationally prioritized candidate CREs underlying asthma risk, experimentally assessed their enhancer activity by massively parallel reporter assay (MPRA) in bronchial epithelial cells (BECs) and further validated a subset by luciferase assays. Combining chromatin interaction data and expression quantitative trait loci, we nominated genes targeted by candidate CREs and prioritized effector genes for AOA and COA. Results:Heritability enrichment analysis suggested a shared role of immune cells in the development of both AOA and COA while highlighting the distinct contribution of lung structural cells in COA. Functional fine-mapping uncovered 21 and 67 credible sets for AOA and COA, respectively, with only 16% shared between the two. Notably, one-third of the loci contained multiple credible sets. Our CRE prioritization strategy nominated 62 and 169 candidate CREs for AOA and COA, respectively. Over 60% of these candidate CREs showed open chromatin in multiple cell lineages, suggesting their potential pleiotropic effects in different cell types. Furthermore, COA candidate CREs were enriched for enhancers experimentally validated by MPRA in BECs. The prioritized effector genes included many genes involved in immune and inflammatory responses. Notably, multiple genes, including TNFSF4, a drug target undergoing clinical trials, were supported by two independent GWAS signals, indicating widespread allelic heterogeneity. Four out of six selected candidate CREs demonstrated allele-specific regulatory properties in luciferase assays in BECs. Conclusions:We present a comprehensive characterization of causal variants, regulatory elements, and effector genes underlying AOA and COA genetics. Our results supported a distinct genetic basis between AOA and COA and highlighted regulatory complexity at many GWAS loci marked by both extensive pleiotropy and allelic heterogeneity.
Asthma, allergic rhinitis, and atopic dermatitis are common, complex traits that are frequently co-morbid and have strong genetic correlation. However, the extent to which genome-wide genetic correlation between traits reflects shared causal variants or risk genes remains unclear. To address this question, we used functional fine-mapping. We generated genomic annotations from primary cells treated with immunomodulatory stimuli, then used these data to identify likely causal variants mediating genetic risk for allergic diseases including adult-onset asthma, childhood-onset asthma, allergic rhinitis, and atopic dermatitis. After identifying likely causal variants, we combined our functional annotations with expression quantitative trait loci and activity-by-contact modeling to predict effector genes. We confirmed a high degree of genetic correlation between GWAS loci for allergic diseases, but on the local level very few of the hundreds of likely causal variants identified by functional fine-mapping were shared between diseases. Instead, we found that each allergic disease was associated with a set of mostly unique variants. Nonetheless, nearly 40% of effector genes predicted to be the regulatory targets of these variants were shared between more than one allergic disease. When we tested candidate regulatory elements containing likely causal variants, we found that regulatory elements demonstrated variable allele-specific enhancer activity depending on the cell type in which they were tested. Overall, our findings suggest a highly pleiotropic gene regulatory network underlying allergic diseases, wherein disease-specific risk variants affect different regulatory elements that converge on the same set of target genes.
Vascular homeostasis and pathophysiology are tightly regulated by mechanical forces generated by hemodynamics. Vascular disorders such as atherosclerotic diseases largely occur at curvatures and bifurcations where disturbed blood flow activates endothelial cells while unidirectional flow at the straight part of vessels promotes endothelial health. Integrated analysis of the endothelial transcriptome, the 3D epigenome, and human genetics systematically identified the SNP-enriched cistrome in vascular endothelium subjected to well-defined atherosclerosis-prone disturbed flow or atherosclerosis-protective unidirectional flow. Our results characterized the endothelial typical- and super-enhancers and underscored the critical regulatory role of flow-sensitive endothelial super-enhancers. CRISPR interference and activation validated the function of a previously unrecognized unidirectional flow-induced super-enhancer that upregulates antioxidant genes NQO1, CYB5B, and WWP2, and a disturbed flow-induced super-enhancer in endothelium which drives prothrombotic genes EDN1 and HIVEP in vascular endothelium. Our results employing multiomics identify the cis-regulatory architecture of the flow-sensitive endothelial epigenome related to atherosclerosis and highlight the regulatory role of super-enhancers in mechanotransduction mechanisms.
Obesity-associated morbidity is exacerbated by abdominal obesity, which can be measured as the waist-to-hip ratio adjusted for the body mass index (WHRadjBMI). Here we identify genes associated with obesity and WHRadjBMI and characterize allele-sensitive enhancers that are predicted to regulate WHRadjBMI genes in women. We found that several waist-to-hip ratio-associated variants map within primate-specific Alu retrotransposons harboring a DNA motif associated with adipocyte differentiation. This suggests that a genetic component of adipose distribution in humans may involve co-option of retrotransposons as adipose enhancers. We evaluated the role of the strongest female WHRadjBMI-associated gene, SNX10 , in adipose biology. We determined that it is required for human adipocyte differentiation and function and participates in diet-induced adipose expansion in female mice, but not males. Our data identify genes and regulatory mechanisms that underlie female-specific adipose distribution and mediate metabolic dysfunction in women.
The mechanisms that underlie the timing of labor in humans are largely unknown. In most pregnancies, labor is initiated at term (≥ 37 weeks gestation), but in a signifiicant number of women spontaneous labor occurs preterm and is associated with increased perinatal mortality and morbidity. The objective of this study was to characterize the cells at the maternal–fetal interface (MFI) in term and preterm pregnancies in both the laboring and non-laboring state in Black women, who have among the highest preterm birth rates in the U.S. Using mass cytometry to obtain high-dimensional single-cell resolution, we identified 31 cell populations at the MFI, including 25 immune cell types and six non-immune cell types. Among the immune cells, maternal PD1 + CD8 T cell subsets were less abundant in term laboring compared to term non-laboring women. Among the non-immune cells, PD-L1 + maternal (stromal) and fetal (extravillous trophoblast) cells were less abundant in preterm laboring compared to term laboring women. Consistent with these observations, the expression of CD274 , the gene encoding PD-L1, was significantly depressed and less responsive to fetal signaling molecules in cultured mesenchymal stromal cells from the decidua of preterm compared to term women. Overall, these results suggest that the PD1/PD-L1 pathway at the MFI may perturb the delicate balance between immune tolerance and rejection and contribute to the onset of spontaneous preterm labor.
In Tibetans, noncoding alleles in EPAS1—whose protein product hypoxia-inducible factor 2α (HIF-2α) drives the response to hypoxia—carry strong signatures of positive selection; however, their functional mechanism has not been systematically examined. Here, we report that high-altitude alleles disrupt the activity of four EPAS1 enhancers in one or more cell types. We further characterize one enhancer (ENH5) whose activity is both allele specific and hypoxia dependent. Deletion of ENH5 results in down-regulation of EPAS1 and HIF-2α targets in acute hypoxia and in a blunting of the transcriptional response to sustained hypoxia. Deletion of ENH5 in mice results in dysregulation of gene expression across multiple tissues. We propose that pleiotropic adaptive effects of the Tibetan alleles in EPAS1 underlie the strong selective signal at this gene.
Genome-wide association studies (GWAS) have identified many disease-associated variants, yet mechanisms underlying these associations remain unclear. To understand obesity-associated variants, we generate gene regulatory annotations in adipocytes and hypothalamic neurons across cellular differentiation stages. We then test variants in 97 obesity-associated loci using a massively parallel reporter assay and identify putatively causal variants that display cell type specific or cross-tissue enhancer-modulating properties. Integrating these variants with gene regulatory information suggests genes that underlie obesity GWAS associations. We also investigate a complex genomic interval on 16p11.2 where two independent loci exhibit megabase-range, cross-locus chromatin interactions. We demonstrate that variants within these two loci regulate a shared gene set. Together, our data support a model where GWAS loci contain variants that alter enhancer activity across tissues, potentially with temporally restricted effects, to impact the expression of multiple genes. This complex model has broad implications for ongoing efforts to understand GWAS.
Various human diseases and pregnancy-related disorders reflect endometrial dysfunction. However, rodent models do not share fundamental biological processes with the human endometrium, such as spontaneous decidualization, and no existing human cell cultures recapitulate the cyclic interactions between endometrial stromal and epithelial compartments necessary for decidualization and implantation. Here we report a protocol differentiating human pluripotent stem cells into endometrial stromal fibroblasts (PSC-ESFs) that are highly pure and able to decidualize. Coculture of PSC-ESFs with placenta-derived endometrial epithelial cells generated organoids used to examine stromal-epithelial interactions. Cocultures exhibited specific endometrial markers in the appropriate compartments, organization with cell polarity, and hormone responsiveness of both cell types. Furthermore, cocultures recapitulate a central feature of the human decidua by cyclically responding to hormone withdrawal followed by hormone retreatment. This advance enables mechanistic studies of the cyclic responses that characterize the human endometrium.
Genome-wide association studies (GWAS) have implicated the IL33 locus in asthma, but the underlying mechanisms remain unclear. Here, we identify a 5 kb region within the GWAS-defined segment that acts as an enhancer-blocking element in vivo and in vitro. Chromatin conformation capture showed that this 5 kb region loops to the IL33 promoter, potentially regulating its expression. We show that the asthma-associated single nucleotide polymorphism (SNP) rs1888909, located within the 5 kb region, is associated with IL33 gene expression in human airway epithelial cells and IL-33 protein expression in human plasma, potentially through differential binding of OCT-1 (POU2F1) to the asthma-risk allele. Our data demonstrate that asthma-associated variants at the IL33 locus mediate allele-specific regulatory activity and IL33 expression, providing a mechanism through which a regulatory SNP contributes to genetic risk of asthma.
Human trophoblast stem cells (hTSC) can be isolated from first trimester placenta but not from term placenta. Here we demonstrate that villous cytotrophoblasts (vCTB) from term placenta can be reprogrammed into induced trophoblastic stem-like cells (iTSC) by introducing sets of transcription factors. The iTSCs express TSC markers such as GATA3, TEAD4 and ELF5, and are multipotent, validated by their differentiation into both extravillous trophoblasts (EVT) and syncytiotrophoblasts (STB) in vitro and in vivo. The iTSC can be passaged indefinitely in vitro without slowing of growth. The transcriptome profile of these cells closely resembles the profile of hTSC isolated from first trimester placentae but different from the term placental vCTB from which they originated. The ability to reprogram cells from term placenta into iTSC will allow study of early gestation events which impact placental function later in gestation, including preeclampsia and spontaneous preterm birth.
Whereas coding variants often have pleiotropic effects across multiple tissues, noncoding variants are thought to mediate their phenotypic effects by specific tissue and temporal regulation of gene expression. Here, we investigated the genetic and functional architecture of a genomic region within the FTO gene that is strongly associated with obesity risk. We show that multiple variants on a common haplotype modify the regulatory properties of several enhancers targeting IRX3 and IRX5 from megabase distances. We demonstrate that these enhancers affect gene expression in multiple tissues, including adipose and brain, and impart regulatory effects during a restricted temporal window. Our data indicate that the genetic architecture of disease-associated loci may involve extensive pleiotropy, allelic heterogeneity, shared allelic effects across tissues, and temporally restricted effects.
There is a life-long relationship between rhinovirus (RV) infection and the development and clinical manifestations of asthma. In this study we demonstrate that cultured primary bronchial epithelial cells from adults with asthma (n = 9) show different transcriptional and chromatin responses to RV infection compared to those without asthma (n = 9). Both the number and magnitude of transcriptional and chromatin responses to RV were muted in cells from asthma cases compared to controls. Pathway analysis of the transcriptionally responsive genes revealed enrichments of apoptotic pathways in controls but inflammatory pathways in asthma cases. Using promoter capture Hi-C we tethered regions of RV-responsive chromatin to RV-responsive genes and showed enrichment of these regions and genes at asthma GWAS loci. Taken together, our studies indicate a delayed or prolonged inflammatory state in cells from asthma cases and highlight genes that may contribute to genetic risk for asthma.
While Mediator plays a key role in eukaryotic transcription, little is known about its mechanism of action. This study combines CRISPR-Cas9 genetic screens, degron assays, Hi-C, and cryoelectron microscopy (cryo-EM) to dissect the function and structure of mammalian Mediator (mMED). Deletion analyses in B, T, and embryonic stem cells (ESC) identified a core of essential subunits required for Pol II recruitment genome-wide. Conversely, loss of non-essential subunits mostly affects promoters linked to multiple enhancers. Contrary to current models, however, mMED and Pol II are dispensable to physically tether regulatory DNA, a topological activity requiring architectural proteins. Cryo-EM analysis revealed a conserved core, with non-essential subunits increasing structural complexity of the tail module, a primary transcription factor target. Changes in tail structure markedly increase Pol II and kinase module interactions. We propose that Mediator’s structural pliability enables it to integrate and transmit regulatory signals and act as a functional, rather than an architectural bridge, between promoters and enhancers.
Over 500 genetic loci have been associated with risk of cardiovascular diseases (CVDs), however most loci are located in gene-distal non-coding regions and their target genes are not known. Here, we generated high-resolution promoter capture Hi-C (PCHi-C) maps in human induced pluripotent stem cells (iPSCs) and iPSC-derived cardiomyocytes (CMs) to provide a resource for identifying and prioritizing the functional targets of CVD associations. We validate these maps by demonstrating that promoters preferentially contact distal sequences enriched for tissue-specific transcription factor motifs and are enriched for chromatin marks that correlate with dynamic changes in gene expression. Using the CM PCHi-C map, we linked 1,999 CVD-associated SNPs to 347 target genes. Remarkably, more than 90% of SNP-target gene interactions did not involve the nearest gene, while 40% of SNPs interacted with at least two genes, demonstrating the importance of considering long-range chromatin interactions when interpreting functional targets of disease loci.
Article Figures and data Abstract eLife digest Introduction Results Discussion Materials and methods Data availability References Decision letter Author response Article and author information Metrics Abstract Over 500 genetic loci have been associated with risk of cardiovascular diseases (CVDs); however, most loci are located in gene-distal non-coding regions and their target genes are not known. Here, we generated high-resolution promoter capture Hi-C (PCHi-C) maps in human induced pluripotent stem cells (iPSCs) and iPSC-derived cardiomyocytes (CMs) to provide a resource for identifying and prioritizing the functional targets of CVD associations. We validate these maps by demonstrating that promoters preferentially contact distal sequences enriched for tissue-specific transcription factor motifs and are enriched for chromatin marks that correlate with dynamic changes in gene expression. Using the CM PCHi-C map, we linked 1999 CVD-associated SNPs to 347 target genes. Remarkably, more than 90% of SNP-target gene interactions did not involve the nearest gene, while 40% of SNPs interacted with at least two genes, demonstrating the importance of considering long-range chromatin interactions when interpreting functional targets of disease loci. https://doi.org/10.7554/eLife.35788.001 eLife digest Our genomes contain around 20,000 different genes that code for instructions to create proteins and other important molecules. When changes, or mutations, occur within these genes, malfunctioning proteins that are damaging to the cell may be produced. Researchers of human genetics have tried to spot the genetic mutations that are associated with illnesses, for example heart diseases. However, they found that most of these mutations are actually located outside of genes, in the ‘non-coding’ areas that make up the majority of our genome. These mutations do not modify proteins directly, which makes it challenging to understand how they may be related to heart conditions. One possibility is that the genetic changes affect regions called enhancers, which control where, when and how much a gene is turned on by physically interacting with it. Mutations in enhancers could lead to a gene producing too much or too little of a protein, which might create problems in the cell. Yet, it is difficult to match an enhancer with the gene or genes it controls. One reason is that a non-coding region can influence a gene placed far away on the DNA strand. Indeed, the long DNA molecule precisely folds in on itself to fit inside its compartment in the cell, which can bring together distant sequences. Montefiori et al. take over 500 non-coding areas, which can carry mutations associated with heart diseases, and use a technique called Hi-C to try to identify which genes these regions may control. The tool can model the 3D organization of the genome, and it was further modified to capture only the regions of the genome that contain genes, and the DNA sequences that interact with them, in human heart cells. This helped to create a 3D map of 347 genes which come in contact with the non-coding areas that carry mutations associated with heart diseases. In fact, deleting those genes often causes heart disorders in mice. In addition, Montefiori et al. reveal that 90% of the non-coding regions examined were influencing genes that were far away. This shows that, despite a common assumption, enhancers often do not regulate the coding sequences they are nearest to on the DNA strand. Pinpointing the genes regulated by the non-coding regions involved in cardiovascular diseases could lead to new ways of treating or preventing these conditions. The 3D map created by Montefiori et al. may also help to visualize how the genetic information is organized in heart cells. This will contribute to the current effort to understand the role of the 3D structure of the genome, especially in different cell types. https://doi.org/10.7554/eLife.35788.002 Introduction A major goal in human genetics research is to understand genetic contributions to complex diseases, specifically the molecular mechanisms by which common DNA variants impact disease etiology. Most genome-wide association studies (GWAS) implicate non-coding variants that are far from genes, complicating interpretation of their mode of action and correct identification of the target gene (Maurano et al., 2012). Mounting evidence suggests that disease variants disrupt the function of cis-acting regulatory elements, such as enhancers, which in turn affects expression of the specific gene or genes that are functional targets of these elements (Wright et al., 2010; Musunuru et al., 2010; Cowper-Sal lari et al., 2012; Smemo et al., 2014; Claussnitzer et al., 2015). However, because cis-acting regulatory elements can be located kilobases (kb) away from their target gene(s), identifying the true functional targets of regulatory elements remains challenging (Smemo et al., 2014). Chromosome conformation capture techniques such as Hi-C (Lieberman-Aiden et al., 2009) enable the genome-wide mapping of long-range chromatin contacts and therefore represent a promising strategy to identify distal gene targets of disease-associated genetic variants. Recently, Hi-C maps have been generated in numerous human cell types including embryonic stem cells and early embryonic lineages (Dixon et al., 2012, 2015), immune cells (Rao et al., 2014), fibroblasts (Jin et al., 2013) and other primary tissue types (Schmitt et al., 2016). However, despite the increasing abundance of Hi-C maps, most datasets are of limited resolution (>40 kb) and do not precisely identify the genomic regions in contact with gene promoters. More recently, promoter capture Hi-C (PCHi-C) was developed which greatly increases the power to detect interactions involving promoter sequences (Schoenfelder et al., 2015; Mifsud et al., 2015). PCHi-C in different cell types identified thousands of enhancer-promoter contacts and revealed extensive differences in promoter architecture between cell types and throughout differentiation (Schoenfelder et al., 2015; Mifsud et al., 2015; Javierre et al., 2016; Freire-Pritchett et al., 2017; Rubin et al., 2017; Siersbæk et al., 2017). These studies collectively demonstrated that genome architecture reflects cell identity, suggesting that disease-relevant cell types are critical for successful interrogation of the gene regulatory mechanisms of disease loci. In support of this notion, several recent studies utilized high-resolution promoter interaction maps to identify tissue-specific target genes of GWAS associations. Javierre et al. generated promoter capture Hi-C data in 17 primary human blood cell types and identified 2604 potentially causal genes for immune- and blood-related disorders, including many genes with unannotated roles in those diseases (Javierre et al., 2016). Similarly, Mumbach et al. interrogated GWAS SNPs associated with autoimmune diseases using HiChIP where they identified ~10,000 promoter-enhancer interactions that linked several hundred SNPs to target genes, most of which were not the nearest gene (Mumbach et al., 2017). Importantly, both studies reported cell-type specificity of SNP-target gene interactions. Cardiovascular diseases, including cardiac arrhythmia, heart failure, and myocardial infarction, continue to be the leading cause of death world-wide. Over 50 GWAS have been conducted for these specific cardiovascular phenotypes alone, with more than 500 loci implicated in cardiovascular disease risk (NHGRI GWAS catalog, https://www.ebi.ac.uk/gwas/), most of which map to non-coding genomic regions. To begin to dissect the molecular mechanisms by which genetic variants contribute to CVD risk, a comprehensive gene regulatory map of human cardiac cells is required. Here, we present high-resolution promoter interaction maps of human iPSCs and iPSC-derived cardiomyocytes (CMs). Using PCHi-C, we identified hundreds of thousands of promoter interactions in each cell type. We demonstrate the physiological relevance of these datasets by functionally interrogating the relationship between gene expression and long-range promoter interactions, and demonstrate the utility of long-range chromatin interaction data to resolve the functional targets of disease-associated loci. Results iPSC-derived cardiomyocytes provide an effective model to study the architecture of CVD genetics We used iPSC-derived CMs (Burridge et al., 2014) as a model to study cardiovascular gene regulation and disease genetics. The CMs generated in this study were 86–94% pure based on cardiac Troponin T protein expression and exhibited spontaneous, uniform beating (Figure 1—figure supplement 1A, Video 1). To demonstrate that iPSCs and CMs recapitulate transcriptional and epigenetic profiles of matched primary cells, we conducted RNA-seq and ChIP-seq for the active enhancer mark H3K27ac in both cell types and compared these data with similar cell types from the Epigenome Roadmap Project (Kundaje et al., 2015). RNA-seq profiles of iPSCs clustered tightly with H1 embryonic stem cells, whereas CMs clustered with both left ventricle (LV) and fetal heart (FH) profiles (Figure 1—figure supplement 1B). Furthermore, we observed that matched cell types exhibited three-fold greater overlap in the number of promoter-distal H3K27ac ChIP-seq peaks than non-matched cell types (Figure 1—figure supplement 1C,D), indicating that both iPSCs and CMs recapitulate tissue-specific epigenetic states of human stem cells and primary cardiomyocytes, respectively. To further validate our system, we analyzed differentially expressed genes between iPSCs and CMs. Among the top 10% of over-expressed genes in CMs were genes directly related to cardiac function including essential cardiac transcription factors (GATA4, MEIS1, TBX5, and TBX20) and differentiation products (TNNT2, MYH7B, MYL7, ACTN2, NPPA, HCN4, and RYR2) (fold-change >1.5, Padj <0.05, Figure 1—figure supplement 2A–C). Gene Ontology (GO) enrichment analysis for genes over-expressed in CMs relative to iPSCs further confirmed the cardiac-specific phenotypes of these cells with top terms relating to the development of the cardiac conduction system and cardiac muscle cell contraction (Figure 1—figure supplement 2D). Promoter-capture Hi-C identifies distal regulatory elements in iPSCs and CMs To comprehensively map long-range regulatory elements in iPSCs and CMs, we performed in-situ Hi-C (Rao et al., 2014) in triplicate iPSC-CM differentiations; importantly, we used the four-cutter restriction enzyme MboI which generates ligation fragments with an average size of 422 bp, enabling enhancer-level resolution of promoter contacts. We enriched iPSC and CM in situ Hi-C libraries for promoter interactions through hybridization with a set of 77,476 biotinylated RNA probes (‘baits’) targeting 22,600 human RefSeq protein-coding promoters (see Materials and methods) and sequenced each library to an average depth of ~413 million (M) paired-end reads. After removing duplicates and read-pairs that did not map to a bait, we obtained an average of 31M and 41M read-pairs per replicate for iPSC and CM, respectively. We used CHiCAGO (Cairns et al., 2016), a computational pipeline which accounts for bias from the sequence capture, to identify significant interactions and further filtered for those significant in at least two out of three replicates (see Materials and methods). Finally, we exclusively focused on interactions that were separated by a distance of at least 10 kb. This criterion addresses the high frequency of close-proximity ligation evets in Hi-C data, which are difficult to distinguish as random Brownian contacts or functional chromatin interactions (Cairns et al., 2016). In total, we identified 350,062 promoter interactions in iPSCs and 401,098 in CMs. A large proportion (~55%) of interactions were shared between the two cell types, indicating that even at high resolution many long-range interactions are stable across cell types (Figure 1A). Approximately 20% of all interactions were between two promoters, demonstrating the high connectivity between genes and supporting the recently suggested role of promoters acting as regulatory inputs for distal genes (Dao et al., 2017; Diao et al., 2017) (Figure 1B). Most interactions were promoter-distal, with a median of ~170 kb between the promoter and the distal-interacting region (Figure 1C). Figure 1 with 3 supplements see all Download asset Open asset General features of promoter interactions. (A) Venn diagram displaying the number of cell-type-specific and shared promoter interactions in each cell type. (B) Proportion of interactions in each distance category: promoter (P)-promoter (both interacting ends overlap a transcription start site (TSS)); P-proximal (non-promoter end overlaps captured region but not the TSS); P-distal (non-promoter end is outside of captured region). Note that all promoter interactions are separated by at least 10 kb. (C) Distribution of the distances spanning each interaction in iPSCs and CMs. The red line depicts the median (170 kb in iPSCs, 164 kb in CMs); the black line depicts the mean (208 kb in iPSCs, 206 kb in CMs). (D) A ~ 2 Mb region of chromosome 8 encompassing the GATA4 gene is shown along with pre-capture (whole genome) Hi-C interaction maps at 40 kb resolution for iPSCs (top) and CMs (bottom). TADs called with TopDom are shown as colored bars (median TAD size = 640 kb in both cell types, mean TAD size = 742 kb in iPSCs and 743 kb in CMs) and significant PCHi-C interactions as colored arcs. (E) Zoomed-in view of the GATA4 locus (promoter highlighted in yellow) in iPSCs (top) and CMs (bottom) along with corresponding RNA-seq data generated as part of this study, and ChIP-seq data for H3K27ac, H3K4me1, H3K27me3 and CTCF from the Epigenome Roadmap Project/ENCODE (H1 and left ventricle for iPSC and CM, respectively). Filtered GATA4 read counts used by CHiCAGO are displayed in blue with the corresponding significant interactions shown as arcs. For clarity, only GATA4 interactions are shown. Gray highlighted regions show interactions overlapping in vivo validated heart enhancers (pink boxes), with representative E11.5 embryos for each enhancer element (Visel et al., 2007). Red arrowhead points to the heart. https://doi.org/10.7554/eLife.35788.003 Video 1 Download asset This video cannot be played in place because your browser does support HTML5 video. You may still download the video for offline viewing. Download as MPEG-4 Download as WebM Download as Ogg Video of iPSC-derived cardiomyocytes exhibiting spontaneous beating at day 20 of the differentiation (day of cell harvesting). https://doi.org/10.7554/eLife.35788.007 To compare the PCHi-C maps with known features of genome organization, we sequenced our pre-capture Hi-C libraries to an average depth of 665M reads per cell type and identified topologically associating domains (TADs) with TopDom (see Materials and methods). TADs are organizational units of chromosomes defined by <1 megabase (Mb) genomic blocks that exhibit high self-interacting frequencies with a very low interaction frequency across TAD boundaries (Dixon et al., 2012; Nora et al., 2012). Notably, this organization is thought to constrain the activity of cis-regulatory elements to target genes within the same TAD, as disruption of TAD boundaries has been shown to lead to aberrant activation of genes in neighboring TADs (Nora et al., 2012; Lupiáñez et al., 2015; Franke et al., 2016; Symmons et al., 2016; Tsujimura et al., 2015). We found that the majority of PCHi-C interactions occurred within TADs (73 and 77% in iPSCs and CMs, respectively; Figure 1D and Figure 1—figure supplement 3A). TAD-crossing interactions (‘inter-TAD’) contained proportionally more promoter-promoter interactions than intra-TAD interactions, and were more likely to overlap promoter-distal CTCF sites; however, they were similarly enriched for looping to distal H3K27ac sites, a mark of active chromatin (Figure 1—figure supplement 3B–D). Inter-TAD interactions had slightly lower CHiCAGO scores, reflecting a lower number of reads supporting these interactions, and spanned greater genomic distances than intra-TAD interactions (Figure 1—figure supplement 3E,F). Additionally, promoters with inter-TAD interactions were preferentially located close to TAD boundaries (Figure 1—figure supplement 3G) and had higher expression levels compared to promoters with intra-TAD interactions, particularly in CMs (Figure 1—figure supplement 3H). These observations are consistent with previous studies which demonstrated that highly expressed genes, specifically housekeeping genes, are enriched at TAD boundaries (Dixon et al., 2012). To illustrate the utility of high-resolution PCHi-C interaction maps, we highlight the GATA4 locus in Figure 1D and E. GATA4 is a master regulator of heart development (Watt et al., 2004; Pikkarainen et al., 2004) and the GATA4 gene is located in a TAD structure that is relatively stable between iPSCs and CMs (Figure 1D). However, PCHi-C identified increased interaction frequencies between the GATA4 promoter and several H3K27ac-marked regions, including four in vivo validated heart enhancers from the Vista enhancer browser (Visel et al., 2007), specifically in CMs and coincident with strong up-regulation of GATA4 (Figure 1—figure supplement 2C). Although TAD-based analyses help define a gene’s cis-regulatory landscape, high-resolution promoter interaction data provides the resolution necessary to precisely map enhancer-promoter interactions in the context of cellular differentiation. To validate the CM interaction map as a resource for cardiovascular disease genetics we next extensively characterized several important aspects of genetic architecture in CMs. We compared CMs with iPSCs in each analysis as a measure of cell-type specificity. These analyses serve as benchmarks that build on established features of genome organization and aid interpretations of the roles that long range interactions play in gene regulation. Promoter interactions are enriched for tissue-specific transcription factor motifs Distal enhancers activate target genes through DNA looping, a mechanism that enables distally bound transcription factors to contact the transcription machinery of target promoters (Pennacchio et al., 2013; Miele and Dekker, 2008; Deng et al., 2012). To assess whether this feature of gene regulation was reflected in the iPSC and CM interactions, we conducted motif analysis using HOMER (Heinz et al., 2010) on the set of promoter-distal interacting sequences in each cell type. We initially focused on interactions for genes differentially expressed between iPSCs and CMs (fold-change >1.5, Padj <0.05). We identified CTCF as the most enriched motif in each case (Figure 2A,B), consistent with the known role of this factor in mediating long-range genomic interactions (Phillips and Corces, 2009; Phillips-Cremins et al., 2013; Nora et al., 2017). Among the other top motifs, we identified the pluripotency factor motifs OCT4-SOX2-TCF-NANOG (OSN) and SOX2 as preferentially enriched in distal sequences looping to genes over-expressed in iPSCs (Figure 2A,C), whereas top motifs in distal sequences looping to genes over-expressed in CMs included TBX20, ESRRB and MEIS1 (Figure 2B,C). TBX20 and MEIS1 transcription factors are important regulators of heart development and function (Cai et al., 2005; Sakabe et al., 2012; Mahmoud et al., 2013) and ESRRB was previously identified as a potential binding partner of TBX20 in adult mouse cardiomyocytes (Shen et al., 2011). We also observed that distal interactions unique to either iPSCs or CMs were similarly enriched for tissue-specific transcription factor motifs (Figure 2D). In line with a recent report that AP-1 contributes to dynamic loop formation during macrophage development (Phanstiel et al., 2017), both iPSC- and CM-specific interactions were enriched for AP-1 motifs (Figure 2D), suggesting that AP-1 transcription factors may represent a previously unrecognized genome organizing complex. Figure 2 Download asset Open asset Transcription factor motif enrichment in distal interacting regions. (A,B) Selected transcription factor (TF) motifs identified using HOMER in the promoter-distal interacting sequences for all over-expressed genes in (A) iPSCs and (B) CMs (fold change > 1.5, Padj < 0.05). ‘% sites’ refers to the percent of distal interactions overlapping the motif; rank is based on p-value significance. (C) To compare motif ranks across gene sets, the inverse of the rank is plotted for selected motifs identified in distal interactions from over- or under-expressed genes in both iPSCs and CMs. (D) The top 50 motifs identified in cell-type-specific interactions. OSN, OCT4-SOX2-TCF-NANOG motif. https://doi.org/10.7554/eLife.35788.008 Long-range promoter interactions are enriched for active cis-regulatory elements and correspond to gene expression dynamics Functionally active cis-regulatory elements are characterized by the presence of specific histone modifications; active enhancers are generally associated with H3K4me1 and H3K27ac (Creyghton et al., 2010; Heintzman et al., 2009), whereas inactive (e.g. poised or silenced) elements are often associated with H3K27me3 (Rada-Iglesias et al., 2011; Erceg et al., 2017). In support of the gene-regulatory function of long-range interactions, we found that the promoter-distal MboI fragments involved in significant promoter interactions were enriched for these three histone modifications in both iPSCs and CMs (Figure 3A–C). When promoters were grouped by expression level, we observed that this enrichment increased with increasing expression for H3K27ac and H3K4me1, and decreased with increasing expression for H3K27me3, consistent with an additive nature of enhancer-promoter interactions (Schoenfelder et al., 2015; Javierre et al., 2016), and validating that PCHi-C enriches for likely functional long-range chromatin contacts. Figure 3 with 1 supplement see all Download asset Open asset Enrichment of promoter interactions to distal regulatory features. (A,B) Proportion of promoter-distal interactions overlapping a histone ChIP-seq peak compared to random control MboI fragments (see Materials and methods). iPSC interactions were overlapped with H1 ESC ChIP-seq data; CM interactions were overlapped with left ventricle ChIP-seq data from the Epigenome Roadmap Project (Supplementary file 10). (C) Fold enrichment of the data presented in (A) and (B). (D) Fold enrichment of promoter-distal interactions based on the expression level of the promoter. Promoters were grouped into five bins according to their average TPM values. Dashed line indicates no enrichment. (E) Fold enrichment of cell-type-specific and shared interactions (columns) to tissue-specific and shared chromatin features (rows). (F) Example of the NPPA gene in iPSCs (top) and CMs (bottom). Gray box highlights CM-specific interactions to CM-specific chromatin marks and an in vivo heart enhancer (Visel et al., 2007). For clarity, only interactions for NPPA are shown. *p<0.00001, #p=0.0017, Z-test. https://doi.org/10.7554/eLife.35788.009 A strong correlation (Pearson correlation coefficient r > 0.7) between the degree of histone modifications and gene expression was first reported nearly 10 years ago (Karlić et al., 2010); however, that analysis only considered histone modifications within 2 kb of promoters. To understand whether this relationship extends beyond promoter-proximal regions, we correlated the number of histone ChIP-seq peaks within 300 kb of promoters with the promoter’s expression level (Figure 3—figure supplement 1A,B). H3K27ac and H3K4me1 both positively correlated with expression level (Spearman’s ρ = 0.22 and 0.16, respectively in iPSC and ρ = 0.23 and 0.24, respectively in CMs, p<2.2−16); in contrast, H3K27me3 negatively correlated with expression level in CMs (Spearman’s ρ = −0.20, p<2.2−16); however, this relationship was not present in iPSCs (Spearman’s ρ = 0.02, p=0.06). Although moderate, these correlations could partially explain why higher expressed genes show stronger enrichment for promoter interactions overlapping histone peaks when using a genome-wide background model (see Materials and methods), and lends support to the notion that active genes are located in generally active genomic environments (Stevens et al., 2017; Gilbert et al., 2004). We next investigated the relationship between cell-type-specific interactions and enrichment for tissue-specific CTCF, H3K27ac, and H3K27me3 marks, hypothesizing that interactions unique to iPSCs or CMs would be most enriched for tissue-specific chromatin features. Indeed, we observed that cell-type-specific interactions preferentially involved H3K27ac peaks from the matched cell type, and were either not enriched (iPSC) or depleted (CM) for H3K27ac marks that were specific to the non-matched cell type (Figure 3E, middle panel). However, the strongest enrichment was for cell-type-specific interactions to overlap chromatin features that were present in both cell types (Figure 3E). Additionally, interactions that were shared between iPSCs and CMs were most enriched for shared chromatin features. These results suggest that all interactions, whether shared or unique to one cell type, preferentially contact regulatory regions that are active in both cell types, whereas cell-type-specific interactions are not likely to occur in regions specifically marked in the non-matched cell type. An example of a gene that encompasses these observations is the atrial natriuretic peptide gene NPPA (Figure 3F) which is specifically expressed in cells of the heart atrium and is upregulated in CMs (Figure 1—figure supplement 2C). NPPA makes numerous cell-type-specific interactions to a distal region that is only marked with active chromatin (H3K27ac and H3K4me1) in CMs; furthermore, functional characterization showed that this region corresponds to an in vivo enhancer recapitulating NPPA’s endogenous expression in the developing heart (Visel et al., 2007). Taken together, these results illuminate the complex relationship between long-range promoter interactions and gene regulation and provide evidence that promoter architecture reflects cell-type-specific gene expression. Dynamic changes in genomic compartmentalization involve a subset of cardiac-specific genes As a final benchmark of our datasets, we analyzed large-scale differences in genome organization between iPSCs and CMs. The first Hi-C studies revealed that the genome is organized in two major compartments, A and B, that correspond to open and closed regions of chromosomes, respectively (Lieberman-Aiden et al., 2009; Rao et al., 2014). Although most compartments are stable across different cell types, some compartments switch states in a cell-type-specific manner which may reflect important gene regulatory changes (Dixon et al., 2015). To assess whether capture Hi-C data, which is more cost-effective for capturing promoter-centered interactions, is able to identify A/B compartments, we compared our capture Hi-C data with pre-capture, genome-wide Hi-C libraries. A/B compartments identified using HOMER (Heinz et al., 2010) were remarkably similar in the whole-genome and PCHi-C datasets (97% correspondence, Figure 4A, top panel, and Figure 4—figure supplements 1 and 2), demonstrating that PCHi-C data contains sufficient information to identify broadly active and inactive regions of the genome. As an example, we highlight a 10 Mb region on chromosome 4 containing the CAMK2D gene locus (Figure 4A). Compartments were relatively stable across this region in iPSCs and CMs; however, the CAMK2D gene itself was located in a dynamic compartment that switched from inactive in iPSCs to active in CMs. Correspondingly, this gene was highly upregulated during differentiation to CMs (Figure 4A, inset). Figure 4 with 3 supplements see all Download asset Open asset A/B compartment switching corresponds to activation of tissue-specific genes. (A) Top panel: 10 Mb region on chromosome four showing A (green) and B (blue) compartments based on the first principle component analysis calculated by HOMER (Heinz et al., 2010) of the whole-genome Hi-C and capture Hi-C interaction data. Bottom panel: zoomed in on the CAMK2D locus; only capture Hi-C A/B compartments shown. Inset: expression level of CAMK2D in iPSCs and CMs across the three replicates. (B) Expression level (TPM) of genes located in the A (green) or B (blue) compartment in each replicate of iPSC (left) or CM (right). (C) Difference in expression level (log2 fold change relative to iPSCs) of genes switching compartments from iPSC to CM or remaining in stable compartments. (D) Gene Ontology analysis of biological processes associated with genes switching from B to A compartments during iPSC-CM differentiation. ***p<2.2 × 10−16, Wilcoxon rank-sum test. https://doi.org/10.7554/eLife.35788.011 We observed this effect on a global level, as genes located in A compartments were expressed at significantly higher levels than genes located in the B compartments in both iPSCs and CMs (Figure 4B). Additionally, genes that switched A/B compartments between cell types were correspondingly up- or down-regulated (Figure 4C). GO analysis of the 1008 genes that switched from B to A compartments during iPSC-CM differentiation revealed enrichment for terms such as ‘cardiovascular sys
RATIONALE:Mutations in the transcription factor TBX20 (T-box 20) are associated with congenital heart disease. Germline ablation of Tbx20 results in abnormal heart development and embryonic lethality by embryonic day 9.5. Because Tbx20 is expressed in multiple cell lineages required for myocardial development, including pharyngeal endoderm, cardiogenic mesoderm, endocardium, and myocardium, the cell type-specific requirement for TBX20 in early myocardial development remains to be explored. OBJECTIVE:Here, we investigated roles of TBX20 in midgestation cardiomyocytes for heart development. METHODS AND RESULTS:Ablation of Tbx20 from developing cardiomyocytes using a doxycycline inducible cTnTCre transgene led to embryonic lethality. The circumference of developing ventricular and atrial chambers, and in particular that of prospective left atrium, was significantly reduced in Tbx20 conditional knockout mutants. Cell cycle analysis demonstrated reduced proliferation of Tbx20 mutant cardiomyocytes and their arrest at the G1-S phase transition. Genome-wide transcriptome analysis of mutant cardiomyocytes revealed differential expression of multiple genes critical for cell cycle regulation. Moreover, atrial and ventricular gene programs seemed to be aberrantly regulated. Putative direct TBX20 targets were identified using TBX20 ChIP-Seq (chromatin immunoprecipitation with high throughput sequencing) from embryonic heart and included key cell cycle genes and atrial and ventricular specific genes. Notably, TBX20 bound a conserved enhancer for a gene key to atrial development and identity, COUP-TFII/Nr2f2 (chicken ovalbumin upstream promoter transcription factor 2/nuclear receptor subfamily 2, group F, member 2). This enhancer interacted with the NR2F2 promoter in human cardiomyocytes and conferred atrial specific gene expression in a transgenic mouse in a TBX20-dependent manner. CONCLUSIONS:Myocardial TBX20 directly regulates a subset of genes required for fetal cardiomyocyte proliferation, including those required for the G1-S transition. TBX20 also directly downregulates progenitor-specific genes and, in addition to regulating genes that specify chamber versus nonchamber myocardium, directly activates genes required for establishment or maintenance of atrial and ventricular identity. TBX20 plays a previously unappreciated key role in atrial development through direct regulation of an evolutionarily conserved COUPT-FII enhancer.
ATAC-seq is a high-throughput sequencing technique that identifies open chromatin. Depending on the cell type, ATAC-seq samples may contain ~20–80% of mitochondrial sequencing reads. As the regions of open chromatin of interest are usually located in the nuclear genome, mitochondrial reads are typically discarded from the analysis. We tested two approaches to decrease wasted sequencing in ATAC-seq libraries generated from lymphoblastoid cell lines: targeted cleavage of mitochondrial DNA fragments using CRISPR technology and removal of detergent from the cell lysis buffer. We analyzed the effects of these treatments on the number of usable (unique, non-mitochondrial) reads and the number and quality of peaks called, including peaks identified in enhancers and transcription start sites. Both treatments resulted in considerable reduction of mitochondrial reads (1.7 and 3-fold, respectively). The removal of detergent, however, resulted in increased background and fewer peaks. The highest number of peaks and highest quality data was obtained by preparing samples with the original ATAC-seq protocol (using detergent) and treating them with CRISPR. This strategy reduced the amount of sequencing required to call a high number of peaks, which could lead to cost reduction when performing ATAC-seq on large numbers of samples and in cell types that contain a large amount of mitochondria.