Genome-wide association studies (GWASs) have identified multiple genetic regions that confer risk for juvenile idiopathic arthritis (JIA). However, identifying the single-nucleotide polymorphisms (SNPs) that drive disease risk has been impeded by the fact that the SNPs used to identify risk loci are in linkage disequilibrium (LD) with hundreds of other SNPs. Since the causal SNPs remain unknown, it is difficult to identify target genes and thus use genetic information to elucidate disease biology and inform patient care. We next used existing genotyping data from 3,939 children with JIA and 14,412 healthy controls to identify SNPs on JIA-risk haplotypes that present within open chromatin in multiple immune cell types and are more common in children with JIA than the controls (p < 0.05) in the genotyping datasets. We identified SNPs within cis-regulatory regions (cis-regulatory elements [CREs]) using precision run-on sequencing data and identified likely target genes using MicroC in both resting and activated CD4+ T cells. We identified 138 SNPs within the PROseq-identified CREs and n = 41 genes with which these CREs physically interacted. Data from Genotype-Tissue Expression (GTEx) and the Database of Immune Cell Expression Quantitative Trait Loci (DICE) corroborated these analyses by showing allelic effects for SNPs within the CREs in the ERAP2/LNPEP and locus. We further corroborated IRF1 allelic effects using a luciferase reporter assay. Our findings significantly reduce the genomic search space for risk-driving variants and target genes and support the roles of IRF1, ERAP2, and LNPEP in driving risk for JIA.
Promoter-proximal pausing of RNA polymerase II is a key regulatory checkpoint in metazoan transcription. Despite extensive study of this process, quantitative methods for comparing pausing dynamics across biological contexts have been lacking. Here we introduce a model-based framework for rigorous comparative analysis of both pause-escape kinetics and pause-site distributions across genes, cell types, and species. An application to available PRO-seq datasets revealed striking differences across perturbations, and comparative analyses across cell types and species highlighted distinct patterns of variation in both pause-escape kinetics and pause-site distributions, with only weak coupling between them. Integration with chromatin and sequence features showed that lower pause-escape rates are associated with stronger promoter-proximal nucleosome occupancy, whereas changes in pause-site dispersion are associated with sequence features such as GC skew. Together, these results establish a quantitative framework for comparative analysis of promoter-proximal pausing and reveal kinetic and distributional dimensions of pausing variation across biological contexts.
GWAS have identified multiple genetic regions that confer risk for juvenile idiopathic arthritis (JIA). However, identifying the single nucleotide polymorphisms (SNPs) that drive disease risk has been impeded by the fact that the SNPs used to identify risk loci are in linkage disequilibrium (LD) with hundreds of other SNPs. Since the causal SNPs remain unknown, it is difficult to identify target genes and thus use genetic information to elucidate disease biology and inform patient care. We next used existing genotyping data from 3,939 children with JIA and 14,412 healthy controls to identify SNPs on JIA risk haplotypes that: present within open chromatin in multiple immune cell types and more common in children with JIA than the controls (p<0.05) in the genotyping data sets. We identified SNPs within cis-regulatory regions (CREs) using precision run-on sequencing data, and identified likely target genes using MicroC in both resting and activated CD4+ T cells. We identified 138 SNPs within the PROseq-identified CREs, and n=41 genes with which these CREs physically interacted. Data from GTEx corroborated these analyses by showing allelic effects for SNPs within the CREs in the ERAP2 and IRF1 risk loci. We further corroborated IRF1 allelic effects using a luciferase reporter assay. Our findings significantly reduce the genomic search space for risk-driving variants and target genes and support the roles of IRF1, ERAP2 and LNPEP in driving risk for JIA.
Promoter-proximal pausing of RNA polymerase (Pol) II is a key regulatory step during transcription. Despite the central role of pausing in gene regulation, we do not understand the evolutionary processes that led to the emergence of Pol II pausing or its transition to a rate-limiting step actively controlled by transcription factors. Here, we analyzed transcription in species across the tree of life. Unicellular eukaryotes display an accumulation of Pol II near transcription start sites, which we propose transitioned to the longer-lived, focused pause observed in metazoans. This transition coincided with the evolution of new subunits in the negative elongation factor (NELF) and 7SK complexes. Depletion of NELF in mammals shifted the promoter-proximal buildup of Pol II from the pause site into the early gene body and compromised transcriptional activation for a set of heat-shock genes. Our work details the evolutionary history of Pol II pausing and sheds light on how new transcriptional regulatory mechanisms evolve.
Enhancer-promoter communication is fundamental to gene regulation in metazoans, yet the mechanisms underlying these interactions remain debated. Two primary models have been proposed: the structural bridge model, in which enhancers and promoters come into close proximity through stable, protein-mediated interactions, and the hub model, in which dynamic clusters of transcription-associated proteins facilitate communication over variable distances. Emerging evidence suggests that although enhancer-promoter pairs do come into close proximity during transcriptional activation, these interactions are highly transient, and the precise distances remain challenging to measure. Moving forward, resolving the distinctions between these models will require novel techniques to more precisely measure the spatial and temporal dynamics of enhancer-promoter interactions. Understanding how enhancers interact with promoters will deepen our understanding of the regulation of gene expression and the molecular underpinnings of transcriptional control.
Introduction: GWAS have identified multiple regions that confer risk for juvenile idiopathic arthritis (JIA). However, identifying the single nucleotide polymorphisms (SNPs) that drive disease risk is impeded by the SNPs that identify risk loci being in linkage disequilibrium (LD) with hundreds of other SNPs. Since the causal SNPs remain unknown, it is difficult to identify target genes and use genetic information to inform patient care. We used genotyping and functional data in primary human monocytes/macrophages to nominate disease-driving SNPs on JIA risk haplotypes and identify their likely target genes. Methods: We identified JIA risk haplotypes using Immunochip data from Hinks et al (Nature Gen 2013) and the meta-analysis from McIntosh et al (Arthritis Rheum 2017). We used genotyping data from 3,939 children with JIA and 14,412 healthy controls to identify SNPs that: (1) were situated within open chromatin in multiple immune cell types and (2) were more common in children with JIA than the controls (p< 0.05). We intersected the chosen SNPs (n=846) with regions of bi-directional transcription initiation characteristic of non-coding regulatory regions detected using dREG to analyze GRO-seq data. Finally, we used MicroC data to identify gene promoters interacting with the regulatory regions harboring the candidate causal SNPs. Results: We identified 190 SNPs that overlap with dREG peaks in monocytes and126 SNPs that overlap with dREG peaks in macrophages. Of these SNPs, 101 were situated within dREG peaks in both monocytes and macrophages, suggesting that these SNPs exert their effects independent of the cellular activation state. MicroC data in monocytes identified 20 genes/transcripts whose promoters interact with the enhancers harboring the SNPs of interest. Conclusion: SNPs in JIA risk regions that are candidate causal variants can be further screened using functional data such as GRO-seq. This process identifies a finite number of candidate causal SNPs, the majority of which are likely to exert their biological effects independent of cellular activation state in monocytes. Three-dimensional chromatin data generated with MicroC identifies genes likely to be influenced by these SNPs. These studies demonstrate the importance of investigations into the role of innate immunity in JIA.
During meiotic prophase I, spermatocytes must balance transcriptional activation with homologous recombination and chromosome synapsis, biological processes requiring extensive changes to chromatin state. We explored the interplay between chromatin accessibility and transcription through prophase I of mammalian meiosis by measuring genome-wide patterns of chromatin accessibility, nascent transcription, and processed mRNA. We find that Pol II is loaded on chromatin and maintained in a paused state early during prophase I. In later stages, paused Pol II is released in a coordinated transcriptional burst mediated by the transcription factors A-MYB and BRDT, resulting in ~3-fold increase in transcription. Transcriptional activity is temporally and spatially segregated from key steps of meiotic recombination: double strand breaks show evidence of chromatin accessibility earlier during prophase I and at distinct loci from those undergoing transcriptional activation, despite shared chromatin marks. Our findings reveal mechanisms underlying chromatin specialization in either transcription or recombination in meiotic cells.
How enhancers control target gene expression over long genomic distances remains an important unsolved problem. Here we investigated enhancer–promoter communication by integrating data from nucleosome-resolution genomic contact maps, nascent transcription and perturbations affecting either RNA polymerase II (Pol II) dynamics or the activity of thousands of candidate enhancers. Integration of new Micro-C experiments with published CRISPRi data demonstrated that enhancers spend more time in close proximity to their target promoters in functional enhancer–promoter pairs compared to nonfunctional pairs, which can be attributed in part to factors unrelated to genomic position. Manipulation of the transcription cycle demonstrated a key role for Pol II in enhancer–promoter interactions. Notably, promoter-proximal paused Pol II itself partially stabilized interactions. We propose an updated model in which elements of transcriptional dynamics shape the duration or frequency of interactions to facilitate enhancer–promoter communication. Perturbation of RNA polymerase II (Pol II) via chemical inhibition or dTAG-induced degradation and analysis using Micro-C and run-on sequencing show that enhancer–promoter contacts are dependent on transcription and stabilized by paused RNA Pol II.
How enhancers control target gene expression over long genomic distances remains an important unsolved problem. Here we studied enhancer-promoter contact architecture and communication by integrating data from nucleosome-resolution genomic contact maps, nascent transcription, and perturbations to transcription-associated proteins and thousands of candidate enhancers. Contact frequency between functionally validated enhancer-promoter pairs was most enriched near the +1 and +2 nucleosomes at enhancers and target promoters, indicating that functional enhancer-promoter pairs spend time in close physical proximity. Blocking RNA polymerase II (Pol II) caused major disruptions to enhancer-promoter contacts. Paused Pol II occupancy and the enzymatic activity of poly (ADP-ribose) polymerase 1 (PARP1) stabilized enhancer-promoter contacts. Based on our findings, we propose an updated model that couples transcriptional dynamics and enhancer-promoter communication.
During meiotic prophase I, germ cells must balance transcriptional activation with meiotic recombination and chromosome synapsis, biological processes requiring extensive changes to chromatin state and structure. Here we explored the interplay between chromatin accessibility and transcription across a detailed time-course of murine male meiosis by measuring genome-wide patterns of chromatin accessibility, nascent transcription, and processed mRNA. To understand the relationship between these parameters of gene regulation and recombination, we integrated these data with maps of double-strand break formation. Maps of nascent transcription show that Pol II is loaded on chromatin and maintained in a paused state early during prophase I. In later stages of prophase I, paused Pol II is released in a coordinated transcriptional burst resulting in ∼3-fold increase in transcription. Release from pause is mediated by the transcription factor A-MYB and the testis-specific bromodomain protein, BRDT. The burst of transcriptional activity is both temporally and spatially segregated from key steps of meiotic recombination: double strand breaks show evidence of chromatin accessibility earlier during prophase I and at distinct loci from those undergoing transcriptional activation, despite shared chromatin marks. Our findings reveal the mechanism underlying chromatin specialization in either transcription or recombination in meiotic cells.
E-protein transcription factors limit group 2 innate lymphoid cell (ILC2) development while promoting T cell differentiation from common lymphoid progenitors. Inhibitors of DNA binding (ID) proteins block E-protein DNA binding in common lymphoid progenitors to allow ILC2 development. However, whether E-proteins influence ILC2 function upon maturity and activation remains unclear. Mice that overexpress ID1 under control of the thymus-restricted proximal Lck promoter (ID1(tg/wT)) have a large pool of primarily thymus-derived ILC2s in the periphery that develop in the absence of E-protein activity. We used these mice to investigate how the absence of E-protein activity affects ILC2 function and the genomic landscape in response to house dust mite (HDM) allergens. ID1(tg/wT) mice had increased KLRG1(-) ILC2s in the lung compared with wild-type (WT; ID1(tg/wT)) mice in response to HDM, but ID1(tg/wT) ILC2s had an impaired capacity to produce type 2 cytokines. Analysis of WT ILC2 accessible chromatin suggested that AP-1 and C/EBP transcription factors but not E-proteins were associated with ILC2 inflammatory gene programs. Instead, E-protein binding sites were enriched at functional genes in ILC2s during development that were later dynamically regulated in allergic lung inflammation, including genes that control ILC2 response to cytokines and interactions with T cells. Finally, ILC2s from ID1(tg/wT) compared with WT mice had fewer regions of open chromatin near functional genes that were enriched for AP-1 factor binding sites following HDM treatment. These data show that E-proteins shape the chromatin landscape during ILC2 development to dictate the functional capacity of mature ILC2s during allergic inflammation in the lung.
To control for the possible contribution of age, gender, and cause of death to gene expression patterns, we normalized for these factors using linear regression (see below), such that the calculated residual expression per gene reflected age-/gender-/cause-of-death-independent gene expression. To this end we first divided the RPKM values of each sample by the sum of RPKM values associated with a given gene across all samples. In this way, we defined the expression value of each gene in a given sample as a fraction of its expression out of the sum of its expression values in all available samples. Hence, when calculating correlation of gene expression, we asked whether changes in the residual relative expression of genes are correlated. For gene “i” in sample “S”, the relative gene expression is calculated as follows:
The advent of animal husbandry and hunting increased human exposure to zoonotic pathogens. To understand how a zoonotic disease may have influenced human evolution, we study changes in human expression of anthrax toxin receptor 2 (ANTXR2), which encodes a cell surface protein necessary for Bacillus anthracis virulence toxins to cause anthrax disease. In immune cells, ANTXR2 is 8-fold down-regulated in all available human samples compared to non-human primates, indicating regulatory changes early in the evolution of modern humans. We also observe multiple genetic signatures consistent with recent positive selection driving a European-specific decrease in ANTXR2 expression in multiple tissues affected by anthrax toxins. Our observations fit a model in which humans adapted to anthrax disease following early ecological changes associated with hunting and scavenging, as well as a second period of adaptation after the rise of modern agriculture.
Quantification of mature-RNA isoform abundance from RNA-seq data has been extensively studied, but much less attention has been devoted to quantifying the abundance of distinct precursor RNAs based on nascent RNA sequencing data. Here we address this problem with a new computational method called Deconvolution of Expression for Nascent RNA sequencing data (DENR). DENR models the nascent RNA read counts at each locus as a mixture of user-provided isoforms. The performance of the baseline algorithm is enhanced by the use of machine-learning predictions of transcription start sites (TSSs) and an adjustment for the typical “shape profile” of read counts along a transcription unit. We show using simulated data that DENR clearly outperforms simple read-count-based methods for estimating the abundances of both whole genes and isoforms. By applying DENR to previously published PRO-seq data from K562 and CD4 + T cells, we find that transcription of multiple isoforms per gene is widespread, and the dominant isoform frequently makes use of an internal TSS. We also identify > 200 genes whose dominant isoforms make use of different TSSs in these two cell types. Finally, we apply DENR and StringTie to newly generated PRO-seq and RNA-seq data, respectively, for human CD4 + T cells and CD14 + monocytes, and show that entropy at the pre-RNA level makes a disproportionate contribution to overall isoform diversity, especially across cell types. Altogether, DENR is the first computational tool to enable abundance quantification of pre-RNA isoforms based on nascent RNA sequencing data, and it reveals high levels of pre-RNA isoform diversity in human cells.
Mitochondrial complex I (CI) is the largest multi-subunit oxidative phosphorylation (OXPHOS) protein complex. Recent availability of a high-resolution human CI structure, and from two non-human mammals, enabled predicting the impact of mutations on interactions involving each of the 44 CI subunits. However, experimentally assessing the impact of the predicted interactions requires an easy and high-throughput method. Here, we created such a platform by cloning all 37 nuclear DNA (nDNA) and 7 mitochondrial DNA (mtDNA)-encoded human CI subunits into yeast expression vectors to serve as both 'prey' and 'bait' in the split murine dihydrofolate reductase (mDHFR) protein complementation assay (PCA). We first demonstrated the capacity of this approach and then used it to examine reported pathological OXPHOS CI mutations that occur at subunit interaction interfaces. Our results indicate that a pathological frame-shift mutation in the MT-ND2 gene, causing the replacement of 126 C-terminal residues by a stretch of only 30 amino acids, resulted in loss of specificity in ND2-based interactions involving these residues. Hence, the split mDHFR PCA is a powerful assay for assessing the impact of disease-causing mutations on pairwise protein-protein interactions in the context of a large protein complex, thus offering a possible mechanistic explanation for the underlying pathogenicity.
Supplemental Dataset S1 of: Barshad et al (2018). Genome Research, 28(7): 952-967.
The bacterial heritage of mitochondria, as well as its independent genome [mitochondrial DNA (mtDNA)] and polycistronic transcripts, led to the view that mitochondrial transcriptional regulation relies on an evolutionarily conserved, prokaryotic-like system that is separated from the rest of the cell. Indeed, mtDNA transcription was previously thought to be governed by a few dedicated direct regulators, namely, the mitochondrial RNA polymerase (POLRMT), two transcription factors (TFAM and TF2BM), one transcription elongation (TEFM), and one known transcription termination factor (mTERF1). Recent findings have, however, revealed that known nuclear gene expression regulators are also involved in mtDNA transcription and have identified novel transcriptional features consistent with adaptation of the mitochondria to the regulatory environment of the precursor of the eukaryotic cell. Finally, whereas mammals follow the human mtDNA transcription pattern, other organisms notably diverge in terms of mtDNA transcriptional regulation. Hence, mtDNA transcriptional regulation is likely more evolutionary diverse than once thought.