Long-range transcriptional activation of gene promoters by abundant enhancers in animal genomes calls for mechanisms to limit inappropriate regulation. DNA elements called insulators serve this purpose by shielding promoters from an enhancer when interposed. Unlike promoters and enhancers, insulators have not been systematically characterized due to lacking high-throughput screening assays, and questions regarding how insulators are distributed and encoded in the genome remain. Here, we establish "insulator-seq" as a plasmid-based massively parallel reporter assay in Drosophila cultured cells to perform a systematic insulator screen of selected genomic loci. Screening developmental gene loci showed that not all insulator protein binding sites effectively block enhancer-promoter communication. Deep insulator mutagenesis identified sequences flexibly positioned around the CTCF insulator protein binding motif that are critical for functionality. The ability to screen millions of DNA sequences without positional effect has enabled functional mapping of insulators and provided further insights into the determinants of insulators.
Regulatory elements, such as enhancers and silencers, control transcription by establishing physical proximity to target gene promoters. Neurons in flies and mammals exhibit long-range three-dimensional genome contacts, proposed to connect genes with distal regulatory elements. However, the relevance of these contacts for neuronal gene transcription and the mechanisms underlying their specificity necessitate further investigation. Here, we precisely disrupt several long-range contacts in fly neurons, demonstrating their importance for megabase-range gene regulation and uncovering a hierarchical process in their formation. We further reveal an essential role for the chromosomal boundary-forming protein Cp190 in anchoring many long-range contacts, highlighting a mechanistic interplay between boundary and loop formation. Finally, we develop an unbiased proteomics-based method to systematically identify factors required for specific long-range contacts. Our findings underscore the essential role of architectural proteins such as Cp190 in cell type-specific genome organization in enabling specialized neuronal transcriptional programs.
Genomic insulators are DNA elements that prevent transcriptional activation of a promoter by an enhancer when interposed. We present a protocol for insulator-seq that enables high-throughput screening of genomic insulators using a plasmid-based massively parallel reporter assay in Drosophila cultured cells. We describe steps for insulator reporter plasmid library generation, transient transfection into cultured cells, and sequencing library preparation and provide a pipeline for data analysis.For complete details on the use and execution of this protocol, please refer to Tonelli et al.1
MAF1 is a nutrient-sensitive, TORC1-regulated repressor of RNA polymerase III (Pol III). MAF1 downregulation leads to increased lipogenesis in Drosophila melanogaster, Caenorhabditis elegans, and mice. However, Maf1−/− mice are lean as increased lipogenesis is counterbalanced by futile pre-tRNA synthesis and degradation, resulting in increased energy expenditure. We compared Chow-fed Maf1−/− mice with Chow- or High Fat (HF)-fed Maf1hep−/− mice that lack MAF1 specifically in hepatocytes. Unlike Maf1−/− mice, Maf1hep−/− mice become heavier and fattier than control mice with old age and much earlier under a HF diet. Liver ChIPseq, RNAseq and proteomics analyses indicate increased Pol III occupancy at Pol III genes, very few differences in mRNA accumulation, and protein accumulation changes consistent with increased lipogenesis. Futile pre-tRNA synthesis and degradation in the liver, as likely occurs in Maf1hep−/− mice, thus seems insufficient to counteract increased lipogenesis. Indeed, RNAseq and metabolite profiling indicate that liver phenotypes of Maf1−/− mice are strongly influenced by systemic inter-organ communication. Among common changes in the three phenotypically distinct cohorts, Angiogenin downregulation is likely linked to increased Pol III occupancy of tRNA genes in the Angiogenin promoter.
Previous studies have identified topologically associating domains (TADs) as basic units of genome organization. We present evidence of a previously unreported level of genome folding, where distant TAD pairs, megabases apart, interact to form meta-domains. Within meta-domains, gene promoters and structural intergenic elements present in distant TADs are specifically paired. The associated genes encode neuronal determinants, including those engaged in axonal guidance and adhesion. These long-range associations occur in a large fraction of neurons but support transcription in only a subset of neurons. Meta-domains are formed by diverse transcription factors that are able to pair over long and flexible distances. We present evidence that two such factors, GAF and CTCF, play direct roles in this process. The relative simplicity of higher-order meta-domain interactions in Drosophila, compared with those previously described in mammals, allowed the demonstration that genomes can fold into highly specialized cell-type-specific scaffolds that enable megabase-scale regulatory associations.
Boundaries in animal genomes delimit contact domains with enhanced internal contact frequencies and have debated functions in limiting regulatory cross-talk between domains and guiding enhancers to target promoters. Most mammalian boundaries form by stalling of chromosomal loop-extruding cohesin by CTCF, but most Drosophila boundaries form CTCF independently. However, how CTCF-independent boundaries form and function remains largely unexplored. Here, we assess genome folding and developmental gene expression in fly embryos lacking the ubiquitous boundary-associated factor Cp190. We find that sequence-specific DNA binding proteins such as CTCF and Su(Hw) directly interact with and recruit Cp190 to form most promoter-distal boundaries. Cp190 is essential for early development and prevents regulatory cross-talk between specific gene loci that pattern the embryo. Cp190 was, in contrast, dispensable for long-range enhancer-promoter communication at tested loci. Cp190 is thus currently the major player in fly boundary formation and function, revealing that diverse mechanisms evolved to partition genomes into independent regulatory domains.
Vertebrate genomes are partitioned into contact domains defined by enhanced internal contact frequency and formed by two principal mechanisms: compartmentalization of transcriptionally active and inactive domains, and stalling of chromosomal loop-extruding cohesin by CTCF bound at domain boundaries. While Drosophila has widespread contact domains and CTCF, it is currently unclear whether CTCF-dependent domains exist in flies. We genetically ablate CTCF in Drosophila and examine impacts on genome folding and transcriptional regulation in the central nervous system. We find that CTCF is required to form a small fraction of all domain boundaries, while critically controlling expression patterns of certain genes and supporting nervous system function. We also find that CTCF recruits the pervasive boundary-associated factor Cp190 to CTCF-occupied boundaries and co-regulates a subset of genes near boundaries together with Cp190. These results highlight a profound difference in CTCF-requirement for genome folding in flies and vertebrates, in which a large fraction of boundaries are CTCF-dependent and suggest that CTCF has played mutable roles in genome architecture and direct gene expression control during metazoan evolution.
Mouse liver regeneration after partial hepatectomy involves cells in the remaining tissue synchronously entering the cell division cycle. We have used this system and H3K4me3, Pol II and Pol III profiling to characterize adaptations in Pol III transcription. Our results broadly define a class of genes close to H3K4me3 and Pol II peaks, whose Pol III occupancy is high and stable, and another class, distant from Pol II peaks, whose Pol III occupancy strongly increases after partial hepatectomy. Pol III regulation in the liver thus entails both highly expressed housekeeping genes and genes whose expression can adapt to increased demand.
RNA polymerase II (Pol II) small nuclear RNA (snRNA) promoters and type 3 Pol III promoters have highly similar structures; both contain an interchangeable enhancer and "proximal sequence element" (PSE), which recruits the SNAP complex (SNAPc). The main distinguishing feature is the presence, in the type 3 promoters only, of a TATA box, which determines Pol III specificity. To understand the mechanism by which the absence or presence of a TATA box results in specific Pol recruitment, we examined how SNAPc and general transcription factors required for Pol II or Pol III transcription of SNAPc-dependent genes (i.e., TATA-box-binding protein [TBP], TFIIB, and TFIIA for Pol II transcription and TBP and BRF2 for Pol III transcription) assemble to ensure specific Pol recruitment. TFIIB and BRF2 could each, in a mutually exclusive fashion, be recruited to SNAPc. In contrast, TBP-TFIIB and TBP-BRF2 complexes were not recruited unless a TATA box was present, which allowed selective and efficient recruitment of the TBP-BRF2 complex. Thus, TBP both prevented BRF2 recruitment to Pol II promoters and enhanced BRF2 recruitment to Pol III promoters. On Pol II promoters, TBP recruitment was separate from TFIIB recruitment and enhanced by TFIIA. Our results provide a model for specific Pol recruitment at SNAPc-dependent promoters.
Initiation of gene transcription by RNA polymerase (Pol) III requires the activity of TFIIIB, a complex formed by Brf1 (or Brf2), TBP (TATA-binding protein), and Bdp1. TFIIIB is required for recruitment of Pol III and to promote the transition from a closed to an open Pol III pre-initiation complex, a process dependent on the activity of the Bdp1 subunit. Here, we present a crystal structure of a Brf2–TBP–Bdp1 complex bound to DNA at 2.7 Å resolution, integrated with single-molecule FRET analysis and in vitro biochemical assays. Our study provides a structural insight on how Bdp1 is assembled into TFIIIB complexes, reveals structural and functional similarities between Bdp1 and Pol II factors TFIIA and TFIIF, and unravels essential interactions with DNA and with the upstream factor SNAPc. Furthermore, our data support the idea of a concerted mechanism involving TFIIIB and RNA polymerase III subunits for the closed to open pre-initiation complex transition.
Overlapping gene arrangements can potentially contribute to gene expression regulation. A mammalian interspersed repeat (MIR) nested in antisense orientation within the first intron of the Polr3e gene, encoding an RNA polymerase III (Pol III) subunit, is conserved in mammals and highly occupied by Pol III. Using a fluorescence assay, CRISPR/Cas9-mediated deletion of the MIR in mouse embryonic stem cells, and chromatin immunoprecipitation assays, we show that the MIR affects Polr3e expression through transcriptional interference. Our study reveals a mechanism by which a Pol II gene can be regulated at the transcription elongation level by transcription of an embedded antisense Pol III gene.
TFIIB-related factor 2 (Brf2) is a member of the family of TFIIB-like core transcription factors. Brf2 recruits RNA polymerase (Pol) III to type III gene-external promoters, including the U6 spliceosomal RNA and selenocysteine tRNA genes. Found only in vertebrates, Brf2 has been linked to tumorigenesis but the underlying mechanisms remain elusive. We have solved crystal structures of a human Brf2-TBP complex bound to natural promoters, obtaining a detailed view of the molecular interactions occurring at Brf2-dependent Pol III promoters and highlighting the general structural and functional conservation of human Pol II and Pol III pre-initiation complexes. Surprisingly, our structural and functional studies unravel a Brf2 redox-sensing module capable of specifically regulating Pol III transcriptional output in living cells. Furthermore, we establish Brf2 as a central redox-sensing transcription factor involved in the oxidative stress pathway and provide a mechanistic model for Brf2 genetic activation in lung and breast cancer.
The genomic loci occupied by RNA polymerase (RNAP) III have been characterized in human culture cells by genome-wide chromatin immunoprecipitations, followed by deep sequencing (ChIP-seq). These studies have shown that only ∼40% of the annotated 622 human tRNA genes and pseudogenes are occupied by RNAP-III, and that these genes are often in open chromatin regions rich in active RNAP-II transcription units. We have used ChIP-seq to characterize RNAP-III-occupied loci in a differentiated tissue, the mouse liver. Our studies define the mouse liver RNAP-III-occupied loci including a conserved mammalian interspersed repeat (MIR) as a potential regulator of an RNAP-III subunit-encoding gene. They reveal that synteny relationships can be established between a number of human and mouse RNAP-III genes, and that the expression levels of these genes are significantly linked. They establish that variations within the A and B promoter boxes, as well as the strength of the terminator sequence, can strongly affect RNAP-III occupancy of tRNA genes. They reveal correlations with various genomic features that explain the observed variation of 81% of tRNA scores. In mouse liver, loci represented in the NCBI37/mm9 genome assembly that are clearly occupied by RNAP-III comprise 50 Rn5s (5S RNA) genes, 14 known non-tRNA RNAP-III genes, nine Rn4.5s (4.5S RNA) genes, and 29 SINEs. Moreover, out of the 433 annotated tRNA genes, half are occupied by RNAP-III. Transfer RNA gene expression levels reflect both an underlying genomic organization conserved in dividing human culture cells and resting mouse liver cells, and the particular promoter and terminator strengths of individual genes.
In the past decade, a number of single-molecule methods have been developed with the aim of investigating single protein and nucleic acid interactions. For the first time we use solid-state nanopore sensing to detect a single E. coli RNAP-DNA transcription complex and single E. coli RNAP enzyme. On the basis of their specific conductance translocation signature, we can discriminate and identify between those two types of molecular translocations and translocations of bare DNA. This opens up a new perspectives for investigating transcription processes at the single-molecule level.
Our view of the RNA polymerase III (Pol III) transcription machinery in mammalian cells arises mostly from studies of the RN5S (5S) gene, the Ad2 VAI gene, and the RNU6 (U6) gene, as paradigms for genes with type 1, 2, and 3 promoters. Recruitment of Pol III onto these genes requires prior binding of well-characterized transcription factors. Technical limitations in dealing with repeated genomic units, typically found at mammalian Pol III genes, have so far hampered genome-wide studies of the Pol III transcription machinery and transcriptome. We have localized, genome-wide, Pol III and some of its transcription factors. Our results reveal broad usage of the known Pol III transcription machinery and define a minimal Pol III transcriptome in dividing IMR90hTert fibroblasts. This transcriptome consists of some 500 actively transcribed genes including a few dozen candidate novel genes, of which we confirmed nine as Pol III transcription units by additional methods. It does not contain any of the microRNA genes previously described as transcribed by Pol III, but reveals two other microRNA genes, MIR886 (hsa-mir-886) and MIR1975 (RNY5, hY5, hsa-mir-1975), which are genuine Pol III transcription units.
Chemokines constitute a protein family that exhibit a variety of biological activities involved in normal and pathological physiological processes. CCL11 (eotaxin), CCL19 (MIP-3beta), CCL22 (MDC), CXCL11 (I-TAC) and CXCL12 (SDF-1alpha) chemokines, modified with the Alexa Fluor 647 fluorescent dye at specific positions along their sequence, were produced by a chemical route and their biological activities were characterized. In a migration assay, fluorescent chemokines were as biologically active as the unmodified forms. All labeled chemokines specifically stained cell lines transfected with the appropriate human chemokine receptors. The specificity of binding was further established by showing that the unlabeled ligands efficiently competed with the labeled chemokines for binding to their respective receptor. A low molecular weight antagonist of CXCR4 prevented binding of labeled CXCL12 to CXCR4 comparably to a neutralizing anti-CXCR4 antibody. Finally, labeled CCL19 was used for the staining of primary cells, illustrating that this reagent can be used for studying CCR7 expression on different cell types. Together, these results demonstrate that fluorescent synthetic chemokines constitute promising ligands for the development of chemokine receptor-binding assays on intact cells, for applications such as cell-based, high throughput screening, and studies of chemokine receptor expression by primary cells.
A recombinant rubella virus E1 (rE1) glycoprotein was produced and some of its chemical and immunological features were characterized. Two animal models were then used to establish that the rE1 glycoprotein and rubella virus particles shared antigenic and immunogenic properties. In the first one, sera from rE1 glycoprotein-immunized BALB/c mice neutralized in vitro rubella virus infection. In the second model, severe combined immune deficient (SCID) mice implanted with tonsil fragments from rubella immune donors and immunized with rE1 glycoprotein produced human anti-rubella virus antibodies. Altogether, these results showed that immunization with rE1 glycoprotein elicited neutralizing anti-rubella virus antibodies. This study thus indicated that the rE1 glycoprotein could constitute a non-replicating rubella vaccine.
Maturity-onset diabetes of the young (MODY) is a subtype of early-onset diabetes mellitus which is characterized by autosomal dominant inheritance. Several genes are known to induce MODY : HNF4A/MODY1, GCK/MODY2, TCF1/MODY3, IPF1/MODY4, TCF2/MODY5 and NEUROD1/MODY6. We studied a Swiss family with 13 diabetic patients over 3 generations. The average age at diagnosis was 35 +/- 15 years (7 subjects before 30). In addition, 2 individuals had an abnormal oral glucose tolerance. The mutation present in this family was located in the DNA binding domain of HNF4A, a strongly conserved region across almost all species, and segregated in all the MODY patients. Identification of this missense mutation allowed for presymptomatic diagnosis in the younger generations and will improve medical follow-up of the predisposed individuals.
PURPOSE:To investigate the molecular pathology underlying BIGH3-related corneal dystrophies (CDs) and to further delineate genotype-phenotype specificity.METHODS:Sixty-one index patients with CDs were subjected to phenotypic and genotypic characterization. The corneal phenotypes of all patients were assessed by biomicroscopy and documented by slit lamp photography. The BIGH3 gene was amplified exon by exon from constitutional DNA to perform single-strand conformation polymorphism (SSCP) analysis, followed by direct bidirectional sequencing of abnormal conformers.RESULTS:The phenotypes of CDs were classified as lattice CD in 30 patients, Groenouw type I in 12 (CDGGI), Avellino in 7 (CDA), Reis-Bückler in 8 (CDRB), and Thiel-Behnke in 4 (CDTB). Fifty occurrences of 16 distinct mutations were identified, including 8 novel mutations responsible for lattice type IIIA in three patients (CDLIIA), intermediate type I/IIIA (CDLI/IIIA) in four patients, and atypical CDL with deep deposits in one patient (CDL-deep).CONCLUSIONS:Disease-causing mutations were identified in 80% of the patients (50/61). All mutations localize in two regions of kerato-epithelin: the amino acid R124 and BIGH3 fasc domain 4. This study also confirms the mutation hot spot at positions R124 and R555 with nearly 50% of the mutations targeting these two amino acids (24/50). In addition the corneal phenotypes induced by changes at R124 and R555 are amino acid specific: R124C in CDLI, R555W and R124S in CDGGI, R124H in CDA, R124L in CRRB, and R555Q in CDTB. In CDLIIIA, CDLI/IIIA, and CDL-deep the genotype-phenotype correlation is domain specific, with all changes occurring at the boundary or within the fasc4 domain.