BackgroundEmerging evidence has shown the common occupancy of dozens to hundreds of transcription factors (TFs) on cis-regulatory elements (CREs), yet the underlying details are largely unknown.ResultsIn this study, leveraging extensive collections of TF ChIP-seq data of more than 1000 TFs in human HepG2 and K562 cells, we located highly focused TF binding sites (FBSs) within CREs as single-nucleosome depleted regions, which accommodate the majority of the total TF binding events. Approximately 25,000 strong FBSs were identified in each cell type. For more than 90% of TFs, including some pioneer factors such as GATA1 and JUN, their binding sites out of FBSs barely show nucleosome depletion. Essential cellular function related motifs and phenotypically causal variants are strongly enriched in the FBSs, but not in their immediate flanking regions within CREs. Most TFs bind to FBSs not containing their canonical motifs.ConclusionOur study revealed the critical connection between highly focused TF binding and the nucleosome depleted status of DNA in vivo. Meanwhile, we constructed high-resolution maps of chromatin accessibility at distal CREs in the two human cells. We propose a model of TF co-binding in vivo and suggest that a short DNA residence time of most TFs underlies the requirement of a large number of TFs for sustained nucleosome depletion at CREs.
The Clusters of Orthologous Genes (COG) database, originally created in 1997, has been updated to reflect the constantly growing collection of completely sequenced prokaryotic genomes. This update increased the genome coverage from 1309 to 2296 species, including 2103 bacteria and 193 archaea, in most cases, with a single representative genome per genus. This set covers all genera of bacteria and archaea that included organisms with 'complete genomes' as per NCBI databases in November 2023. The number of COGs has been expanded from 4877 to 4981, primarily by including protein families involved in bacterial protein secretion. Accordingly, COG pathways and functional groups now include secretion systems of types II through X, as well as Flp/Tad and type IV pili. These groupings allow straightforward identification and examination of the prokaryotic lineages that encompass-or lack-a particular secretion system. Other developments include improved annotations for the rRNA and tRNA modification proteins, multi-domain signal transduction proteins, and some previously uncharacterized protein families. The new version of COGs is available at https://www.ncbi.nlm.nih.gov/research/COG, as well as on the NCBI FTP site https://ftp.ncbi.nlm.nih.gov/pub/COG/, which also provides archived data from previous COG releases.
Histones are key epigenetic factors that regulate the accessibility and compaction of eukaryotic genomes, affecting DNA replication and repair, and gene expression. Recent studies have demonstrated that histone missense mutations can perturb normal histone function, promoting the development of phenotypically distinguishable cancers. However, most histone mutations observed in cancer patients remain enigmatic in their potential to promote cancer development. To assess the oncogenic potential of histone missense mutations, we have gathered whole-exome sequencing data for the tumors of about 12 000 patients. Histone mutations occurred in about 16% of cancer patients, although specific cancer types showed substantially higher rates. Using genomic, structural, and biophysical analyses, we found several predominant modes of action by which histone mutations may alter function. Namely, cancer missense mutations primarily affected histone acidic patch residues and protein-binding interfaces in a cancer-specific manner and targeted interaction interfaces with specific DNA repair proteins. Consistent with this finding, we observed a high tumor mutational burden in patients with histone mutations affecting interactions with proteins involved in maintaining genome integrity. We identified potential cancer driver mutations in several histone genes, including mutations on histone H4-a highly conserved histone without previously documented driver mutations.
Histones are key epigenetic factors for regulating the accessibility and compaction of eukaryotic genomes, affecting the replication, repair, and expression of DNA. Recent studies have demonstrated that histone missense mutations can perturb normal histone function, promoting the development of phenotypically distinguishable cancers. However, most histone mutations observed in cancer patients remain enigmatic in their potential to promote cancer development. To assess the oncogenic potential of histone missense mutations, we have gathered whole-exome sequencing data for the tumors of over 12,000 patients. Overall, histone mutations occurred in about 16% of cancer patients, although specific cancer types showed substantially higher rates. Using a combination of genomic, structural, and biophysical analyses, we found several predominant modes of action, where cancer missense mutations in histones affected acidic patches and protein binding interfaces in a cancer-specific manner and targeted interaction sites with specific DNA repair proteins. Consistent with this finding, we observed a high tumour mutational burden in patients with histone mutations affecting interactions of DNA repair proteins. We also identified potential cancer driver mutations in several histone genes, including histone H4-a highly conserved histone without previously documented driver mutations. ### Competing Interest Statement The authors have declared no competing interest.
The cost and complexity of generating a complete reference genome means that many organisms lack an annotated reference. An alternative is to use a de novo reference transcriptome. This technology is cost-effective but is susceptible to off-target RNA contamination. In this manuscript, we present GTax, a taxonomy-structured database of genomic sequences that can be used with BLAST to detect and remove foreign contamination in RNA sequencing samples before assembly. In addition, we use a de novo transcriptome assembly of Solanum lycopersicum (tomato) to demonstrate that removing foreign contamination in sequencing samples reduces the number of assembled chimeric transcripts.
Mouse FOXA1 and GATA4 are prototypes of pioneer factors, initiating liver cell development by binding to the N1 nucleosome in the enhancer of the ALB1 gene. Using cryoelectron microscopy (cryo-EM), we determined the structures of the free N1 nucleosome and its complexes with FOXA1 and GATA4, both individually and in combination. We found that the DNA-binding domains of FOXA1 and GATA4 mainly recognize the linker DNA and an internal site in the nucleosome, respectively, whereas their intrinsically disordered regions interact with the acidic patch on histone H2A-H2B. FOXA1 efficiently enhances GATA4 binding by repositioning the N1 nucleosome. In vivo DNA editing and bioinformatics analyses suggest that the co-binding mode of FOXA1 and GATA4 plays important roles in regulating genes involved in liver cell functions. Our results reveal the mechanism whereby FOXA1 and GATA4 cooperatively bind to the nucleosome through nucleosome repositioning, opening chromatin by bending linker DNA and obstructing nucleosome packing.
Wrapping of DNA into nucleosomes restricts accessibility to DNA and may affect the recognition of binding motifs by transcription factors. A certain class of transcription factors, the pioneer transcription factors, can specifically recognize their DNA binding sites on nucleosomes, initiate local chromatin opening, and facilitate the binding of co-factors in a cell-type-specific manner. For the majority of human pioneer transcription factors, the locations of their binding sites, mechanisms of binding, and regulation remain unknown. We have developed a computational method to predict the cell-type-specific ability of transcription factors to bind nucleosomes by integrating ChIP-seq, MNase-seq, and DNase-seq data with details of nucleosome structure. We have demonstrated the ability of our approach in discriminating pioneer from canonical transcription factors and predicted new potential pioneer transcription factors in H1, K562, HepG2, and HeLa-S3 cell lines. Last, we systematically analyzed the interaction modes between various pioneer transcription factors and detected several clusters of distinctive binding sites on nucleosomal DNA.
Housekeeping genes are considered to be regulated by common enhancers across different tissues. Here we report that most of the commonly expressed mouse or human genes across different cell types, including more than half of the previously identified housekeeping genes, are associated with cell type-specific enhancers. Furthermore, the binding of most transcription factors (TFs) is cell type-specific. We reason that these cell type specificities are causally related to the collective TF recruitment at regulatory sites, as TFs tend to bind to regions associated with many other TFs and each cell type has a unique repertoire of expressed TFs. Based on binding profiles of hundreds of TFs from HepG2, K562, and GM12878 cells, we show that 80% of all TF peaks overlapping H3K27ac signals are in the top 20,000-23,000 most TF-enriched H3K27ac peak regions, and approximately 12,000-15,000 of these peaks are enhancers (nonpromoters). Those enhancers are mainly cell type-specific and include those linked to the majority of commonly expressed genes. Moreover, we show that the top 15,000 most TF-enriched regulatory sites in HepG2 cells, associated with about 200 TFs, can be predicted largely from the binding profile of as few as 30 TFs. Through motif analysis, we show that major enhancers harbor diverse and clustered motifs from a combination of available TFs uniquely present in each cell type. We propose a mechanism that explains how the highly focused TF binding at regulatory sites results in cell type specificity of enhancers for housekeeping and commonly expressed genes.
The Mediator complex is central to transcription by RNA polymerase II (Pol II) in eukaryotes. In budding yeast (Saccharomyces cerevisiae), Mediator is recruited by activators and associates with core promoter regions, where it facilitates preinitiation complex (PIC) assembly, only transiently before Pol II escape. Interruption of the transcription cycle by inactivation or depletion of Kin28 inhibits Pol II escape and stabilizes this association. However, Mediator occupancy and dynamics have not been examined on a genome-wide scale in yeast grown in nonstandard conditions. Here we investigate Mediator occupancy following heat shock or CdCl2 exposure, with and without depletion of Kin28. We find that Pol II occupancy shows similar dependence on Mediator under normal and heat shock conditions. However, although Mediator association increases at many genes upon Kin28 depletion under standard growth conditions, little or no increase is observed at most genes upon heat shock, indicating a more stable association of Mediator after heat shock. Unexpectedly, Mediator remains associated upstream of the core promoter at genes repressed by heat shock or CdCl2 exposure whether or not Kin28 is depleted, suggesting that Mediator is recruited by activators but is unable to engage PIC components at these repressed targets. This persistent association is strongest at promoters that bind the HMGB family member Hmo1, and is reduced but not eliminated in hmo1Δ yeast. Finally, we show a reduced dependence on PIC components for Mediator occupancy at promoters after heat shock, further supporting altered dynamics or stronger engagement with activators under these conditions.
Additional file 3: Table S1. Proteins—HMGN1 and HMGN2 partners in both cell types.
Histone tails harbor a vast number of post-translational modifications which are recognized by regulatory proteins to mediate epigenetic processes. Many recurrent mutations in histones have been identified which drive cancer development, and their effects are not well characterized. Here, we investigate the effects of post-translational modifications and mutations on histone tails using molecular dynamics simulations in conjunction with bioinformatics approaches. We have compiled a dataset of histone mutations from >10,000 tumor samples and predicted the potential oncogenic mutations.
Multiple next-generation-sequencing (NGS)-based studies are enabled by the availability of a reference genome of the target organism. Unfortunately, several organisms remain unannotated due to the cost and complexity of generating a complete (or close to complete) reference genome. These unannotated organisms, however, can also be studied if a de novo reference transcriptome is assembled from whole transcriptome sequencing experiments. This technology is cost effective and widely used but is susceptible to off-target RNA contamination. In this manuscript, we present GTax, a taxonomy structured database of genomic sequences that can be used with BLAST to detect and remove foreign contamination in RNA sequencing samples before assembly. In addition, we investigate the effect of foreign RNA contamination on a de novo transcriptome assembly of Solanum lycopersicum (tomato). Our study demonstrates that removing foreign contamination in sequencing samples reduces the number of assembled chimeric transcripts.
Background Nucleosomal binding proteins, HMGN, is a family of chromatin architectural proteins that are expressed in all vertebrate nuclei. Although previous studies have discovered that HMGN proteins have important roles in gene regulation and chromatin accessibility, whether and how HMGN proteins affect higher order chromatin status remains unknown. Results We examined the roles that HMGN1 and HMGN2 proteins play in higher order chromatin structures in three different cell types. We interrogated data generated in situ, using several techniques, including Hi-C, Promoter Capture Hi-C, ChIP-seq, and ChIP-MS. Our results show that HMGN proteins occupy the A compartment in the 3D nucleus space. In particular, HMGN proteins occupy genomic regions involved in cell-type-specific long-range promoter-enhancer interactions. Interestingly, depletion of HMGN proteins in the three different cell types does not cause structural changes in higher order chromatin, i.e., in topologically associated domains (TADs) and in A/B compartment scores. Using ChIP-seq combined with mass spectrometry, we discovered protein partners that are directly associated with or neighbors of HMGNs on nucleosomes. Conclusions We determined how HMGN chromatin architectural proteins are positioned within a 3D nucleus space, including the identification of their binding partners in mononucleosomes. Our research indicates that HMGN proteins localize to active chromatin compartments but do not have major effects on 3D higher order chromatin structure and that their binding to chromatin is not dependent on specific protein partners.
Cytosine methylation at the 5-carbon position is an essential DNA epigenetic mark in many eukaryotic organisms. Although countless structural and functional studies of cytosine methylation have been reported in both prokaryotes and eukaryotes, our understanding of how it influences the nucleosome assembly, structure, and dynamics remains obscure. Here we investigated the effects of cytosine methylation at CpG sites on nucleosome dynamics and stability. By applying long molecular dynamics simulations (five microsecond long trajectories, 60 microseconds in total), we generated extensive atomic level conformational full nucleosome ensembles. Our results revealed that methylation induces pronounced changes in geometry for both linker and nucleosomal DNA, leading to a more curved, under-twisted DNA, shifting the population equilibrium of sugar-phosphate backbone geometry. These conformational changes are associated with a considerable enhancement of interactions between methylated DNA and the histone octamer, doubling the number of contacts at some key arginines. H2A and H3 tails play important roles in these interactions, especially for DNA methylated nucleosomes. This, in turn, prevents a spontaneous DNA unwrapping of 3-4 helical turns for the methylated nucleosome with truncated histone tails, otherwise observed in the unmethylated system on several microsecond time scale.
BACKGROUND:Disproportionate risks of COVID-19 in congregate care facilities including long-term care homes, retirement homes, and shelters both affect and are affected by SARS-CoV-2 infections among facility staff. In cities across Canada, there has been a consistent trend of geographic clustering of COVID-19 cases. However, there is limited information on how COVID-19 among facility staff reflects urban neighborhood disparities, particularly when stratified by the social and structural determinants of community-level transmission.OBJECTIVE:This study aimed to compare the concentration of cumulative cases by geography and social and structural determinants across 3 mutually exclusive subgroups in the Greater Toronto Area (population: 7.1 million): community, facility staff, and health care workers (HCWs) in other settings.METHODS:We conducted a retrospective, observational study using surveillance data on laboratory-confirmed COVID-19 cases (January 23 to December 13, 2020; prior to vaccination rollout). We derived neighborhood-level social and structural determinants from census data and generated Lorenz curves, Gini coefficients, and the Hoover index to visualize and quantify inequalities in cases.RESULTS:The hardest-hit neighborhoods (comprising 20% of the population) accounted for 53.87% (44,937/83,419) of community cases, 48.59% (2356/4849) of facility staff cases, and 42.34% (1669/3942) of other HCW cases. Compared with other HCWs, cases among facility staff reflected the distribution of community cases more closely. Cases among facility staff reflected greater social and structural inequalities (larger Gini coefficients) than those of other HCWs across all determinants. Facility staff cases were also more likely than community cases to be concentrated in lower-income neighborhoods (Gini 0.24, 95% CI 0.15-0.38 vs 0.14, 95% CI 0.08-0.21) with a higher household density (Gini 0.23, 95% CI 0.17-0.29 vs 0.17, 95% CI 0.12-0.22) and with a greater proportion working in other essential services (Gini 0.29, 95% CI 0.21-0.40 vs 0.22, 95% CI 0.17-0.28).CONCLUSIONS:COVID-19 cases among facility staff largely reflect neighborhood-level heterogeneity and disparities, even more so than cases among other HCWs. The findings signal the importance of interventions prioritized and tailored to the home geographies of facility staff in addition to workplace measures, including prioritization and reach of vaccination at home (neighborhood level) and at work.
The integrity of histone genes is a cautiously guarded attribute in most human cells, demonstrating the integral role that they play in cellular function. Not only are histones capable of allowing the compaction of DNA within the nucleus, but they are also involved in DNA replication, repair and expression. As histones are indispensable for cellular function, mutations that alter histone structure can prove detrimental to cell viability. Recent publications have demonstrated that some histone mutations may be capable of promoting the development of phenotypically distinguishable cancers.
Histones have a long history of research in a wide range of species, leaving a legacy of complex nomenclature in the literature. Community-led discussions at the EMBO Workshop on Histone Variants in 2011 resulted in agreement amongst experts on a revised systematic protein nomenclature for histones, which is based on a combination of phylogenetic classification and historical symbol usage. Human and mouse histone gene symbols previously followed a genome-centric system that was not applicable across all vertebrate species and did not reflect the systematic histone protein nomenclature. This prompted a collaboration between histone experts, the Human Genome Organization (HUGO) Gene Nomenclature Committee (HGNC) and Mouse Genomic Nomenclature Committee (MGNC) to revise human and mouse histone gene nomenclature aiming, where possible, to follow the new protein nomenclature whilst conforming to the guidelines for vertebrate gene naming. The updated nomenclature has also been applied to orthologous histone genes in chimpanzee, rhesus macaque, dog, cat, pig, horse and cattle, and can serve as a framework for naming other vertebrate histone genes in the future.
An essential questions of gene regulation is how large number of enhancers and promoters organize into gene regulatory loops. Using transcription-factor binding enrichment as an indicator of enhancer strength, we identified a portion of H3K27ac peaks as potentially strong enhancers and found a universal pattern of promoter and enhancer distribution: At actively transcribed regions of length of ∼200-300 kb, the numbers of active promoters and enhancers are inversely related. Enhancer clusters are associated with isolated active promoters, regardless of the gene's cell-type specificity. As the number of nearby active promoters increases, the number of enhancers decreases. At regions where multiple active genes are closely located, there are few distant enhancers. With Hi-C analysis, we demonstrate that the interactions among the regulatory elements (active promoters and enhancers) occur predominantly in clusters and multiway among linearly close elements and the distance between adjacent elements shows a preference of ∼30 kb. We propose a simple rule of spatial organization of active promoters and enhancers: Gene transcriptions and regulations mainly occur at local active transcription hubs contributed dynamically by multiple elements from linearly close enhancers and/or active promoters. The hub model can be represented with a flower-shaped structure and implies an enhancer-like role of active promoters.