Strong evolutionary selection has maintained CpG-dense islands (CGIs) at the promoters of constitutively expressed genes throughout the vertebrate genome, suggesting an important role in regulating DNA topology. Here, using Twist-seq, a psoralen-based approach for quantitative genome-wide profiling of DNA supercoiling, we reveal distinct topological states across human gene promoters. We show that CGI promoters accumulate elevated levels of negative supercoiling relative to non-CGI promoters and define localised topological domains at highly transcribed genes. Integrating genome-wide analyses with reaction-diffusion modelling and coarse-grained molecular dynamics simulations, we find that this behaviour is encoded by the intrinsic physical properties of CGI DNA. The GC-rich sequence context promotes nucleosome depletion and focuses torsional stress onto embedded AT-rich pockets, driving localised DNA melting and plectoneme-tip bubble formation within promoter-proximal nucleosome-free regions. This provides an energetically favourable pathway for redistributing transcription-induced torsional stress through transient strand separation and writhe, consistent with increased ssDNA formation at CGI promoters observed by ssDNA-seq. We propose that CGIs function as sequence-encoded topological sinks that buffer supercoiling while maintaining a promoter architecture permissive for transcription initiation, thereby preserving promoter integrity and genome stability.
CENP-B, a centromeric protein known for its role in binding the B box sequence of centromeric DNA, has long been recognized as important, though not essential, for kinetochore attachment and chromosome segregation. Here, we identify an unexpected, non-centromeric role for CENP-B. We demonstrate that CENP-B binds to specific non-centromeric sites along chromosome arms, predominantly at promoters, and depletion of CENP-B leads to dysregulated gene expression. Binding is enriched in G2 phase cells and, importantly, occurs independently of the canonical B box motif. Instead, CENP-B binding in chromosome arms is defined by regions of negatively supercoiled DNA containing repetitive sequences, such as multiple CCAAT boxes, that are prone to forming secondary structures. Consistently, we find that CENP-B binds to hairpin DNA in vitro via its DNA binding domain. The chromosome arm binding pattern is conserved across cell types and is particularly prominent in the promoters of transcriptionally active replication-dependent histone genes. These findings reveal a previously unrecognized centromere-independent binding activity of CENP-B.
Abstract Cell-to-cell transcriptional heterogeneity, or noise, is an intrinsic property of the transcriptome with implications for development, disease progression, and aging. Bulk RNA-seq masks this variability by averaging gene expression across cells, whereas single-cell RNA sequencing (scRNA-seq) resolves it. Nevertheless, separating biological noise from technical variance remains challenging, particularly across platforms with different chemistries. We benchmarked two widely adopted technologies, Evercode WT (SPLiT-seq, Parse Biosciences) and Chromium (10x Genomics), on human lymphoblastoid nuclei. Evercode WT achieved targeted sequencing depth and nuclei number far more reliably, and its random-hexamer priming yielded more intronic reads and non-coding RNA genes; Chromium recovered more cells and detected polyadenylated transcripts and cell-line markers more sensitively. Despite these opposing biases, the platforms showed comparable gene detection and strongly correlated expression profiles. Using datasets from both platforms, we defined a noise metric detrended from mean expression and showed that per-gene estimates were reproducible across chemistries. Noise was lower in G2M than in G1 and was most strongly associated with gene length rather than exonic length. Expression of genes with CpG-island promoters was less variable than that of those without. This study establishes a platform-independent basis for quantifying transcriptional noise and a framework for selecting an appropriate scRNA-seq platform. Highlights Parse Evercode WT demonstrates superior predictability in targeted cell recovery and sequencing depth estimations compared to 10x Chromium. Distinct biases: Parse captures intronic sequences; 10x targets polyadenylated mRNA. Detrended transcriptional noise is reproducible across both barcoding chemistries. Noise scales with gene length, not exonic length, and is lower at CGI promoters.
Lamina-associated domains (LADs) are megabase-sized genomic regions anchored to the nuclear lamina (NL). Factors controlling the interactions of the genome with the NL have largely remained elusive. Here, we identified DNA topoisomerase 2 beta (TOP2B) as a regulator of these interactions. TOP2B binds predominantly to inter-LAD (iLAD) chromatin and its depletion results in a partial loss of genomic partitioning between LADs and iLADs, suggesting that this enzyme might protect specific iLADs from interacting with the NL. TOP2B depletion affects LAD interactions with lamin B receptor (LBR) more than with lamins. LBR depletion phenocopies the effects of TOP2B depletion, despite the different positioning of the two proteins in the genome. This suggests a complementary mechanism for organizing the genome at the NL. Indeed, co-depletion of TOP2B and LBR causes partial LAD/iLAD inversion, reflecting changes typical of oncogene-induced senescence. We propose that a coordinated axis controlled by TOP2B in iLADs and LBR in LADs maintains the partitioning of the genome between the NL and the nuclear interior.
Autophagy is a conserved cellular degradation process. While autophagy-related proteins were shown to influence the signaling and trafficking of some receptor tyrosine kinases, the relevance of this during cancer development is unclear. Here, we identify a role for autophagy in regulating platelet-derived growth factor receptor alpha (PDGFRA) signaling and levels. We find that PDGFRA can be targeted for autophagic degradation through the activity of the autophagy cargo receptor p62. As a result, short-term autophagy inhibition leads to elevated levels of PDGFRA but an unexpected defect in PDGFA-mediated signaling due to perturbed receptor trafficking. Defective PDGFRA signaling led to its reduced levels during prolonged autophagy inhibition, suggesting a mechanism of adaptation. Importantly, PDGFA-driven gliomagenesis in mice was disrupted when autophagy was inhibited in a manner dependent on Pten status, thus highlighting a genotype-specific role for autophagy during tumorigenesis. In summary, our data provide a mechanism by which cells require autophagy to drive tumor formation.
Maintaining chromatin integrity at the repetitive non-coding DNA sequences underlying centromeres is crucial to prevent replicative stress, DNA breaks and genomic instability. The concerted action of transcriptional repressors, chromatin remodelling complexes and epigenetic factors controls transcription and chromatin structure in these regions. The histone chaperone complex ATRX/DAXX is involved in the establishment and maintenance of centromeric chromatin through the deposition of the histone variant H3.3. ATRX and DAXX have also evolved mutually-independent functions in transcription and chromatin dynamics. Here, using paediatric glioma and pancreatic neuroendocrine tumor cell lines, we identify a novel ATRX-independent function for DAXX in promoting genome stability by preventing transcription-associated R-loop accumulation and DNA double-strand break formation at centromeres. This function of DAXX required its interaction with histone H3.3 but was independent of H3.3 deposition and did not reflect a role in the repression of centromeric transcription. DAXX depletion mobilized BRCA1 at centromeres, in line with BRCA1 role in counteracting centromeric R-loop accumulation. Our results provide novel insights into the mechanisms protecting the human genome from chromosomal instability, as well as potential perspectives in the treatment of cancers with DAXX alterations.
Classical observations suggest a connection between 3D gene structure and function, but testing this hypothesis has been challenging due to technical limitations. To explore this, we developed epigenetic highly predictive heteromorphic polymer (e-HiP-HoP), a model based on genome organization principles to predict the 3D structure of human chromatin. We defined a new 3D structural unit, a “topos,” which represents the regulatory landscape around gene promoters. Using GM12878 cells, we predicted the 3D structure of over 10,000 active gene topoi and stored them in the 3DGene database. Data mining revealed folding motifs and their link to Gene Ontology features. We computed a structural diversity score and identified influential nodes—chromatin sites that frequently interact with gene promoters, acting as key regulators. These nodes drive structural diversity and are tied to gene function. e-HiP-HoP provides a framework for modeling high-resolution chromatin structure and a mechanistic basis for chromatin contact networks that link 3D gene structure with function.
Human centromeres appear as constrictions on mitotic chromosomes and form a platform for kinetochore assembly in mitosis. Biophysical experiments led to a suggestion that repetitive DNA at centromeric regions form a compact scaffold necessary for function, but this was revised when neocentromeres were discovered on non-repetitive DNA. To test whether centromeres have a special chromatin structure we have analysed the architecture of a neocentromere. Centromere formation is accompanied by RNA pol II recruitment and active transcription to form a decompacted, negatively supercoiled domain enriched in open chromatin fibres. In contrast, centromerisation causes a spreading of repressive epigenetic marks to surrounding regions, delimited by H3K27me3 polycomb boundaries and divergent genes. This flanking domain is transcriptionally silent and partially remodelled to form compact chromatin, similar to satellite-containing DNA sequences, and exhibits genomic instability. We suggest transcription disrupts chromatin to provide a foundation for kinetochore formation whilst compact pericentromeric heterochromatin generates mechanical rigidity.
Classical observations have long suggested there is a link between 3D gene structure and transcription1–4. However, due to the many factors regulating gene expression, and to the challenge of visualizing DNA and chromatin dynamics at the same time in living cells, this hypothesis has been difficult to quantitatively test experimentally. Here we take an orthogonal approach and use computer simulations, based on the known biophysical principles of genome organisation5–7, to simultaneously predict 3D structure and transcriptional output of human chromatin genome wide. We validate our model by quantitative comparison with Hi-C contact maps, FISH, GRO-seq and single-cell RNA-seq data, and provide the 3DGene resource to visualise the panoply of structures adopted by any active human gene in a population of cells. We find transcription strongly correlates with the formation of protein-mediated microphase separated clusters of promoters and enhancers, associated with clouds of chromatin loops, and show that gene noise is a consequence of structural heterogeneity. Our results also indicate that loop extrusion by cohesin does not affect average transcriptional patterns, but instead impacts transcriptional noise. These findings provide a functional role for intranuclear microphase separation, and an evolutionary mechanism for loop extrusion halted at CTCF sites, to modulate transcriptional noise.
The ARID1A subunit of SWI/SNF chromatin remodeling complexes is a potent tumor suppressor. Here, a degron is applied to detect rapid loss of chromatin accessibility at thousands of loci where ARID1A acts to generate accessible minidomains of nucleosomes. Loss of ARID1A also results in the redistribution of the coactivator EP300. Co-incident EP300 dissociation and lost chromatin accessibility at enhancer elements are highly enriched adjacent to rapidly downregulated genes. In contrast, sites of gained EP300 occupancy are linked to genes that are transcriptionally upregulated. These chromatin changes are associated with a small number of genes that are differentially expressed in the first hours following loss of ARID1A. Indirect or adaptive changes dominate the transcriptome following growth for days after loss of ARID1A and result in strong engagement with cancer pathways. The identification of this hierarchy suggests sites for intervention in ARID1A-driven diseases.
Human centromeric chromatin is assembled from CENP-A nucleosomes1 and repetitive α-satellite DNA sequences2 and provides a foundation for kinetochore assembly in mitosis. Biophysical experiments led to a hypothesis that the repetitive DNA sequences form a highly folded chromatin scaffold necessary for function, but this idea was revised when fully functional evolutionary new centromeres (ENCs) or neocentromeres were found to form on non-repetitive DNA. To understand if centromeres have a special chromatin structure we have genetically isolated a single human chromosome harbouring a neocentromere and investigated its organisation. The centromere core is enriched in RNA pol II, active epigenetic marks and remodelled by transcription to form a negatively supercoiled ‘open’ chromatin domain. In contrast, centromerisation causes a spreading of repressive epigenetic marks to flanking regions, delimited by H3K27me3 polycomb boundaries and divergent genes. The flanking domain is partially remodelled to form ‘compact’ chromatin, with characteristics similar to satellite-containing pericentromeric chromatin, but exhibits low level genomic instability. We provide a model for centromere chromatin structure and suggest that open chromatin provides a foundation for a stable kinetochore whilst pericentromeric heterochromatin generates surrounding mechanical rigidity.
Cells coordinate interphase-to-mitosis transition, but recurrent cytogenetic lesions appear at common fragile sites (CFSs), termed CFS expression, in a tissue-specific manner after replication stress, marking regions of instability in cancer. Despite such a distinct defect, no model fully provides a molecular explanation for CFSs. We show that CFSs are characterized by impaired chromatin folding, manifesting as disrupted mitotic structures visible with molecular fluorescence in situ hybridization (FISH) probes in the presence and absence of replication stress. Chromosome condensation assays reveal that compaction-resistant chromatin lesions persist at CFSs throughout the cell cycle and mitosis. Cytogenetic and molecular lesions are marked by faulty condensin loading at CFSs, a defect in condensin-I-mediated compaction, and are coincident with mitotic DNA synthesis (MIDAS). This model suggests that, in conditions of exogenous replication stress, aberrant condensin loading leads to molecular defects and CFS expression, concomitantly providing an environment for MIDAS, which, if not resolved, results in chromosome instability.
Centromeres are highly specialized genomic loci that function during mitosis to maintain genome stability. Formed primarily on repetitive α-satellite DNA sequence characterisation of native centromeric chromatin structure has remained challenging. Fortuitously, neocentromeres are formed on a unique DNA sequence and represent an excellent model to interrogate centromeric chromatin structure. This review uncovers the specific findings from independent neocentromere studies that have advanced our understanding of canonical centromere chromatin structure.
Interphase to mitosis transition is integral to cellular life and requires extensive large-scale chromatin remodelling. Cells coordinate compaction but recurrent cytogenetic lesions appear at common fragile sites (CFSs) in a tissue-specific manner following replication stress, marking regions of genome instability in cancer. Current views suggest lesion formation is dependent on late replication, but instead we report that replication stress causes changes in replication dynamics and origin firing efficiency. We show that CFSs are characterised by impaired compaction, manifested either as cytogenetic abnormalities or as disrupted mitotic figures on cytogenetically normal chromosomes. Chromosome condensation assays reveal that compaction-resistant chromatin lesions persist at CFSs throughout the cell cycle, due to faulty condensin loading at CFS region and a defect in condensin I mediated compaction. Our data supports a model for CFS formation where aberrant replication dynamics lead to faulty condensin I recruitment and mitotic misfolding, mitotic DNA synthesis, and subsequent chromosomal instability.
Higher eukaryotic chromosomes are organized into topologically constrained functional domains; however, the molecular mechanisms required to sustain these complex interphase chromatin structures are unknown. A stable matrix underpinning nuclear organization was hypothesized, but the idea was abandoned as more dynamic models of chromatin behavior became prevalent. Here, we report that scaffold attachment factor A (SAF-A), originally identified as a structural nuclear protein, interacts with chromatin-associated RNAs (caRNAs) via its RGG domain to regulate human interphase chromatin structures in a transcription-dependent manner. Mechanistically, this is dependent on SAF-A’s AAA+ ATPase domain, which mediates cycles of protein oligomerization with caRNAs, in response to ATP binding and hydrolysis. SAF-A oligomerization decompacts large-scale chromatin structure while SAF-A loss or monomerization promotes aberrant chromosome folding and accumulation of genome damage. Our results show that SAF-A and caRNAs form a dynamic, transcriptionally responsive chromatin mesh that organizes large-scale chromosome structures and protects the genome from instability.
Transitions in DNA structure have the capacity to regulate genes, but have been poorly characterised in eukaryotes due to a lack of appropriate techniques. One important example is DNA supercoiling, which can directly regulate transcription initiation, elongation and coordinated expression of neighbouring genes. DNA supercoiling is the over- or under-winding of the DNA double helix, which occurs as a consequence of polymerase activity and is modulated by topoisomerase activity [5]. To map the distribution of DNA supercoiling in nuclei, we developed biotinylated 4,5,8-trimethylpsoralen (bTMP) pull-down to preferentially enrich for under-wound DNA. Here we describe in detail the experimental design, quality controls and analyses associated with the study by Naughton et al. [13] that characterised for the first time the large-scale distribution of DNA supercoiling in human cells (GEO: GSE43488 and GSE43450GSE43488GSE43450).
New approaches using biotinylated-psoralen as a probe for investigating DNA structure have revealed new insights into the relationship between DNA supercoiling, transcription and chromatin compaction. We explore a hypothesis that divergent RNA transcription generates negative supercoiling at promoters facilitating initiation complex formation and subsequent promoter clearance.
New approaches using biotinylated-psoralen as a probe for investigating DNA structure have revealed new insights into the relationship between DNA supercoiling, transcription and chromatin compaction. We explore a hypothesis that divergent RNA transcription generates negative supercoiling at promoters facilitating initiation complex formation and subsequent promoter clearance.
DNA supercoiling is an inherent consequence of twisting DNA and is critical for regulating gene expression and DNA replication. However, DNA supercoiling at a genomic scale in human cells is uncharacterized. To map supercoiling, we used biotinylated trimethylpsoralen as a DNA structure probe to show that the human genome is organized into supercoiling domains. Domains are formed and remodeled by RNA polymerase and topoisomerase activities and are flanked by GC-AT boundaries and CTCF insulator protein-binding sites. Underwound domains are transcriptionally active and enriched in topoisomerase I, 'open' chromatin fibers and DNase I sites, but they are depleted of topoisomerase II. Furthermore, DNA supercoiling affects additional levels of chromatin compaction as underwound domains are cytologically decondensed, topologically constrained and decompacted by transcription of short RNAs. We suggest that supercoiling domains create a topological environment that facilitates gene activation, providing an evolutionary purpose for clustering genes along chromosomes.