Enhancers have been proposed to act as privileged loading sites for cohesin, raising the idea that they actively fold the genome to engage distal target promoters for transcription. Supporting this idea, NIPBL/MAU2, which is required for cohesin loading, binds at enhancers in mouse embryonic stem cells. However, we find that driving cohesin recruitment near an enhancer strongly inhibits transcription from its target distal promoter, indicating that strong focal cohesin loading at enhancers is not compatible with their long-range regulatory functions. Quantitative experiments and biophysical modeling further indicate that cohesin loading at enhancers does not make major contributions to genome-wide cohesin binding and chromosome folding patterns. Instead, cohesin must load throughout the genome to extrude it, regardless of enhancer proximity, with the major determinants of cohesin traffic being extrusion barriers such as transcription and clustered CTCF sites. These findings indicate that enhancer function is largely ancillary to the general mechanisms of chromosome folding, informing further study of the relationship between genome architecture and transcriptional regulation.
Genome folding is not static but emerges from dynamic processes that control transcription, replication, recombination, and repair. DNA loop extrusion by cohesin is central to genome organization, yet it remains unclear how cells can tune extrusion kinetics to achieve precise and functional chromosome folding patterns. Here, we show that extrusion rate acts as a tunable biophysical parameter in cells, quantitatively dialed by the respective dosage of the cohesin cofactors NIPBL and PDS5. Modulation of extrusion rate can offset changes in cohesin lifetime to buffer steady-state chromosome structure and transcriptional states, even in the face of abnormal extrusion dynamics. These findings provide a long-sought mechanistic basis for the genetic interactions between cohesin cofactors and for the molecular origin of haploinsufficiency in cohesinopathies, such as Cornelia de Lange syndrome.
Three-dimensional genome organization constrains the regulatory interactions that govern vital cellular processes. Chromatin loops are key features of genome folding, yet it is unclear how genetic and epigenetic variation influences differential loop formation across individuals. Loops primarily form between two CTCF binding proteins, which recognize a specific motif at loop anchors. CTCF binding site motifs are frequently altered by base substitutions, structural variation, and 5-methylcytosine (m5C) CpG methylation, yet no study has comprehensively profiled this variation across diverse individuals. Moreover, existing approaches relying on binary loop calls fail to capture subtle changes in genetic and epigenetic features, as well as CTCF occupancy, that drive variation in loop strength. Here, we combined high-resolution Hi-C, Fiber-seq, near telomere-to-telomere phased assemblies, and m5C methylation maps across five lymphoblastoid cell lines to quantify how genetic and epigenetic variation shape genome folding. We used DiffHiC to identify 367 differential pixels and found that sequence variation, chromatin accessibility, and m5C CpG methylation are each significantly associated with differential chromatin contacts. Next, we developed CLASH (Chromatin Loop Across-sample Score Harmonizer) to harmonize loop calls across samples and enable robust comparisons of loop strengths across individuals. CLASH substantially improved loop calls and loop score calibration with respect to the classification boundary over existing methods and confirmed a significant relationship between CTCF occupancy and loop strength. We then characterized independent contributions of sequence and epigenetic variation to differential loop formation, demonstrating that 57% of sequence variation- and 40% of methylation-associated effects on loop formation acted through CTCF occupancy. Together, we present a multimodal dataset and computational approach to facilitate the study of 3D genome structure across human populations.
Enhancers are critical genetic elements controlling transcription from promoters, yet how they convey regulatory information across large genomic distances remains unclear. In this study, we engineered pluripotent stem cells in which cohesin loop extrusion can be inducibly disrupted without confounding cell cycle defects. Transcriptional dysregulation is cell type specific, and not all loci with distal enhancers depend equally on cohesin extrusion. Using comparative genome editing, we demonstrated that enhancer-promoter communication over just 20 kb can require cohesin. However, promoter-proximal elements can support long-range, cohesin-independent enhancer action-even across strong CCCTC-binding factor (CTCF) insulators. Lastly, transcriptional dynamics and the emergence of embryonic cell types remain largely robust despite disrupted extrusion. Beyond establishing strategies to study cohesin in enhancer biology, our work provides mechanistic insight into cell type specificity and genomic context specificity.
Chromatin has a complex 3D structure and diverse binding proteins that coordinate the genome's most essential functions. Many microscopy and genomics technologies that map chromatin proteins and modifications rely on the diffusion of antibodies (Abs) to target epitopes within whole nuclei. Here, we reveal a critical flaw in such methods that arises when Abs become trapped at the edge of nuclear structures and fail to reach internally positioned epitopes. This "Ab-trapping" results in artifactual peripheral signal that fundamentally distorts the apparent positions of chromatin features across the genome and nucleus. Using computational modeling and experimental validation, we demonstrate that Ab-trapping is caused by a combination of three compounding factors-high epitope abundance, high Ab affinity, and low Ab diffusion rates. Ab-trapping can thus systematically misrepresent the localization of many prevalent chromatin features like histone modifications, transcription factors, nucleolar proteins, and protein tags. We also show that this artifact manifests in multiple technologies, including immunofluorescence microscopy, more recent CUT&Tag-seq, and likely any method relying on Ab diffusion. Finally, we outline readily implementable strategies to identify and mitigate Ab-trapping. Combined, our work presents a previously unrecognized yet prevalent artifact in Ab-based chromatin mapping methods and the means to resolve it.
Mammalian genomes display complex three-dimensional organization which is crucial for downstream processes like gene regulation. Local features of genome organization are largely driven by loop extrusion and manifest as boundaries, dots, and flames in genome contact maps. Still, the rational design of DNA sequences that produce desired folding patterns has not been demonstrated. Here, we present Akita Semifreddo, a framework that enables the rational in silico design of DNA sequences with programmable 3D folding outcomes. This combines a computationally efficient "half-frozen" version of the AkitaV2 genome folding model with the Ledidi sequence optimizer. We systematically demonstrate that this approach spans the full repertoire of known local folding features. We show that ~2 kb synthetic sequences can be designed to induce boundaries, dots, and flames at desired strengths, with CTCF motif configurations consistent with their known mechanistic bases. We further demonstrate that weak boundaries can be designed through transcription-associated sequence features alone, without introducing CTCF motifs, and that strong boundaries can be suppressed by introducing SINE B2 retroelement-like sequences. Collectively, these results reveal a many-to-one relationship between DNA sequence and folding outcomes and uncover the biological basis of sequence features leveraged by our model. In short, Akita Semifreddo provides a platform for dissecting the sequence grammar of three-dimensional chromatin architecture and engineering synthetic regulatory landscapes.
Quantitative interpretation of ChIP-seq data is instrumental to derive insight into chromatin and transcription factor biology. Here we developed ChIP-FRiP, an end-to-end pipeline enabling systematic comparison of pairwise protein positioning, and applied it to the study of cohesin. In mammalian interphase, loop extruding cohesin complexes are positioned by CTCF barriers to generate locus-specific 3D genome folding patterns. Many aspects of our understanding of cohesin loop extrusion come from interpreting the amount of cohesin ChIP-seq signal at CTCF barriers, which has been reported to change variably after perturbing cohesin co-factors, such as NIPBL, PDS5A/B, and WAPL. Using ChIP-FRiP to homogeneously process 140 cohesin ChIP-seq datasets from 13 publicly available studies, we observed substantial variation attributable to technical effects, obscuring biological interpretability. To better understand how technical considerations, such as antibody specificity, influence apparent cohesin binding patterns, we integrated technical aspects of ChIP-seq into biophysical simulations of loop extrusion. Leveraging a simple biochemical model for background ChIP-seq signal, we derived a strategy to estimate and correct for the background using paired spike-in ChIP-seq data from wild-type and depletion conditions. Our results establish a framework for reliable comparative analysis, demonstrating that accurate background correction is requisite for interpreting the roles of cohesin cofactors in cohesin positioning.
Central to genome function, enhancers are non-coding sequences that can control transcription from promoters hundreds of kilobases away. Yet the physical basis of this long-range communication remains unclear. A prevalent view is that enhancers activate promoters when the two elements come into spatial proximity through the 3D folding of chromatin. However, activation by spatial proximity alone has struggled to explain several core features of enhancer function. Here, we propose that the molecular motor cohesin transmits long-range enhancer action by forming bridges between enhancers and promoters during loop extrusion. In this view, rare and transient bridges carry regulatory communication, rather than mere spatial proximity. We develop a quantitative model that predicts transcriptional output from cohesin-bridging dynamics and validate it by engineering cells in which strategically positioned CTCF sites rewire loop extrusion trajectories. The model explains how enhancer action scales with genomic distance, and how it can be either facilitated or insulated by CTCF sites across two orders of magnitude-behaviors incompatible with proximity-based models. Finally, our framework reveals that CTCF sites can block enhancers bidirectionally, by either blocking or releasing cohesin loops, resolving longstanding paradoxes between their effects on transcriptional regulation and genome folding. Together, our results establish cohesin bridging as a mode of enhancer-promoter communication that can be modulated by genomic context to achieve selective and tunable transcriptional control over long genomic distances.
Three-dimensional genome folding plays roles in gene regulation and disease. In this review, we compare and contrast recent deep learning models for predicting genome contact maps. We survey preprocessing, architecture, training, evaluation, and interpretation methods, highlighting the capabilities and limitations of different models. In each area, we highlight challenges, opportunities, and potential future directions for genome-folding models.
Changes in gene regulation were a major driver of the divergence of archaic hominins (AHs)—Neanderthals and Denisovans—and modern humans (MHs). The three-dimensional (3D) folding of the genome is critical for regulating gene expression; however, its role in recent human evolution has not been explored because the degradation of ancient samples does not permit experimental determination of AH 3D genome folding. To fill this gap, we apply novel deep learning methods for inferring 3D genome organization from DNA sequence to Neanderthal, Denisovan, and diverse MH genomes. Using the resulting 3D contact maps across the genome, we identify 167 distinct regions with diverged 3D genome organization between AHs and MHs. We show that these 3D-diverged loci are enriched for genes related to the function and morphology of the eye, supra-orbital ridges, hair, lungs, immune response, and cognition. Despite these specific diverged loci, the 3D genome of AHs and MHs is more similar than expected based on sequence divergence, suggesting that the pressure to maintain 3D genome organization constrained hominin sequence evolution. We also find that 3D genome organization constrained the landscape of AH ancestry in MHs today: regions more tolerant of 3D variation are enriched for introgression in modern Eurasians. Finally, we identify loci where modern Eurasians have inherited novel 3D genome folding from AH ancestors, which provides a putative molecular mechanism for phenotypes associated with these introgressed haplotypes. In summary, our application of deep learning to predict archaic 3D genome organization illustrates the potential of inferring molecular phenotypes from ancient DNA to reveal previously unobservable biological differences. 1 Highlights The 3D genome organization of archaic hominins can be inferred from sequence to facilitate comparisons to modern humans. Loci with 3D genome folding divergence between humans and Neanderthals highlight functional differences in the eye, supra-orbital ridges, hair, lungs, immune response, and cognition. 3D genome organization constrained recent human evolution. Tolerance to variation in 3D genome organization shaped the landscape of Neanderthal ancestry in modern humans. Neanderthal introgression contributed novel 3D genome folding patterns to Eurasians.
Antibodies (Ab) are essential for detecting specific epitopes in microscopy and genomics, but can produce artifacts leading to erroneous interpretations. Here, we characterize a novel artifact, Ab-trapping, in which antibodies bind at the periphery of a cellular structure and do not diffuse further into its interior. This causes anomalous peripheral staining for multiple critical targets, including endogenous or ectopically expressed nuclear proteins like nucleolar proteins, histone variants and their modifications like H3K9me2. Ab-trapping can affect any assay relying on Ab diffusion, including immunofluorescence microscopy and recent genomics approaches like CUT&Tag. Critically, computational modeling and experimental validation reveal that Ab-trapping is caused by high epitope abundance, high Ab affinity, and low diffusion rates. Consequently, its effects can be mitigated by using alternative Abs and optimizing incubation conditions. Ab-trapping is therefore a considerable artifact that should be considered when designing experiments and interpreting results. ### Competing Interest Statement The authors have declared no competing interest.
Cohesin drives genome organization via loop extrusion, orchestrated by the dynamic exchange of multiple essential accessory proteins. Although these regulators bind the core cohesin complex only transiently, their disruption can dramatically alter loop-extrusion dynamics and chromosome morphology. Still, a quantitative theory of cohesin regulation and its interplay with genome folding is still elusive. Here, we derive a chemical-reaction network model of loop-extrusion regulation from first principles that is fully specified by available in vivo measurements. This “bursty extrusion model” untangles the distinct roles of regulators, whose exchange coincides with intermittent periods of motor activity. By incorporating bursty extrusion in polymer simulations, we reveal how variations in regulatory protein abundance can alter chromatin architecture across length and timescales. Our results are corroborated by in vivo and in vitro observations, bridging the gap between cohesin-regulator dynamics at the molecular scale and their genome-wide consequences on chromosome organization.
Cross-species mosaic genomes provide insight into synthetic chromosome design
The spatial folding of the genome shapes gene regulation by controlling which loci interact, yet inferring the mechanisms behind these 3D structures from contact maps remains difficult. Cohesin-mediated loop extrusion is a key organizer of domains and loops, but existing methods either predict contacts without mechanistic insight or simulate extrusion with limited scalability. We present the differentiable loop extrusion model (dLEM), a scalable framework that reformulates extrusion as a smooth, trainable process. dLEM represents extrusion through position-specific velocity profiles for leftward and rightward cohesin movement. Fitting dLEM to chromosome conformation capture data yields a one-dimensional, interpretable description of extrusion dynamics that aligns with genomic and epigenomic features. dLEM parameters also capture architectural changes under CTCF and WAPL perturbations, enabling genome-wide prediction of extrusion disruptions. Extending our observations, we demonstrate that dLEM can be seamlessly incorporated into deep learning models to infer extrusion parameters directly from sequence and chromatin features, reducing model complexity by nearly three orders of magnitude while preserving predictive accuracy. Indeed, when incorporated into deep dLEM, dLEM acts as a biophysically-motivated layer for long range genomic communication, and together they provide a predictive, inter-pretable framework linking 1D genomic features to 3D chromatin folding and its response to sequence and chromatin state, with dLEM’s mechanistic modeling enabling prediction of trans factor perturbation effects. ### Competing Interest Statement The authors have declared no competing interest. National Science Foundation, NSF 2238125 National Human Genome Research Institute, NIH R01 HG 009299-6A1 National Eye Institute, NIH R01 EY 030546-01A1
Genome folding is not static, but emerges from dynamic processes that control transcription, replication, recombination, and repair. DNA loop extrusion by cohesin is central to genome organization, yet it remains unclear how cells can tune extrusion kinetics to achieve precise and functional chromosome folding patterns. Here we discover extrusion rate acts as a tunable biophysical parameter in cells, quantitatively dialed by the respective dosage of the cohesin cofactors NIPBL and PDS5. Modulation of extrusion rate can offset changes in cohesin lifetime to buffer steady-state chromosome structure and transcriptional states, even in the face of abnormal extrusion dynamics. These findings provide a long-sought mechanistic basis for the genetic interactions between cohesin cofactors and the molecular origin of haploinsufficiency in cohesinopathies, such as Cornelia de Lange syndrome.
Interphase mammalian genomes are folded in 3D with complex locus-specific patterns that impact gene regulation. CTCF (CCCTC-binding factor) is a key architectural protein that binds specific DNA sites, halts cohesin-mediated loop extrusion, and enables long-range chromatin interactions. There are hundreds of thousands of annotated CTCF-binding sites in mammalian genomes; disruptions of some result in distinct phenotypes, while others have no visible effect. Despite their importance, the determinants of which CTCF sites are necessary for genome folding and gene regulation remain unclear. Here, we update and utilize Akita, a convolutional neural network model, to extract the sequence preferences and grammar of CTCF contributing to genome folding. Our analyses of individual CTCF sites reveal four predictions: (i) only a small fraction of genomic sites are impactful; (ii) impact is highly dependent on sequences flanking the core CTCF binding motif; (iii) core and flanking nucleotides contribute largely additively to the overall impact of a site; (iv) sites created as combinations of different core and flanking sequences have impacts proportional to the product of their average impacts, i.e. they are broadly compatible. Our analysis of collections of CTCF sites make two predictions for multi-motif grammar: (i) insulation strength depends on the number of CTCF sites within a cluster, and (ii) pattern formation is governed by the orientation and spacing of these sites, rather than any inherent specialization of the CTCF motifs themselves. In sum, we present a framework for using neural network models to probe the sequences instructing genome folding and provide a number of predictions to guide future experimental inquiries.
In mammalian interphase cells, genomes are folded by cohesin loop extrusion limited by directional CTCF barriers. This process enriches cohesin at barriers, isolates neighboring topologically associating domains, and elevates contact frequency between convergent CTCF barriers across the genome. However, recent in vivo measurements present a puzzle: reported CTCF residence times on chromatin are in the range of a few minutes, whereas cohesin lifetimes are much longer. Can the observed features of genome folding result from relatively transient barriers? To address this question, we develop a dynamic barrier model, where CTCF sites switch between bound and unbound states. Using this model, we investigate how barrier dynamics would impact observables for a range of experimental genomic and imaging data sets, including ChIP-seq, Hi-C, and microscopy. We find the interplay of CTCF and cohesin binding timescales influence the strength of each of these features, leaving a signature of barrier dynamics even in the population-averaged snapshots offered by genomic data sets. First, in addition to barrier occupancy, barrier bound times are crucial for instructing features of genome folding. Second, the ratio of boundary to extruder lifetime greatly alters simulated ChIP-seq and simulated Hi-C. Third, large-scale changes in chromosome morphology observed experimentally after increasing extruder lifetime require dynamic barriers. By integrating multiple sources of experimental data, our biophysical model argues that CTCF barrier bound times effectively approach those of cohesin extruder lifetimes. Together, we demonstrate how models that are informed by biophysically measured protein dynamics broaden our understanding of genome folding.
Motivation Genomic intervals are one of the most prevalent data structures in computational genome biology, and used to represent features ranging from genes, to DNA binding sites, to disease variants. Operations on genomic intervals provide a language for asking questions about relationships between features. While there are excellent interval arithmetic tools for the command line, they are not smoothly integrated into Python, one of the most popular general-purpose computational and visualization environments. Results Bioframe is a library to enable flexible and performant operations on genomic interval dataframes in Python. Bioframe extends the Python data science stack to use cases for computational genome biology by building directly on top of two of the most commonly-used Python libraries, numpy and pandas . The bioframe API enables flexible name and column orders, and decouples operations from data formats to avoid unnecessary conversions, a common scourge for bioinformaticians. Bioframe achieves these goals while maintaining high performance and a rich set of features. Availability and implementation Bioframe is open-source under MIT license, cross-platform, and can be installed from the Python package index. The source code is maintained by Open2C on Github at https://github.com/open2c/bioframe .
Chromosome conformation capture (3C) technologies reveal the incredible complexity of genome organization. Maps of increasing size, depth, and resolution are now used to probe genome architecture across cell states, types, and organisms. Larger datasets add challenges at each step of computational analysis, from storage and memory constraints to researchers’ time; however, analysis tools that meet these increased resource demands have not kept pace. Furthermore, existing tools offer limited support for customizing analysis for specific use cases or new biology. Here we introduce cooltools ( https://github.com/open2c/cooltools ), a suite of computational tools that enables flexible, scalable, and reproducible analysis of high-resolution contact frequency data. Cooltools leverages the widely-adopted cooler format which handles storage and access for high-resolution datasets. Cooltools provides a paired command line interface (CLI) and Python application programming interface (API), which respectively facilitate workflows on high-performance computing clusters and in interactive analysis environments. In short, cooltools enables the effective use of the latest and largest genome folding datasets.
The field of 3D genome organization produces large amounts of sequencing data from Hi-C and a rapidly-expanding set of other chromosome conformation protocols (3C+). Massive and heterogeneous 3C+ data require high-performance and flexible processing of sequenced reads into contact pairs. To meet these challenges, we present pairtools–a flexible suite of tools for contact extraction from sequencing data. Pairtools provides modular command-line interface (CLI) tools that can be flexibly chained into data processing pipelines. The core operations provided by pairtools are parsing of.sam alignments into Hi-C pairs, sorting and removal of PCR duplicates. In addition, pairtools provides auxiliary tools for building feature-rich 3C+ pipelines, including contact pair manipulation, filtration, and quality control. Benchmarking pairtools against popular 3C+ data pipelines shows advantages of pairtools for high-performance and flexible 3C+ analysis. Finally, pairtools provides protocol-specific tools for restriction-based protocols, haplotype-resolved contacts, and single-cell Hi-C. The combination of CLI tools and tight integration with Python data analysis libraries makes pairtools a versatile foundation for a broad range of 3C+ pipelines.