The structure and dynamics of individual chromosomes are significantly modulated by their local environment within the nucleus. It has long been established that heterochromatin (B compartment loci) associates with the lamina meshwork that lines the inner nuclear envelope. However, the extent to which these interactions perturb the chromosomal organization within a territory remains unclear. Using computer simulations and published imaging data, we further characterize the interaction between chromatin and the lamina and find striking consequences of lamina-association on the chromosomal structural ensemble. We find that lamina-associating chromosomes have their compartmentalization polarized in a direction perpendicular to the nuclear surface. Further, lamina-associated chromosomes exhibit shorter average distances between euchromatin loci (A compartment) than chromosomes without lamina contact. Energy landscape analysis of our simulations reveal that the sequestration of heterochromatin to the lamina allows for euchromatin to form more spatial contacts with other segments of euchromatin. The compaction of euchromatin due to lamina-association also diminishes the mobility of these regions. These findings suggest a mechanism for bringing together euchromatic segments that are separated by large genomic distances within a chromosome, potentially to share transcriptional machinery and enhance transcription for large chromosomes, which are often found at the nuclear periphery.
DNA-binding response regulators (DBRRs) are a broad class of proteins that operate in tandem with their partner kinase proteins to form two-component signal transduction systems in bacteria. Typical DBRRs are composed of two domains where the conserved N-terminal domain accepts transduced signals and the evolutionarily diverse C-terminal domain binds to DNA. These domains are assumed to be functionally independent, and hence recombination of the two domains should yield novel DBRRs of arbitrary input/output response, which can be used as biosensors. This idea has been proved to be successful in some cases; yet, the error rate is not trivial. Improvement of the success rate of this technique requires a deeper understanding of the linker-domain and inter-domain residue interactions, which have not yet been thoroughly examined. Here, we studied residue coevolution of DBRRs of the two main subfamilies (OmpR and NarL) using large collections of bacterial amino acid sequences to extensively investigate the evolutionary signatures of linker-domain and inter-domain residue interactions. Coevolutionary analysis uncovered evolutionarily selected linker-domain and inter-domain residue interactions of known experimental structures, as well as previously unknown inter-domain residue interactions. We examined the possibility of these inter-domain residue interactions as contacts that stabilize an inactive conformation of the DBRR where DNA binding is inhibited for both subfamilies. The newly gained insights on linker-domain/inter-domain residue interactions and shared inactivation mechanisms improve the understanding of the functional mechanism of DBRRs, providing clues to efficiently create functional DBRR-based biosensors. Additionally, we show the feasibility of applying coevolutionary landscape models to predict the functionality of domain-swapped DBRR proteins. The presented result demonstrates that sequence information can be used to filter out bioengineered DBRR proteins that are predicted to be nonfunctional due to a high negative predictive value.
Transcription factors (TFs) control the rate of transcription of genetic information by binding to specific DNA sequences. The time needed for a TF to find its specific target sites is a bottleneck to the genetic response mechanism. While TF target site search is a well studied problem, the effect of genome 3D architecture on the TF target search times is poorly understood. Here, we use accurate and cell-specific 3D structural ensembles of human chromosomes to investigate how the spatial organization of binding sites on chromosomes influences the dynamics of TFs. We use Chromatin Immuno-Precipitation data to map the position of binding sites for several TFs on chromosomal structures and simulate the dynamics of individual TF within chromosomal territories. We find that the distribution of binding sites along chromosomes cooperates with the 3D folding of the chromatin fiber to induce dynamics in which TFs tend to visit sites distributed sequentially along the genome. In this way, genome 3D architecture appears to reduce the time each TF spends in the unbound state while commuting from one target site to the other. At the same time, genome 3D architecture further reduces the flux of TFs between binding sites already well separated along the genome, effectively isolating distant clusters of binding sites. We compare the TF traffic patterns generated by the 3D structures of human chromosomes with those generated by several alternative structural models characterized by increasing randomness. Finally, we study the effect of lengthwise compaction and phase separation, known architectural features of the human genome, in TFs target search. In short, our analysis demonstrates that genome architecture regulates the traffic of TFs within chromosomal territories and reduces the time each TF spends commuting between binding sites.
Chromosomes endure mechanical stresses throughout the cell cycle; for example, resulting from the pulling of chromosomes by spindle fibers during mitosis or deformation of the nucleus during cell migration. The response to physical stress is closely related to chromosome structure and function. Micromechanical studies of mitotic chromosomes have revealed them to be remarkably extensible objects and informed early models of mitotic chromosome organization. We use a data-driven, coarse-grained polymer modeling approach to explore the relationship between the spatial organization of individual chromo-somes and their emergent mechanical properties. In particular, we investigate the mechanical properties of our model chromo-somes by axially stretching them. Simulated stretching led to a linear force-extension curve for small strain, with mitotic chromosomes behaving about 10-fold stiffer than interphase chromosomes. Studying their relaxation dynamics, we found that chromosomes are viscoelastic solids with a highly liquid-like, viscous behavior in interphase that becomes solid-like in mitosis. This emergent mechanical stiffness originates from lengthwise compaction, an effective potential capturing the activity of loop-extruding SMC complexes. Chromosomes denature under large strains via unraveling, which is characterized by open-ing of large-scale folding patterns. By quantifying the effect of mechanical perturbations on the chromosome's structural fea-tures, our model provides a nuanced understanding of in vivo mechanics of chromosomes.
In recent years, much effort has been devoted to understanding the three-dimensional (3D) organization of the genome and how genomic structure mediates nuclear function. The development of experimental techniques that combine DNA proximity ligation with high-throughput sequencing, such as Hi-C, have substantially improved our knowledge about chromatin organization. Numerous experimental advancements, not only utilizing DNA proximity ligation but also high-resolution genome imaging (DNA tracing), have required theoretical modeling to determine the structural ensembles consistent with such data. These 3D polymer models of the genome provide an understanding of the physical mechanisms governing genome architecture. Here, we present an overview of the recent advances in modeling the ensemble of 3D chromosomal structures by employing the maximum entropy approach combined with polymer physics. Particularly, we discuss the minimal chromatin model (MiChroM) along with the "maximum entropy genomic annotations from biomarkers associated with structural ensembles" (MEGABASE) model, which have been remarkably successful in the accurate modeling of chromosomes consistent with both Hi-C and DNA-tracing data.
Micromechanical studies of mitotic chromosomes have revealed them to be remarkably extensible objects and informed early models of mitotic chromosome organization. We use a data-driven, coarsegrained polymer modeling approach, capable of generating ensembles of chromosome structures that are quantitatively consistent with experiments, to explore the relationship between the spatial organization of individual chromosomes and their emergent mechanical properties. In particular, we investigate the mechanical properties of our model chromosomes by axially stretching them. Simulated stretching led to a linear force-extension curve for small strain, with mitotic chromosomes behaving about ten-fold stiffer than interphase chromosomes. Studying the relaxation dynamics we found that chromosomes are viscoelastic solids, with a highly liquid-like, viscous behavior in interphase that becomes solid-like in mitosis. This emergent mechanical stiffness in our model originates from lengthwise compaction, an effective potential capturing the activity of loop-extruding SMC complexes. Chromosomes denature under large strains via unraveling, which is characterized by opening of large-scale folding patterns. By quantifying the effect of mechanical perturbations on the chromosome’s structural features, our model provides a nuanced understanding of in vivo mechanics of chromosomes.
We introduce the Nucleome Data Bank, a web-based platform to simulate and analyze the three-dimensional organization of genomes. The Nucleome Data Bank enables physics-based simulation of chromosomal structural dynamics through the MEGABASE + MiChroM computational pipeline. The input of the pipeline consists of epigenetic information sourced from the Encode database; the output consists of the trajectories of chromosomal motions that accurately predict Hi-C and FISH data, as well as multiple observations of chromosomal dynamics in vivo . As an intermediate step, users can also generate chromosomal sub-compartment annotations directly from the same epigenetic input, without the use of any DNA-DNA proximity ligation data. Additionally, the Nucleome Data Bank freely hosts both experimental and computational structural genomics data. Besides being able to perform their own genome simulations and download the hosted data, users can also analyze and visualize the same data through custom-designed web-based tools. In particular, the one-dimensional genetic and epigenetic data can be overlaid onto accurate three-dimensional structures of chromosomes, to study the spatial distribution of genetic and epigenetic features. The Nucleome Data Bank aims to be a shared resource to biologists, biophysicists, and all genome scientists. The Nucleome Data Bank (NDB) is available at https://ndb.rice.edu .
We explore the energetic frustration patterns associated with the binding between the SARS-CoV-2 spike protein and the ACE2 receptor protein in a broad selection of animals. Using energy landscape theory and the concept of energy frustration—theoretical tools originally developed to study protein folding—we are able to identify interactions among residues of the spike protein and ACE2 that result in COVID-19 resistance. This allows us to identify whether or not a particular animal is susceptible to COVID-19 from the protein sequence of ACE2 alone. Our analysis predicts a number of experimental observations regarding COVID-19 susceptibility, demonstrating that this feature can be explained, at least partially, on the basis of theoretical means.
Direct coupling analysis (DCA) is a global statistical approach that uses information encoded in protein sequence data to predict spatial contacts in a three-dimensional structure of a folded protein. DCA has been widely used to predict the monomeric fold at amino acid resolution and to identify biologically relevant interaction sites within a folded protein. Going beyond single proteins, DCA has also been used to identify spatial contacts that stabilize the interaction in protein complex formation. However, extracting this higher order information necessary to predict dimer contacts presents a significant challenge. A DCA evolutionary signal is much stronger at the single protein level (intraprotein contacts) than at the protein-protein interface (interprotein contacts). Therefore, if DCA-derived information is to be used to predict the structure of these complexes, there is a need to identify statistically significant DCA predictions. We propose a simple Z-score measure that can filter good predictions despite noisy, limited data. This new methodology not only improves our prediction ability but also provides a quantitative measure for the validity of the prediction.
Response regulators are one group of signal transduction proteins found in bacteria, called two-component signaling proteins. The DNA-binding response regulators are homodimers upon activation, with each monomer composed of receiver (REC) and effector (EFF) domains. While the three-dimensional structure of the REC domain is conserved across all DNA-binding response regulator subfamilies, the DNA-binding EFF domain is structurally distinct among three main subfamilies: OmpR, NarL and NtrC. Despite a lot of efforts of study, it is not fully understood how the sequences of response regulators have diverged to maintain stable interactions between the REC and EFF domains, as well as forming dimeric interfaces. To elucidate the molecular mechanisms that maintain the interactions simultaneously, we used Direct Coupling Analysis (DCA) to quantify the co-evolution between residues belonging to each of the three subfamilies of the DNA-binding response regulators. In spite of the significant sequence and structural differences, DCA revealed that all response regulators exhibit two interaction regions between the REC and EFF domains. The analysis suggests two types of interactions between REC and EFF domains: those responsible for monomeric conformational changes that plausibly stabilize inactive monomer conformations and those stabilizing dimeric conformations. We hypothesize that these interactions play a physiologically important role in regulation of DNA-binding of the protein.
Using computer simulations, we generate cell-specific 3D chromosomal structures and compare them to recently published chromatin structures obtained through microscopy. We demonstrate using machine learning and polymer physics simulations that epigenetic information can be used to predict the structural ensembles of multiple human cell lines. Theory predicts that chromosome structures are fluid and can only be described by an ensemble, which is consistent with the observation that chromosomes exhibit no unique fold. Nevertheless, our analysis of both structures from simulation and microscopy reveals that short segments of chromatin make two-state transitions between closed conformations and open dumbbell conformations. Finally, we study the conformational changes associated with the switching of genomic compartments observed in human cell lines. The formation of genomic compartments resembles hydrophobic collapse in protein folding, with the aggregation of denser and predominantly inactive chromatin driving the positioning of active chromatin toward the surface of individual chromosomal territories.
We develop a simple, coarse-grained approach for simulating the folding of the Beet Western Yellow Virus (BWYV) pseudoknot toward the goal of creating a transferable model that can be used to study other small RNA molecules. This approach combines a structure-based model (SBM) of RNA with an electrostatic scheme that has previously been shown to correctly reproduce ionic condensation in the native basin. Mg2+ ions are represented explicitly, directly incorporating ion-ion correlations into the system, and K+ is represented implicitly, through the mean-field generalized Manning counterion condensation theory. Combining the electrostatic scheme with a SBM enables the electrostatic scheme to be tested beyond the native basin. We calibrate the SBM to reproduce experimental BWYV unfolding data by eliminating overstabilizing backbone interactions from the molecular contact map and by strengthening base pairing and stacking contacts relative to other native contacts, consistent with the experimental observation that relative helical stabilities are central determinants of the RNA unfolding sequence. We find that this approach quantitatively captures the Mg2+ dependence of the folding temperature and generates intermediate states that better approximate those revealed by experiment. Finally, we examine how our model captures Mg2+ condensation about the BWYV pseudoknot and a U-tail variant, for which the nine 3' end nucleotides are replaced with uracils, and find our results to be consistent with experimental condensation measurements. This approach can be easily transferred to other RNA molecules by eliminating and strengthening the same classes of contacts in the SBM and including generalized Manning counterion condensation.
This chapter argues that the energy landscape theory, which has proven useful in understanding protein folding, and provides a powerful bridge between the top-down and bottom-up perspectives when thinking about chromosome structure and function. It describes a physical theory motivated by symmetry principles that explains how local interactions between genomic loci can lead to the conformations of human chromosomes in interphase. The principles learned using the energy landscape theory allow the construction of a force field that predicts detailed Hi-C data for many distinct human chromosomes, using epigenetic marks found through binding assays. While some argue that information theory is logically prior to physics, many others would hold that physics is prior to information theory. Genome-wide ligation assays measure the frequency at which any given pair of genomic segments in the chromosomes comes into close spatial proximity inside the nucleus.
Protein assemblies consisting of structural maintenance of chromosomes (SMC) and kleisin subunits are essential for the process of chromosome segregation across all domains of life. Prokaryotic condensin belonging to this class of protein complexes is composed of a homodimer of SMC that associates with a kleisin protein subunit called ScpA. While limited structural data exist for the proteins that comprise the (SMC)-kleisin complex, the complete structure of the entire complex remains unknown. Using an integrative approach combining both crystallographic data and coevolutionary information, we predict an atomic-scale structure of the whole condensin complex, which our results indicate being composed of a single ring. Coupling coevolutionary information with molecular-dynamics simulations, we study the interaction surfaces between the subunits and examine the plausibility of alternative stoichiometries of the complex. Our analysis also reveals several additional configurational states of the condensin hinge domain and the SMC-kleisin interaction domains, which are likely involved with the functional opening and closing of the condensin ring. This study provides the foundation for future investigations of the structure-function relationship of the various SMC-kleisin protein complexes at atomic resolution.
Cohesin extrusion is thought to play a central role in establishing the architecture of mammalian genomes. However, extrusion has not been visualized in vivo, and thus, its functional impact and energetics are unknown. Using ultra-deep Hi-C, we show that loop domains form by a process that requires cohesin ATPases. Once formed, however, loops and compartments are maintained for hours without energy input. Strikingly, without ATP, we observe the emergence of hundreds of CTCF-independent loops that link regulatory DNA. We also identify architectural "stripes,'' where a loop anchor interacts with entire domains at high frequency. Stripes often tether super-enhancers to cognate promoters, and in B cells, they facilitate Igh transcription and recombination. Stripe anchors represent major hotspots for topoisomerase-mediated lesions, which promote chromosomal translocations and cancer. In plasmacytomas, stripes can deregulate Igh-translocated oncogenes. We propose that higher organisms have coopted cohesin extrusion to enhance transcription and recombination, with implications for tumor development.
Point mutations that preserve the overall fold of a protein can change the stability. Designing enzymes that remain active at severe conditions, such as high temperature or low pH, is just one of the many possible applications of a method that can reliably predict stability changes upon mutation. Experimental measurements of stability changes upon mutation have been collected in databases, and, in this work, we use two different models to predict these changes. The first method is Direct Coupling Analysis (DCA), which takes as input a multiple sequence alignment of proteins that share the same fold and yields a probability model for sequences in that family. This probability model can be used to obtain an Energy function based on the sequence. The second method is the Associative Memory, Water Mediated, Structure and Energy Model, which is a coarse-grained protein folding model that has been used for structure prediction and mechanistic studies of proteins. We show that both models can present good correlations with experimental data in several protein families. This work is partially supported by the Center for Theoretical Biological Physics sponsored by the NSF (PHY-1427654 ) and by the National Institute of General Medical Sciences (R01-GM44557).
Selecting amino acids to design novel protein-protein interactions that facilitate catalysis is a daunting challenge. We propose that a computational coevolutionary landscape based on sequence analysis alone offers a major advantage over expensive, time-consuming bruteforce approaches currently employed. Our coevolutionary landscape allows prediction of single amino acid substitutions that produce functional interactions between non-cognate, interspecies signaling partners. In addition, it can also predict mutations that maintain segregation of signaling pathways across species. Specifically, predictions of phosphotransfer activity between the Escherichia coli histidine kinase EnvZ to the non-cognate receiver Spo0F from Bacillus subtilis were compiled. Twelve mutations designed to enhance, suppress, or have a neutral effect on kinase phosphotransfer activity to a non-cognate partner were selected. We experimentally tested the ability of the kinase to relay phosphate to the respective designed Spo0F receiver proteins against the theoretical predictions. Our key finding is that the coevolutionary landscape theory, with limited structural data, can significantly reduce the search-space for successful prediction of single amino acid substitutions that modulate phosphotransfer between the two-component His-Asp relay partners in a predicted fashion. This combined approach offers significant improvements over large-scale mutations studies currently used for protein engineering and design.
Short Abstract — Actin is one of the most abundant proteins found in eukaryotic cells. It plays a crucial role in cell motility, division, and forms the cytoskeleton. The actin filament is a helical polymer and its crystal structure has yet to be solved. Low resolution atomic-models of the filament, constructed using crystallographic data and known G-actin crystal structure, provide useful, but limited, clues as to the interactions involved in filament formation and regulation. Current statistical modeling tools can predict residue pair coevolution from genomic information. Here, we utilize both structural and genomic data to elucidate the functional interactions of actin.
Inside the cell nucleus, genomes fold into organized structures that are characteristic of cell type. Here, we show that this chromatin architecture can be predictedde novousing epigenetic data derived from ChIP-Seq. We exploit the idea that chromosomes encode a one-dimensional sequence of chromatin structural types. Interactions between these chromatin types determine the three-dimensional (3D) structural ensemble of chromosomes through a process similar to phase separation. First, a recurrent neural network is used to infer the relation between the epigenetic marks present at a locus, as assayed by ChIP-Seq, and the genomic compartment in which those loci reside, as measured by DNA-DNA proximity ligation (Hi-C). Next, types inferred from this neural network are used as an input to an energy landscape model for chromatin organization (MiChroM) in order to generate an ensemble of 3D chromosome conformations. After training the model, dubbed MEGABASE (Maximum Entropy Genomic Annotation from Biomarkers Associated to Structural Ensembles), on odd numbered chromosomes, we predict the chromatin type sequences and the subsequent 3D conformational ensembles for the even chromosomes. We validate these structural ensembles by using ChIP-Seq tracks alone to predict Hi-C maps as well as distances measured using 3D FISH experiments. Both sets of experiments support the hypothesis of phase separation being the driving process behind compartmentalization. These findings strongly suggest that epigenetic marking patterns encode sufficient information to determine the global architecture of chromosomes and thatde novostructure prediction for whole genomes may be increasingly possible.