DNA superhelicity and transcription are intimately related because changes to DNA topology can influence gene expression and vice versa. Information is transferred through the modulation of local DNA torsional stress, where the expression of one gene may influence the superhelical level of neighbouring genes, either promoting or repressing their expression. In this work, we introduce a one-dimensional physical model that simulates supercoiling-mediated regulation. This TORCphysics model takes as input a genome architecture represented either by a plasmid or chromosomal DNA sequence with ends constrained under specific biological conditions and computes the molecule's output. Our findings demonstrate that the expression profiles of genes are directly influenced by the gene circuit design, including gene location, the positions of topological barriers, promoter sequences, and topoisomerase activity. The novelty that TORCphysics offers is versatility, where users can define distinct activity models for different types of proteins and protein-binding sites. The aim of this research is to establish a flexible framework for developing physical simulations of gene circuits to deepen our comprehension of the intricate mechanisms involved in gene regulation.
Closing each strand of a DNA duplex upon itself fixes its linking number L. This topological condition couples together the secondary and tertiary structures of the resulting ccDNA topoisomer, a constraint that is not present in otherwise identical nicked or linear DNAs. Fixing L has a range of structural, energetic and functional consequences. Here we consider how L having different integer values (that is, different superhelicities) affects ccDNA molecules. The approaches used are primarily theoretical, and are developed from a historical perspective. In brief, processes that either relax or increase superhelicity, or repartition what is there, may either release or require free energy. The energies involved can be substantial, sufficient to influence many events, directly or indirectly. Here two examples are developed. The changes of unconstrained superhelicity that occur during nucleosome attachment and release are examined. And a simple theoretical model of superhelically driven DNA structural transitions is described that calculates equilibrium distributions for populations of identical topoisomers. This model is used to examine how these distributions change with superhelicity and other factors, and applied to analyze several situations of biological interest.
Many cellular processes occur out of equilibrium. This includes site-specific unwinding in supercoiled DNA, which may play an important role in gene regulation. Here, we use the Convex Lens-induced Confinement (CLiC) single-molecule microscopy platform to study these processes with high-throughput and without artificial constraints on molecular structures or interactions. We use two model DNA plasmid systems, pFLIP-FUSE and pUC19, to study the dynamics of supercoiling-induced secondary structural transitions after perturbations away from equilibrium. We find that structural transitions can be slow, leading to long-lived structural states whose kinetics depend on the duration and direction of perturbation. Our findings highlight the importance of out-of-equilibrium studies when characterizing the complex structural dynamics of DNA and understanding the mechanisms of gene regulation.
Physiologically, MYC levels must be precisely set to faithfully amplify the transcriptome, but in cancer MYC is quantitatively misregulated. Here, we study the variation of MYC amongst single primary cells (B-cells and murine embryonic fibroblasts, MEFs) for the repercussions of variable cellular MYC-levels and setpoints. Because FUBPs have been proposed to be molecular “cruise controls” that constrain MYC expression, their role in determining basal or activated MYC-levels was also examined. Growing cells remember low and high-MYC setpoints through multiple cell divisions and are limited by the same expression ceiling even after modest MYC-activation. High MYC MEFs are enriched for mRNAs regulating inflammation and immunity. After strong stimulation, many cells break through the ceiling and intensify MYC expression. Lacking FUBPs, unstimulated MEFs express levels otherwise attained only with stimulation and sponsor MYC chromatin changes, revealed by chromatin marks. Thus, the FUBPs enforce epigenetic setpoints that restrict MYC expression.
DNA unwinding is an important cellular process involved in DNA replication, transcription and repair. In cells, molecular crowding caused by the presence of organelles, proteins, and other molecules affects numerous internal cellular structures. Here, we visualize plasmid DNA unwinding and binding dynamics to an oligonucleotide probe as functions of ionic strength, crowding agent concentration, and crowding agent species using single-molecule CLiC microscopy. We demonstrate increased probe-plasmid interaction over time with increasing concentration of 8kDa polyethylene glycol (PEG), a crowding agent. We show decreased probe-plasmid interactions as ionic strength is increased without crowding. However, when crowding is introduced via 10% 8kDa PEG, interactions between plasmids and oligos are enhanced. This is beyond what is expected for normal in vitro conditions, and may be a critically important, but as of yet unknown, factor in DNA's proper biological function in vivo. Our results show that crowding has a strong effect on the initial concentration of unwound plasmids. In the dilute conditions used in these experiments, crowding does not impact probe-plasmid interactions once the site is unwound.
R-loops are abundant three-stranded nucleic-acid structures that form in cis during transcription. Experimental evidence suggests that R-loop formation is affected by DNA sequence and topology. However, the exact manner by which these factors interact to determine R-loop susceptibility is unclear. To investigate this, we developed a statistical mechanical equilibrium model of R-loop formation in superhelical DNA. In this model, the energy involved in forming an R-loop includes four terms-junctional and base-pairing energies and energies associated with superhelicity and with the torsional winding of the displaced DNA single strand around the RNA:DNA hybrid. This model shows that the significant energy barrier imposed by the formation of junctions can be overcome in two ways. First, base-pairing energy can favor RNA:DNA over DNA:DNA duplexes in favorable sequences. Second, R-loops, by absorbing negative superhelicity, partially or fully relax the rest of the DNA domain, thereby returning it to a lower energy state. In vitro transcription assays confirmed that R-loops cause plasmid relaxation and that negative superhelicity is required for R-loops to form, even in a favorable region. Single-molecule R-loop footprinting following in vitro transcription showed a strong agreement between theoretical predictions and experimental mapping of stable R-loop positions and further revealed the impact of DNA topology on the R-loop distribution landscape. Our results clarify the interplay between base sequence and DNA superhelicity in controlling R-loop stability. They also reveal R-loops as powerful and reversible topology sinks that cells may use to nonenzymatically relieve superhelical stress during transcription.
We directly visualize the topology-mediated interactions between an unwinding site on a supercoiled DNA plasmid and a specific probe molecule designed to bind to this site, as a function of DNA supercoiling and temperature. The visualization relies on containing the DNA molecules within an enclosed array of glass nanopits using the Convex Lens-induced Confinement (CLiC) imaging method. This method traps molecules within the focal plane while excluding signal from out-of-focus probes. Simultaneously, the molecules can freely diffuse within the nanopits, allowing for accurate measurements of exchange rates, unlike other methods which could introduce an artifactual bias in measurements of binding kinetics. We demonstrate that the plasmid’s structure influences the binding of the fluorescent probes to the unwinding site through the presence, or lack, of other secondary structures. With this method, we observe an increase in the binding rate of the fluorescent probe to the unwinding site with increasing temperature and negative supercoiling. This increase in binding is consistent with the results of our numerical simulations of the probability of site-unwinding. The temperature dependence of the binding rate has allowed us to distinguish the effects of competing higher order DNA structures, such as Z-DNA, in modulating local site-unwinding, and therefore binding.
We directly visualize the weak and slow interactions between specific unwinding sites on supercoiled DNA, and site-specific probes designed to bind to these sites, as a function of DNA supercoiling and temperature. We use Convex Lens-induced Confinement microscopy to confine the DNA molecules within a sealed, glass array of nanoscale pits, embedded in a coverslip. Throughout our wide-field observations of single-molecule binding and diffusive trajectories, the DNA molecules are free to explore all possible configurations, which we show is influential on the observed dynamics. The increase in binding rate that we observe, with both temperature and supercoiling, is consistent with Z-DNA formation playing a key role in governing DNA structural mechanics. This alternate structure suppresses supercoil-induced DNA unwinding in DNA at low temperatures, for a wide range of negative superhelicities. The new single-molecule methodology that we present may be used to visualize a wide range of molecular interactions that are challenging or impossible to access with other methods such as optical or magnetic tweezers. These include interactions that are highly dependent on molecular topology, proceed over many seconds to minutes, or are rare.
DNA in cells is predominantly B-form double helix. Though certain DNA sequences in vitro may fold into other structures, such as triplex, left-handed Z form, or quadruplex DNA, the stability and prevalence of these structures in vivo are not known. Here, using computational analysis of sequence motifs, RNA polymerase II binding data, and genomewide potassium permanganate-dependent nuclease footprinting data, we map thousands of putative non-B DNA sites at high resolution in mouse B cells. Computational analysis associates these non-B DNAs with particular structures and indicates that they form at locations compatible with an involvement in gene regulation. Further analyses support the notion that non-B DNA structure formation influences the occupancy and positioning of nucleosomes in chromatin. These results suggest that non-B DNAs contribute to the control of a variety of critical cellular and organismal processes.
It is well established that gene regulation can be achieved through activator and repressor proteins that bind to DNA and switch particular genes on or off, and that complex metabolic networks determine the levels of transcription of a given gene at a given time. Using three complementary computational techniques to study the sequence-dependence of DNA denaturation within DNA minicircles, we have observed that whenever the ends of the DNA are constrained, information can be transferred over long distances directly by the transmission of mechanical stress through the DNA itself, without any requirement for external signalling factors. Our models combine atomistic molecular dynamics (MD) with coarse-grained simulations and statistical mechanical calculations to span three distinct spatial resolutions and timescale regimes. While they give a consensus view of the non-locality of sequence-dependent denaturation in highly bent and supercoiled DNA loops, each also reveals a unique aspect of long-range informational transfer that occurs as a result of restraining the DNA within the closed loop of the minicircles.
SUMMARY:Supercoiling imposes stress on a DNA molecule that can drive susceptible sequences into alternative non-B form structures. This phenomenon occurs frequently in vivo and has been implicated in biological processes, such as replication, transcription, recombination and translocation. SIST is a software package that analyzes sequence-dependent structural transitions in kilobase length superhelical DNA molecules. The numerical algorithms in SIST are based on a statistical mechanical model that calculates the equilibrium probability of transition for each base pair in the domain. They are extensions of the original stress-induced duplex destabilization (SIDD) method, which analyzes stress-driven DNA strand separation. SIST also includes algorithms to analyze B-Z transitions and cruciform extrusion. The SIST pipeline has an option to use the DZCBtrans algorithm, which analyzes the competition among these three transitions within a superhelical domain.AVAILABILITY AND IMPLEMENTATION:The package and additional documentation are freely available at https://bitbucket.org/benhamlab/sist_codes.CONTACT:dzhabinskaya@ucdavis.edu.
Where possible, developments enabling the establishment of cell lines with predictable, long-term stable expression capacity are based on single-copy integrations at safe genomic loci with predictable properties. Robust performance could be assigned to lentiviral transduction systems anchoring single LV-units at sites with adequate transcription potential. In the case of gene therapeutic vectors it is essential that the expression interval can be safely terminated following individual requirements, which has mostly been achieved by lox-mediated excision ("floxing"). To extend the spectrum of possible applications we replaced the common, phage-derived Cre/loxP-setup by modules derived from the yeast "Flp/FRT" site-specific recombination system. This change enables a variety of additional options, for instance by "multiplexing" strategies, which rely on a variety of heterospecific FRT-site variants (F'). If we provide lentiviral LTRs with a "twin-site", here an FF3 fusion, the presence of Flp-recombinase will effectively excise the expression cassette, leaving behind a single neutral, genomically anchored FF3 unit. This tag serves to identify the integration locus and to apply sequence- and structural (SIDD-) analyses to predict its functions. Candidate loci are then used to accommodate, at the given site, other genes of interest by "Recombinase-Mediated Twin Site Targeting" (RMTT), a contemporary extension of existing cassette exchange (RMCE-) routines. Supported by the fact that FF3 twins remain accessible within the host genome, RMTT provides access to certified cell lines as it complies with recently defined stringent genomic safe harbor criteria. Our discussion- and outlook-sections will cover lentiviral targeting strategies and current possibilities to enable their fine-tuning.
Although the right-handed double helical B-form DNA is most common under physiological conditions, DNA is dynamic and can adopt a number of alternative structures, such as the four-stranded G-quadruplex, left-handed Z-DNA, cruciform and others. Active transcription necessitates strand separation and can induce such non-canonical forms at susceptible genomic sequences. Therefore, it has been speculated that these non-B DNA motifs can play regulatory roles in gene transcription. Such conjecture has been supported in higher eukaryotes by direct studies of several individual genes, as well as a number of large-scale analyses. However, the role of non-B DNA structures in many lower organisms, in particular proteobacteria, remains poorly understood and incompletely documented. In this study, we performed the first comprehensive study of the occurrence of B DNA-non-B DNA transition-susceptible sites (non-B DNA motifs) within the context of the operon structure of the Escherichia coli genome. We compared the distributions of non-B DNA motifs in the regulatory regions of operons with those from internal regions. We found an enrichment of some non-B DNA motifs in regulatory regions, and we show that this enrichment cannot be simply explained by base composition bias in these regions. We also showed that the distribution of several non-B DNA motifs within intergenic regions separating divergently oriented operons differs from the distribution found between convergent ones. In particular, we found a strong enrichment of cruciforms in the termination region of operons; this enrichment was observed for operons with Rho-dependent, as well as Rho-independent terminators. Finally, a preference for some non-B DNA motifs was observed near transcription factor-binding sites. Overall, the conspicuous enrichment of transition-susceptible sites in these specific regulatory regions suggests that non-B DNA structures may have roles in the transcriptional regulation of specific operons within the E. coli genome.
A vast literature has explored the genetic interactions among the cellular components regulating gene expression in many organisms. Early on, in the absence of any biochemical definition, regulatory modules were conceived using the strict formalism of genetics to designate the modifiers of phenotype as either cis- or trans-acting depending on whether the relevant genes were embedded in the same or separate DNA molecules. This formalism distilled gene regulation down to its essence in much the same way that consideration of an ideal gas reveals essential thermodynamic and kinetic principles. Yet just as the anomalous behavior of materials may thwart an engineer who ignores their non-ideal properties, schemes to control and manipulate the genetic and epigenetic programs of cells may falter without a fuller and more quantitative elucidation of the physical and chemical characteristics of DNA and chromatin in vivo.
The susceptibility to recombination of a plasmid inserted into a chromosome varies with its genomic position. This recombination position effect is known to correlate with the average G+C content of the flanking sequences. Here we propose that this effect could be mediated by changes in the susceptibility to superhelical duplex destabilization that would occur. We use standard nonparametric statistical tests, regression analysis and principal component analysis to identify statistically significant differences in the destabilization profiles calculated for the plasmid in different contexts, and correlate the results with their measured recombination rates. We show that the flanking sequences significantly affect the free energy of denaturation at specific sites interior to the plasmid. These changes correlate well with experimentally measured variations of the recombination rates within the plasmid. This correlation of recombination rate with superhelical destabilization properties of the inserted plasmid DNA is stronger than that with average G+C content of the flanking sequences. This model suggests a possible mechanism by which flanking sequence base composition, which is not itself a context-dependent attribute, can affect recombination rates at positions within the plasmid.
Although a variety of possible functions have been proposed for inverted repeat sequences (IRs), it is not known which of them might occur in vivo. We investigate this question by assessing the distributions and properties of IRs in the Saccharomyces cerevisiae (SC) genome. Using the IRFinder algorithm we detect 100,514 IRs having copy length greater than 6 bp and spacer length less than 77 bp. To assess statistical significance we also determine the IR distributions in two types of randomization of the S. cerevisiae genome. We find that the S. cerevisiae genome is significantly enriched in IRs relative to random. The S. cerevisiae IRs are significantly longer and contain fewer imperfections than those from the randomized genomes, suggesting that processes to lengthen and/or correct errors in IRs may be operative in vivo. The S. cerevisiae IRs are highly clustered in intergenic regions, while their occurrence in coding sequences is consistent with random. Clustering is stronger in the 3′ flanks of genes than in their 5′ flanks. However, the S. cerevisiae genome is not enriched in those IRs that would extrude cruciforms, suggesting that this is not a common event. Various explanations for these results are considered.
We present an interactive web application for visualizing genomic data of prokaryotic chromosomes. The tool (GeneWiz browser) allows users to carry out various analyses such as mapping alignments of homologous genes to other genomes, mapping of short sequencing reads to a reference chromosome, and calculating DNA properties such as curvature or stacking energy along the chromosome. The GeneWiz browser produces an interactive graphic that enables zooming from a global scale down to single nucleotides, without changing the size of the plot. Its ability to disproportionally zoom provides optimal readability and increased functionality compared to other browsers. The tool allows the user to select the display of various genomic features, color setting and data ranges. Custom numerical data can be added to the plot allowing, for example, visualization of gene expression and regulation data. Further, standard atlases are pre-generated for all prokaryotic genomes available in GenBank, providing a fast overview of all available genomes, including recently deposited genome sequences. The tool is available online from http://www.cbs.dtu.dk/services/gwBrowser. Supplemental material including interactive atlases is available online at http://www.cbs.dtu.dk/services/gwBrowser/suppl/.
Numerical models of mesoscale DNA dynamics relevant to in vivo scenarios require methods that incorporate important features of the intracellular environment, while maintaining computational tractability. Because the explicit inclusion of ions leads to electrostatic calculations that scale as the square of the number of charged particles, such models typically handle these calculations using low-potential, mean-field approaches, rather than by considering the discrete interactions of ions. This allows approximation of the long-range, screened self-repulsion of DNA, but is unable to capture detailed electrostatic phenomena, such as short-range attractions mediated by ion-ion correlations. Here, we develop a dynamical model of explicitly double-stranded, sequence-specific DNA in a bulk environment consisting of other polyions and explicitly represented counterions and coions. DNA is represented as two interwound chains of charged Stokes spheres, and ions as free, monovalently charged Stokes spheres. Brownian dynamics simulations performed at salt concentrations of 0.1, 1, 10, and 100 mM demonstrate this model captures anticipated behaviors of the system, including increasing compaction of the polyion by the ionic atmosphere with increasing ionic strength. The decay of the distance dependence of the ion concentrations as one moves away from the polyion approaches their equilibrium values in quantitative agreement with predictions of Poisson-Boltzmann theory. The simulation results also demonstrate quantitative agreement with experimental measurements of the persistence length of B-DNA, which increases significantly at low ionic strengths. The model also captures behaviors intimating the importance of explicitly representing ionic and polyionic structure. These include penetration of the polyion interior by both coions and counterions, and counterion-mediated accumulation of coions near the surface of the polyion. Such phenomena are likely to play an important role in the formation of alternative DNA secondary structures, suggesting the present methods will prove valuable to dynamic models of superhelical stress-induced DNA structural transitions.
Stress-induced DNA duplex destabilization (SIDD) analysis exploits the known structural and energetic properties of DNA to predict sites that are susceptible to strand separation under negative superhelical stress. When this approach was used to calculate the SIDD profile of the entire Escherichia coli K12 genome, it was found that strongly destabilized sites occur preferentially in intergenic regions that are either known or inferred to contain promoters, but rarely occur in coding regions. Here, we investigate whether the genes grouped in different functional categories have characteristic SIDD properties in their upstream flanks. We report that strong SIDD sites in the E. coli K12 genome are statistically significantly overrepresented in the upstream regions of genes encoding transcriptional regulators. In particular, the upstream regions of genes that directly respond to physiological and environmental stimuli are more destabilized than are those regions of genes that are not involved in these responses. Moreover, if a pathway is controlled by a transcriptional regulator whose gene has a destabilized 5' flank, then the genes (operons) in that pathway also usually contain strongly destabilized SIDD sites in their 5' flanks. We observe this statistically significant association of SIDD sites with upstream regions of genes functioning in transcription in 38 of 43 genomes of free-living bacteria, but in only four of 18 genomes of endosymbionts or obligate parasitic bacteria. These results suggest that strong SIDD sites 5' to participating genes may be involved in transcriptional responses to environmental changes, which are known to transiently alter superhelicity. We propose that these SIDD sites are active and necessary participants in superhelically mediated regulatory mechanisms governing changes in the global pattern of gene expression in prokaryotes in response to physiological or environmental changes.