Transcription factor IIH (TFIIH) is a protein assembly essential for transcription initiation and nucleotide excision repair (NER). Yet, understanding of the conformational switching underpinning these diverse TFIIH functions remains fragmentary. TFIIH mechanisms critically depend on two translocase subunits, XPB and XPD. To unravel their functions and regulation, we build cryo-EM based TFIIH models in transcription- and NER-competent states. Using simulations and graph-theoretical analysis methods, we reveal TFIIH’s global motions, define TFIIH partitioning into dynamic communities and show how TFIIH reshapes itself and self-regulates depending on functional context. Our study uncovers an internal regulatory mechanism that switches XPB and XPD activities making them mutually exclusive between NER and transcription initiation. By sequentially coordinating the XPB and XPD DNA-unwinding activities, the switch ensures precise DNA incision in NER. Mapping TFIIH disease mutations onto network models reveals clustering into distinct mechanistic classes, affecting translocase functions, protein interactions and interface dynamics.
Transcription-coupled repair is essential for the removal of DNA lesions from the transcribed genome. The pathway is initiated by CSB protein binding to stalled RNA polymerase II. Mutations impairing CSB function cause severe genetic disease. Yet, the ATP-dependent mechanism by which CSB powers RNA polymerase to bypass certain lesions while triggering excision of others is incompletely understood. Here we build structural models of RNA polymerase II bound to the yeast CSB ortholog Rad26 in nucleotide-free and bound states. This enables simulations and graph-theoretical analyses to define partitioning of this complex into dynamic communities and delineate how its structural elements function together to remodel DNA. We identify an allosteric pathway coupling motions of the Rad26 ATPase modules to changes in RNA polymerase and DNA to unveil a structural mechanism for CSB-assisted progression past less bulky lesions. Our models allow functional interpretation of the effects of Cockayne syndrome disease mutations.
Suboptimal path analysis in a protein structural or dynamical network becomes increasingly popular for identifying critical residues involved in allosteric communication and regulation. Several software packages have been developed for calculating suboptimal paths, including NetworkView, WISP, and CNAPATH (Bio3D). Although these packages work well for biological systems of moderate sizes, they either dramatically slow down or are subjected to accuracy issues when applied to large systems such as supramolecular complexes. In this work, we develop a new method called SOAN, which implements a modified version of Yen's algorithm for finding loopless k-shortest paths. Instead of searching the entire protein network, SOAN builds up a subgraph for path calculations based on an initial evaluation of the optimal path and its neighbouring nodes. We test our method on four systems of increasing size and compare it to the NetworkView, WISP and CNAPATH methods. The result shows that SOAN is approximately five times faster than NetworkView and orders of magnitude faster than CNAPATH and WISP. In terms of accuracy, SOAN is comparable to CNAPATH and WISP and superior to NetworkView. We also discuss the influence of SOAN input parameters on performance and suggest optimal values.
DNA replication origins serve as sites of replicative helicase loading. In all eukaryotes, the six-subunit origin recognition com-plex (Orc1-6; ORC) recognizes the replication origin. During late M-phase of the cell-cycle, Cdc6 binds to ORC and the ORC-Cdc6 complex loads in a multistep reaction and, with the help of Cdt1, the core Mcm2-7 helicase onto DNA. A key intermediate is the ORC-Cdc6-Cdt1-Mcm2-7 (OCCM) complex in which DNA has been already inserted into the central channel of Mcm2-7. Until now, it has been unclear how the origin DNA is guided by ORC-Cdc6 and inserted into the Mcm2-7 hexamer. Here, we truncated the C-terminal winged-helix-domain (WHD) of Mcm6 to slow down the loading reaction, thereby capturing two loading intermediates prior to DNA insertion in budding yeast. In "semi-attached OCCM," the Mcm3 and Mcm7 WHDs latch onto ORC-Cdc6 while the main body of the Mcm2-7 hexamer is not connected. In "pre-insertion OCCM," the main body of Mcm2-7 docks onto ORC-Cdc6, and the origin DNA is bent and positioned adjacent to the open DNA entry gate, poised for insertion, at the Mcm2-Mcm5 interface. We used molecular simulations to reveal the dynamic transition from pre -loading conformers to the loaded conformers in which the loading of Mcm2-7 on DNA is complete and the DNA entry gate is fully closed. Our work provides multiple molecular insights into a key event of eukaryotic DNA replication.
Proofreading by replicative DNA polymerases is a fundamental mechanism ensuring DNA replication fidelity. In proofreading, mis-incorporated nucleotides are excised through the 3′-5′ exonuclease activity of the DNA polymerase holoenzyme. The exonuclease site is distal from the polymerization site, imposing stringent structural and kinetic requirements for efficient primer strand transfer. Yet, the molecular mechanism of this transfer is not known. Here we employ molecular simulations using recent cryo-EM structures and biochemical analyses to delineate an optimal free energy path connecting the polymerization and exonuclease states of E. coli replicative DNA polymerase Pol III. We identify structures for all intermediates, in which the transitioning primer strand is stabilized by conserved Pol III residues along the fingers, thumb and exonuclease domains. We demonstrate switching kinetics on a tens of milliseconds timescale and unveil a complete pol-to-exo switching mechanism, validated by targeted mutational experiments.
Advances in cryoelectron microscopy (cryo-EM) have revolutionized the structural investigation of large macromolecular assemblies. In this review, we first provide a broad overview of modeling methods used for flexible fitting of molecular models into cryo-EM density maps. We give special attention to approaches rooted in molecular simulations-atomistic molecular dynamics and Monte Carlo. Concise descriptions of the methods are given along with discussion of their advantages, limitations, and most popular alternatives. We also describe recent extensions of the widely used molecular dynamics flexible fitting (MDFF) method and discuss how different model-building techniques could be incorporated into new hybrid modeling schemes and simulation workflows. Finally, we provide two illustrative examples of model-building and refinement strategies employing MDFF, cascade MDFF, and RosettaCM. These examples come from recent cryo-EM studies that elucidated transcription preinitiation complexes and shed light on the functional roles of these assemblies in gene expression and gene regulation.
DNA polymerase III (PolIII) is a high-fidelity enzyme that can synthesize over 100,000 basepairs (bp) per binding event, with an error rate of ∼1 per million. On the rare occasion that a misincorporation does occur, PolIII utilizes a 3'-5' exonuclease to remove the misincorporated nucleotide so that synthesis may continue. Combined with its exonuclease activity, PolIII's error rate reaches ∼1 per billion. Although recent cryo-EM data has elucidated structural information on PolIII during DNA synthesis (polymerase mode) and error correction (editing mode), details on the primer strand separation and its translocation to the exonuclease active site remain elusive. Using novel path-sampling methodologies we have constructed a transition pathway between the polymerase mode and editing mode. Evenly spaced configurations along the transition path were further sampled using molecular dynamics (MD). We then utilized time-independent component analysis (tICA) to distinguish between the slowest motions in the transition between polymerase and editing modes. Additionally, we've identified several key intermediate states highlighting crucial roles of the polymerase thumb domain in primer strand separation and stabilization. Our initial results show that global motions coupled with intimate protein-nucleic acid contacts provide PolIII with a powerful mechanism for the correction of misincorporated nucleotides.
Transcription preinitiation complexes (PICs) are vital assemblies whose function underlies the expression of protein-encoding genes. Cryo-EM advances have begun to uncover their structural organization. Nevertheless, functional analyses are hindered by incompletely modeled regions. Here we integrate all available cryo-EM data to build a practically complete human PIC structural model. This enables simulations that reveal the assembly's global motions, define PIC partitioning into dynamic communities and delineate how structural modules function together to remodel DNA. We identify key TFIIE-p62 interactions that link core-PIC to TFIIH.p62 rigging interlaces p34, p44 and XPD while capping the DNA-binding and ATP-binding sites of XPD. PIC kinks and locks substrate DNA, creating negative supercoiling within the Pol II cleft to facilitate promoter opening. Mapping disease mutations associated with xeroderma pigmentosum, trichothiodystrophy and Cockayne syndrome onto defined communities reveals clustering into three mechanistic classes that affect TFIIH helicase functions, protein interactions and interface dynamics.
Regulation of gene-expression by specific targeting of protein-nucleic acid interactions has been a long-standing goal in medicinal chemistry. Transcription factors are considered "undruggable" because they lack binding sites well suited for binding small-molecules. In order to overcome this obstacle, we are interested in designing small molecules that bind to the corresponding promoter sequences and either prevent or modulate transcription factor association via an allosteric mechanism. To achieve this, we must design small molecules that are both sequence-specific and able to target G/C base pair sites. A thorough understanding of the relationship between binding affinity and the structural aspects of the small molecule-DNA complex would greatly aid in rational design of such compounds. Here we present a comprehensive analysis of sequence-specific DNA association of a synthetic minor groove binder using long timescale molecular dynamics. We show how binding selectivity arises from a combination of structural factors. Our results provide a framework for the rational design and optimization of synthetic small molecules in order to improve site-specific targeting of DNA for therapeutic uses in the design of selective DNA binders targeting transcription regulation.
a. Department of Chemistry and Center for Diagnostics and Therapeutics, Georgia State University, Atlanta, GA 30302, USA, iivanov@gsu.edu b. Department of Molecular Biosciences and Chemistry of Life Processes Institute, Northwestern University, Evanston, IL 60208, USA, yuanhe@northwestern.edu c. Molecular Biophysics and Integrated Bioimaging, Lawrence Berkeley National Laboratory, Berkeley, CA 94720, USA, setsutakawa@lbl.gov d. Department of Molecular and Cellular Oncology, The University of Texas M. D. Anderson Cancer Center, Houston, TX 77030 USA, jatainer@gmail.com
Thymine DNA glycosylase (TDG) is a pivotal enzyme with dual roles in both genome maintenance and epigenetic regulation. TDG is involved in cytosine demethylation at CpG sites in DNA. Here we have used molecular modeling to delineate the lesion search and DNA base interrogation mechanisms of TDG. First, we examined the capacity of TDG to interrogate not only DNA substrates with 5-carboxyl cytosine modifications but also G:T mismatches and nonmismatched (A:T) base pairs using classical and accelerated molecular dynamics. To determine the kinetics, we constructed Markov state models. Base interrogation was found to be highly stochastic and proceeded through insertion of an arginine-containing loop into the DNA minor groove to transiently disrupt Watson-Crick pairing. Next, we employed chain-of-replicas path-sampling methodologies to compute minimum free energy paths for TDG base extrusion. We identified the key intermediates imparting selectivity and determined effective free energy profiles for the lesion search and base extrusion into the TDG active site. Our results show that DNA sculpting, dynamic glycosylase interactions, and stabilizing contacts collectively provide a powerful mechanism for the detection and discrimination of modified bases and epigenetic marks in DNA.