The DNA mismatch repair protein MutS recognizes mispaired bases in DNA and initiates repair in an ATP-dependent manner. Understanding of the allosteric coupling between DNA mismatch recognition and two asymmetric nucleotide binding sites at opposing sides of the MutS dimer requires identification of the relevant MutS.mmDNA.nucleotide species. Here, we use native mass spectrometry to detect simultaneous DNA mismatch binding and asymmetric nucleotide binding to Escherichia coli MutS. To resolve the small differences between macromolecular species bound to different nucleotides, we developed a likelihood based algorithm capable to deconvolute the observed spectra into individual peaks. The obtained mass resolution resolves simultaneous binding of ADP and AMP.PNP to this ABC ATPase in the absence of DNA. Mismatched DNA regulates the asymmetry in the ATPase sites; we observe a stable DNA-bound state containing a single AMP.PNP cofactor. This is the first direct evidence for such a postulated mismatch repair intermediate, and showcases the potential of native MS analysis in detecting mechanistically relevant reaction intermediates.
Motivation: Macromolecular crystal structures in the Protein Data Bank (PDB) are a key source of structural insight into biological processes. These structures, some >30 years old, were constructed with methods of their era. With PDB_REDO, we aim to automatically optimize these structures to better fit their corresponding experimental data, passing the benefits of new methods in crystallography on to a wide base of non-crystallographer structure users. Results: We developed new algorithms to allow automatic rebuilding and remodeling of main chain peptide bonds and side chains in crystallographic electron density maps, and incorporated these and further enhancements in the PDB_REDO procedure. Applying the updated PDB_REDO to the oldest, but also to some of the newest models in the PDB, corrects existing modeling errors and brings these models to a higher quality, as judged by standard validation methods. Availability and Implementation: The PDB_REDO database and links to all software are available at http://www.cmbi.ru.nl/pdb_redo. Contact: r.joosten@nki.nl; a.perrakis@nki.nl Supplementary Information: Supplementary data are available at Bioinformatics online.
The human LINE‐1 endonuclease (L1‐EN) contributes in defining the genomic integration sites of the abundant human L1 and Alu retrotransposons. LINEs have been considered as possible vehicles for gene delivery and understanding the mechanism of L1‐EN could help engineering them as genetic tools. We tested the in vitro activity of point mutants in three L1‐EN residues—Asp145, Arg155, Ile204—that are key for DNA cleavage, and determined their crystal structures. The L1‐EN structure remains overall unaffected by the mutations, which change the enzyme activity but leave DNA cleavage sequence specificity mostly unaffected. To better understand the mechanism of L1‐EN, we performed molecular dynamics simulations using as model the structures of wild type EN‐L1, of two βB6‐βB5 loop exchange mutants we have described previously to be important for DNA recognition, of the R155A mutant from this study, and of the homologous TRAS1 endonuclease: all confirm a rigid scaffold. The simulations crucially indicate that the βB6‐βB5 loop shows an anticorrelated motion with the surface loops βA6‐βA5 and βB3‐αB1. The latter loop harbors N118, a residue that alters DNA cleavage specificity in homologous endonucleases, and implies that the plasticity and correlated motion of these loops has a functional importance in DNA recognition and binding. To further explore how these loops are possibly involved in DNA binding, we docked computationally two DNA substrates to our structure, one involving a flipped‐out nucleotide downstream the scissile phosphodiester; and one not. The models for both scenarios are feasible and agree with the hypotheses derived from the dynamic simulations. The reduced cleavage activity we have observed for the I204Y mutant above however, favors the flipped out nucleotide model. Proteins 2009. © 2008 Wiley‐Liss, Inc.
The automated building of a protein model into an electron density map remains a challenging problem. In the ARP/wARP approach, model building is facilitated by initially interpreting a density map with free atoms of unknown chemical identity; all structural information for such chemically unassigned atoms is discarded. Here, this is remedied by applying restraints between free atoms, and between free atoms and a partial protein model. These are based on geometric considerations of protein structure and tentative (conditional) assignments for the free atoms. Restraints are applied in the REFMAC5 refinement program and are generated on an ad hoc basis, allowing them to fluctuate from step to step. A large set of experimentally phased and molecular replacement structures showcases individual structures where automated building is improved drastically by the conditional restraints. The concept and implementation we present can also find application in restraining geometries, such as hydrogen bonds, in low-resolution refinement.
Automatic iterative model (re-) building, as implemented in ARP/wARP and its new control system flex-wARP, is particularly well suited to follow structure solution by molecular replacement. More than 100 molecular-replacement solutions automatically solved by the BALBES software were submitted to three standard protocols in flex-wARP and the results were compared with final models from the PDB. Standard metrics were gathered in a systematic way and enabled the drawing of statistical conclusions on the advantages of each protocol. Based on this analysis, an empirical estimator was proposed that predicts how good the final model produced by flex-wARP is likely to be based on the experimental data and the quality of the molecular-replacement solution. To introduce the differences between the three flex-wARP protocols ( keeping the complete search model, converting it to atomic coordinates but ignoring atom identities or using the electron-density map calculated from the molecular-replacement solution), two examples are also discussed in detail, focusing on the evolution of the models during iterative rebuilding. This highlights the diversity of paths that the flex-wARP control system can employ to reach a nearly complete and accurate model while actually starting from the same initial information.
ARP/wARP is a software suite to build macromolecular models in X-ray crystallography electron density maps. Structural genomics initiatives and the study of complex macromolecular assemblies and membrane proteins all rely on advanced methods for 3D structure determination. ARP/wARP meets these needs by providing the tools to obtain a macromolecular model automatically, with a reproducible computational procedure. ARP/wARP 7.0 tackles several tasks: iterative protein model building including a high-level decision-making control module; fast construction of the secondary structure of a protein; building flexible loops in alternate conformations; fully automated placement of ligands, including a choice of the best-fitting ligand from a 'cocktail'; and finding ordered water molecules. All protocols are easy to handle by a nonexpert user through a graphical user interface or a command line. The time required is typically a few minutes although iterative model building may take a few hours.
Coot [1] is a molecular graphics application for macromolecular model building against X-ray data.Coot provides a modern interface drawing from usability paradigms of popular desktop applications.In so doing, it has become increasingly popular [2], particularly in the UK and parts of Europe.Coot provides several tools which can be used to build, refine and protein structures and other models.In the last year, there has been more focus on developments to better handle lower resolution data, these include the fitting of alpha helices and beta strands for model building and the addition of extra restraints and modification of restraints when refining.Also to be discussed are the tools for validation for the detection of model-building errors and feature outliers.[1] Emsley & Cowtan ( 2004) "Coot: Model-Building Tools for Molecular Graphics", Acta Cryst D 60, 2126-2132.[2] But not close to the popularity of a more established application.
One of the most cumbersome and time-demanding tasks in completing a protein model is building short missing regions or ;loops'. A method is presented that uses structural and electron-density information to build the most likely conformations of such loops. Using the distribution of angles and dihedral angles in pentapeptides as the driving parameters, a set of possible conformations for the C(alpha) backbone of loops was generated. The most likely candidate is then selected in a hierarchical manner: new and stronger restraints are added while the loop is built. The weight of the electron-density correlation relative to geometrical considerations is gradually increased until the most likely loop is selected on map correlation alone. To conclude, the loop is refined against the electron density in real space. This is started by using structural information to trace a set of models for the C(alpha) backbone of the loop. Only in later steps of the algorithm is the electron-density correlation used as a criterion to select the loop(s). Thus, this method is more robust in low-density regions than an approach using density as a primary criterion. The algorithm is implemented in a loop-building program, Loopy, which can be used either alone or as part of an automatic building cycle. Loopy can build loops of up to 14 residues in length within a couple of minutes. The average root-mean-square deviation of the C(alpha) atoms in the loops built during validation was less than 0.4 A. When implemented in the context of automated model building in ARP/wARP, Loopy can increase the completeness of the built models.
The Structural Proteomics In Europe ( SPINE) consortium contained a workpackage to address the automated X- ray analysis of macromolecules. The aim of this workpackage was to increase the throughput of three- dimensional structures while maintaining the high quality of conventional analyses. SPINE was able to bring together developers of software with users from the partner laboratories. Here, the results of a workshop organized by the consortium to evaluate software developed in the member laboratories against a set of bacterial targets are described. The major emphasis was on molecular- replacement suites, where automation was most advanced. Data processing and analysis, use of experimental phases and model construction were also addressed, albeit at a lower level.
Structure determination and functional characterization of macromolecular complexes requires the purification of the different subunits in large quantities and their assembly into a functional entity. Although isolation and structure determination of endogenous complexes has been reported, much progress has to be made to make this technology easily accessible. Co- expression of subunits within hosts such as Escherichia coli and insect cells has become more and more amenable, even at the level of high- throughput projects. As part of SPINE ( Structural Proteomics In Europe), several laboratories have investigated the use co- expression techniques for their projects, trying to extend from the common binary expression to the more complicated multi- expression systems. A new system for multi- expression in E. coli and a database system dedicated to handle co- expression data are described. Results are also reported from various case studies investigating different methods for performing co- expression in E. coli and insect cells.
Progress towards structure determination that is both high-throughput and high-value is dependent on the development of integrated and automatic tools for electron-density map interpretation and for the analysis of the resulting atomic models. Advances in map-interpretation algorithms are extending the resolution regime in which fully automatic tools can work reliably, but at present human intervention is required to interpret poor regions of macromolecular electron density, particularly where crystallographic data is only available to modest resolution [for example, I/sigma(I) < 2.0 for minimum resolution 2.5 A]. In such cases, a set of manual and semi-manual model-building molecular-graphics tools is needed. At the same time, converting the knowledge encapsulated in a molecular structure into understanding is dependent upon visualization tools, which must be able to communicate that understanding to others by means of both static and dynamic representations. CCP4 mg is a program designed to meet these needs in a way that is closely integrated with the ongoing development of CCP4 as a program suite suitable for both low- and high-intervention computational structural biology. As well as providing a carefully designed user interface to advanced algorithms of model building and analysis, CCP4 mg is intended to present a graphical toolkit to developers of novel algorithms in these fields.
New procedures are outlined that enable ARP/wARP to automatically build protein models with diffraction data extending to about 2.5 A. An overview of ongoing research is given and possible future advances are discussed.
Group leader: Victor Lamzin Staff scientist: Andrea Schmidt Postdoctoral fellows: Olga Kirillova, Gerrit Langer*, Tilo Strutz* PhD student: Petrus Zwart* Technicians: Venkataraman Parthasarathy*, Babu Pothineni*, Katja Schirwitz Visitors: Serge Cohen, Zbigniew Dauter, Francisco Fernandez Perez*, Christian Jelsch, Mattheos Kakaris, Joergen Koepke, Olga V. Koroleva, Richard J. Morris, Peter Østergaard, Tatiana V. Pegasova, Anastassis Perrakis, Denis V. Rebrikov, Wojciech Rypniewski, Jozef Sevcic , Elena V. Stepanova, Clemens Vonrhein
The design of a new versatile control system that will underlie future releases of the automated model-building package ARP/wARP is presented. A sophisticated expert system is under development that will transform ARP/wARP from a very useful model-building aid to a truly automated package capable of delivering complete, well refined and validated models comparable in quality to the result of intensive manual checking, rebuilding, hypothesis testing, refinement and validation cycles of an experienced crystallographer. In addition to the presentation of this control system, recent advances, ideas and future plans for improving the current model-building algorithms, especially for completing partially built models, are presented. Furthermore, a concept for integrating validation routines into the iterative model-building process is also presented.
Glia cell missing (GCM) transcription factors form a small family of transcriptional regulators in metazoans. The prototypical Drosophila GCM protein directs the differentiation of neuron precursor cells into glia cells, whereas mammalian GCM proteins are involved in placenta and parathyroid development. GCM proteins share a highly conserved 150 amino acid residue region responsible for DNA binding, known as the GCM domain. Here we present the crystal structure of the GCM domain from murine GCMa bound to its octameric DNA target site at 2.85 Å resolution. The GCM domain exhibits a novel fold consisting of two domains tethered together by one of two structural Zn ions. We observe the novel use of a β‐sheet in DNA recognition, whereby a five‐ stranded β‐sheet protrudes into the major groove perpendicular to the DNA axis. The structure combined with mutational analysis of the target site and of DNA‐contacting residues provides insight into DNA recognition by this new type of Zn‐containing DNA‐binding domain.
Glial cells missing (GCM) proteins form a small family of transcriptional regulators involved in different developmental processes. They contain a DNA-binding domain that is highly conserved from flies to mice and humans and consists of approximately 150 residues. The GCM domain of the mouse GCM homolog a was expressed in bacteria. Extended X-ray absorption fine structure and particle-induced X-ray emission analysis techniques showed the presence of two Zn atoms with four-fold coordination and cysteine/histidine residues as ligands. Zn atoms can be removed from the GCM domain by the Zn chelator phenanthroline only under denaturating conditions. This suggests that the Zn ions are buried in the interior of the GCM domain and that their removal abolishes DNA-binding because it impairs the structure of the GCM domain. Our results define the GCM domain as a new type of Zn-coordinating, sequence-specific DNA-binding domain.