It is difficult to imagine where the signaling community would be today without the Protein Data Bank. This visionary resource, established in the 1970s, has been an essential partner for sharing information between academics and industry for over 3 decades. We describe here the history of our journey with the protein kinases using cAMP-dependent protein kinase as a prototype. We summarize what we have learned since the first structure, published in 1991, why our journey is still ongoing, and why it has been essential to share our structural information. For regulation of kinase activity, we focus on the cAMP-binding protein kinase regulatory subunits. By exploring full-length macromolecular complexes, we discovered not only allostery but also an essential motif originally attributed to crystal packing. Massive genomic data on disease mutations allows us to now revisit crystal packing as a treasure chest of possible protein:protein interfaces where the biological significance and disease relevance can be validated. It provides a new window into exploring dynamic intrinsically disordered regions that previously were deleted, ignored, or attributed to crystal packing. Merging of crystallography with cryo-electron microscopy, cryo-electron tomography, NMR, and millisecond molecular dynamics simulations is opening a new world for the signaling community where those structure coordinates, deposited in the Protein Data Bank, are just a starting point!
Vaccinia virus interferes with early events of the activation pathway of the transcriptional factor NF-kB by binding to numerous host TIR-domain containing adaptor proteins. We have previously determined the X-ray structure of the A46 C-terminal domain; however, the structure and function of the A46 N-terminal domain and its relationship to the C-terminal domain have remained unclear. Here, we biophysically characterize residues 1-83 of the N-terminal domain of A46 and present the X-ray structure at 1.55 Å. Crystallographic phases were obtained by a recently developed ab initio method entitled ARCIMBOLDO_BORGES that employs tertiary structure libraries extracted from the Protein Data Bank; data analysis revealed an all β-sheet structure. This is the first such structure solved by this method which should be applicable to any protein composed entirely of β-sheets. The A46(1-83) structure itself is a β-sandwich containing a co-purified molecule of myristic acid inside a hydrophobic pocket and represents a previously unknown lipid-binding fold. Mass spectrometry analysis confirmed the presence of long-chain fatty acids in both N-terminal and full-length A46; mutation of the hydrophobic pocket reduced the lipid content. Using a combination of high resolution X-ray structures of the N- and C-terminal domains and SAXS analysis of full-length protein A46(1-240), we present here a structural model of A46 in a tetrameric assembly. Integrating affinity measurements and structural data, we propose how A46 simultaneously interferes with several TIR-domain containing proteins to inhibit NF-κB activation and postulate that A46 employs a bipartite binding arrangement to sequester the host immune adaptors TRAM and MyD88.
Computational docking is a useful tool for predicting macromolecular complexes, which are often difficult to determine experimentally. Here, we present the DOT2 software suite, an updated version of the DOT intermolecular docking program. DOT2 provides straightforward, automated construction of improved biophysical models based on molecular coordinates, offering checkpoints that guide the user to include critical features. DOT has been updated to run more quickly, allow flexibility in grid size and spacing, and generate an infinitive complete list of favorable candidate configurations. Output can be filtered by experimental data and rescored by the sum of electrostatic and atomic desolvation energies. We show that this rescoring method improves the ranking of correct complexes for a wide range of macromolecular interactions and demonstrate that biologically relevant models are essential for biologically relevant results. The flexibility and versatility of DOT2 accommodate realistic models of complex biological systems, improving the likelihood of a successful docking outcome. © 2013 Wiley Periodicals, Inc.
ABSTRACTProtein–DNA interactions are essential for many biological processes. X‐ray crystallography can provide high‐resolution structures, but protein‐DNA complexes are difficult to crystallize and typically contain only small DNA fragments. Thus, there is a need for computational methods that can provide useful predictions to give insights into mechanisms and guide the design of new experiments. We used the program DOT, which performs an exhaustive, rigid‐body search between two macromolecules, to investigate four diverse protein–DNA interactions. Here, we compare our computational results with subsequent experimental data on related systems. In all cases, the experimental data strongly supported our structural hypotheses from the docking calculations: a mechanism for weak, nonsequence‐specific DNA binding by a transcription factor, a large DNA‐binding footprint on the surface of the DNA‐repair enzyme uracil‐DNA glycosylase (UNG), viral and host DNA‐binding sites on the catalytic domain of HIV integrase, and a three‐DNA‐contact model of the linker histone bound to the nucleosome. In the case of UNG, the experimental design was based on the DNA‐binding surface found by docking, rather than the much smaller surface observed in the crystallographic structure. These comparisons demonstrate that the DOT electrostatic energy gives a good representation of the distinctive electrostatic properties of DNA and DNA‐binding proteins. The large, favourably ranked clusters resulting from the dockings identify active sites, map out large DNA‐binding sites, and reveal multiple DNA contacts with a protein. Thus, computational docking can not only help to identify protein–DNA interactions in the absence of a crystal structure, but also expand structural understanding beyond known crystallographic structures. Proteins 2013; 81:2106–2118. © 2013 Wiley Periodicals, Inc.
X-ray crystallography provides excellent structural data on protein-DNA interfaces, but crystallographic complexes typically contain only small fragments of large DNA molecules. We present a new approach that can use longer DNA substrates and reveal new protein-DNA interactions even in extensively studied systems. Our approach combines rigid-body computational docking with hydrogen/deuterium exchange mass spectrometry (DXMS). DXMS identifies solvent-exposed protein surfaces; docking is used to create a 3-dimensional model of the protein-DNA interaction. We investigated the enzyme uracil-DNA glycosylase (UNG), which detects and cleaves uracil from DNA. UNG was incubated with a 30 bp DNA fragment containing a single uracil, giving the complex with the abasic DNA product. Compared with free UNG, the UNG-DNA complex showed increased solvent protection at the UNG active site and at two regions outside the active site: residues 210-220 and 251-264. Computational docking also identified these two DNA-binding surfaces, but neither shows DNA contact in UNG-DNA crystallographic structures. Our results can be explained by separation of the two DNA strands on one side of the active site. These non-sequence-specific DNA-binding surfaces may aid local uracil search, contribute to binding the abasic DNA product and help present the DNA product to APE-1, the next enzyme on the DNA-repair pathway.
The drastic reduction of novel folds in proteins newly determined by x-ray crystallography suggests that large macromolecules are built from domains with already known structures. We have developed a novel integrative protocol that combines experimentally-measured topographic surfaces of single molecules with atomic coordinates of molecular constituents of large proteins or assemblies. Topographic surfaces are obtained using high-resolution atomic force microscopy (AFM) imaging. The present integrative method is based on real-space docking of macromolecular constituents beneath the experimental topographic surface. Assembly of molecular constituents is performed using a combinatorial approach. Only steric clashes between assembled constituents are computed; assemblies having more than a given threshold of bumps are eliminated. The goodness of fit is obtained by a score named E-factor which determines the agreement between the experimental topographic surface with that of the assembled constituents. A proof of concept has been determined on three different systems: Immunoglobulin G, Tobacco mosaic virus, and Aquaporin Z. Results demonstrated that partial topographic surface is adequate for complete macromolecular reconstruction. This protocol may be extremely useful for "difficult proteins" such as membrane proteins, partially unfolded proteins, and hard-to-produce proteins.
TNT (Tronrud et al., 1987) is a computer program package that optimizes the parameters of a molecular model given a set of observations and indicates the location of errors that it cannot correct. Its authors presume the principal set of observations to be the structure factors observed in a single-crystal diffraction experiment. To complement such a data set, which for most macromolecules has limitations, stereochemical restraints such as standard bond lengths and angles are also used as observations. A molecule is parameterized as a set of atoms, each with a position in space, an isotropic B factor and an occupancy. The complete model also includes an overall scale factor, which converts the arbitrary units of the measured structure factors to e Å , and a two-parameter model of the electron density of the bulk solvent. Because a TNT model of a macromolecule does not allow anisotropic B factors, TNT cannot be used to finish the refinement of any structure that diffracts to high enough resolution to justify the use of these parameters. If one has a crystal that diffracts to 1.4 Å or better, the final model should probably include these parameters and TNT cannot be used. One may still use TNT in the early stages of such a refinement because one usually begins with only isotropic B’s. At the other extreme of resolution, TNT begins to break down with data sets limited to only about 3.5 Å data. This breakdown occurs for two reasons. First, at 3.5 Å resolution, the maps can no longer resolve -sheet strands or -helices. The refinement of a model against data of such low resolution requires strong restraints on dihedral angles and hydrogen bonds – tasks for which TNT is not well suited. Second, the errors in an initial model constructed with only 3.5 Å data are usually of such a magnitude and quality that the function minimizer in TNT cannot correct them.
TM0077 from Thermotoga maritima is a member of the carbohydrate esterase family 7 and is active on a variety of acetylated compounds, including cephalosporin C. TM0077 esterase activity is confined to short‐chain acyl esters (C2–C3), and is optimal around 100°C and pH 7.5. The positional specificity of TM0077 was investigated using 4‐nitrophenyl‐β‐ D ‐xylopyranoside monoacetates as substrates in a β‐xylosidase‐coupled assay. TM0077 hydrolyzes acetate at positions 2, 3, and 4 with equal efficiency. No activity was detected on xylan or acetylated xylan, which implies that TM0077 is an acetyl esterase and not an acetyl xylan esterase as currently annotated. Selenomethionine‐substituted and native structures of TM0077 were determined at 2.1 and 2.5 Å resolution, respectively, revealing a classic α/β‐hydrolase fold. TM0077 assembles into a doughnut‐shaped hexamer with small tunnels on either side leading to an inner cavity, which contains the six catalytic centers. Structures of TM0077 with covalently bound phenylmethylsulfonyl fluoride and paraoxon were determined to 2.4 and 2.1 Å, respectively, and confirmed that both inhibitors bind covalently to the catalytic serine (Ser188). Upon binding of inhibitor, the catalytic serine adopts an altered conformation, as observed in other esterase and lipases, and supports a previously proposed catalytic mechanism in which Ser hydroxyl rotation prevents reversal of the reaction and allows access of a water molecule for completion of the reaction. Proteins 2012. © 2012 Wiley Periodicals, Inc.
The TNT refinement package was created in the late 1970s and its development continued for about 25 years. Its design included many features not present in its contemporaries and allowed for the testing and addition of many novel tools during its lifespan. These include the implementation of stereochemical restraints to molecules in other asymmetric units (both bonded and non-bonded), space-group-optimized fast Fourier transforms, which allowed rapid crystallographic calculations, and a quick and easy method to model the scattering of bulk solvent, which allowed the use of low-resolution data in refinement. TNT is no longer being developed by its original authors but its code has been included in Global Phasing, Ltd's program Buster, where considerable improvements have been made.
Identifying conserved pockets on the surfaces of a family of proteins can provide insight into conserved geometric features and sites of protein-protein interaction. Here we describe mapping and comparison of the surfaces of aligned crystallographic structures, using the protein kinase family as a model. Pockets are rapidly computed using two computer programs, FADE and Crevasse. FADE uses gradients of atomic density to locate grooves and pockets on the molecular surface. Crevasse, a new piece of software, splits the FADE output into distinct pockets. The computation was run on 10 kinase catalytic cores aligned on the alphaF-helix, and the resulting pockets spatially clustered. The active site cleft appears as a large, contiguous site that can be subdivided into nucleotide and substrate docking sites. Substrate specificity determinants in the active site cleft between serine/threonine and tyrosine kinases are visible and distinct. The active site clefts cluster tightly, showing a conserved spatial relationship between the active site and alphaF-helix in the C-lobe. When the alphaC-helix is examined, there are multiple mechanisms for anchoring the helix using spatially conserved docking sites. A novel site at the top of the N-lobe is present in all the kinases, and there is a large conserved pocket over the hinge and the alphaC-beta4 loop. Other pockets on the kinase core are strongly conserved but have not yet been mapped to a protein-protein interaction. Sites identified by this algorithm have revealed structural and spatially conserved features of the kinase family and potential conserved intermolecular and intramolecular binding sites.
Cyclic nucleotides (cAMP and cGMP) regulate multiple intracellular processes and are thus of a great general interest for molecular and structural biologists. To study the allosteric mechanism of different cyclic nucleotide binding (CNB) domains, we compared cAMP-bound and cAMP-free structures (PKA, Epac, and two ionic channels) using a new bioinformatics method: local spatial pattern alignment. Our analysis highlights four major conserved structural motifs: 1) the phosphate binding cassette (PBC), which binds the cAMP ribose-phosphate, 2) the "hinge," a flexible helix, which contacts the PBC, 3) the beta(2,3) loop, which provides precise positioning of an invariant arginine from the PBC, and 4) a conserved structural element consisting of an N-terminal helix, an eight residue loop and the A-helix (N3A-motif). The PBC and the hinge were included in the previously reported allosteric model, whereas the definition of the beta(2,3) loop and the N3A-motif as conserved elements is novel. The N3A-motif is found in all cis-regulated CNB domains, and we present a model for an allosteric mechanism in these domains. Catabolite gene activator protein (CAP) represents a trans-regulated CNB domain family: it does not contain the N3A-motif, and its long range allosteric interactions are substantially different from the cis-regulated CNB domains.
Structures of set of serine-threonine and tyrosine kinases were investigated by the recently developed bioinformatics tool Local Spatial Patterns (LSP) alignment. We report a set of conserved motifs comprised mostly of hydrophobic residues. These residues are scattered throughout the protein sequence and thus were not previously detected by traditional methods. These motifs traverse the conserved protein kinase core and play integrating and regulatory roles. They are anchored to the F-helix, which acts as an organizing “hub” providing precise positioning of the key catalytic and regulatory elements. Consideration of these discovered structures helps to explain previously inexplicable results.
Protein kinases are a large family of enzymes heavily involved in signal transduction, regulation of metabolism, and control of cell growth and differentiation. These functions require precise recognition of widely diverse signals and substrates, and very detailed control of protein kinase activity. Large molecules interact primarily through recognition of surface features. Comparison of surfaces is complicated by both sequence diversity and conformational variability, including multiple possible rotameric states of side chains. We used a recently developed method of protein surface comparison to compare different serine/threonine and tyrosine kinases. As we have shown, two hydrophobic cores inside a protein kinase molecule are connected by a unique formation, called the “spine”. It exists only in the active conformation of protein kinases and is dynamically disassembled during the inactivation process. Detection of such structures by any other method was not possible as the residues which comprise the spine do not form any sequence or 3D motifs in a traditional sense.
Protein kinases represent a large protein superfamily which regulates numerous processes in living cells. Structures of several serine‐threonine and tyrosine kinases were investigated by recently developed bioinformatics method: Local Spatial Patterns alignment. Previously this method proved to be an effective tool, which is capable to detect highly conserved spatial patterns of amino acid residues. These residues do not form traditional motifs in terms of sequence or protein fold and can not be detected by traditional methods. An example of that was discovery of “the spine” in protein kinases, an important regulatory element. In this work we report detailed analysis of the protein kinase catalytic core. A set of conserved ensembles comprised mostly by hydrophobic residues was detected. The formations traverse protein kinase molecules and play integrating and regulatory roles. They are anchored to the F‐helix located in the middle of the large lobe, which acts as an organizing “hub” providing exact positioning of the key catalytic and regulatory elements. This work was supported by National Science Foundation Grant NSF‐DBI 99111196 and National Institute of General Medical Sciences Grants GM70996 (to L.F.T.E.); GM19301 and National Science Foundation Grant DBI0217951 (to S.S.T.).
The two isoforms (RI and RII) of the regulatory (R) subunit of cAMP-dependent protein kinase or protein kinase A (PKA) are similar in sequence yet have different biochemical properties and physiological functions. To further understand the molecular basis for R-isoform-specificity, the interactions of the RIIβ isoform with the PKA catalytic (C) subunit were analyzed by amide H/2H exchange mass spectrometry to compare solvent accessibility of RIIβ and the C subunit in their free and complexed states. Direct mapping of the RIIβ-C interface revealed important differences between the intersubunit interfaces in the type I and type II holoenzyme complexes. These differences are seen in both the R-subunits as well as the C-subunit. Unlike the type I isoform, the type II isoform complexes require both cAMP-binding domains, and ATP is not obligatory for high affinity interactions with the C-subunit. Surprisingly, the C-subunit mediates distinct, overlapping surfaces of interaction with the two R-isoforms despite a strong homology in sequence and similarity in domain organization. Identification of a remote allosteric site on the C-subunit that is essential for interactions with RII, but not RI subunits, further highlights the considerable diversity in interfaces found in higher order protein complexes mediated by the C-subunit of PKA.
Proteins that undergo cooperative unfolding events display EX1 kinetic signatures in hydrogen exchange mass spectra. The hallmark bimodal isotope pattern observed for EX1 kinetics is distinct from the binomial isotope pattern for uncorrelated exchange (EX2), the normal exchange regime for folded proteins. Detection and characterization of EX1 kinetics is simple when the cooperative unit is large enough that the isotopic envelopes in the bimodal pattern are resolved in the m/z scale but become complicated in cases where the unit is small or there is a mixture of EX1 and EX2 kinetics. Here we describe a data interpretation method involving peak width analysis that makes characterization of EX1 kinetics simple and rapid. The theoretical basis for EX1 and EX2 isotopic signatures and the effects each have on peak width are described. Modeling of EX2 widening and analysis of empirical data for proteins and peptides containing purely EX2 kinetics showed that the amount of widening attributable to stochastic forwardand back exchange in a typical experiment is small and can be quantified. Proteins and peptides with both obvious and less obvious EX1 kinetics were analyzed with the peak width method. Such analyses provide the half-life for the cooperative unfolding event and the relative number of residues involved. Automated analysis of peak width was performed with custom Excel macros and the DEX software package. Peak width analysis is robust, capable of automation, and provides quick interpretation of the key information contained in EX1 kinetic events. (J Am Soc Mass Spectrom 2006, 17, 1498–1509) © 2006 American Society for Mass Spectrometry
J. Ben Rosen合作论文数California Institute for Telecommunications and Information Technology2