ALK (anaplastic lymphoma kinase) is an RTK (receptor tyrosine kinase) of the IRK (insulin receptor kinase) superfamily, which share an YXXXYY autophosphorylation motif within their A-loops (activation loops). A common activation and regulatory mechanism is believed to exist for members of this superfamily typified by IRK and IGF1RK (insulin-like growth factor receptor kinase-1). Chromosomal translocations involving ALK were first identified in anaplastic large-cell lymphoma, a subtype of non-Hodgkin's lymphoma, where aberrant fusion of the ALK kinase domain with the NPM (nucleophosmin) dimerization domain results in autophosphosphorylation and ligand-independent activation. Activating mutations within the full-length ALK kinase domain, most commonly R1275Q and F1174L, which play a major role in neuroblastoma, were recently identified. To provide a structural framework for understanding these mutations and to guide structure-assisted drug discovery efforts, the X-ray crystal structure of the unphosphorylated ALK catalytic domain was determined in the apo, ADP- and staurosporine-bound forms. The structures reveal a partially inactive protein kinase conformation distinct from, and lacking, many of the negative regulatory features observed in inactive IGF1RK/IRK structures in their unphosphorylated forms. The A-loop adopts an inhibitory pose where a short proximal A-loop helix (alphaAL) packs against the alphaC helix and a novel N-terminal beta-turn motif, whereas the distal portion obstructs part of the predicted peptide-binding region. The structure helps explain the reported unique peptide substrate specificity and the importance of phosphorylation of the first A-loop Tyr1278 for kinase activity and NPM-ALK transforming potential. A single amino acid difference in the ALK substrate peptide binding P-1 site (where the P-site is the phosphoacceptor site) was identified that, in conjunction with A-loop sequence variation including the RAS (Arg-Ala-Ser)-motif, rationalizes the difference in the A-loop tyrosine autophosphorylation preference between ALK and IGF1RK/IRK. Enzymatic analysis of recombinant R1275Q and F1174L ALK mutant catalytic domains confirms the enhanced activity and transforming potential of these mutants. The transforming ability of the full-length ALK mutants in soft agar colony growth assays corroborates these findings. The availability of a three-dimensional structure for ALK will facilitate future structure-function and rational drug design efforts targeting this receptor tyrosine kinase.
A novel aminoacyl-tRNA synthetase that contains an iron-sulfur cluster in the tRNA anticodon-binding region and efficiently charges tRNA with tryptophan has been found in Thermotoga maritima. The crystal structure of TmTrpRS (tryptophanyl-tRNA synthetase; TrpRS; EC 6.1.1.2) reveals an iron-sulfur [4Fe-4S] cluster bound to the tRNA anticodon-binding (TAB) domain and an L-tryptophan ligand in the active site. None of the other T. maritima aminoacyl-tRNA synthetases (AARSs) contain this [4Fe-4S] cluster-binding motif (C-x₂₂-C-x₆-C-x₂-C). It is speculated that the iron-sulfur cluster contributes to the stability of TmTrpRS and could play a role in the recognition of the anticodon.
The crystal structures of two homologous endopeptidases from cyanobacteria Anabaena variabilis and Nostoc punctiforme were determined at 1.05 and 1.60 A resolution, respectively, and contain a bacterial SH3-like domain (SH3b) and a ubiquitous cell-wall-associated NlpC/P60 (or CHAP) cysteine peptidase domain. The NlpC/P60 domain is a primitive, papain-like peptidase in the CA clan of cysteine peptidases with a Cys126/His176/His188 catalytic triad and a conserved catalytic core. We deduced from structure and sequence analysis, and then experimentally, that these two proteins act as gamma-D-glutamyl-L-diamino acid endopeptidases (EC 3.4.22.-). The active site is located near the interface between the SH3b and NlpC/P60 domains, where the SH3b domain may help define substrate specificity, instead of functioning as a targeting domain, so that only muropeptides with an N-terminal L-alanine can bind to the active site.
The TM1249 gene (purH) of Thermotoga maritima encodes a bifunctional enzyme with 5-aminoimidazole-4-carboxamide ribonucleotide (AICAR) transformylase (AICAR Tfase; EC 2.1.2.3) and inosine 5′-monophosphate (IMP) cyclohydrolase (IMPCH; EC 3.5.4.10) activities. These activities represent the final two steps of the de novo purine biosynthesis pathway. PurH contains two functional regions, with the IMPCH activity performed by the N-terminal domain and the AICAR Tfase activity by the C-terminal region, which is composed of two structural domains. The AICAR Tfase region catalyzes the transfer of the formyl group from the cofactor 10-formyl-tetrahydrofolate (10-f-THF) to the substrate AICAR, while the IMPCH domain catalyzes the intramolecular cyclization of 5-formyl-AICAR (FAICAR) to IMP [Fig. 1(A)]. Crystal structure of PurH from Thermotoga maritima. A: The reactions carried out by PurH. B: Stereo ribbon diagram of the TM1249 monomer, color-coded from N-terminus (blue) to C-terminus (red), showing the domain organization. Helices (H1–H18) and β-strands (β1–β21) are labeled. C: Diagram showing the secondary structural elements of TM1249 superimposed on its primary sequence. The α-helices, 310-helices, β-strands, β-turns, and γ-turns are indicated. The β-hairpins are depicted as red loops. The four β-sheets are labeled as A, B, C, and D, respectively. The TM1249 monomer has a molecular weight of 49.7 kDa (residues 1–452) and a calculated isoelectric point of 5.49. Crystal structures for the avian1 and human2 forms of this enzyme have been previously reported. In bacteria and eukaryotes, the IMPCH and AICAR Tfase activities reside on a single, bifunctional enzyme. A similar bifunctional protein coded by a purH gene has not been found in archaea. Instead, these activities have been identified in a nonhomologous IMPCH3 and a separate AICAR Tfase.4 Previously, we reported the crystal structures of PurS (TM1244)5 and smPurL (TM1246)6 from Thermotoga maritima. These two proteins, along with TM1245, form the formylglycinamide ribonucleotide amidotransferase complex (EC 6.3.5.3), which catalyzes an earlier step in the de novo purine biosynthesis pathway. Here, we report the crystal structure of PurH from Thermotoga maritima (TM1249), a bacterial form of this bifunctional enzyme, determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG).7 Because this enzyme serves an important role in nucleotide biosynthesis in bacteria and eukaryotes, it is a potential target for the development of antibacterial and anticancer agents. The TM1249 structure offers new insights about the mechanistic features of this essential biochemical reaction. PurH from Thermotoga maritima (TIGR: TM1249; Swiss-Prot: Q9X0X6) was amplified by polymerase chain reaction (PCR) from genomic DNA using PfuTurbo DNA polymerase (Stratagene) and primers corresponding to the predicted 5′ and 3′ ends. The PCR product was cloned into plasmid pMH4, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The cloning junctions were confirmed by DNA sequencing. Protein expression was performed in a modified Terrific Broth using the Escherichia coli strain GeneHogs (Invitrogen). At the end of fermentation, lysozyme was added to the culture to a final concentration of 250 μg/mL, and the cells were harvested. After one freeze/thaw cycle, the cells were sonicated in lysis buffer [50 mM Tris pH 7.9, 50 mM NaCl, 10 mM imidazole, 1 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the lysate was clarified by centrifugation at 32,500 g for 30 min. The soluble fraction was passed over nickel-chelating resin (GE Healthcare) pre-equilibrated with lysis buffer, the resin was washed with wash buffer [50 mM Tris pH 7.9, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP], and the protein was eluted with elution buffer [20 mM Tris pH 7.9, 300 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP]. The eluate was diluted 10-fold with buffer Q [20 mM Tris pH 7.9, 5% (v/v) glycerol, 0.25 mM TCEP] containing 50 mM NaCl and loaded onto a RESOURCE Q column (GE Healthcare) pre-equilibrated with the same buffer. The protein was eluted with a linear gradient of 50–500 mM NaCl in buffer Q, buffer exchanged with crystallization buffer [20 mM Tris pH 7.9, 150 mM NaCl, 0.25 mM TCEP], and concentrated for crystallization assays to 20 mg/mL by centrifugal ultrafiltration (Millipore). TM1249 was crystallized using the nanodroplet vapor diffusion method8 with standard JCSG crystallization protocols.7 The crystallization reagent that produced the crystal used for the structure solution contained 20% (w/v) polyethylene glycol (PEG) 6000 and 0.1 M citrate pH 5.0. PEG 200 was added as a cryoprotectant to a final concentration of 10% (v/v). Initial screening for diffraction was carried out using the Stanford Automated Mounting system (SAM)9 at the Stanford Synchrotron Radiation Laboratory (SSRL, Menlo Park, CA). The crystal was indexed in triclinic space group P1 (Table I).10, 11 Molecular weight and oligomeric state of TM1249 were determined using a 1 cm × 30 cm Superdex 200 column (GE Healthcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl, and 0.02% (w/v) sodium azide. Diffraction data were collected at the Advanced Light Source (ALS, Berkeley, CA) on beamline 5.0.1 at 100K using a Quantum 210 CCD detector (ADSC). Data were integrated and reduced using Denzo12 and then scaled with the program SCALEPACK.12 Primary phasing was accomplished using the JCSG molecular replacement (MR) protocol.13 A conserved functional domain analysis14 of the amino acid sequence of TM1249 showed that it contains an N-terminal IMPCH domain and a C-terminal AICAR Tfase region. MR search models were constructed from the avian homolog (PDB accession code: 1g8m) using WHATIF,15 based on FFAS16 alignments. To accommodate a possible difference in the relative positioning of the domains in TM1249 compared with the avian homolog, the N-terminal domain (residues 1–156) and the C-terminal region (residues 187–452) were treated as independent entities in the search model. The hinge region (residues 157–186), which separates the N-terminal domain and the C-terminal region, was omitted from the search. After initial phasing with MOLREP17 from the CCP4 suite,11 a preliminary model was traced from the MR-phased electron density maps using ARP/wARP.18 Further refinement was carried out using REFMAC519 and XtalView.20 Data collection, model, and refinement statistics are summarized in Table I. Analysis of the stereochemical quality of the model was accomplished using AutoDepInputTool,21 MolProbity,22 SFcheck 4.0,23 and WHATIF 5.0.15 Figure 1(C) was adapted from an analysis using PDBsum,24 and all other figures were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors of TM1249 have been deposited in the PDB and are accessible under the code 1zcz. The three-dimensional structure of TM1249 [Fig. 1(B)] was determined to 1.88 Å resolution by MR. The refined structure includes a dimer (residues 1–452 for chains A and B), two K+ ions, two PEG 200 molecules, and 636 water molecules in the asymmetric unit. The two PEG molecules are symmetrically positioned near the noncrystallographic twofold axis relating the two monomers of the dimer, and each is within hydrogen bonding distance of Ala 161 in the hinge region that links the N-terminal domain and the C-terminal region. Electron density was not observed for residues from the expression and purification tag. The Matthews' coefficient (Vm)25 is 2.3 Å3/Da, and the estimated solvent content is 46.1%. A Ramachandran plot produced by MolProbity22 shows that 98.3% and 100% of the residues are in favored and allowed regions, respectively. The peptide bond between Ser 355 and Asn 356 (in chain A) is in the cis conformation, with ϕ = −54.8° and ψ = −78.0° for Ser 355, and ϕ = −64.3° and ψ = 135.5° for Asn 356. These residues surround a cavity that forms the proposed AICAR Tfase active site. The K+ ion bound to each monomer is located ∼4 Å from the side-chain hydroxyl of Ser 355 [Fig. 2(A)]. A bound K+ ion and the cis peptide bond are also observed in the avian1 and human structures,2 suggesting that these conserved features are important determinants of AICAR Tfase activity. The biologically relevant dimer of PurH from Thermotoga maritima, and a shift in the relative orientation of the two functional regions of TM1249 compared to the avian enzyme. A: Stereo ribbon diagram of the dimer of TM1249 showing the locations of the IMPCH and AICAR Tfase regions. The active sites of IMPCH and AICAR Tfase are labeled with cyan and orange asterisks, respectively. Bound K+ atoms are shown as spheres (grey) and PEG molecules as sticks (pink) (B) (Left) The relative shift in orientation between the N-terminal IMPCH domain of TM1249 (blue) and that of the avian enzyme (grey) (PDB accession code: 1m9n). (Right) FATCAT-optimized alignment of the N- and C-terminal regions of the avian enzyme and TM1249 after implementing a ∼90° twist in the hinge region of TM1249. A monomer of the avian protein is shown in grey, the N-terminal domain of TM1249 is in green, and the C-terminal domains of TM1249 are in red. The TM1249 monomer contains 21 β-strands (β1–β21) in four β-sheets, 14 α-helices (H2, H3, H6–H15, H17, and H18), and four 310-helices (H1, H4, H5, and H16) [Fig. 1(B)]. The total β-strand, α-helical, and 310-helical content is 22.1%, 35.2%, and 2.7%, respectively. TM1249 comprises three structural domains: an N-terminal domain, which functions as an IMPCH, and two C-terminal domains, which function as an AICAR Tfase [Fig. 1(B)]. The N-terminal IMPCH domain (residues 1–156), which adopts a Rossmann fold topology, has been classified in the SCOP25, 26 database as a methylglyoxal synthase-like fold. Compared with the canonical nucleotide-binding Rossmann fold, which contains a six-stranded β-sheet, this α/β fold lacks the last β-strand and, therefore, consists of a β-sheet formed by five parallel β-strands. The second (residues 187–301) and third (residues 328–452) structural domains are related by tandem duplication and, along with residues 302–327, compose the AICAR Tfase region (residues 187–452) [Fig. 1(B)]. When aligned with each other, these two domains show an 18% sequence identity and an RMSD of 1.4 Å over 112 aligned Cα atoms. These domains adopt a three-layered α+β fold comprised of a mixed six-stranded β-sheet and four α-helices arranged as βαββαβαβαβ, which is similar to the cytidine deaminase-like fold in SCOP.25, 26 In addition, both domains are preceded by β-hairpins (1st: residues 164–182; 2nd: residues 307–318) that connect the three domains. These β-hairpins are involved in extensive inter-subunit interactions and contribute significantly to the dimeric interface, which buries a surface area of 5270 Å2 from each monomer [Fig. 2(A)]. The dimerization interface is very extensive, running along a noncrystallographic twofold axis that covers the full length of the long axis of the molecule. As in the avian1 and human structures,2 the AICAR Tfase active site of TM1249 is formed by the dimerization of the AICAR Tfase regions. Moreover, analytical size exclusion chromatography coupled with static light scattering also indicates that the dimer is the biologically relevant form. Unexpectedly, β-strands β6 from the two monomers form a two-stranded antiparallel β-sheet that packs against α-helix H9 from the IMPCH domain in the TM1249 dimer [Fig 2(A)]. This feature is not present in the avian1 or human structures.2 The top of this β-sheet forms the bottom of a hydrophobic circular cleft delineated by β7, β14, and β15. This hydrophobic patch is located at the hinge region that separates the two functional halves of TM1249. In the crystal structure of TM1249, a PEG 200 molecule partially shields this region from the solvent [Fig. 2(A)]. As anticipated, a structural comparison of TM1249 using the DALI server26 revealed several highly similar structures with statistically significant Z-scores. The avian enzyme (PDB: 1g8m, 1m9n) was found as the top hit (Z = 33.6). Individually, the IMPCH domain and AICAR Tfase region in TM1249 show topological similarity with the corresponding regions in the avian enzyme. The RMSD for the structural alignment of the N-terminal IMPCH domains of TM1249 (residues 1–156) and the avian enzyme (residues 1–197) is 2.2 Å over 154 aligned residues with 42% sequence identity. For the C-terminal AICAR Tfase regions, the RMSD is 1.9 Å over 281 aligned residues with 36% sequence identity. Although these proteins share a similar architecture for the individual regions, the arrangement of the N-terminal IMPCH domain with respect to the C-terminal AICAR Tfase region is significantly different between TM1249 and the avian or human enzyme. Consequently, the RMSD for a structural alignment of TM1249 and the avian enzyme reported by DALI is 12.3 Å over 389 aligned Cα atoms with 38% sequence identity [Fig. 2(B), left]. A similar alignment of these two proteins can be performed using the FATCAT server27 with an RMSD of 13.4 Å (opt-RMSD of 6.3 Å) over 383 Cα atoms when treating the protein chains as rigid bodies. However, the FATCAT server can further optimize this structural alignment by incorporating twists around the hinge regions of the protein chains. When conformational flexibility (one twist) is assumed, these structures can be superimposed with an opt-RMSD of 2.1 Å over 442 Cα atoms [Fig. 2(B), right], revealing the striking difference in the arrangement of the N-terminal IMPCH domain with respect to the C-terminal AICAR Tfase region between the avian and T. maritima structures. The relative arrangement of the two functional regions of this bifunctional enzyme in different species is determined by the structure of the hinge region. In the TM1249 dimer, the two-stranded anti-parallel β-sheet, comprised of strands β6 [Fig. 2(A)] from the symmetry-related monomers, forms a rigid linker that dictates that both functional regions remain on the same side of the dimerization axis, as opposed to the 90° twist that is observed in the avian and human structures. In the avian protein, the linker region is more flexible and is not constrained by the β-sheet formation. Although TM1249 and numerous homologs share a significant degree of overall sequence identity (i.e., >30%), their structures suggest an important role for the hinge region (157–186 in TM1249) in the relative positioning of the two functional regions. Therefore, an NCBI-BLAST29 search was performed using residues from the hinge region of TM1249, which, surprisingly, revealed only a few examples of similar sequences. For example, the hinge region of TM1249 shows 73% sequence identity with the corresponding region in AICAR Tfase IMPCH from Erythrobacter sp. NAP1. This suggests that the interactions between the IMPCH and AICAR Tfase regions in Thermotoga maritima and Erythrobacter sp. NAP1 are likely to be similar and that the two bacterial proteins adopt a similar arrangement of the two regions. The IMPCH and AICAR Tfase active sites of TM1249 were inferred from sequence and structural comparison to the avian enzyme. The IMPCH active sites are located in clefts on the two monomers [Fig. 3(A)]. One boundary of this cleft, which is most likely adjacent to the ribose-phosphate moiety of the substrate and product, is at one end of a β-sheet comprised of β1, β2, β4, and β5. The N-terminal ends of H6 and H7 form a second boundary of this active site cleft for interaction with the 5-aminoimidazole-4-carboxamide moiety of the substrate. The active site region contains the highly conserved residues Asp 94, Tyr 88, Lys 63, Thr 64, Thr 34, Ser 31, Lys 11, and Ser 7, which adopt side-chain conformations similar to those in the avian homolog [Fig. 3(A)]. The active sites of PurH from Thermotoga maritima. A: Stereo diagram of residues in the putative IMPCH active site of TM1249 superimposed on the avian structure (PDB ID: 1m9n). The carbon atoms on the side chains of TM1249 are shown in green, while the carbon atoms on the active site ligand (XMP) and on the side chains of the avian enzyme are in cyan. Residues from the avian enzyme are shown in parentheses. The side chain of Lys 11 in TM1249 is disordered and, consequently, only the β carbon atom is shown. The location of XMP, an inhibitor of IMPCH, from the avian crystal structure is shown. B: Stereo diagram of residues in the putative AICAR Tfase active site of TM1249 superimposed on the avian structure, with the same color and labeling scheme used in (A). The AICAR substrate is from the avian structure. The two AICAR Tfase active sites are situated in large clefts at the dimer interface [Figs. 2(A), 3(B)]. A structural comparison of TM1249 to the avian enzyme reveals similarities in their AICAR Tfase active sites [Fig. 3(B)]. As in the avian enzyme, the AICAR Tfase active site of TM1249 is situated between the two cytidine deaminase-like domains of one monomer and the C-terminal cytidine deaminase-like domain of the second monomer.28 The substrate-bound avian structure reveals several conserved residues that are involved in AICAR binding and in the activation of the 5-amino group of AICAR for formyl group transfer. These residues are conserved in TM1249 and include Phe 401, Arg 376, and His 452 from one monomer, and Glu 275, His 224, Lys 223, Tyr 171, and Arg 170 from the other. In TM1249, difference electron density was observed near one of the putative AICAR-binding sites adjacent to Lys 223 and His 224. This difference electron density was not modeled in the deposited coordinates but can likely be attributed to citrate from the crystallization buffer. The PurH family contains hundreds of homologs in bacteria and eukaryotes, and the TM1249 structure provides a more suitable template for the bacterial homologs. Models for these homologs can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=1ZCZA. The two independent active sites in this bacterial bifunctional enzyme are very similar to those in its eukaryotic counterparts. The most striking structural difference between TM1249 and the avian homolog is the ∼90° rotation in the relative orientation of the N-terminal IMPCH domain and C-terminal AICAR region. TM1249 shares a >30% sequence identity with both the avian and human homologs, and, not surprisingly, the architecture of the individual IMPCH and AICAR Tfase regions shows a substantial degree of structural similarity across this broad phylogenetic spectrum. However, a prediction of the relative positioning of the two conserved functional regions is more challenging, and the structures can reveal significant differences in the interactions between the regions. These are likely related to differences in coupling between IMPCH and AICAR Tfase activities in TM1249 versus the eukaryotic enzymes. The JCSG has developed The Open Protein Structure Annotation Network (TOPSAN), a wiki-based community project to collect, share, and distribute information about protein structures determined at PSI centers. TOPSAN offers a combination of automatically generated, as well as comprehensive, expert-curated annotations, provided by JCSG personnel and members of the research community. Additional information about TM1249 is available at https://www.topsan.org/1ZCZ. Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory (SSRL) and the Advanced Light Source (ALS). The SSRL is a national user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U.S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institute of General Medical Sciences or the National Institutes of Health.
Methionine is one of 10 essential dietary amino acids in mammals; however, most bacteria, fungi, and plants are able to synthesize methionine de novo from aspartic acid. While the methionine biosynthetic pathway shares many common steps in microorganisms and plants, the pathway diverges in their activation of homoserine,1-4 the initial step of methionine biosynthesis. Bacteria, fungi, and archaea acylate homoserine, but most plants phosphorylate the γ-hydroxyl of homoserine for activation.5 The acylation reactions in bacterial and fungal systems utilize very similar substrates and most likely follow similar reaction mechanisms.6 Upon acylation, the acylhomoserine is shuttled to cystathionine γ-synthase for condensation with cysteine to form cystathionine. Subsequently, the cystathionine is hydrolyzed to form homocysteine, which is then methylated to produce methionine.7 In bacteria, two distinct families of enzymes are responsible for the activation of homoserine. Succinylation of homoserine in some bacteria, such as Escherichia coli and Bacillus cereus, is catalyzed by homoserine O-succinyltransferase (HTS; E.C 2.3.1.46), while other bacteria, such as Haemophilus influenzae, Pseudomonas aeruginosa, and Mycobacterium tuberculosis, acetylate homoserine via homoserine O-acetyltransferase (HTA; E.C 2.3.1.31). HTAs and HTSs fall into two distinct families based on primary sequence, but membership in a family does not necessarily predict substrate preference or specificity.4 While a representative member of the HTA family from Haemophilus influenzae has been structurally characterized,8 no structures from the HTS family are currently available. The metA gene of B. cereus (strain ATCC 10987) encodes an HTS with a molecular weight of 35,338 Da (residues 1–301) and a calculated isoelectric point of 5.69. Here, we report the crystal structure of the B. cereus HTS, which was determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG),9 which is part of the National Institute of General Medical Sciences (NIGMS) Protein Structure Initiative (PSI). The B. cereus metA gene (GenBank: NP_981826, gi|42784579) was amplified by polymerase chain reaction (PCR) from genomic DNA using PfuTurbo (Stratagene) and primers corresponding to the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pSpeedET, which encodes an expression and purification tag followed by a tobacco etch virus (TEV) protease cleavage site (MGSDKIHHHHHHENLYFQG) at the amino terminus of the full-length protein. The cloning junctions were confirmed by DNA sequencing. Protein expression was performed in a selenomethionine-containing medium using the Escherichia coli strain GeneHogs (Invitrogen). At the end of fermentation, lysozyme was added to the culture to a final concentration of 250 μg/mL, and the cells were harvested. After one freeze/thaw cycle, the cells were sonicated in lysis buffer [50 mM HEPES pH 8.0, 50 mM NaCl, 10 mM imidazole, 1 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the lysate was clarified by centrifugation at 32,500 × g for 30 min. The soluble fraction was loaded onto nickel-chelating resin (GE Healthcare) pre-equilibrated with Lysis Buffer, the resin was washed with wash buffer [50 mM HEPES pH 8.0, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP], and the protein was eluted with elution buffer [20 mM HEPES pH 8.0, 300 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP]. The eluate was buffer exchanged with HEPES crystallization buffer [20 mM HEPES pH 8.0, 200 mM NaCl, 40 mM imidazole, 1 mM TCEP] using a PD-10 column (GE Healthcare) and treated with 1 mg TEV protease per 10 mg eluted protein. The digested eluate was passed over nickel-chelating resin (GE Healthcare) pre-equilibrated with HEPES crystallization buffer, and the resin was washed with the same buffer. The flow-through and wash fractions were combined and concentrated for crystallization assays to 14.6 mg/mL by centrifugal ultrafiltration (Millipore). The protein was crystallized using the nanodroplet vapor diffusion method10 with standard JCSG crystallization protocols.9 The crystallization reagent contained 1.6 M (NH4)2SO4 and 0.1 M Tris pH 8.0. Glycerol was added as a cryoprotectant to a final concentration of 15% (v/v). Screening for diffraction was carried out using the Stanford Automated Mounting system (SAM)11 at the Stanford Synchrotron Radiation Laboratory (SSRL, Stanford, CA). The crystals were indexed in tetragonal space group P4122 (Table I). To determine its oligomeric state, B. cereus HTS was analyzed using a 1 cm × 30 cm Superdex 200 column (GE Healthcare) coupled with miniDAWN static light scattering and Optilab differential refractive index detectors (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl, and 0.02% (w/v) sodium azide. The molecular weight was calculated using ASTRA 5.1.5 software (Wyatt Technology). Multi-wavelength anomalous diffraction (MAD) data were collected at SSRL on beamline 1–5 at wavelengths corresponding to the high energy remote (λ1), inflection point (λ2), and peak (λ3) of a selenium MAD experiment using the BLU-ICE14 data collection environment. The data sets were collected at 100K using an ADSC Q315 CCD detector. The MAD data were integrated and reduced using MOSFLM15 and then scaled with the program SCALA.12 The heavy atom substructure was solved by SHELXD16, 17 and refined with autoSHARP.18 Automatic model building was performed with RESOLVE.19 Model completion and refinement were performed with the λ1 data set using COOT20 and REFMAC5,12 respectively. Data and refinement statistics are summarized in Table I. Analysis of the stereochemical quality of the model was accomplished using AutoDepInputTool,21 MolProbity,22 SFcheck 4.0,23 and WHATIF 5.0.24 Protein quaternary structure analysis was performed using the PQS and PISA servers.25, 26 Figure 1(B) was adapted from an analysis using PDBsum,27 and all others were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors for HTS from B. cereus at 2.4 Å resolution have been deposited in the PDB and are accessible under the code 2ghr Crystal structure of HTS from B. cereus. (A) Stereo ribbon diagram of HTS monomer color-coded from N-terminus (blue) to C-terminus (red). Helices (H1–H11) and β-strands (β1–β11) are indicated. The disordered region (75–86) is depicted by a grey dashed line. β-sheets are indicated by a red A′, A, and B. (B) Diagram showing the secondary structural elements of HTS superimposed on its primary sequence. The α-helices, 310-helices, β-strands, β-bulges, and γ-turns are indicated. The β-hairpin is depicted as a red loop. Disordered regions are depicted by dashed regions with the corresponding sequence shown below. The crystal structure of HTS from B. cereus (Fig. 1) was determined to 2.4 Å resolution using the MAD method. Data collection, model, and refinement statistics are summarized in Table I. The final model includes one protein molecule (residues 17–74 and 87–297), one SO ion, and 78 water molecules in the asymmetric unit. No electron density was observed for residues 0 (glycine residue of protease cleavage site), 1–16, 75–86, or 298–301. The side chains of Glu18, Pro45, Thr46, Arg93, Glu215, Lys265, and Lys272 were in regions of poor electron density and were not modeled. The Matthews' coefficient (Vm)28 for HTS is 2.44 Å3/Da, and the estimated solvent content is 54%. The Ramachandran plot produced by MolProbity22 shows that 96.9 and 99.6% of the residues are in favored and allowed regions, respectively. Pro45 was the only outlier and is located in a region of poor electron density. Analysis of the crystallographic packing of HTS indicates that a crystallographic dimer is the biologically relevant form [Fig. 2(A)]. This finding is consistent with results from analytical size exclusion chromatography in combination with static light scattering, which gave a molecular weight of approximately 70,000 Da. HTS is a single-domain protein with a Rossmann fold topology. The core of the protein is a parallel β-sheet sandwiched by α-helices. HTS is composed of 11 β-strands, 7 α-helices (H1-H2, H4-H5, H7, H9, H11), and four 310-helices (H3, H6, H8, H10) [Fig. 1(A,B)]. The total β-sheet, α-helical, and 310-helical content is 28.6, 26.0, and 4.5%, respectively. The N-terminus of the monomer is in an extended conformation and forms the primary dimerization interface via β-strand exchange [Fig. 2(A)]. Dimerization buries approximately 2274 Å2 of solvent accessible surface area per monomer. This mode of dimerization via N-terminal β-strand exchange is not observed in structurally similar or mechanistically related HTAs. HTS dimer and structural comparisons with GMP synthase from Thermoplasma acidophilum (PDB accession code 2a9v). (A) Stereo view of the HTS dimer showing the N-terminal β-strand exchange between the monomers, colored green and orange. The N- and C-termini are labeled. (B) Stereo overlay of the HTS monomer (grey) and GMP synthase from T. acidophilum (light green) with the proposed His-Glu-Cys catalytic triad in ball-and-stick representation. (C) Close-up view of the active sites of HTS and GMP synthase colored as per (B). Active site catalytic triad residues are depicted as ball-and-sticks and colored by atom type (N blue, S orange, O red, and C light green or grey). HTS residues are labeled, and corresponding GMP synthase residues are in parentheses. HTA from H. influenzae (PDB accession code 2b61) also consists of a Rossmann-type fold, but with an additional predominantly α-helical sub-domain (residues 168–281) that caps the active site. The helical sub-domain also forms the primary dimerization interface. The Rossmann fold regions of HTA and HTS can be aligned structurally; however, the helical cap sub-domain in HTA is not present in HTS and the dimer assembly is, therefore, very different. Further structural characterization of HTS family members will be necessary to determine whether dimerization via N-terminal β-strand exchange is a general property of bacterial HTSs or specific to B. cereus. Based on previous biochemical characterization of E. coli HTS and structural alignments with H. influenzae HTA, the active site of B. cereus HTS can be tentatively identified. Enzymatic studies identified Cys142 in E. coli HTS (Cys142 in B. cereus HTS) as the critical residue for catalysis.7 In B. cereus, this residue is part of a previously unidentified catalytic triad consisting of His235, Glu237, and Cys142. This triad is not only conserved in the E. coli HTS, but also in other annotated bacterial HTS enzymes. The putative catalytic residues, His235 and Glu237, are located between β-strand 11 and helix H9. The catalytic cysteine, Cys142, is found between β-strand 5 and helix H5. While the residues involved in the putative catalytic triad are not adjacent to each other in the primary sequence, they are in spatial proximity in the tertiary structure. Another study using proteomics and mass spectrometry identified Lys47 (B. cereus HTS) as the key catalytic residue and not Cys142.29 Lys47 is approximately 15 Å from Cys142 and is located at the N-terminus of helix H1 on the surface of the protein. While we cannot exclude Lys47 as playing a catalytic role in the B. cereus HTS, the triad of His235, Glu237, and Cys142 appears more likely to form the catalytic core of the active site. Furthermore, mechanistically related enzymes, such as the HTAs, utilize a His-Asp-Ser triad that occurs in a similar location to the HTS catalytic triad, as determined by structural comparison with H. influenza HTA. To date, the enzymatic and kinetic characterization of HTAs6 and HTSs7 point to a common ping-pong mechanism for acetate or succinate transfer. This mechanism necessitates a nucleophilic attack by the enzyme to form an acetyl- or succinyl-enzyme intermediate followed by transfer of the acyl group to homoserine to yield the O-acylhomoserine product. Usually, the initial nucleophilic attack occurs via an active site cysteine (Cys142 in B. cereus HTS) or a serine residue. The biochemical analysis identifying Cys142 as the key catalytic residue, coupled with kinetic analysis, demonstrated a ping-pong mechanism for trans-succinylation for HTS7 and is consistent with the putative active site identified here. Sequence and structural comparisons of remote homologs of HTS reveal mechanistically diverse enzymes that include ligases, isomerases, hydrolases, and transferases with many variants on the canonical Rossmann α/β fold. The closest homolog in the PDB to HTS, based on sequence identity, is the tetrameric GMP synthase from Thermoplasma acidophilum (PDB accession code 2a9v). An alignment of one monomer from each protein using the program DALI30 shows an essentially identical fold with an RMSD of 2.3 Å over 189 Cα atoms. Although the two proteins exhibit a low sequence identity of 18%, the structural alignment reveals a strong conservation of the overall fold and active site residues [Fig. 2(B,C)]. The utility of the classical Rossmann fold is further diversified here by varying the protein oligomeric state. For example, based on DALI Z-score, the five structures most related to B. cereus HTS are either monomeric (PDB accession code 1o1y) or tetrameric (PDB accession codes 2a9v, 1qdl, 1l9x, and 1gpm) enzymes. Clearly, the Rossmann fold has been modified by evolution to select for diverse substrates and to perform different reactions by altering sequence/structure and oligomerization state while conserving the same overall fold. The structure of HTS from B. cereus is the first structure from the HTS protein family (Pfam accession code PF04204). This family contains over 100 bacterial homologs by Pfam definition and over 500 by PSI-BLAST. Models for HTS homologs can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key = 2ghr. A primary goal of the PSI is to rapidly expand the coverage of protein fold space and to target large protein superfamilies for structural characterization. During target selection, we focus on proteins that are essential for fundamental biological processes and have a broad phylogenetic distribution (“Central Machinery of Life”). The biosynthesis of methionine is a fundamental process carried out by plants, archaea, fungi, and bacteria. The structure of HTS from B. cereus is the first example for this enzyme family. The crystal structure presented here provides important information as to the function and mechanism of HTS enzymes. Moreover, mechanistic and mutagenesis studies should confirm and elucidate residues important for catalysis and for substrate specificity. Furthermore, detailed structural and biochemical knowledge of HTS and the succinylation of homoserine may aid in the design of effective inhibitors against pathogenic bacteria which utilize HTS during methionine biosynthesis. Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory (SSRL). The SSRL is a national user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences).
Proteins: Structure, Function, and BioinformaticsVolume 69, Issue 2 p. 415-421 Structure Note Crystal structure of NMA1982 from Neisseria meningitidis at 1.5 Å resolution provides a structural scaffold for nonclassical, eukaryotic-like phosphatases S. Sri Krishna, S. Sri Krishna Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, California Sri Krishna and Lutz Tautz contributed equally to this work.Search for more papers by this authorLutz Tautz, Lutz Tautz Burnham Institute for Medical Research, La Jolla, California Sri Krishna and Lutz Tautz contributed equally to this work.Search for more papers by this authorQingping Xu, Qingping Xu Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDaniel McMullan, Daniel McMullan Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMitchell D. Miller, Mitchell D. Miller Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorPolat Abdubek, Polat Abdubek Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEileen Ambing, Eileen Ambing Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorTamara Astakhova, Tamara Astakhova Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorHerbert L. Axelrod, Herbert L. Axelrod Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDennis Carlton, Dennis Carlton Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorHsiu-Ju Chiu, Hsiu-Ju Chiu Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorThomas Clayton, Thomas Clayton Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorMichael DiDonato, Michael DiDonato Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLian Duan, Lian Duan Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorMarc-André Elsliger, Marc-André Elsliger Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorSlawomir K. Grzechnik, Slawomir K. Grzechnik Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorJoanna Hale, Joanna Hale Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Hampton, Eric Hampton Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorGye Won Han, Gye Won Han Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJustin Haugen, Justin Haugen Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLukasz Jaroszewski, Lukasz Jaroszewski Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorKevin K. Jin, Kevin K. Jin Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorHeath E. Klock, Heath E. Klock Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMark W. Knuth, Mark W. Knuth Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Koesema, Eric Koesema Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorAndrew T. Morse, Andrew T. Morse Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorTomas Mustelin, Tomas Mustelin Burnham Institute for Medical Research, La Jolla, CaliforniaSearch for more papers by this authorEdward Nigoghossian, Edward Nigoghossian Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSilvya Oommachen, Silvya Oommachen Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorRon Reyes, Ron Reyes Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorChristopher L. Rife, Christopher L. Rife Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorHenry van den Bedem, Henry van den Bedem Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDana Weekes, Dana Weekes Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, CaliforniaSearch for more papers by this authorAprilfawn White, Aprilfawn White Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorKeith O. Hodgson, Keith O. Hodgson Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorJohn Wooley, John Wooley Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorAshley M. Deacon, Ashley M. Deacon Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorAdam Godzik, Adam Godzik Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorScott A. Lesley, Scott A. Lesley Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, California The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorIan A. Wilson, Corresponding Author Ian A. Wilson [email protected] Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaJCSG, The Scripps Research Institute, BCC206, 10550 North Torrey Pines Road, La Jolla, CA 92037===Search for more papers by this author S. Sri Krishna, S. Sri Krishna Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, California Sri Krishna and Lutz Tautz contributed equally to this work.Search for more papers by this authorLutz Tautz, Lutz Tautz Burnham Institute for Medical Research, La Jolla, California Sri Krishna and Lutz Tautz contributed equally to this work.Search for more papers by this authorQingping Xu, Qingping Xu Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDaniel McMullan, Daniel McMullan Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMitchell D. Miller, Mitchell D. Miller Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorPolat Abdubek, Polat Abdubek Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEileen Ambing, Eileen Ambing Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorTamara Astakhova, Tamara Astakhova Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorHerbert L. Axelrod, Herbert L. Axelrod Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDennis Carlton, Dennis Carlton Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorHsiu-Ju Chiu, Hsiu-Ju Chiu Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorThomas Clayton, Thomas Clayton Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorMichael DiDonato, Michael DiDonato Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLian Duan, Lian Duan Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorMarc-André Elsliger, Marc-André Elsliger Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorSlawomir K. Grzechnik, Slawomir K. Grzechnik Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorJoanna Hale, Joanna Hale Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Hampton, Eric Hampton Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorGye Won Han, Gye Won Han Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJustin Haugen, Justin Haugen Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLukasz Jaroszewski, Lukasz Jaroszewski Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorKevin K. Jin, Kevin K. Jin Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorHeath E. Klock, Heath E. Klock Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMark W. Knuth, Mark W. Knuth Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Koesema, Eric Koesema Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorAndrew T. Morse, Andrew T. Morse Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorTomas Mustelin, Tomas Mustelin Burnham Institute for Medical Research, La Jolla, CaliforniaSearch for more papers by this authorEdward Nigoghossian, Edward Nigoghossian Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSilvya Oommachen, Silvya Oommachen Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorRon Reyes, Ron Reyes Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorChristopher L. Rife, Christopher L. Rife Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorHenry van den Bedem, Henry van den Bedem Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorDana Weekes, Dana Weekes Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, CaliforniaSearch for more papers by this authorAprilfawn White, Aprilfawn White Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorKeith O. Hodgson, Keith O. Hodgson Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorJohn Wooley, John Wooley Joint Center for Structural Genomics (JCSG) Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorAshley M. Deacon, Ashley M. Deacon Joint Center for Structural Genomics (JCSG) Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, CaliforniaSearch for more papers by this authorAdam Godzik, Adam Godzik Joint Center for Structural Genomics (JCSG) Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorScott A. Lesley, Scott A. Lesley Joint Center for Structural Genomics (JCSG) Genomics Institute of the Novartis Research Foundation, San Diego, California The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorIan A. Wilson, Corresponding Author Ian A. Wilson [email protected] Joint Center for Structural Genomics (JCSG) The Scripps Research Institute, La Jolla, CaliforniaJCSG, The Scripps Research Institute, BCC206, 10550 North Torrey Pines Road, La Jolla, CA 92037===Search for more papers by this author First published: 16 July 2007 https://doi.org/10.1002/prot.21314Citations: 6 Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinkedInRedditWechat REFERENCES 1 Alonso A,Sasin J,Bottini N,Friedberg I,Friedberg I,Osterman A,Godzik A,Hunter T,Dixon J,Mustelin T. Protein tyrosine phosphatases in the human genome. Cell 2004; 117: 699–711. 2 Dixon JE. Structure and catalytic properties of protein tyrosine phosphatases. Ann N Y Acad Sci 1995; 766: 18–22. 3 Bateman A,Coin L,Durbin R,Finn RD,Hollich V,Griffiths-Jones S,Khanna A,Marshall M,Moxon S,Sonnhammer EL,Studholme DJ,Yeats C,Eddy SR. The Pfam protein families database. Nucleic Acids Res 2004; 32: D138–D141. 4 Yooseph S,Sutton G,Rusch DB,Halpern AL,Williamson SJ,Remington K,Eisen JA,Heidelberg KB,Manning G,Li W,Jaroszewski L,Cieplak P,Miller CS,Li H,Mashiyama ST,Joachimiak MP,van Belle C,Chandonia JM,Soergel DA,Zhai Y,Natarajan K,Lee S,Raphael BJ,Bafna V,Freidman R,Brenner SE,Godzik A,Eisenberg D,Dixon JE,Taylor SS,Strausberg RL,Frazier M,Venter JC. The Sorcerer II global ocean sampling expedition: expanding the universe of protein families. PLoS Biol 2007; 5: e16. 5 Jaroszewski L,Rychlewski L,Li Z,Li W,Godzik A. FFAS03: a server for profile–profile sequence alignments. Nucleic Acids Res 2005; 33: W284–W288. 6 Ginalski K,Pas J,Wyrwicz LS,von Grotthuss M,Bujnicki JM,Rychlewski L. ORFeus: detection of distant homology using sequence profiles and predicted secondary structure. Nucleic Acids Res 2003; 31: 3804–3807. 7 Lesley SA,Kuhn P,Godzik A,Deacon AM,Mathews I,Kreusch A,Spraggon G,Klock HE,McMullan D,Shin T,Vincent J,Robb A,Brinen LS,Miller MD,McPhillips TM,Miller MA,Scheibe D,Canaves JM,Guda C,Jaroszewski L,Selby TL,Elsliger MA,Wooley J,Taylor SS,Hodgson KO,Wilson IA,Schultz PG,Stevens RC. Structural genomics of the Thermotoga maritima proteome implemented in a high-throughput structure determination pipeline. Proc Natl Acad Sci USA 2002; 99: 11664–11669. 8 Santarsiero BD,Yegian DT,Lee CC,Spraggon G,Gu J,Scheibe D,Uber DC,Cornell EW,Nordmeyer RA,Kolbe WF,Jin J,Jones AL,Jaklevic JM,Schultz PG,Stevens RC. An approach to rapid protein crystallization using nanodroplets. J Appl Crystallogr 2002; 35: 278–281. 9 Cohen AE,Ellis PJ,Miller MD,Deacon AM,Phizackerley RP. An automated system to mount cryo-cooled protein crystals on a synchrotron beamline, using compact sample cassettes and a small-scale robot. J Appl Crystallogr 2002; 35: 720–726. 10 Collaborative Computional Project Number 4. The CCP4 suite: programs for protein crystallography. Acta Crystallogr D Biol Crystallogr 1994; 50: 760–763. 11 Tickle IJ,Laskowski RA,Moss DS. Error estimates of protein structure coordinates and deviations from standard geometry by full-matrix refinement of gammaB- and betaB2-crystallin. Acta Crystallogr D Biol Crystallogr 1998; 54: 243–252. 12 McPhillips TM,McPhillips SE,Chiu HJ,Cohen AE,Deacon AM,Ellis PJ,Garman E,Gonzalez A,Sauter NK,Phizackerley RP,Soltis SM,Kuhn P. Blu-Ice and the distributed control system: software for data acquisition and instrument control at macromolecular crystallography beamlines. J Synchrotron Radiat 2002; 9: 401–406. 13 Leslie AGW. Recent changes to the MOSFLM package for processing film and image plate data. Joint CCP4+ESF-EAMCB Newsletter on Protein Crystallography 1992; 26. 14 Terwilliger TC,Berendzen J. Automated MAD and MIR structure solution. Acta Crystallogr D Biol Crystallogr 1999; 55 (Part 4): 849–861. 15 Perrakis A,Morris R,Lamzin VS. Automated protein model building combined with iterative structure refinement. Nat Struct Biol 1999; 6: 458–463. 16 Emsley P,Cowtan K. Coot: model-building tools for molecular graphics. Acta Crystallogr D Biol Crystallogr 2004; 60: 2126–2132. 17 Kleywegt GJ. Validation of protein crystal structures. Acta Crystallogr D Biol Crystallogr 2000; 56: 249–265. 18 Yang H,Guranovic V,Dutta S,Feng Z,Berman HM,Westbrook JD. Automated and accurate deposition of structures solved by X-ray diffraction to the Protein Data Bank. Acta Crystallogr D Biol Crystallogr 2004; 60: 1833–1839. 19 Davis IW,Murray LW,Richardson JS,Richardson DC MOLPROBITY: structure validation and all-atom contact analysis for nucleic acids and their complexes. Nucleic Acids Res 2004; 32: W615–W619. 20 Vaguine AA,Richelle J,Wodak SJ. SFCHECK: a unified set of procedures for evaluating the quality of macromolecular structure-factor data and their agreement with the atomic model. Acta Crystallogr D Biol Crystallogr 1999; 55: 191–205. 21 Vriend G. WHAT IF: a molecular modeling and drug design program. J Mol Graph 1990; 8: 52–56, 29. 22 Henrick K,Thornton JM. PQS: a protein quaternary structure file server. Trends Biochem Sci 1998; 23: 358–361. 23 Matthews BW. Solvent content of protein crystals. J Mol Biol 1968; 33: 491–497. 24 Murzin AG,Brenner SE,Hubbard T,Chothia C. SCOP: a structural classification of proteins database for the investigation of sequences and structures. J Mol Biol 1995; 247: 536–540. 25 Holm L,Sander C. Dali: a network tool for protein structure comparison. Trends Biochem Sci 1995; 20: 478–480. Citing Literature Volume69, Issue21 November 2007Pages 415-421 ReferencesRelatedInformation
Transcriptional regulators play a crucial role in the adaptation of microorganisms to diverse environmental challenges.1-3 Most microbial transcriptional regulators contain an effector binding regulatory domain and a DNA-binding domain that interacts with a specific operator DNA to either prevent (transcriptional repressors) or stimulate (transcriptional activators) transcription of a nearby gene(s).4 Prokaryotic transcriptional regulators have been classified into a number of families based on amino acid sequence similarity and domain architecture.4-8 The tetracycline repressor (TetR) family of proteins exhibits a high degree of sequence similarity at the N-terminal DNA-binding domain (∼50 amino acids), which adopts a helix-turn-helix (HTH) motif. In contrast, the regulatory domain is more variable, possibly reflecting the need to specifically accommodate different effectors.4, 9 TM1030 from Thermotoga maritima, a hyperthermophilic bacterium that typically thrives in high temperature ecosystems, is a 200 amino acid protein with a molecular weight of 24 kDa and an isoelectric point of 6.25. The N-terminal DNA-binding domain of TM1030 shows sequence similarity to members of the TetR family, but no significant similarity is found for the regulatory C-terminal region (∼150 amino acids). Here, we present the crystal structure of a ligand-bound form of TM1030, which was determined to 2.3 Å resolution, using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG)10 as part of the National Institute of General Medical Sciences (NIGMS)-funded Protein Structure Initiative (PSI). The TM1030 gene (GenBank: AAD36107.1, GI: 4981571, Swiss-Prot: Q9×0C0) from Thermotoga maritima was amplified by polymerase chain reaction (PCR) from genomic DNA using PfuTurbo (Stratagene) and primers corresponding to the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pMH1, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The TM1030 gene uses an alternate start codon (GUG) that results in a valine at position 1 when expressed as a fusion with the expression and purification tag. The cloning junctions were confirmed by DNA sequencing. Protein expression was performed in a selenomethionine-containing medium using the Escherichia coli methionine auxotrophic strain DL41. At the end of fermentation, lysozyme was added to the culture to a final concentration of 250 μg/mL, and the cells were harvested. After one freeze/thaw cycle, the cells were sonicated in lysis buffer [50 mM Tris pH 8.0, 50 mM NaCl, 10 mM imidazole, 0.25 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the lysate was clarified by centrifugation at 32,500g for 30 min. The soluble fraction was passed over nickel-chelating resin (GE Healthcare) pre-equilibrated with Lysis Buffer, the resin was washed with Wash Buffer [50 mM potassium phosphate, pH 7.8, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP], and the protein was eluted with elution buffer [20 mM Tris pH 8.0, 300 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP]. The eluate was diluted ten-fold with Buffer Q [20 mM Tris pH 7.9, 50 mM NaCl, 5% (v/v) glycerol, 0.25 mM TCEP] and applied to a RESOURCE Q column (GE Healthcare) pre-equilibrated with the same buffer. The flow-through fraction, which contained TM1030, was further purified on a Superdex 200 column (GE Healthcare), with isocratic elution in Crystallization Buffer [20 mM Tris pH 7.9, 150 mM NaCl, 0.25 mM TCEP]. The protein was concentrated for crystallization assays to 15 mg/mL by centrifugal ultrafiltration (Millipore) and crystallized using the nanodroplet vapor diffusion method11 with standard JCSG crystallization protocols.10 The crystallization reagent contained 30% (w/v) polyethylene glycol (PEG) 8000, 0.2M Mg(NO3)2, and 0.1M citrate pH 4.5. Ethylene glycol was added as a cryoprotectant to a final concentration of 5% (v/v). Initial screening for diffraction was carried out using the Stanford Automated Mounting system (SAM)12 at the Stanford Synchrotron Radiation Laboratory (SSRL, Stanford, CA). The crystals were indexed in monoclinic space group P21 (Table I). The molecular weight and oligomeric state of TM1030 were determined using a 1 cm × 30 cm Superdex 200 column (GE Healthcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl, and 0.02% (w/v) sodium azide. Multi-wavelength anomalous diffraction (MAD) data sets were collected at 100 K using a charge-coupled device detector (ADSC Q315) on SSRL beamline 11-1 using the BLU-ICE13 data collection environment (Table I). Data were collected at wavelengths corresponding to the high energy remote (λ1) and inflection (λ2) of a selenium MAD experiment. Data were indexed and reduced with Mosflm16 and scaled using SCALA from the CCP4 suite.14 Diffraction data statistics are summarized in Table I. The selenium substructure was solved using SOLVE.17 Refinement of the Se sites resulted in a mean figure of merit of 0.39 to a resolution of 2.5 Å. Phase extension to 2.3 Å was performed using RESOLVE,17 with a solvent content of 0.5 and a starting two-fold noncrystallographic symmetry (NCS) matrix derived from the substructure solution. Automatic model building was performed with RESOLVE, resulting in a dimer model containing 288 residues (72%), with 89 (22%) of the side chains fitted. This initial model was rebuilt using iterative ARP/wARP runs,18 which built 354 residues (88%), with 345 residues docked into the sequence (86%). Model completion and refinement were performed with the remote (λ1) data set using COOT19 and REFMAC5.20 Refinement statistics are summarized in Table I. Analysis of the stereochemical quality of the structure was accomplished using AutoDepInputTool,21 MolProbity,15 SFcheck 4.0,14 and WHATIF 5.0.22 Protein quaternary structure analysis was performed using the PQS server.23 Figure 1 was adapted from an analysis using PDBsum,24 and all other figures were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors for TM1030 at 2.3 Å resolution have been deposited in the PDB and are accessible under the code 1zkg. Stereo ribbon diagram of the crystal structure of TM1030 monomer. A: The DNA-binding domain and the regulatory domain are colored in green and violet, respectively. The helices, as well as the N- and C-termini, are labeled. B: Schematic diagram showing the secondary structural elements in TM1030 superimposed on its primary sequence. The α-helices and 310-helix (H6A) are indicated. The crystal structure of TM1030 (Fig. 1) was determined to 2.3 Å resolution using the MAD method (Table I). The asymmetric unit includes two TM1030 subunits, two unknown ligands (UNLs) and 56 water molecules. Electron density was not observed for residues from the expression and purification tag for both subunits and residue Val 1 of subunit B. The Matthews' coefficient (Vm)25 for TM1030 is 2.6 Å3/Da, and the estimated solvent content is 52.2%. The Ramachandran plot,26 as produced by Molprobity,27 shows that 97.2 and 99.8% of the main chain torsion angles are in the favored and allowed regions, respectively. The only outlier is the surface exposed residue R74 (subunit A), which is poorly defined in the electron density map. TM1030 is an all-helical protein, comprised of 10 α-helices (H1-H7, H7A-H9) and a 310-helix (H6A), and adopts a two-domain architecture similar to TetR (Fig. 1). The N-terminal DNA-binding domain is composed of the first three α-helices. The H2 and H3 α-helices of this domain form a canonical HTH motif. The regulatory domain is made up of an antiparallel helical bundle (H4-H5 and H7-H9) and helix H6 that is packed nearly orthogonal to the long axis of this helical bundle. A DALI28 search revealed structural similarity to several microbial transcriptional regulators. The top 13 hits in the search belong to the TetR family and include proteins from Pseudomonas aeruginosa (PDB accession code: 2gen, 2fbq, 2fd5), Salmonella typhimurium (1t33), Staphylococcus aureus (1jty), Mycobacterium tuberculosis (1t56), Rhodococcus sp. (2gfn, 2g7g), Bacillus cereus (1sgm, 2fx0, 1zk8, 2fq4), and Streptomyces coelicolor (1ui5). The Z-scores for the structural alignments of TM1030 with these top hits were in the range of 14.1–7.4 and the corresponding RMSDs are in the range of 3.0–6.8 Å where at least 75% of the Cα atoms (of the total 200 amino acids) were included. Notably, these structures share less than 21% sequence identity to TM1030, and nine of these structures were determined at PSI-funded Structural Genomics (SG) centers. The N-terminal DNA-binding domain of TM1030 displays remarkable structural similarity to the homologous domains in all TetR-like proteins, while differences are much greater in the C-terminal regulatory domain. In particular, the relative orientation of the α-helices in the regulatory domain that mediates homodimerization4 differs significantly among members of the TetR family. A pair of α-helices (H8 and H9) in the TM1030 regulatory domain is involved in mediating most of the inter-subunit interaction. The inter-subunit interactions are mostly hydrophobic (V145, I149, F153, W156, F157, F161, V164, V189, M190, I193, and L194) and are further stabilized by four salt-bridges [D144(A) – R192(B), K152(A) – E186(B), D144(B) – R192(A), and K152(B) – E186(A)], and three hydrogen bonds [E163(A) – E163(B), K196(A) – T199(B), and K196(B) – T199(A)]. An analysis using size exclusion chromatography coupled with static light scattering supports the assignment of TM1030 as a dimer in solution. Furthermore, the biologically relevant homodimerization in the TetR family is mediated by similar helix-to-helix contacts,4, 29-31 suggesting that the dimer observed for TM1030 is functionally relevant. The Midwest Center for Structural Genomics (MCSG) has also determined the crystal structure of a TM1030 construct to 2.0 Å resolution (PDB code 1z77). A structural superposition revealed a significant global conformational difference between the two structures (Fig. 2), despite very similar crystallization conditions (Table II). Two modes of structural alignments were explored using as anchors either the conserved N-terminal DNA-binding domain or the C-terminal homodimerization-mediating α-helices (H8 and H9). The corresponding RMSD values for the N-terminal or C-terminal based structural alignments for all 200 Cα atoms of TM1030 (JCSG, subunits A and B) with TM1030 (MCSG) are in the range of 3.3–5.0 and 5.1–6.4 Å, respectively. In spite of the large structural difference, the global fold and individual structural elements are mostly retained in both structures, with noticeable differences confined to the lengths of α-helices H1, H3, H4, and H6A. Moreover, the inter-subunit interactions mediated by α-helices H8 and H9 in TM1030 (JCSG) are largely retained in TM1030 (MCSG), even though the biological dimer in TM1030 (JCSG) is formed from subunits related by twofold NCS, while the TM1030 (MCSG) subunits are related by exact crystallographic symmetry. The calculated total buried surface area between the monomers in TM1030 (JCSG) and TM1030 (MCSG) is also quite comparable [1497 Å2 (JCSG) vs. 1401 Å2 (MCSG)]. In addition, the residues involved in inter-subunit interactions are largely unperturbed in these structures suggesting that the conformational changes between the two TM1030 structures are unlikely to be caused by the inter-subunit/crystal-packing interactions. When the structural superpositions are restricted to individual domains of TM1030 [Fig. 2(B,C)], regions of large conformational differences are mostly confined to the regulatory domain (RMSD of 1.75 Å for 153 Cα atoms) rather than in the DNA-binding domain (RMSD of 0.5 Å for 47 Cα atoms). Structure comparison of TM1030 (JCSG) and TM1030 (MCSG). Structure alignment of A: full-length TM1030 monomer. The alignment was optimized over residues from the C-terminal homodimerization mediating α-helices (H8-H9). B: The DNA-binding domain. C: The regulatory domain. JCSG and MCSG structures are shown in red and green, respectively. The α-helices, as well as the N- and C-termini, are indicated. While searching for a plausible basis for the conformational differences in the regulatory domain, we identified a ∼12 Å deep cavity in each of the TM1030 subunits that is located within the helical bundle of the regulatory domain [Fig. 3(A)]. The binding pocket, whose total volume is approximately 2000 Å3, is predominantly lined by hydrophobic residues. The TM1030 cavity has a 10–17 Å wide opening, that is formed by residues from α-helices H6A, H7, H7A, and H8 [Figs. 1(A) and 3(A)], which is likely to serve as an entrance to this putative binding pocket. The location of the cavity is in proximity to the ligand-binding pocket of TetR, but the volume of the cavity and the nature of the residues lining the cavity in the two proteins are quite different. Interestingly, each TM1030 (JCSG) subunit contains a semi-circular region of positive electron density in both the omit Fo − Fc and 2Fo − Fc maps in the ligand-binding pocket, indicative of a bound ligand (Fig. 3(B,C)]. However, such density was not observed in the TM1030 (MCSG) structure. The residues within 4 Å of this electron density in TM1030 (JCSG) are T59, L62, and F66 from α-helix H4; W85, I86, and K89 from α-helix H5; S124, Q125, and F128 from α-helix H7 and helix H7A; and F158, F161, E162, and Y165 from α-helix H8. The bound ligand is surrounded by hydrophobic, polar, and electrostatic groups including a cluster of aromatic rings. No biologically relevant ligand that could fit such electron density was added during the protein purification or crystallization of TM1030 (JCSG). Consideration of the shape and length of the density suggests it might represent a lipid molecule acquired in vivo within the E. coli host cells used for heterologous expression. However, the electron density is not consistent with any lipid with a head group or a carboxylate, such as palmitic acid. On the other hand, the density can be modeled by a relatively short fragment of PEG, in particular heptaethylene glycol [Fig. 3(D)], that probably originates from the crystallization solution (see e.g. Koepke et al., 2003; Zhu et al., 2006).32, 33 Nevertheless, as the precise identity of the ligand molecule has yet to be determined, we modeled and deposited it in the PDB as an UNL. The absence of a ligand in TM1030 (MCSG) together with the difference in conformation strongly suggests that this structure represents the apo-form of TM1030. A DNA-binding model for TM1030 based on TetR-tetO complex shows that the two DNA-binding domains from apo-TM1030 fit into the major groove of dsDNA, whereas the ligand-bound form of TM1030 is not in a favorable conformation to bind dsDNA [Fig. 3(E)], suggesting that TM1030 is most likely a transcription repressor. Ligand-binding pocket and DNA-binding model. A: Surface representation of the ligand-binding pocket. B: The omit Fo − Fc electron density map corresponding to a region in the ligand-binding pocket. C: The residues within 4 Å of the bound ligand in TM1030 (JCSG). D: The omit Fo − Fc electron density map. A heptaethylene glycol molecule has been modeled to fit the electron density, but coordinates for the ligand are deposited in the PDB as a UNL. The electron density is contoured at 2.5 σ. E: Computational model of TM1030-operator dsDNA complex. The model is based on the crystal structure of TetR-operator dsDNA complex (PDB accession code: 1qpi). Apo-TM1030 and ligand-bound TM1030 are colored in blue and grey, respectively. N- and C-termini of TM1030 are indicated. One of the fundamental means by which bacteria adapt to varying environmental conditions is based on the regulation of gene expression at the transcriptional level.1-3, 34 Structural information regarding these transcriptional regulators is crucial to our understanding of how transcriptional regulatory networks control the microbial responses to different environmental challenges, including multidrug resistance, solvent tolerance, stress response, and pathogenesis. The efforts of the PSI-funded SG centers have resulted so far in the determination of nine TetR-like protein structures. Although these TetR-like structures share a high degree of overall fold similarity, their structures, particularly those of the regulatory domains, are very divergent and cannot be readily predicted. The ability to create novel binding sites for various effectors within the regulatory domain of proteins is perhaps driven by mutations that have little effect on the overall three-dimensional structure, but exert a large effect on the plasticity of effector binding sites. A detailed structural analysis, including identification of a biologically relevant ligand for TM1030, will offer insights into the structural basis of its transcription repressor function. Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory (SSRL), the Advanced Light Source (ALS), and the Advanced Photon Source (APS). The SSRL is a national user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U.S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory. Use of the Argonne National Laboratory Structural Biology Center beamlines at the APS was supported by the U. S. Department of Energy, Office of Biological and Environmental Research, under Contract No. W-31-109-ENG-38.
Bt DyP from Bacteroides thetaiotaomicron (strain VPI‐5482) and TyrA from Shewanella oneidensis are dye‐decolorizing peroxidases (DyPs), members of a new family of heme‐dependent peroxidases recently identified in fungi and bacteria. Here, we report the crystal structures of BtDyP and TyrA at 1.6 and 2.7 Å, respectively. BtDyP assembles into a hexamer, while TyrA assembles into a dimer; the dimerization interface is conserved between the two proteins. Each monomer exhibits a two‐domain, α+β ferredoxin‐like fold. A site for heme binding was identified computationally, and modeling of a heme into the proposed active site allowed for identification of residues likely to be functionally important. Structural and sequence comparisons with other DyPs demonstrate a conservation of putative heme‐binding residues, including an absolutely conserved histidine. Isothermal titration calorimetry experiments confirm heme binding, but with a stoichiometry of 0.3:1 (heme:protein). Proteins 2007. © 2007 Wiley‐Liss, Inc.
Glutathione S-transferases (GSTs) comprise a diverse superfamily of enzymes found in organisms from all kingdoms of life. GSTs are involved in diverse processes, notably small-molecule biosynthesis or detoxification, and are frequently also used in protein engineering studies or as biotechnology tools. Here, we report the high-resolution X-ray structure of Atu5508 from the pathogenic soil bacterium Agrobacterium tumefaciens (atGST1). Through use of comparative sequence and structural analysis of the GST superfamily, we identified local sequence and structural signatures, which allowed us to distinguish between different GST classes. This approach enables GST classification based on structure, without requiring additional biochemical or immunological data. Consequently, analysis of the atGST1 crystal structure suggests a new GST class, distinct from previously characterized GSTs, which would make it an attractive target for further biochemical studies.
Qingping Xu, S. Sri Krishna, Daniel McMullan, Robert Schwarzenbacher, Mitchell D. Miller, Polat Abdubek, Sanjay Agarwalla, Eileen Ambing, Tamara Astakhova, Herbert L. Axelrod, Jaume M. Canaves, Dennis Carlton, Hsiu-Ju Chiu, Thomas Clayton, Michael DiDonato, Lian Duan, Marc-André Elsliger, Julie Feuerhelm, Slawomir K. Grzechnik, Joanna Hale, Eric Hampton, Gye Won Han, Justin Haugen, Lukasz Jaroszewski, Kevin K. Jin, Heath E. Klock, Mark W. Knuth, Eric Koesema, Andreas Kreusch, Peter Kuhn, Andrew T. Morse, Edward Nigoghossian, Linda Okach, Silvya Oommachen, Jessica Paulsen, Kevin Quijano, Ron Reyes, Christopher L. Rife, Glen Spraggon, Raymond C. Stevens, Henry van den Bedem, Aprilfawn White, Guenter Wolf, Keith O. Hodgson, John Wooley, Ashley M. Deacon, Adam Godzik, Scott A. Lesley, and Ian A. Wilson* Joint Center for Structural Genomics Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, California Genomics Institute of the Novartis Research Foundation, San Diego, California Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The Scripps Research Institute, La Jolla, California
Qingping Xu, Robert Schwarzenbacher, S. Sri Krishna, Daniel McMullan, Sanjay Agarwalla, Kevin Quijano, Polat Abdubek, Eileen Ambing, Herbert Axelrod, Tanya Biorac, Jaume M. Canaves, Hsiu-Ju Chiu, Marc-André Elsliger, Carina Grittini, Slawomir K. Grzechnik, Michael DiDonato, Joanna Hale, Eric Hampton, Gye Won Han, Justin Haugen, MichaelHornsby, Lukasz Jaroszewski, Heath E. Klock, Mark W. Knuth, Eric Koesema, Andreas Kreusch, Peter Kuhn, Mitchell D. Miller, Kin Moy, Edward Nigoghossian, Jessica Paulsen, Ron Reyes, Chris Rife, Glen Spraggon, Raymond C. Stevens, Henry van den Bedem, Jeff Velasquez, Aprilfawn White, Guenter Wolf, Keith O. Hodgson, John Wooley, Ashley M. Deacon, Adam Godzik, Scott A. Lesley, and Ian A. Wilson* The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California San Diego, La Jolla, California The Genomics Institute of the Novartis Research Foundation, San Diego, California The Scripps Research Institute, La Jolla, California
The TM1553 gene of Thermotoga maritima encodes a lipoprotein with a molecular weight of 39,409 Da (residues 1–352) and a calculated isoelectric point of 5.6. Sequence analysis reveals that TM1553 belongs to a predominantly prokaryotic family of proteins that are homologous to the ApbE family of periplasmic lipoproteins. ApbE is involved in thiamine (vitamin B1) biosynthesis and has been proposed to carry out the conversion of aminoimidazole ribotide (AIR) to 4-amino-5-hydroxymethyl-2-methyl pyrimidine (HMP).1, 2 Although the precise biochemical function of ApbE is not known, mutagenesis studies have indicated that ApbE is important for Fe-S cluster metabolism.3, 4 The exact role played by ApbE in either of these activities remains unclear. Herein, we report the crystal structure of TM1553, the first structural representative of the ApbE family, which was determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG).5 The TM1553 gene (TIGR: TM1553; Swiss-Prot: Q9X1N9), was amplified by polymerase chain reaction (PCR) from genomic DNA using Pfu Turbo and primer pairs encoding the predicted 5- and 3-ends (residues 40-352; cloned without the N-terminal transmembrane helix 1-39). The PCR product was cloned into plasmid pMH4, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the protein. The cloning junctions were confirmed by DNA sequencing. Protein expression was performed in a selenomethionine-containing medium using the Escherichia coli methionine auxotrophic strain DL41. Lysozyme was added to the culture at the end of fermentation to a final concentration of 250 μg/mL. Bacteria were lysed by sonication after a freeze/thaw procedure in Lysis Buffer [50 mM Tris pH 8.0, 50 mM NaCl, 10 mM imidazole, 1 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the cell debris was pelleted by centrifugation at 32,500 × g for 30 min. The soluble fraction was applied to a nickel-chelating resin (GE Healthcare) preequilibrated with Lysis Buffer. The resin was washed with Wash Buffer [50 mM Tris pH 8.0, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP], and the target protein was eluted with Elution Buffer [20 mM Tris pH 8.0, 300 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP]. The eluate was diluted 10-fold with Buffer Q [20 mM Tris, pH 7.9, 50 mM NaCl, 5% (v/v) glycerol, 1 mM TCEP] and applied to a RESOURCE Q column (GE Healthcare) preequilibrated with the same buffer. The flow-through fraction which contained the target protein was buffer exchanged with Crystallization Buffer [20 mM Tris pH 7.9, 150 mM NaCl, 1 mM TCEP] and concentrated for crystallization assays to 14 mg/mL by centrifugal ultrafiltration (Millipore). Molecular weight and oligomeric state of TM1553 were determined using a 1.0 × 30 cm Superdex 200 column (GE Healthcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl, and 0.02% (w/v) sodium azide. The protein was crystallized using the nanodroplet vapor diffusions method6 with standard JCSG crystallization protocols.5 The crystallization reagent contained 10% (v/v) 2-methyl-2,4-pentanediol (MPD), 0.1 M citrate pH 5.0. Twenty percent (v/v) MPD (final concentration) was added as a cryoprotectant. The crystals were indexed in the orthorhombic space group P212121 (Table I). Multiwavelength anomalous diffraction (MAD) data were collected at the Advanced Light Source (ALS, Berkeley, CA) on beamline 8.2.1 at wavelengths corresponding to the inflection (λ1) and low energy remote (λ2) of a MAD experiment. The datasets were collected at 100 K using an ADSC CCD detector. The MAD data were integrated and reduced using Mosflm9 and then scaled with the program SCALA from the CCP4 suite.7 Data statistics are summarized in Table I. The structure was determined with 1.58 Å selenium MAD data using the CCP4 suite7 and SOLVE/RESOLVE.10 Automatic model building was performed with iterative ARP/wARP runs.11 Model completion and refinement were performed with dataset λ1 using O12 and REFMAC5.13 Refinement statistics are summarized in Table I. Analysis of the stereochemical quality of the model was accomplished using AutoDepInputTool,14 MolProbity,15 SFcheck 4.0,16 and WHATIF 5.0.17 Protein quaternary structure analysis used the PQS server.18 Figures were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors for TM1553 at 1.58 Å resolution have been deposited in the Protein Data Bank (PDB) and are accessible under the code 1vrm. The PSI-BLAST19 program was used to detect homologs of DUF375 proteins in the NCBI nonredundant protein sequence database (March 8, 2005; 2,354,365 sequences; 800,120,167 total letters). Initially, a PSI-BLAST search (inclusion threshold 0.001) was performed until profile convergence, using as a query one of the DUF375 family members, hypothetical protein MTH727 from Methanothermobacter thermautotrophicus (gi|2621816). Subsequently, collected sequences were subjected to further transitive PSI-BLAST searches to identify other distantly related proteins. Additional searches were performed with the meta profile alignment method Meta-BASIC,20 which is available via the GRDB system (http://basic.bioinfo.pl/meta). Meta-BASIC combines the use of sequence profiles and secondary structure predictions (meta profiles) for the query sequence and given protein families with various scoring systems, and meta profile alignment algorithms to detect distant similarity between proteins, even if the structure of the reference protein is not known. Specifically, the consensus sequence of DUF375 was compared with all 7,418 PfamA families21 and with 10,128 proteins (clustered at 90% sequence identity) extracted from the PDB.22 The crystal structure of TM1553 [Fig. 1(A)] was determined to 1.58 Å resolution using the MAD method. Data collection, model, and refinement statistics are summarized in Table I. The final model includes a monomer (residues 44–352), three MPD molecules, an unknown ligand (UNL), and 404 water molecules in the asymmetric unit. No electron density was observed for residues 40–43 or for the expression and purification tag. The Matthews' coefficient (Vm)23 for TM1553 is 2.37 Å3/Da, and the estimated solvent content is 47.8%. The Ramachandran plot produced by MolProbity15 shows that 98.3% and 100.0% of the residues are in favored and allowed regions, respectively. Crystal structure of TM1553. A: Ribbon drawing where the N-terminal domain (cyan), helical domain (pink), linker region (slate-blue), and C-terminal domain (yellow) are colored to illustrate the domain organization. Helices (H1–H12) and β-strands (β1–β15) are labeled, and the unknown ligand is represented as orange spheres. B: Diagram showing the secondary structural elements in TM1553 superimposed on its primary sequence. The α-helices, 310-helices, β-strands (indicated by red A, B, C, and D to represent the β-sheet designation), β-bulges, and γ-turns are indicated. The β-hairpins are depicted as red loops. Dashed lines indicate that no electron density was observed for that region. TM1553 is composed of 15 β-strands (β1–β15), nine α-helices (H1–H4, H6, H8, H10–H12), and four 310-helices (H5, H7, H9, H11′) [Fig. 1(A,B)]. The total β-sheet, α-helical, and 310-helical content is 26.2, 32.4, and 3.6%, respectively. The TM1553 monomer contains three structural domains. The N-terminal (40–87, 191–226) and C-terminal (254–352) domains are extremely similar in topology and adopt a fold described previously as the tunneling fold (T-fold).24, 25 The T-fold core is composed of a four-stranded, antiparallel β-sheet (strand order 1234) with a pair of antiparallel α-helices placed between the second and third β-strands (ββααββ) [Figs. 1(A) and 2(A)]. This fold is present in a uricase and in a group of tetrahydrobiopterin biosynthesis enzymes whose known members form large, tunnel-shaped, homo-oligomeric barrels of different sizes [Fig. 2(A,B)]. When the two T-fold domains of TM1553 are structurally aligned, the root-mean-square deviation (RMSD) between 65 Cα atoms from each domain is 2.6 Å, indicating that TM1553 may have arisen from an ancestral gene duplication event. Although this alignment indicates only 8% sequence identity, it reveals a strong conservation of hydrophobic and polar residues between the domains. The two T-fold domains are connected by a short linker region (227–253) composed of a 310-helix (H9) and a β-hairpin (β8–β9). In addition to the T-fold domains, TM1553 contains a predominantly helical domain (88–190) that is inserted between α-helices H1 and H8 of the N-terminal T-fold domain [Fig. 1(A)]. This domain is composed of three short β-strands, four α-helices, and two 310-helices. A DALI26 structure similarity search of this domain did not return any significant hits, although it has a remote structural resemblance to members of the SCOP 3-helical bundle fold.27 A: Structural diagram of the monomeric unit of uricase (PDB 1r51), which is also composed of a tandem duplication of the T-fold. The N- and C-terminal domains are colored green and gray, respectively. B: Tunnel-shaped oligomeric complex of uricase. The active site is composed of residues from different subunits (colored gray, beige, green, and slate-blue) of the enzyme. The ligand molecules bound at the active sites are shown in CPK. C: Stereo diagram of the structure alignment of the monomeric unit of uricase and TM1553. TM1553 is colored blue and uricase gray. Uricase does not contain the additional helical domain. The ligands at the active sites of uricase (shown in CPK) and TM1553 (shown as orange spheres) are in completely different locations. D: Stereo diagram of the putative active site of TM1553. The structure of TM1553 is colored gray and the side-chains of residues within a 4 Å shell of the UNL are displayed in ball-and-stick representation. The electron density of a 2Fo-Fc omit map contoured at 1σ (blue) around the UNL (orange spheres) in the active site is shown. Water molecules in the active site are not displayed. So far, the structures of GTP cyclohydrolase I, 6-pyruvoyl tetrahydropterin synthase, 7,8-dihydroneopterin aldolase, 7,8-dihydroneopterin triphosphate epimerase, and uricase are known to adopt the T-fold. With the exception of uricase, whose monomeric unit also contains a tandem duplication of the T-fold unit, all these structures are homo-oligomers of a single T-fold domain. In the uricase and TM1553 structures, the duplicated T-fold domains are arranged in tandem so that β-strand 7 of the N-terminal β-sheet hydrogen bonds with β-strand 10 of the C-terminal β-sheet to form an eight-stranded, antiparallel β-sheet [Fig. 1(A)]. TM1553 and uricase (PDB 1r51) can be structurally aligned with an RMSD of 4.0 Å and a sequence identity of 8% over 141 Cα atoms [Fig. 2(C)]. This structural superposition reveals a similar overall fold in both structures, with the α-helices present on the concave surface of the curved, eight-stranded β-sheet. The T-fold domains of uricase contain long, N-terminal β-strands that are involved in oligomerization. TM1553, however, contains shorter β-strands as compared with uricase, and lacks the corresponding extensions of the β-strands that are involved in oligomerization. Therefore, it is unlikely that TM1553 would form an oligomeric complex, like the other members of the T-fold. Indeed, analysis of the crystallographic packing of TM1553 using the PQS server18 indicates that a monomer is the biologically relevant form [Fig. 1(A)]. This finding is also consistent with results from analytical size exclusion chromatography in combination with static light scattering. This observation is noteworthy because all previously determined structures that adopt the T-fold form large homo-oligomeric assemblies. In fact, it has been suggested that the T-fold may not be stable in isolation and must be associated with identical subunits to form barrel-shaped, homo-oligomeric complexes in order to be functional.24 TM1553 represents the first structure of a stable and functional enzyme that possesses T-fold domains, but does not oligomerize. Another interesting difference among the current members of the T-fold family and TM1553 is the location of the active site. All known members of the T-fold for which structures are available possess multiple symmetry-related active sites in their tunnel-shaped, homo-oligomeric complexes. Typically, each active site is located at the interface between T-fold monomers. In uricase, which contains a tandem duplication of the T-fold domain, the functional active site is also formed at the interface between two monomers upon oligomerization [Fig. 2(B)]. Surprisingly, TM1553 possesses a putative active site that is contained within the monomeric unit and is located at the interface between the two T-fold units and the helical insert domain [Fig. 1(A)]. An analysis of the TM1553 structure using the CastP server28 reveals a deep cavity of 1,300 Å3 at this interface. Location of additional electron density in this cavity that corresponds to a noncovalently bound small molecule, as well as sequence conservation between TM1553 homologs in this region, point to its role as the active site of this protein. Based on the electron density, different purines, nucleosides, and nucleotides where modeled and adenosine monophosphate proved to be the best fit. However, mass spectrometry failed to reveal the precise chemical identity of the ligand, and the “ligand” electron density was not sufficiently resolved to model any known small molecules with confidence. Therefore, it was modeled as an UNL. This UNL is coordinated by the side-chains of Phe130, Val134, Leu138, Asp188, Asp221, Ala258, Ser260, Glu264, His276, Ile277, Pro280, Asp303, Ser306, and Thr307, the main-chain of Ala129, Asp131, Gly190, Gly191, Thr259, and Leu278, and seven water molecules [Fig. 2(D)]. Of these, Asp188, Thr259, Ser260, His276, Pro280, Asp303, and Thr307 are particularly well conserved among homologs of TM1553, suggesting that they are essential for the enzyme's function. A DALI26 structural similarity search using the complete protein as a query did not find any significant hits to other protein structures. A search using the program ProSMoS (Grishin NV, unpublished) revealed similarities between the N- and C-terminal domains of TM1553 and members of the T-fold. A subsequent DALI26 search using the individual T-fold domains of TM1553 as a query then found links to members of the T-fold of SCOP. Pfam domains of unknown function (DUFs) are clusters of related protein sequences for which no fold or functional assignment could be made based on similarities to other functionally annotated protein sequences in the Pfam database.21 The crystal structure of TM1553 provides a structural template for DUF375 that encompasses many uncharacterized proteins from bacterial and archaeal species, including sequences belonging to an uncharacterized cluster of orthologs COG2122. The evolutionary relationship between ApbE (PF02424 in Pfam database) and DUF375 (PF04040 in Pfam database) families was indicated by the meta profile alignment method, Meta-BASIC, with a Z-score of about 25 (n.b. predictions with Z-score >12 have <5% probability of being incorrect).20 In addition, a PSI-BLAST19 search against the NCBI nonredundant protein sequence database (E-value threshold 0.001) initiated with one of the DUF375 family members, hypothetical protein MMP1236 from Methanococcus maripaludis (gi|45358799), found the putative thiamine biosynthesis lipoprotein ApbE (gi|29377698) from Enterococcus faecalis with an E-value of 4e-04 in the second iteration. In the next iteration, about 200 ApbE family proteins were detected with statistically significant E-values, including COG1477 (membrane-associated lipoprotein involved in thiamine biosynthesis), as well as TM1553, the only ApbE protein for which a crystal structure has been determined. A global multiple sequence alignment of DUF375 and representative ApbE family sequences was generated using the PCMA program29 and then adjusted manually according to TM1553 structure (gi|15644301, PDB 1vrm) (Fig. 3). The alignment reveals a strong conservation of hydrophobic residues and a highly conserved ligand binding site, indicating that DUF375, which belongs to the superfamily of ApbE-like proteins, is likely to possess its active site at a position similar to that observed in the ApbE structure. Multiple sequence alignment for DUF375 proteins (top) and ApbE representatives (bottom). Sequences are labeled according to the NCBI gene identification (gi) number and an abbreviation of the species name: Af, Archaeoglobus fulgidus; Bf, Burkholderia fungorum; Bj, Bradyrhizobium japonicum; Ca, Clostridium acetobutylicum; Dd, Desulfovibrio desulfuricans; De, Dehalococcoides ethenogenes; Dv, Desulfovibrio vulgaris; Ef, Enterococcus faecalis; Fn, Fusobacterium nucleatum; Ma, Methanosarcina acetivorans; Mb, Methanosarcina barkeri; Mj, Methanocaldococcus jannaschii; Mk, Methanopyrus kandleri; Ml, Melittangium lichenicola; Mm, Methanococcus maripaludis; Mo, Mesorhizobium loti; Ms, Methanobrevibacter smithii; Mt, Methanothermobacter thermautotrophicus; Mu, Methanococcoides burtonii; Ma, Methanosarcina mazei; Pg, Porphyromonas gingivalis; Ps, Polaromonas sp. JS666; Sp, Silicibacter pomeroyi; Tm, T. maritima; Ub, uncultured bacterium 582. Gi numbers for DUF375 sequences are colored according to taxonomy in green (bacterial) and red (archaeal). The first and last residue numbers are indicated before and after each sequence with the total sequence length of each protein indicated in square brackets at the end of the alignment. The numbers of excluded residues are specified in parentheses. Residue conservation is denoted with the following scheme: uncharged, highlighted in yellow; charged or polar, highlighted in gray; small, red. Locations of predicted (gi|2621816) and observed (gi|15644301; PDB 1vrm) secondary structure elements (E, β-strand; H, α-helix) are marked above the sequences. The TM1553 protein family (PF02424) contains more than 190 homologous sequences mostly of bacterial and archaeal origin. Models for TM1553 homologs can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=TM1553. The TM1553 crystal structure reported herein represents the first structure determined for the ApbE-like protein superfamily. The structure includes an unknown ligand that suggests a putative active site location. TM1553 reveals unexpected similarities to enzymes from the tetrahydrobiopterin biosynthesis pathway and offers a structural template for the so far uncharacterized DUF375 family in addition to the ApbE-like proteins. The evolutionary relationship between TM1553 and the enzymes from the tetrahydrobiopterin biosynthesis pathway, however, remains unclear. The information presented herein, in combination with further biochemical and biophysical studies, should yield valuable insights into the functional role of this enzyme. This work was supported by NIH Protein Structure Initiative grants from the National Institute of General Medical Sciences (www.nigms.nih.gov). Portions of this research were performed at the Stanford Synchrotron Radiation Laboratory (SSRL) and the Advanced Light Source (ALS). The SSRL is a national user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U.S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory.
Qingping Xu, Robert Schwarzenbacher, Daniel McMullan, Polat Abdubek, Sanjay Agarwalla, Eileen Ambing, Herbert Axelrod, Tanya Biorac, Jaume M. Canaves, Hsiu-Ju Chiu, Ashley M. Deacon, Michael DiDonato, Marc-André Elsliger, Adam Godzik, Carina Grittini, Slawomir K. Grzechnik, Joanna Hale, Eric Hampton, Gye Won Han, Justin Haugen, Michael Hornsby, Lukasz Jaroszewski, Heath E. Klock, Eric Koesema, Andreas Kreusch, Peter Kuhn, Scott A. Lesley, Mitchell D. Miller, Kin Moy, Edward Nigoghossian, Jessica Paulsen, Kevin Quijano, Ron Reyes, Chris Rife, Glen Spraggon, Raymond C. Stevens, Henry van den Bedem, Jeff Velasquez, Aprilfawn White, Guenter Wolf, Keith O. Hodgson, John Wooley, and Ian A. Wilson* The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, California The Genomics Institute of the Novartis Research Foundation, San Diego, California The Scripps Research Institute, La Jolla, California
The TM0604 gene of Thermotoga maritima encodes a single-stranded DNA-binding protein (SSB) with a molecular weight of 16,166 Da (residues 1–141) and a calculated isoelectric point of 4.8. SSB plays essential roles in DNA replication, recombination, and repair.1-3 SSB, also known as helix-destabilizing protein, binds tightly to single-stranded DNA (ssDNA) as a homotetramer. SSB from E. coli can bind long ssDNAs in two main binding modes, which occlude 35 or 65 nucleotides per tetramer, depending on how many SSB subunits from the tetramer interact with ssDNA. In the SSB35 binding mode, two SSB subunits per tetramer interact with ssDNA and self-assemble into long clusters. In the SSB65 binding mode, all four SSB subunits interact with ssDNA and do not form clusters.4 Closely related variants of SSB are encoded in the genome of a variety of large, self-transmissible plasmids. The eukaryotic mitochondrial proteins that bind ssDNA are structurally and evolutionarily related to prokaryotic SSB and are probably involved in mitochondrial DNA replication.5 Here, we report the crystal structure of TM0604, which was determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG).6 Protein Production and Crystallization. SSB from Thermotoga maritima (TIGR: TM0604, Swiss-Prot: Q9WZ73) was amplified by PCR from genomic DNA using PfuTurbo (Stratagene) and primer pairs encoding the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pMH2T7, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The cloning junctions were confirmed by sequencing. Protein expression was performed in a modified Terrific Broth using the Escherichia coli strain DL41. Lysozyme was added to the culture at the end of fermentation to a final concentration of 250 μg/mL. Bacteria were lysed by sonication after a freeze/thaw procedure in Lysis Buffer [50 mM Tris pH 7.9, 50 mM NaCl, 10 mM imidazole, 0.25 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the cell debris was pelleted by centrifugation at 3400 × g for 60 min. The soluble fraction was applied to a nickel-chelating resin (GE Heathcare) pre-equilibrated with Lysis Buffer. The resin was washed with Wash Buffer [50 mM potassium phosphate pH 7.8, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP], and the target protein was eluted with Elution Buffer [20 mM Tris pH 7.9, 300 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP]. The eluate was buffer exchanged with Buffer Q [20 mM Tris pH 7.9, 5% (v/v) glycerol, 0.25 mM TCEP] containing 50 mM NaCl and applied to a RESOURCE Q column (GE Heathcare) pre-equilibrated with the same buffer. The target protein was eluted using a linear gradient of 50 to 500 mM NaCl in Buffer Q. The appropriate RESOURCE Q fractions were pooled and further purified using a Superdex 200 column (GE Healthcare) with elution in Crystallization Buffer [20 mM Tris pH 7.9, 150 mM NaCl, 0.25 mM TCEP]. The appropriate Superdex 200 fractions were pooled and concentrated for crystallization assays to 10 mg/mL by centrifugal ultrafiltration (Millipore). Molecular weight and oligomeric state of the target protein were determined using a 1.0 × 30 cm Superdex 200 column (GE Heathcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl. The protein was crystallized using the nanodroplet vapor diffusion method7 with standard JCSG crystallization protocols.6 The crystallization reagent contained 25% MPD, 0.1M Tris pH 8.0. An additional 10% MPD (35% final concentration) was added as a cryoprotectant. The crystals were indexed in the orthorhombic space group F222 (Table I). Native diffraction data were collected to 2.3 Å resolution at the Advanced Light Source (ALS, Berkeley, CA) on beamline 5.0.3. The dataset was collected at 100 K using an ADSC Q4 CCD detector. Data were integrated and reduced using Mosflm8 and then scaled with the program SCALA from the CCP4 suite.9 Data statistics are summarized in Table I. Diffraction was anisotropic, resulting in incomplete data in the higher resolution shells. Intensity falloff is greatest in the b* axis direction. Because the data are incomplete to 2.3 Å, we define the nominal resolution as 2.6 Å, which is the resolution of a dataset that is 100% complete and has the same number of reflections as observed in the current dataset.10 There are 739 observed reflections between 2.6 and 2.3 Å (39% complete for this shell). The structure was determined with the JCSG molecular replacement pipeline11 using E. coli SSB protein (PDB: 1sru, sequence identity 34%) as the search model. Refinement was carried out using RESOLVE,12 REFMAC5,13 and XFIT.14 Refinement statistics are summarized in Table I. Analysis of the stereochemical quality of the model was accomplished using the AutoDepInputTool,15 MolProbity,16 SFcheck 4.0,9 and WHATIF 5.0.17 Figure 1(B) was adapted from an analysis using PDBsum,18 and all others were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors of TM0604 have been deposited within the PDB and are accessible under the code 1z9f. Crystal structure of SSB from Thermotoga maritima. (A) Stereo ribbon diagram of TM0604 monomer color-coded from N-terminus (blue) to C-terminus (red). The N-terminus corresponding to the actual start of SSB is labeled. The α-helix H1 and β-strands (β1–β6) are labeled. Disordered regions are indicated by dashed lines. (B) Diagram showing the secondary structural elements of TM0604 superimposed on its primary sequence. The α-helix, β-strands, β-turns, and disordered regions, with corresponding sequence in brackets, are indicated. The three-dimensional structure of TM0604 [Fig. 1(A)] was determined to a nominal resolution of 2.60 Å by the molecular replacement (MR) method using the coordinates of the E. coli SSB protein (PDB: 1sru, sequence identity of 34%)19 as the search model. Data collection, model, and refinement statistics are summarized in Table I. The final model contains only 88 of the 141 residues and includes one monomer (residues 1–23, 26–37, 49–85, and 93–108), one residue from the His-tag, and 14 water molecules. No electron density was observed for residues 24 and 25, 38 to 48, 86 to 92, 109 to 141, and the remaining residues of the expression and purification tag. The Matthews' coefficient (Vm)20 is 1.94 Å3/Da, and the estimated solvent content is 36.2%. The Ramachandran plot, produced by MolProbity,16 shows that 93.8%, 97.5%, and 2.5% of the residues are in favored, allowed, and disallowed regions, respectively. The outlier residues are Met1 (ϕ = 172.4° and ψ = −57.8°) and Ser2 (ϕ = 104.7° and ψ = 19.9°), which are located at the N-terminus, but have good electron density. The TM0604 monomer contains six β-strands (β1–β6) and one α-helix (H1) [Fig. 1(A,B)], with about 37% of the structure being disordered. The total β-strand and α-helical content is 34.1% and 8.4%, respectively. The Structural Classification of Proteins database (SCOP) classifies this protein as an oligonucleotide-binding fold (OB-fold).21 The OB-fold comprises a closed, five-stranded β-barrel architecture (β1, β3–β6) that packs, in the case of TM0604, against one α-helix (H1) and an additional β-strand (β2) [Fig. 1(A)]. Analytical size exclusion chromatography in combination with static light scattering indicates the oligomeric state to be a tetramer. A tetramer is also consistent with the analysis of crystallographic packing using the PQS server,22 as well as what has been reported for SSB from E. coli (PDB: 1eyg) [Fig. 2(A,B)]. (A) Stereo ribbon diagram of a superposition of the TM0604 tetramer (blue) and an SSB-ssDNA complex from E. coli (gray). (B) Stereo view of a model of a TM0604-ssDNA complex in surface representation with ssDNA depicted as sticks. The model is based on the SSB-ssDNA complex from E. coli. The surface is colored according to electrostatic potential in a range where red is negative (−84 kT/e) and blue is positive (+ 84 kT/e). A structural similarity search, performed with the coordinates of TM0604 using the DALI23 server, reveals the structure of SSB from Mycobacterium tuberculosis (PDB: 1ue1),24 with an RMSD of 1.3 Å over 86 aligned Cα atoms and a sequence identity of 30%. A superposition of TM0604 with SSB from E. coli (PDB: 1sru)19 reveals an RMSD of 1.5 Å over 79 aligned Cα atoms, whereas an alignment with an SSB-ssDNA complex from E. coli (PDB: 1eyg)4 gives an RMSD of 1.6 Å over 79 aligned Cα atoms, with a sequence identity of 34% in both cases. A structural comparison between tetramers of TM0604 and the SSB-ssDNA complex from E. coli reveals a similar tetrameric architecture and a dramatic decrease of disorder in the loops between 23 to 26, 37 to 49, and 85 to 93 upon binding to ssDNA25 [Fig. 2(A)]. From a number of residues implicated in critical interactions with ssDNA in E. coli SSB (Arg3, Gly15, Trp40, Trp54, Phe60, Lys62 Tyr70, Lys73, Trp88, Met109, and Met111), only Phe60, Tyr70, Lys73, Trp88, and Met109 are strictly conserved in TM0604. Nevertheless, a superposition of TM0604 with the SSB-ssDNA complex shows good agreement of surface topography and electrostatic potential and suggests a very similar ssDNA binding mode for TM0604 [Fig. 2(B)]. Currently, the SSB protein family contains more than 400 sequence homologues mainly of bacterial, eukaryotic, and viral origin. Models for TM0604 homologues can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=TM0604. The TM0604 structure reported here represents an SSB protein, whose structure has been determined by X-ray crystallography. The information reported here, in combination with further biochemical and biophysical studies, will yield valuable insights regarding the role of this protein in DNA replication, recombination, and repair. This work was supported by NIH Protein Structure Initiative grants P50-GM 62411 and U54 GM074898 from the National Institute of General Medical Sciences (http://www.nigms.nih.gov). Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory (SSRL) and the Advanced Light Source (ALS). The SSRL is a national user facility operated by Stanford University on behalf of the U. S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U. S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory.
The hyperthermophilic bacterium Thermotoga maritima has been the target of choice at the Joint Center for Structural Genomics (JCSG) for pipeline development and genome-wide structural annotation.1 In order to increase the fold coverage for this bacterium, the structure determination of the TM1367 gene product, which encodes a hypothetical protein with a molecular weight of 13,791 Da (residues 1–124) and a calculated isoelectric point of 4.86, was undertaken. This protein is a representative of the PFam2 domain of unknown function 369 (DUF369). Currently, no sequence-based functional annotation has been made for this protein, since profile-based sequence similarity methods, such as PSI-BLAST3 and Meta-BASIC,4 found similarities only to other hypothetical proteins, mainly of archaeal and bacterial origin. Here, we report the crystal structure of TM1367, which was determined using the semiautomated, high-throughput pipeline of JCSG,1 and propose that TM1367 is a novel member of the cyclophilin (peptidylprolyl isomerase) fold. The crystal structure of TM1367 [Fig. 1(A)] was determined to 1.90 Å resolution using the multi-wavelength anomalous diffraction (MAD) method. Data collection, model, and refinement statistics are summarized in Table I. The final refined structure includes three TM1367 monomers (residues 1–124 for chains A and B, and residues 1–123 for chain C), three His-tag residues for chains A and B, five His-tag residues for chain C, three PEG-200 ligands, one Ni2+ ion, and 243 water molecules in the asymmetric unit. The Matthews' coefficient (Vm)5 for TM1367 is 2.31 Å3/Da, and the estimated solvent content is 43.4%. The Ramachandran plot, produced by MolProbity,6 shows that 98.4 and 100.0% of the residues are in favored and allowed regions, respectively. A: Crystal structure of TM1367. Stereo ribbon diagram of TM1367 monomer color-coded from N-terminus (blue) to C-terminus (red). Helices (H1 and H2) and β-strands (β1–β8) are labeled. The structure includes His residues from the purification tag (three in chains A and B and five in chain C). B: Diagram showing the secondary structural elements in TM1367 superimposed on its primary sequence. The helices, β-strands of sheet A, and β-bulges are indicated. The β-hairpins are depicted as red loops. TM1367 is composed of eight β-strands (β1–β8), one α-helix (H1), and one 310 helix (H2) [Fig. 1(A,B)]. The total β-strand, α-helical, and 310-helical content is 41.9, 7.3, and 4.8%, respectively. The TM1367 monomer is comprised of a single structural domain (1–124) and belongs to the cyclophilin (peptidylprolyl isomerase) fold of the SCOP database.7 SCOP defines the core of this fold as an eight-stranded, antiparallel, closed β-barrel [Fig. 1(A)]. In addition to the β-strands, TM1367 and other members of the cyclophilin (peptidylprolyl isomerase) fold contain an α-helix (H1), which is located between β-strands β2 and β3, and a 310 helix (H2) between β-strands β7 and β8. Unlike other enzymes based on a β-barrel architecture, which generally have their active sites inside the β-barrel, members of this fold have both ends of the β-barrel closed due to the presence of the helices, and the active site is located on the exterior of the β-barrel. The interior of the closed β-barrel of TM1367 is extremely hydrophobic, similar to other members of the cyclophilin (peptidylprolyl isomerase) fold. A DALI8 structural similarity search, using the TM1367 structure as a query, found similarities to human cyclophilin A (PDB: 2cpl, Z = 10.8) and the C-terminal domain of a protein of unknown function from Bacillus cereus (PDB: 1x7f, Z = 8.3), whose structure was determined by the Midwest Center for Structural Genomics. The structural alignment of TM1367 with human cyclophilin A superimposes 115 Cα atoms with an RMSD of 2.7 Å (sequence identity 10%) [Fig. 2(A,B)], while the C-terminal domain of the protein of unknown function from B. cereus can be aligned over 100 Cα atoms with an RMSD of 2.9 Å and a sequence identity of 13%. A: Structural superposition of TM1367 (blue) and the human cyclophilin A (PDB: 2cpl, gray). Residues corresponding to the calcineurin binding site of human cyclophilin A are shown in ball-and-stick representation. B: Residues corresponding to PPlase active site of cyclophilin A are shown in ball-and-stick representation. TM1367 residues are shown in brackets. Cyclophilin A (CyPA) belongs to a group of cytosolic enzymes called peptidylprolyl cis-trans isomerases or rotamases (PPIase, E.C. 5.2.1.8). These enzymes mediate the conversion of Xaa-Pro peptide bonds between trans and cis conformations in Pro-containing polypeptides and have been proposed to play an important role in catalyzing the refolding of partly-denatured proteins in vivo. Based on structural and biochemical evidence, it has been suggested that the mechanism of the PPIase activity of CyPA depends on its ability to recognize a specific conformation of the peptide rather than the sequence.9-11 Cyclosporin A (CsA) is a fungally produced “11-residue” cyclic peptide that can bind CyPA with a very high affinity and inhibit its enzymatic function. It is a highly effective immunosuppressant whose main clinical use is in the prevention of organ-transplant rejection and in the treatment of autoimmune disorders mediated by T cells. Apart from the PPIase activity of CyPA, the CyPA-CsA binary complex binds to and inhibits the Ca2+-activated Ser/Thr phosphatase calcineurin. Binding of the CyPA-CsA complex to calcineurin, which plays an essential role in the antigen-induced proliferation of T cells, prevents it from dephosphorylating NFATp (nuclear factor of activated T cells), which suppresses T cell proliferation. It has been argued that the presence of the CyPA-CsA binary complex, rather than the inhibition of the rotamase activity, interferes with T cell proliferation.12, 13 The calcineurin binding site of human CyPA is comprised of Arg69, Asn71, Glu81, Lys82, Glu84, Pro105, Ser147, and Arg1489, 14 [Fig. 2(A)]. This site is not conserved in TM1367, and most of the residues of the calcineurin binding site correspond to insertions in human CyPA, as compared to TM1367 [Fig. 2(A)]. However, this is not totally unexpected since calcineurin inhibition by the CyPA-CsA complex has so far been reported only in mammals.15 CyPA contains a large number of isozymes, including cyclophilins B and C, that vary in their organism/tissue distribution.16 For example, while CyPA is found in the cytosol, CyPB and CyPC are found in the endoplasmic reticulum (ER) and mitochondrial matrices, respectively. The sequences at the ligand binding site of these different forms of cyclophilin-type PPIases are highly conserved.9 In human CyPA, the CsA binding site (also the rotamase active site) is comprised of Arg55, Phe60, Gln63, Thr73, Asn102, Phe113, Trp121, and His126 [Fig. 2(B)]. The corresponding side-chain residues in TM1367 (shown in brackets) are Asn37, Glu41, Tyr44, Pro70, Cys76, and Val95. Thr73 and Trp121 of CyPA are present in insert regions and have no corresponding residues in TM1367 [Fig. 2(B)]. It has been proposed that Arg55 of human CyPA, which is crucial for enzymatic function, hydrogen bonds with the imino nitrogen atom of the Xaa-Pro peptide bond, thus facilitating the cis-trans isomerization by deconjugating and, hence, weakening the peptidyl-prolyl amide bond.9 An Arg55 to Ala mutation in the human CyPA reduces its enzymatic activity 100-fold. TM1367, which exhibits a remarkable structural similarity to CyPA, does not have even a single residue of the PPIase active site conserved and, thus, it represents a novel member of the cyclophilin (peptidylprolyl isomerase) fold. Also, significant structural differences are observed in the loop regions that surround the active site. TM1367 contains a PEG-200 molecule bound in the location corresponding to the active site of human CyPA [Fig. 3(A)]. This PEG-200 molecule is stacked between the indole side-chains of Trp39 and Trp69. Other residues that contribute to this PEG binding site include Glu42, Tyr44, Pro91, Ala92, and Val95 [Fig. 3(A)]. The carboxamide side-chain of Asn37, which corresponds to the catalytic Arg55 of the human CyPA structure, points away from the PEG site and, hence, the biological function of this site is unclear. A: Structure of TM1367 with a PEG-200 molecule in the putative active site. Contact residues of chain B in a shell 4 Å from the PEG molecule are shown in ball-and-stick. B: Structural superposition of TM1367 (blue) and the C-terminal domain of a protein of unknown function from B. cereus (PDB: 1×7f, green), which is a seven- rather than eight-stranded, closed β-barrel. TM1367 also exhibits a strong structural similarity to the C-terminal domain of a protein of unknown function from B. cereus. However, this protein forms a seven-stranded β-barrel [Fig. 3(B)] rather than an eight-stranded β-barrel and lacks the corresponding N-terminal β-strand in the TM1367 structure. Despite lacking one β-strand of the β-barrel, the MCSG hypothetical protein superimposes remarkably well with CyPA and TM1367. Further, Arg272 of this protein aligns structurally with the catalytic Arg55 of human CyPA, although the other residues of the CyPA active site are not conserved. This Arg272 residue, which adopts a side-chain conformation that differs from CyPA Arg55, is conserved in all known homologues of the B. cereus protein. Analysis of the crystallographic packing of TM1367 using the PQS server17 indicates that a monomer is the biologically relevant form [Fig. 1(A)]. This finding is also consistent with results from analytical size exclusion chromatography in combination with static light scattering. The TM1367 protein presently contains about 25 close sequence homologues mainly of bacterial and archaeal origin. Structural models for these sequences are accessible at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=TM1367. The TM1367 crystal structure reveals unexpected similarities to members of the cyclophilin (peptidylprolyl isomerase) fold of the SCOP database and to the C-terminal domain of a hypothetical protein of unknown function from B. cereus. In the absence of statistically significant sequence similarity, pronounced structural similarity such as a high DALI Z-score (Z > 9)18 and the conservation of subtle structural features between protein structures19 have been used to argue for a divergent relationship. We hypothesize such a divergent evolutionary relationship between TM1367, CyPA, and the B. cereus hypothetical protein on the basis of their extensive structural similarity. However, residues of the PPIase active site and the calcineurin binding site are not conserved in TM1367, and the possibility of a convergent relationship cannot be ruled out. At this stage, the available sequence and structure information cannot define the function of TM1367. It will be of interest to investigate whether TM1367 has rotamase activity, binds CsA, or acts as a regulator of calcineurin. Further biochemical and biophysical studies, in combination with additional sequence and structural information for proteins of this fold, should yield valuable insights into the precise functional role and evolutionary origin of TM1367. TM1367 from Thermotoga maritima (TIGR: TM1367, Swiss-Prot: Q9X187) was amplified by PCR from genomic DNA using PfuTurbo (Stratagene) and primer pairs encoding the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pMH4, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The cloning junctions were confirmed by sequencing. Protein expression was performed in a selenomethionine-containing medium using the Escherichia coli strain GeneHogs® (Invitrogen). Lysozyme was added to the culture at the end of fermentation to a final concentration of 250 μg/mL. Bacteria were lysed by sonication after a freeze/thaw procedure in Lysis Buffer [50 mM Tris pH 7.9, 50 mM NaCl, 10 mM imidazole, 1 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the cell debris was pelleted by centrifugation at 32,500g for 30 min. The soluble fraction was applied to nickel-chelating resin (GE Healthcare) pre-equilibrated with Lysis Buffer. The resin was washed with Wash Buffer [50 mM Tris pH 7.9, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP], and the target protein was eluted with Elution Buffer [20 mM Tris pH 7.9, 300 mM imidazole, 10% (v/v) glycerol, 1 mM TCEP]. The eluate was diluted 10-fold with Buffer Q [20 mM Tris, pH 7.9, 50 mM NaCl, 5% (v/v) glycerol, 1 mM TCEP] and applied to a RESOURCE Q column (GE Healthcare) pre-equilibrated with the same buffer. The flow-through fraction, which contained the target protein, was buffer exchanged with Crystallization Buffer (20 mM Tris pH 7.9, 150 mM NaCl, 1 mM TCEP) and concentrated for crystallization assays to 15 mg/mL by centrifugal ultrafiltration (Millipore). Molecular weight and oligomeric state of the target protein were determined using a 1.0 × 30 cm Superdex 200 column (GE Healthcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 8.0, 150 mM NaCl, 0.02% (w/v) sodium azide. The protein was crystallized using the nanodroplet vapor diffusion method20 with standard JCSG crystallization protocols.1 The crystallization reagent contained 50% polyethylene glycol 200 (PEG-200) 200, 0.2 M NaCl, 0.1 M phosphate/citrate pH 4.2. The crystals were indexed in orthorhombic space group P21212 (Table I). Multi-wavelength anomalous diffraction data were collected at SSRL (Stanford, CA) on beamline 9-2 at wavelengths corresponding to the high energy remote (λ1) and the inflection point (λ2) of a selenium MAD using the BLU-ICE21 data collection environment (Table I). All data sets were collected at 100K using a Mar 325 CCD detector. Data were integrated and reduced using Mosflm22 and then scaled with the program SCALA from the CCP4 suite.23 Data statistics are summarized in Table I. The initial structure was determined with the 1.90 Å selenium MAD data (λ1,2) using the CCP4 suite23 and SOLVE/RESOLVE.24 Model building and refinement were performed using O25 and REFMAC5.23 Refinement statistics are summarized in Table I. The final model includes three monomers (residues 1–124 for chains A and B and 1–123 for chain C), three residues of the His-tag for chains A and B, five residues of the His-tag for chain C, three PEG-200 ligands, one Ni2+ ion, and 243 water molecules in the asymmetric unit. Analysis of the stereochemical quality of the model was accomplished using AutoDepInputTool,26 MolProbity,6 SFcheck 4.0,27 and WHATIF 5.0.28 Protein quaternary structure analysis was performed using the PQS server.17 Figures were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors for TM1367 at 1.90 Å resolution have been deposited in the PDB and are accessible under the code 1zx8. This work was supported by NIH Protein Structure Initiative grants P50-GM62411 and U54-GM074898 from the National Institute of General Medical Sciences (http://www.nigms.nih.gov). Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory (SSRL) and the Advanced Light Source (ALS). The SSRL is a national user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the U.S. Department of Energy under Contract No. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory.
Irimpan I. Mathews, S. Sri Krishna, Robert Schwarzenbacher, Daniel McMullan, Lukasz Jaroszewski, Mitchell D. Miller, Polat Abdubek, Sanjay Agarwalla, Eileen Ambing, Herbert L. Axelrod, Jaume M. Canaves, Dennis Carlton, Hsiu-Ju Chiu, Thomas Clayton, Michael DiDonato, Lian Duan, Marc-André Elsliger, Slawomir K. Grzechnik, Joanna Hale, Eric Hampton, Justin Haugen, Kevin K. Jin, Heath E. Klock, Eric Koesema, John S. Kovarik, Andreas Kreusch, Peter Kuhn, Inna Levin, Andrew T. Morse, Edward Nigoghossian, Linda Okach, Silvya Oommachen, Jessica Paulsen, Kevin Quijano, Ron Reyes, Christopher L. Rife, Glen Spraggon, Raymond C. Stevens, Henry van den Bedem, Aprilfawn White, Guenter Wolf, Qingping Xu, Keith O. Hodgson, John Wooley, Ashley M. Deacon, Adam Godzik, Scott A. Lesley, and Ian A. Wilson* Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California Joint Center for Structural Genomics Burnham Institute for Medical Research, La Jolla, California Center for Research in Biological Systems, University of California, San Diego, La Jolla, California Genomics Institute of the Novartis Research Foundation, San Diego, California The Scripps Research Institute, La Jolla, California San Diego Supercomputer Center, La Jolla, California
The TM1246 gene (smPurL) of Thermotoga maritima encodes a phosphoribosylformylglycinamidine synthase II (FGAM; EC 6.3.5.3) with a molecular weight of 65,834 Da (residues 1–603) and a calculated isoelectric point of 5.38. This enzyme is part of the de novo purine biosynthesis subsystem, where it forms a complex with purS (TM1244) and purQ (TM1245) to form formylglycinamide ribonucleotide amidotransferase (FGAR-AT). FGAR-AT catalyzes the adenosine 5′-triphosphate (ATP)-dependent conversion of FGAR and glutamine to formylglycinamidine ribonucleotide (FGAM), adenosine 5′-diphosphate (ADP), Pi, and glutamate in the fourth step of the purine biosynthetic pathway (EC 6.3.5.3).1-3 In Gram-positive bacteria, archaebacteria, and the Gram-negative T. maritima, FGAR-AT is a complex of three proteins: PurS, PurL (designated as smPurL), and PurQ. In eukaryotes and other Gram-negative bacteria, FGAR-AT is a multidomain protein with a molecular mass of about 140 kDa (designated as lgPurL). Based on iterative sequence similarity searches, it has been proposed that smPurL together with lgPurL, aminoimidazole ribonucleotide synthetase (PurM), Ni-Fe hydrogenase maturation protein (HypE), selenophosphate synthetase (SelD), and thiamine monophosphate kinase (ThiL) form a new structural superfamily of ATP-dependent enzymes.4 The structures of FGAR-AT (lgPurL) from Salmonella typhimurium [Protein Data Bank (PDB) 1t3t], PurM protein from Escherichia coli (PDB 1cli), and ThiL from Aquifex aeolicus (PDB 1vqv) are known.4-6 Herein, we report the crystal structure of TM1246 (smPurL), determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG).7 The crystal structure of TM1246 [Fig. 1(A)] was determined to 2.15 Å resolution using the multi-wavelength anomalous dispersion (MAD) method. Data collection, model, and refinement statistics are summarized in Table I. The final model includes one monomer (residues 2–186, 203–603), one chloride ion, and 221 water molecules in the asymmetric unit. The Matthews' coefficient (Vm)8 for TM1246 is 2.14 Å3/Da and the estimated solvent content is 42.0%. A main-chain torsion angle analysis by the program MolProbity9 shows that 97% and 100% of the residues are in the favored and allowed regions of the Ramachandran plot, respectively. Crystal structure of smPurL from T. maritima. A: Stereo ribbon diagram of TM1246 monomer color-coded from N-terminus (blue) to C-terminus (red) showing the domain organization. Helices H1–H24 and β-strands (β1–β25) are indicated. The disordered region is depicted by a dashed line with the start and end residues labeled. B: Diagram showing the secondary structure elements of TM1246 superimposed on its primary sequence. The α-helices, 310-helices, β-strands, β-bulges, and γ-turns are indicated. The four β-sheets are indicated by a red A, B, C, and E. β-Hairpins are depicted as red loops. Disordered regions are depicted by a dashed line with the corresponding sequence shown below. The TM1246 monomer contains 25 β-strands (β1–β25) in four β-sheets (A, B, C, E), one β-hairpin (D), 18 α-helices (H1–H7, H10–H11, H13, H15–H17, H19–H21, H23–H24), and seven 310-helices (H8–H9, H12, H14, H18, H22–H23′) [Fig. 1(A,B)]. The total β-strand, α-helical, and 310-helical content is 27.6, 31.7, and 3.1%, respectively. TM1246 comprises four α+β domains [Figs. 1(A) and 2(A)]. The first (A1: 2–166) and third domains (A2: 362–507) are related by pseudo-twofold symmetry and pack against each other to form the central structural unit of TM1246 [Fig. 2(A)]. These domains adopt a two-layered α+β fold whose core is composed of a mixed four-stranded β-sheet (strand order 1423) and two α-helices arranged in a βαβαββ motif. These domains resemble the N-terminal domain of PurM and belong to the “Bacillus chorismate mutase-like” fold.10 The other domains of smPurL (B1: 167–345 and B2: 508–603) have a curved β-sheet and adopt a three-layered α+β fold composed of a four-stranded antiparallel β-sheet and three α-helices arranged in a βαβαβαβ motif. This fold has some resemblance to the ferredoxin fold and has been classified in the SCOP database10 under the “PurM C-terminal domain-like” fold. A linker peptide (residues 346–361) connects the two PurM-like units [Figs. 1(A) and 2(A)]. The domains of smPurL are likely a result of a tandem duplication of an ancestral PurM-like subunit and are arranged like the PurM dimer (A1B1-A2B2) [Figs. 1(A) and 2(A)]. Domain arrangement and structural alignment of smPurL and lgPurL. A: Domain arrangement of PurM, ThiL, smPurL, and lgPurL proteins. PurM and ThiL are homodimers. smPurL (gray) and the central domain of lgPurL (blue) have the same domain arrangement as a PurM dimer. lgPurL has two additional domains as compared with smPurL. The N-terminal domain homologous to the PurS protein is colored green and the C-terminal glutaminase domain is colored red. B: Stereo ribbon diagram of a superposition of smPurL (gray) and residues 183–960 of lgPurL from S. typhimurium (blue). The lgPurL N-terminal domain of unknown function and the C-terminal glutaminase domain are colored green and red, respectively. These extra domains of lgPurL correspond to the PurS (TM1244) and PurQ (TM1245) proteins in T. maritima. The structures were aligned using the DALI server. The crystallographic packing of the TM1246 structure, as well as analytical size exclusion chromatography in combination with static light scattering, indicates that a monomer is the biologically-relevant form. A search performed with the coordinates of TM1246 using the DALI server11 showed structural similarity to residues 183–960 of formylglycinamide synthetase from S. typhimurium (PDB: 1t3t) (Z = 28.6).6 The root-mean-square deviation (RMSD) for this structural alignment is 2.8 Å over 553 aligned Cα atoms with 21% sequence identity [Fig. 2(B)]. The smPurL is also homologous to the structures of PurM and ThiL.4-6, 12 The structural alignment of PurM from E. coli4 (PDB: 1cli) with the N-terminal domains of smPurL superimposes 229 Cα atoms with an RMSD of 3.2 Å, whereas the ThiL protein from Aquifex aeolicus (PDB: 1vqv) can be aligned over 210 Cα atoms with an RMSD of 3.6 Å. Because TM1246 represents the first structure of a smPurL, it is interesting to compare it with the lgPurL structure. Although the overall fold between TM1246 and the FGAM synthetase domain of lgPurL is very similar, two large insertions are found in the lgPurL structure [Fig. 2(A,B)]. These insertions, which occur on A1 and B2, are placed in close structural proximity, creating a new surface on one side of the enzyme [Fig. 2(A,B)]. The very compact nature of the T. maritima enzyme structure seems to suggest that this represents a minimal version of the FGAM synthetase domain. The active site of the smPurL from T. maritima was inferred from sequence and structural comparison to the lgPurL from S. typhimurium.6 The putative active site is located in the cleft formed by the “Bacillus chorismate mutase-like” fold and the succeeding “PurM C-terminal domain-like” fold [Figs. 2(B) and 3(A)]. Unlike lgPurL, which binds two sulfate ions in its putative active site, the smPurL structure does not contain any bound ligands. The secondary structural elements around the active site of smPurL superimpose very well with the lgPurL structure. Seven of the nine residues that interact with the sulfate ions in the active site region of lgPurL are conserved in smPurL and adopt similar side-chain conformations [Fig. 3(A)]. A: Stereo diagram of a close-up of the putative active site of TM1246 superimposed on the lgPurL structure shown in ribbon representation. The sulfate ions are from the lgPurL structure. The start and end regions of the glycine-rich loop in both the structures are labeled in red. B: Stereo diagram of a close-up of the lgPurL ADP-binding site superimposed on the smPurL structure shown in ribbon representation. The ADP moiety and Mg2+ ions are from the lgPurL structure. In A and B, residues are numbered according to TM1246 structure (PDB 1vk3) and the equivalent residues of lgPurL (PDB 1t3t) are shown in parentheses. The lgPurL structure contains an auxiliary ADP-binding site that is related to the active site by pseudo- twofold symmetry [Figs. 2(B) and 3(B)]. The sequence and structural conservation at this region is less pronounced than for the putative active site. Recent biochemical studies on the Bacillus subtilis smPurL have indicated that ADP is required for the assembly of the PurSLQ complex.5 However, in the TM1246 structure, no ADP or bound anions are observed at this site. Of the five highly conserved residues of the lgPurL protein family that are involved in interactions with the ADP moiety and the Mg2+ ions (K649, E718, N722, D884, and D887 of PDB 1t3t), only E425 (E718 in lgPurL) is conserved in TM1246 [Fig. 3(B)]. None of the residues that are involved in forming the hydrophobic pocket in lgPurL are conserved in TM1246, although these regions are structurally similar. The structural differences occur mainly around the regions that accommodate the base ring and the sugar moiety of ADP. smPurL contains a glycine-rich loop that is structurally disordered (residues 187–202) and is located close to the active site [Fig. 3(A)]. Interestingly, the equivalent region of lgPurL (448–466) is also disordered and is positioned to cover the active site upon binding of the ATP moiety. It is likely that this loop will become ordered in the ATP-bound form of the enzyme. As noted before, the lgPurL structure has two extra domains as compared with the smPurL structure: an N-terminal domain of unknown function [Fig. 2(A,B), colored green] and a C-terminal glutaminase domain [Fig. 2(A,B), colored red]. The structural study of lgPurL revealed that the N-terminal domain of unknown function is structurally homologous to a dimer of PurS.6 It is likely that the PurS and PurQ domains of T. maritima (TM1244 and TM1245) occupy similar positions to the additional N- and C-terminal domains of lgPurL [Fig. 2(A,B)]. smPurL is the last remaining enzyme in the purine biosynthetic pathway to have its structure determined. The smPurL family contains hundreds of sequence homologs. Models for TM1246 homologs can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=TM1246. The crystal structure of TM1246 represents a smPurL protein. The information reported herein, in combination with the structure of lgPurL and further biochemical and biophysical studies, will yield valuable insights regarding the role of this protein in purine biosynthesis. Phosphoribosylformylglycinamidine synthase II from T. maritima (TIGR: TM1246, Swissprot: Q9X0X3) was amplified by polymerase chain reaction (PCR) from genomic DNA using PfuTurbo (Stratagene) and primer pairs encoding the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pMH2T7, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The cloning junctions were confirmed by sequencing. Protein expression was performed in a selenomethionine-containing medium using the E. coli methionine auxotrophic strain DL41. Lysozyme was added to the culture at the end of fermentation to a final concentration of 250 μg/mL. Bacteria were lysed by sonication after a freeze/thaw procedure in Lysis Buffer [50 mM Tris pH 7.9, 50 mM NaCl, 10 mM imidazole, 0.25 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the cell debris was pelleted by centrifugation at 3,400 × g for 60 min. The soluble fraction was applied to nickel-chelating resin (GE Healthcare) pre-equilibrated with Lysis Buffer. The resin was washed with Wash Buffer [50 mM potassium phosphate pH 7.8, 300 mM NaCl, 40 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP], and the target protein was eluted with Elution Buffer [20 mM Tris pH 7.9, 300 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP]. The eluate was buffer exchanged with Buffer Q [20 mM Tris pH 7.9, 5% (v/v) glycerol, 0.25 mM TCEP] containing 50 mM NaCl and was applied to a RESOURCE Q column (GE Healthcare) preequilibrated with the same buffer. The target protein was eluted using a linear gradient of 50–500 mM NaCl in Buffer Q. The appropriate fractions were pooled, buffer exchanged with Crystallization Buffer [20 mM Tris pH 7.9, 150 mM NaCl, 0.25 mM TCEP], and concentrated for crystallization assays to 15 mg/mL by centrifugal ultrafiltration (Millipore). Molecular weight and oligomeric state of the target protein were determined using a 1.0 × 30 cm Superdex 200 column (GE Healthcare) in combination with static light scattering (Wyatt Technology). The mobile phase consisted of 20 mM Tris pH 7.9, 150 mM NaCl. The protein was crystallized using the nanodroplet vapor diffusion method13 with standard JCSG crystallization protocols.7 The crystallization reagent contained 30% polyethylene glycol (PEG)-200, 7% PEG-4000, 0.1 M Tris pH 8.0. No additional cryoprotectant was required. The crystals were indexed in orthorhombic space group P212121 (Table I). MAD data were collected at the Advanced Light Source (ALS, Berkeley, CA) on beamline 8.2.1 at wavelengths corresponding to the low energy remote (λ1) and the peak (λ2) of a selenium MAD experiment. In addition, a second crystal was used to collect a 2.15 Å high-resolution dataset (λ0) on beamline 8.2.2. The datasets were collected at 100 K using ADSC CCD detectors. Data were integrated and reduced using Mosflm14 and then scaled with the program SCALA from the CCP4 suite.15 Data statistics are summarized in Table I. The initial structure was determined with the 2.20 Å selenium MAD data (λ1, 2) using the CCP4 suite15 and SOLVE/RESOLVE.16 Model building and refinement were performed on the high-resolution data set (λ0) using O17 and REFMAC5.15 Refinement statistics are summarized in Table I. The final model includes one protein monomer, one chloride ion, and 221 water molecules in the asymmetric unit. No electron density was observed for residues 1, 187–202, and the expression and purification tag. Analysis of the stereochemical quality of the model was performed using AutoDepInputTool (http://deposit.pdb.org/adit/), MolProbity,9 SFcheck 4.0,18 and WHAT IF 5.0.19 Figure 1(B) was adapted from an analysis using PDBsum (http://www.biochem.ucl.ac.uk/bsm/pdbsum/) and all others were prepared with PyMOL (DeLano Scientific). Atomic coordinates and experimental structure factors of TM1246 have been deposited within the PDB and are accessible under the code 1vk3. This work was supported by the National Institutes of Health Protein Structure Initiative grants P50 GM62411 and U54 GM074898 awarded by the National Institute of General Medical Sciences (www.nigms.nih.gov). Portions of this research were performed at the Stanford Synchrotron Radiation Laboratory (SSRL) and the Advanced Light Source (ALS). The SSRL is a national user facility operated by Stanford University on behalf of the United States Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). The ALS is supported by the Director, Office of Science, Office of Basic Energy Sciences, Materials Sciences Division, of the United States Department of Energy under contract no. DE-AC03-76SF00098 at Lawrence Berkeley National Laboratory.
Proteins: Structure, Function, and BioinformaticsVolume 65, Issue 3 p. 771-776 Structure Note Crystal structure of 2-phosphosulfolactate phosphatase (ComB) from Clostridium acetobutylicum at 2.6 Å resolution reveals a new fold with a novel active site Michael DiDonato, Michael DiDonato The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorS. Sri Krishna, S. Sri Krishna The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorRobert Schwarzenbacher, Robert Schwarzenbacher The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorDaniel McMullan, Daniel McMullan The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSanjay Agarwalla, Sanjay Agarwalla The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorScott M. Brittain, Scott M. Brittain The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMitchell D. Miller, Mitchell D. Miller The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorPolat Abdubek, Polat Abdubek The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEileen Ambing, Eileen Ambing The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorHerbert L. Axelrod, Herbert L. Axelrod The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJaume M. Canaves, Jaume M. Canaves The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorHsiu-Ju Chiu, Hsiu-Ju Chiu The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAshley M. Deacon, Ashley M. Deacon The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorLian Duan, Lian Duan The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorMarc-André Elsliger, Marc-André Elsliger The Joint Center for Structural GenomicsSearch for more papers by this authorAdam Godzik, Adam Godzik The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorSlawomir K. Grzechnik, Slawomir K. Grzechnik The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorJoanna Hale, Joanna Hale The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Hampton, Eric Hampton The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorJustin Haugen, Justin Haugen The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLukasz Jaroszewski, Lukasz Jaroszewski The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorKevin K. Jin, Kevin K. Jin The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorHeath E. Klock, Heath E. Klock The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMark W. Knuth, Mark W. Knuth The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Koesema, Eric Koesema The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorAndreas Kreusch, Andreas Kreusch The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorPeter Kuhn, Peter Kuhn The Joint Center for Structural GenomicsSearch for more papers by this authorScott A. Lesley, Scott A. Lesley The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorInna Levin, Inna Levin The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAndrew T. Morse, Andrew T. Morse The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorEdward Nigoghossian, Edward Nigoghossian The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLinda Okach, Linda Okach The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSilvya Oommachen, Silvya Oommachen The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJessica Paulsen, Jessica Paulsen The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorKevin Quijano, Kevin Quijano The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorRon Reyes, Ron Reyes The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorChristopher L. Rife, Christopher L. Rife The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorGlen Spraggon, Glen Spraggon The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorRaymond C. Stevens, Raymond C. Stevens The Joint Center for Structural GenomicsSearch for more papers by this authorHenry van den Bedem, Henry van den Bedem The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAprilfawn White, Aprilfawn White The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorGuenter Wolf, Guenter Wolf The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorQingping Xu, Qingping Xu The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorKeith O. Hodgson, Keith O. Hodgson The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJohn Wooley, John Wooley The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorIan A. Wilson, Corresponding Author Ian A. Wilson wilson@scripps.edu The Joint Center for Structural GenomicsJCSG, The Scripps Research Institute, BCC206, 10550 North Torrey Pines Road, La Jolla, CA 92037===Search for more papers by this author Michael DiDonato, Michael DiDonato The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorS. Sri Krishna, S. Sri Krishna The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorRobert Schwarzenbacher, Robert Schwarzenbacher The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorDaniel McMullan, Daniel McMullan The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSanjay Agarwalla, Sanjay Agarwalla The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorScott M. Brittain, Scott M. Brittain The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMitchell D. Miller, Mitchell D. Miller The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorPolat Abdubek, Polat Abdubek The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEileen Ambing, Eileen Ambing The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorHerbert L. Axelrod, Herbert L. Axelrod The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJaume M. Canaves, Jaume M. Canaves The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorHsiu-Ju Chiu, Hsiu-Ju Chiu The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAshley M. Deacon, Ashley M. Deacon The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorLian Duan, Lian Duan The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorMarc-André Elsliger, Marc-André Elsliger The Joint Center for Structural GenomicsSearch for more papers by this authorAdam Godzik, Adam Godzik The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorSlawomir K. Grzechnik, Slawomir K. Grzechnik The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorJoanna Hale, Joanna Hale The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Hampton, Eric Hampton The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorJustin Haugen, Justin Haugen The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLukasz Jaroszewski, Lukasz Jaroszewski The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorKevin K. Jin, Kevin K. Jin The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorHeath E. Klock, Heath E. Klock The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorMark W. Knuth, Mark W. Knuth The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorEric Koesema, Eric Koesema The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorAndreas Kreusch, Andreas Kreusch The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorPeter Kuhn, Peter Kuhn The Joint Center for Structural GenomicsSearch for more papers by this authorScott A. Lesley, Scott A. Lesley The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorInna Levin, Inna Levin The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAndrew T. Morse, Andrew T. Morse The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorEdward Nigoghossian, Edward Nigoghossian The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorLinda Okach, Linda Okach The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorSilvya Oommachen, Silvya Oommachen The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJessica Paulsen, Jessica Paulsen The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorKevin Quijano, Kevin Quijano The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorRon Reyes, Ron Reyes The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorChristopher L. Rife, Christopher L. Rife The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorGlen Spraggon, Glen Spraggon The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorRaymond C. Stevens, Raymond C. Stevens The Joint Center for Structural GenomicsSearch for more papers by this authorHenry van den Bedem, Henry van den Bedem The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorAprilfawn White, Aprilfawn White The Joint Center for Structural Genomics The Genomics Institute of the Novartis Research Foundation, San Diego, CaliforniaSearch for more papers by this authorGuenter Wolf, Guenter Wolf The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorQingping Xu, Qingping Xu The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorKeith O. Hodgson, Keith O. Hodgson The Joint Center for Structural Genomics The Scripps Research Institute, La Jolla, CaliforniaSearch for more papers by this authorJohn Wooley, John Wooley The Joint Center for Structural Genomics The University of California, San Diego, La Jolla, CaliforniaSearch for more papers by this authorIan A. Wilson, Corresponding Author Ian A. Wilson wilson@scripps.edu The Joint Center for Structural GenomicsJCSG, The Scripps Research Institute, BCC206, 10550 North Torrey Pines Road, La Jolla, CA 92037===Search for more papers by this author First published: 22 August 2006 https://doi.org/10.1002/prot.20978Citations: 2Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinked InRedditWechat Citing Literature Volume65, Issue315 November 2006Pages 771-776 RelatedInformation
In order to extend the structural coverage of eukaryotic members in the Protein Family database (Pfam),1 we selected 400 open reading frames (ORFs) from available cDNA libraries of the Mouse genome. One of these, the Mouse gene DcpS (GI: 16740816, PF05652), belongs to the family of scavenger mRNA decapping enzymes. DcpS is a scavenger pyrophosphatase that hydrolyzes the residual cap structure following mRNA degradation in the 3′–5′ exoribonuclease pathway.2 The association of DcpS with 3′–5′ exonuclease exosome components suggests that these two activities are linked within a coupled exonucleolytic, decay-dependent decapping pathway.3 The family contains a histidine triad (HIT) motif4 with three histidines (His274, His276, and His278 in Mouse DcpS) separated by hydrophobic residues. The central histidine within the DcpS HIT motif is critical for decapping activity and defines the HIT motif as a new mRNA decapping domain, making DcpS the first member of the HIT family of proteins with a defined biological function. Mouse DcpS has a molecular weight of 37,921 Da (residues 1–351) and a calculated isoelectric point of 5.4. Here, we report the crystal structure of DcpS from Mouse determined using the semiautomated, high-throughput pipeline of the Joint Center for Structural Genomics (JCSG).5 The crystal structure of Mouse DcpS [Fig. 1(A)] was determined to 1.83-Å resolution by the molecular replacement (MR) method using a search model constructed from human DcpS [Protein Data Bank (PDB) code: 1st0], with a sequence identity of 89%.6 Data collection, model, and refinement statistics are summarized in Table I. The final model includes two monomers, [(chain A residues 38–337); (chain B residues 39–337)], six 1,2-ethanediol (ethylene glycol) molecules, and 800 water molecules. The Matthews' coefficient (Vm) is 2.56 Å3/Da, and the estimated solvent content is 52.0%. The Ramachandran plot, produced by MolProbity,9 shows that 98.43% and 1.67% of the residues are in favored and allowed regions, respectively. Crystal structure of mRNA decapping enzyme (DcpS) from Mouse. (A) Stereo ribbon diagram of Mouse DcpS color-coded from N-terminus (blue) to C-terminus (red) showing the domain organization with helices H1–H13 and β-strands β1–β14, as well as β-sheets A, B, C, and D. (B) Diagram showing the secondary structure elements in Mouse DcpS (chain A) superimposed on its primary sequence. The β-sheet strands are indicated by a red A, B, C, and D. β-bulges and γ-turns are indicated. β-hairpins are depicted as red loops. Disordered regions are depicted by a dashed line with the corresponding sequence in brackets. The HIT sequence motif is indicated by a black line above the three histidines. The DcpS monomer consists of 14 β-strands (β1–β14), nine α-helices (H1, H3–H6, H9, and H11–H13), and four 310-helices (H2, H7–H8, and H10) [Fig. 1(A, B)]. The total β-strand, α-helical, and 310-helical content is 29.6%, 33.3%, and 5.8%, respectively. DcpS consists of an N-terminal domain (residues 33–130) followed by a short linker region (residues 131–147) and a larger C-terminal domain (residues 148–337). The N-terminal domain is elongated and consists of two α-helices (H1, H2) and two antiparallel β-sheets, A and B, comprised of β-strands β1, β2, β3, and β6 with 1234 topology, and β4 and β5, respectively. The C-terminal domain is formed by 10 α-helices (H4–H13) and 2 antiparallel β-sheets, C and D, comprised of β-strands β7 and β14, and β8– β13 with 123546 topology respectively [Fig. 1(A and B). DcpS forms an intimate homodimer primarily through domain-swapped interactions of two strands (β4, β5) and 1 helix (H1) from the N-terminal domain [Fig. 2(A)]. The dimer interface in the N-terminal domain is then formed from α-helix H1, the 310-helices H2 and H7, the two β-strands β5 and β6, and two short loop regions between residues 54 and 61 and 83 and 88. This crossover interaction creates a tightly-assembled domain in which the two β-sheets, A and B, form two symmetry-related, continuous, antiparallel β-sheets [Fig. 2(A)]. The dimer interface in the C-terminal domain involves interactions of α-helix H11, β-strand β14, and the loop region between residues 282 and 294. The dimer interface is 60% hydrophobic but includes 4 pairs of salt bridges and hydrogen bonds between residues Lys59 and Glu84, Glu102 and Arg121, Arg151 and Glu302, and Glu292 and Arg263. The interface buries an accessible surface area of 3936 Å2 per monomer. Comparison of Mouse and human DcpS. (A) Ribbon diagram of the Mouse DcpS dimer. Chain A is in gray and chain B is in pink. The dimer is symmetric with a Cα–Cα distance between Asp110 and Trp174 of ∼28 Å. Residues 131 and 147 flanking the linker region are labeled. (B) Ribbon diagram of the human DcpS/mGpppG complex. Chain A is shown in gray and chain B in purple. The dimer has a symmetric bottom and an asymmetric top and features an open side and a closed side with a 36 Å and 6 Å Cα–Cα gap between Asp111 and Trp175, respectively. The mGpppG nucleotides bound to the active sites are shown in ball-and-stick configuration, with carbon atoms colored in yellow, phosphorous in purple, oxygen in red, and nitrogen in blue. The movement of the N-terminal domain is facilitated by a conformational change in the linker region around helix E3. Superposition of the DcpS active sites in the relaxed empty state with the open (C) and closed (D) mGpppG bound state. The mGpppG-interacting residues from human DcpS (gray, residues labeled in brackets) and their counterparts in Mouse DcpS (blue) are shown in ball-and-stick configuration. A structural alignment, performed with the coordinates of DcpS using the FATCAT server,10 indicates significant structural similarity to the recently solved crystal structure of the human DcpS/7-methyl-guanosine-5′-triphosphate-5′-guanosine (mGpppG) complex (PDB code: 1st0).6 The Cα root-mean-square deviation (RMSD) values for individual alignments of the N-terminal (residues 39–130) and C-terminal domains (residues 148–337) are 1.23 Å and 0.82 Å, respectively. The higher RMSD between the N-terminal domains is due to conformational changes upon substrate binding in the human DcpS/mGpppG structure. Interestingly, a superposition of both structures as a whole requires the introduction of a pivot point at residue 141, indicative of a rigid movement of the N-terminal domain in the human DcpS/mGpppG complex. The overall structural alignment shows an RMSD of 0.89 Å for 287 equivalent Cα positions. In contrast to the symmetric dimer observed in the Mouse DcpS structure, the DcpS/mGpppG complex presents an asymmetric dimer with an open and closed active site6 due to this large movement of the N-terminal domain upon substrate binding. While the C-terminal domain remains symmetric, the N-terminal domain shifts by a ∼30° rotation and a ∼20° twist toward one of the C-terminal domains in the dimer [Fig. 2(B)]. This asymmetric movement of the N-terminal domain is facilitated by conformational changes in the linker region (residues 131–147). This rearrangement has a profound effect on the two active sites that are located in the interface between the N- and C-terminal domains, where it simultaneously opens one active site, while closing the other one. This rearrangement alters the Cα–Cα distances between Asp111 and Trp175 in the human DcpS/mGpppG complex to 6 Å in the closed side and 36 Å in the open side of the dimer. In contrast, Mouse DcpS features a 28-Å gap (Cα–Cα) between Asp110 and Trp174, indicative of a relaxed state, where both sides of the dimer are neither open nor closed. Interestingly, both active sites of human DcpS contain an mGpppG molecule, indicating that closure and catalysis in one active site may have an allosteric effect for substrate recognition in the other. The HIT triad in Mouse DcpS is represented by residues His274, His276, and His278, which form part of the active site. His276 is crucial for the catalytic activity and serves as the nucleophile attacking the cap γ-phosphate in mGpppG.11 The human DcpS/mGpppG complex contains a His277Asn mutation that inactivated the decapping activity in order to facilitate crystallization of the substrate complex. A superposition of Mouse DcpS with the human DcpS/mGpppG complex [Fig. 2(C and D)] shows that all of the active site residues are well conserved. Several active site residues in the linker region [Arg144(145), Gln145(146)] and in the N-terminal region [Arg53(54), Glu84(85), and Lys127(128)] undergo large conformational changes upon substrate binding. Different rotamers are found in the open active site for residues Glu184(185), Arg187(188), Asp204(205), Tyr216(217), His267(268), Ser271(272), Tyr272(273), and Arg321(322) [Fig. 2(C)] and in the closed active site for residues Glu184(185), Arg187(188), Asp204(205), Lys206(207), Tyr216(217), His267(268), Ser271(272), Tyr272(273), and Pro287(288) (human DcpS residue numbers are in parentheses) [Fig. 2(D)]. The structural similarity between Mouse and human DcpS [Fig. 2(C and D)] indicates that both mRNA decapping enzymes share the same mechanism and have very similar functional properties. According to the Fold and Function Assignment System (FFAS),12 the DcpS family has homologous sequences only in eukaryotic proteomes. Models for DcpS homologues can be accessed at http://www1.jcsg.org/cgi-bin/models/get_mor.pl?key=16740816. The 16740816 structure reported here represents a scavenger mRNA decapping enzyme (DcpS) from Mouse, whose structure has been determined by X-ray crystallography. The information reported here, in combination with the human DcpS structure (PDB code: 1st0)6 and further biochemical and biophysical studies, will yield valuable insights into the functional determinants of this protein and extend homology modeling of eukaryotic proteins. The scavenger mRNA decapping enzyme DcpS from Mouse (GI: 16740816; IMAGE: 4511427; SwissProt: Q9DAR7) was amplified by polymerase chain reaction (PCR) from a clone obtained from the IMAGE consortium using PfuTurbo (Stratagene) and primer pairs encoding the predicted 5′- and 3′-ends. The PCR product was cloned into plasmid pMH4, which encodes an expression and purification tag (MGSDKIHHHHHH) at the amino terminus of the full-length protein. The cloning junctions were confirmed by sequencing. Protein expression was performed in a modified Terrific Broth [24 g/L yeast extract, 12 g/L tryptone, 1% (v/v) glycerol, 50 mM NaCl, 50 mM 3-(N-morpholino)propanesulfonic acid (MOPS), pH 7.6] using the Escherichia coli strain GeneHogs® (Invitrogen). Lysozyme was added to the culture at the end of fermentation to a final concentration of 250 μg/mL. Bacteria were lysed by sonication after a freeze-thaw procedure in Lysis Buffer [50 mM Tris pH 7.9, 50 mM NaCl, 10 mM imidazole, 0.25 mM Tris(2-carboxyethyl)phosphine hydrochloride (TCEP)], and the cell debris was pelleted by centrifugation at 3400 × g for 60 min. The soluble fraction was applied to a nickel-resin (Amersham Biosciences) pre-equilibrated with Lysis Buffer. The nickel-resin was washed with Wash Buffer [50 mM potassium phosphate, pH 7.8, 40 mM imidazole, 300 mM NaCl, 10% (v/v) glycerol, 0.25 mM TCEP], and the protein was eluted with Elution Buffer [20 mM Tris, pH 7.9, 300 mM imidazole, 10% (v/v) glycerol, 0.25 mM TCEP]. Buffer exchange was performed to remove imidazole from the eluate, and the protein in Buffer Q [20 mM Tris, pH 7.9, 5% (v/v) glycerol, 0.25 mM TCEP] containing 50 mM NaCl was applied to a Resource Q column (Amersham Biosciences) pre-equilibrated with the same buffer. The protein was eluted using a linear gradient of 50–500 mM NaCl in Buffer Q. The appropriate fractions were pooled, further purified using a Superdex 200 size exclusion column (SEC; Amersham Biosciences) with elution in Crystallization Buffer [20 mM Tris, pH 7.9, 150 mM NaCl, 0.25 mM TCEP], and concentrated for crystallization assays to 16 mg/mL by centrifugal ultrafiltration (Millipore). The protein was crystallized using the nanodroplet vapor diffusion method13 with standard JCSG crystallization protocols.5 The crystallization reagent contained 20% polyethylene glycol (PEG)-3350 and 0.2 M potassium fluoride at pH 7.2. 15% ethylene glycol (final concentration) was included as a cryoprotectant. The crystals were indexed in the monoclinic space group P21 (Table I). Native diffraction data were collected at the Advanced Photon Source (APS, Chicago, IL) on beamline 31-ID (Table I). The data set was collected at 100 K using a MAR charge-coupled device (CCD) detector. Data were integrated and reduced using MOSFLM14 and then scaled with the program SCALA from the CCP4 suite.7 Data statistics are summarized in Table I. The structure was determined by molecular replacement using the program MOLREP from the CCP4 suite.7 A homology model based on the FFAS10 alignment between Mouse and human DcpS (PDB code: 1st0)6 was constructed with the modeling program WHAT IF15 and used as a search model. Structure refinement was performed using REFMAC57 and O.16 Refinement statistics are summarized in Table I. The final model includes a protein dimer (residues A38–337, B39–337), six 1,2-ethanediol, and 800 water molecules in the asymmetric unit. No electron density was observed for residues 1–37, 70–75, and 182–185 in chain A, residues 1–38 and 70–76 in chain B, or the expression or purification tag. The side-chain atoms of residues 104, 137, 144, 186, 328, 332, and 335 in chain A and residues 69, 91, 110, 144, and 332 in chain B were not visible in the electron density maps and were omitted from the model. Analysis of the stereochemical quality of the model was accomplished using the AutoDepInputTool (http://deposit.pdb.org/adit/), MolProbity,9 SFcheck 4.0,7 and WHAT IF 5.0.15 Protein quaternary structure analysis was performed using the PQS server (http://pqs.ebi.ac.uk/). Figure 1(B) was adapted from PDBsum (http://www.biochem.ucl.ac.uk/bsm/pdbsum/), and all others were prepared with PYMOL (DeLano Scientific). Atomic coordinates and experimental structure factors of Mouse DcpS have been deposited within the PDB and are accessible under the code 1vlr. Portions of this research were carried out at the Stanford Synchrotron Radiation Laboratory, a National user facility operated by Stanford University on behalf of the U.S. Department of Energy, Office of Basic Energy Sciences. The SSRL Structural Molecular Biology Program is supported by the Department of Energy, Office of Biological and Environmental Research, and by the National Institutes of Health (National Center for Research Resources, Biomedical Technology Program, and the National Institute of General Medical Sciences). Additional research was conducted at the Northeastern Collaborative Access Team beamlines of the Advanced Photon Source, supported by award RR-15301 from the National Center for Research Resources at the National Institute of Health. Use of the Advanced Photon Source is supported by the U.S. Department of Energy, Office of Basic Energy Sciences, under contract No. W-31-109-ENG-38.
Gye Won Han, Robert Schwarzenbacher, Rebecca Page, Lukasz Jaroszewski, Polat Abdubek, Eileen Ambing, Tanya Biorac, Jaume M. Canaves, Hsiu-Ju Chiu, Xiaoping Dai, Ashley M. Deacon, Michael DiDonato, Marc-André Elsliger, Adam Godzik, Carina Grittini, Slawomir K. Grzechnik, Joanna Hale, Eric Hampton, Justin Haugen, Michael Hornsby, Heath E. Klock, Eric Koesema, Andreas Kreusch, Peter Kuhn, Scott A. Lesley, Inna Levin, Daniel McMullan, Timothy M. McPhillips, Mitchell D. Miller, Andrew Morse, Kin Moy, Edward Nigoghossian, Jie Ouyang, Jessica Paulsen, Kevin Quijano, Ron Reyes, Eric Sims, Glen Spraggon, Raymond C. Stevens, Henry van den Bedem, Jeff Velasquez, Juli Vincent, Frank von Delft, Xianhong Wang, Bill West, Aprilfawn White, Guenter Wolf, Qingping Xu, Olga Zagnitko, Keith O. Hodgson, John Wooley, and Ian A. Wilson* The Joint Center for Structural Genomics Stanford Synchrotron Radiation Laboratory, Stanford University, Menlo Park, California The San Diego Supercomputer Center, La Jolla, California The Genomics Institute of the Novartis Research Foundation, San Diego, California The University of California, San Diego, La Jolla, California The Scripps Research Institute, La Jolla, California