Proteins: Structure, Function, and BioinformaticsVolume 79, Issue 4 p. 1329-1336 Structure Note Structural architecture of Galdieria sulphuraria DCN1L E. Sethe Burgie, E. Sethe Burgie Department of Genetics, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorCraig A. Bingman, Craig A. Bingman Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorShin-ichi Makino, Shin-ichi Makino Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorGary E. Wesenberg, Gary E. Wesenberg Department of Mathematics, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorXiaokang Pan, Xiaokang Pan Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorBrian G. Fox, Brian G. Fox Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorGeorge N. Phillips Jr., Corresponding Author George N. Phillips Jr. [email protected] Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Department of Biochemistry, University of Wisconsin-Madison, 433 Babcock Drive, Madison, WI 53706-1544===Search for more papers by this author E. Sethe Burgie, E. Sethe Burgie Department of Genetics, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorCraig A. Bingman, Craig A. Bingman Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorShin-ichi Makino, Shin-ichi Makino Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorGary E. Wesenberg, Gary E. Wesenberg Department of Mathematics, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorXiaokang Pan, Xiaokang Pan Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorBrian G. Fox, Brian G. Fox Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Search for more papers by this authorGeorge N. Phillips Jr., Corresponding Author George N. Phillips Jr. [email protected] Department of Biochemistry, Center for Eukaryotic Structural Genomics, University of Wisconsin-Madison, Madison Wisconsin 53706-1544Department of Biochemistry, University of Wisconsin-Madison, 433 Babcock Drive, Madison, WI 53706-1544===Search for more papers by this author First published: 11 November 2010 https://doi.org/10.1002/prot.22937Citations: 4Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Citing Literature Volume79, Issue4April 2011Pages 1329-1336 RelatedInformation
The identification of sequence-based protein domains and their boundaries is often a prelude to structure determination. An accurate prediction of disordered regions, secondary structures and low complexity segments of target protein sequences can improve the efficiency of selection in structural genomics and also aid in design of constructs for directed structural biology studies. At the Center for Eukaryotic Structural Genomics (CESG) we have developed DomainView, a web tool to visualize and analyze predicted protein domains, disordered regions, secondary structures and low complexity segments of target protein sequences for selection of experimental protein structure attempts. DomainView consists of a relational database and a web graphical-user interface. The database was developed based on MySQL, which stores data from target protein sequences and their domains, disordered regions, secondary structures and low complexity segments. The program of the web user interface is a Perl CGI script. When a user searches for a target protein sequence, the script displays the combinational information about the domains and other features of that target sequence graphically on a web page by querying the database. The graphical representation for each feature is linked to a web page showing more detailed annotation information or to a new window directly running the corresponding prediction program to show further information about that feature.
The region surrounding a protein, known as the surface of interaction or molecular surface, can provide valuable insight into its function. Unfortunately, due to the complexity of both their geometry and their surface fields, study of these surfaces can be slow and difficult and important features may be hard to identify. Here, we describe our GRaphical Abstracted Protein Explorer, or GRAPE, a web server that allows users to explore abstracted representations of proteins. These abstracted surfaces effectively reduce the level of detail of the surface of a macromolecule, using a specialized algorithm that removes small bumps and pockets, while preserving large-scale structural features. Scalar fields, such as electrostatic potential and hydropathy, are smoothed to further reduce visual complexity. This entirely new way of looking at proteins complements more traditional views of the molecular surface. GRAPE includes a thin 3D viewer that allows users to quickly flip back and forth between both views. Abstracted views provide a fast way to assess both a molecule's shape and its different surface field distributions. GRAPE is freely available at http://grape.uwbacter.org.
chipD is a web server that facilitates design of DNA oligonucleotide probes for high-density tiling arrays, which can be used in a number of genomic applications such as ChIP-chip or gene-expression profiling. The server implements a probe selection algorithm that takes as an input, in addition to the target sequences, a set of parameters that allow probe design to be tailored to specific applications, protocols or the array manufacturer's requirements. The algorithm optimizes probes to meet three objectives: (i) probes should be specific; (ii) probes should have similar thermodynamic properties; and (iii) the target sequence coverage should be homogeneous and avoid significant gaps. The output provides in a text format, the list of probe sequences with their genomic locations, targeted strands and hybridization characteristics. chipD has been used successfully to design tiling arrays for bacteria and yeast. chipD is available at http://chipd.uwbacter.org/.
The gene LOC791917 Danio rerio (zebrafish) encodes a protein annotated in the UniProt knowledgebase1 as the “middle domain of eukaryotic initiation factor 4G domain containing protein b” (MIF4Gdb). Its molecular weight is 25.8 kDa, and it comprises 222 amino acid residues. BLAST searches revealed homologues of D. rerio MIF4Gdb in many eukaryotes including humans.2 The homologues and MIF4Gdb were identified as members of the Pfam family, MIF4G (PF02854), which is named after the middle domain of eukaryotic initiation factor 4G (eIF4G).3-5 eIF4G is a component of eukaryotic translational initiation complex, and contains binding sites for other initiation factors, suggesting its critical role in translational initiation.6 The MIF4G domain also occurs in several other proteins involved in RNA metabolism, including the Nonsense-mediated mRNA decay 2 protein (NMD2/UPF2), and the nuclear cap-binding protein 80-kD subunit (CBP80).5 Sequence and structure analysis of the MIF4G domains in many proteins indicates that the domain assumes all helical fold and has tandem repeated motifs.5,7 The zebrafish protein described here has homology to domains of other proteins variously referred to as NIC-containing proteins (NMD2, eIF4G, CBP80). The biological function of D. rerio MIF4Gdb has not yet been experimentally characterized, and the annotation is based on amino acid sequence comparison. D. rerio MIF4Gdb did not share more than 25% sequence identity with any protein for which the three-dimensional structure is known and was selected as a target for structure determination by the Center for Eukaryotic Structural Genomics (CESG). Here, we report the crystal structure of D. rerio MIF4Gdb (UniGene code Dr.79360, UniProt code {"type":"entrez-protein","attrs":{"text":"Q5EAQ1","term_id":"82178873"}}Q5EAQ1, CESG target number GO.79294).
Os09g0567400 codes for a hypothetical protein from Oryza sativa that is annotated as the "Histidine-containing phosphotransfer (Hpt) protein". Hpt domain is a protein module with a histidine residue mediating phosphotransfer reaction in the histidine-aspartate phosphorelay system. We report here the crystal structure and analysis of Os09g0567400.
ABSTRACT DesA3 (Rv3229c) from Mycobacterium tuberculosis is a membrane-bound stearoyl coenzyme A Δ 9 desaturase that reacts with the oxidoreductase Rv3230c to produce oleic acid. This work provides evidence for a mechanism used by mycobacteria to regulate this essential enzyme activity. DesA3 expressed as a fusion with either a C-terminal His 6 or c-myc tag had consistently higher activity and stability than native DesA3 having the native C-terminal sequence of LAA, which apparently serves as a binding determinant for a mycobacterial protease/degradation system directed at DesA3. Fusion of only the last 12 residues of native DesA3 to the C terminus of green fluorescent protein (GFP) was sufficient to make GFP unstable. Furthermore, the comparable C-terminal sequence from the Mycobacterium smegmatis DesA3 homolog Msmeg_1886 also conferred instability to the GFP fusion. Systematic examination revealed that residues with charged side chains, large nonpolar side chains, or no side chain at the last two positions were most important for stabilizing the construct, while lesser effects were observed at the third-from-last position. Using these rules, a combinational substitution of the last three residues of DesA3 showed that either DKD or LEA gave the best enhancement of stability for the modified GFP in M. smegmatis . Moreover, upon mutagenesis of LAA at the C terminus in native DesA3 to either of these tripeptides, the modified enzyme had enhanced catalytic activity and stability. Since many proteases are conserved within bacterial families, it is reasonable that M. tuberculosis will use a similar C-terminal degradation system to posttranslationally regulate the activity of DesA3 and other proteins. Application of these rules to the M. tuberculosis genome revealed that ∼10% the proteins encoded by essential genes may be susceptible to C-terminal proteolysis. Among these, an annotation is known for less than half, underscoring a general lack of understanding of proteins that have only temporal existence in a cell.
Soluble N‐ethylmaleimide‐sensitive factor attachment protein gamma (γ‐SNAP) is a member of an eukaryotic protein family involved in intracellular membrane trafficking. The X‐ray structure of Brachydanio rerio γ‐SNAP was determined to 2.6 Å and revealed an all‐helical protein comprised of an extended twisted‐sheet of helical hairpins with a helical‐bundle domain on its carboxy‐terminal end. Structural and conformational differences between multiple observed γ‐SNAP molecules and Sec17, a SNAP family protein from yeast, are analyzed. Conformational variation in γ‐SNAP molecules is matched with great precision by the two lowest frequency normal modes of the structure. Comparison of the lowest‐frequency modes from γ‐SNAP and Sec17 indicated that the structures share preferred directions of flexibility, corresponding to bending and twisting of the twisted sheet motif. We discuss possible consequences related to the flexibility of the SNAP proteins for the mechanism of the 20S complex disassembly during the SNAP receptors recycling. Proteins 2008. © 2007 Wiley‐Liss, Inc.
X-ray crystallography typically uses a single set of coordinates and B factors to describe macromolecular conformations. Refinement of multiple copies of the entire structure has been previously used in specific cases as an alternative means of representing structural flexibility. Here, we systematically validate this method by using simulated diffraction data, and we find that ensemble refinement produces better representations of the distributions of atomic positions in the simulated structures than single-conformer refinements. Comparison of principal components calculated from the refined ensembles and simulations shows that concerted motions are captured locally, but that correlations dissipate over long distances. Ensemble refinement is also used on 50 experimental structures of varying resolution and leads to decreases in R-free values, implying that improvements in the representation of flexibility observed for the simulated structures may apply to real structures. These gains are essentially independent of resolution or data-to-parameter ratio, suggesting that even structures at moderate resolution can benefit from ensemble refinement.
Determination of a protein structure requires a series of decisions and processes, starting with target selection, through cloning, expression, purification, and finally structure determination. Structural genomics projects may distribute these steps among several different groups of researchers. Although this division may achieve a lower cost per solved structure, it creates a unique set of challenges for integrating and passing information on the progress of a given target across several functional divisions. Laboratory information management systems (LIMS) are essential for gathering this information, but may not display the progress of a given target in an intuitive way. In addition, structural genomics projects funded by the Protein Structure Initiative (PSI) are obliged to disseminate data regularly to the TargetDB and PepcDB data repositories, and this requires the creation of specialized views of the data. We report here how the flow of a target through a structural genomics pipeline and reports to TargetDB and PepcDB can be abstracted as directed acyclic graphs or trees. To implement this kind of display, we created software that tracks the flow of activity leading toward protein structure determination and prepares XML reports as input to TargetDB and PepcDB. The target tracing software consists of a set of Perl CGI scripts that integrate with the Graphviz visualization system to provide a graphical, user-friendly Web interface. The database reporting software, also coded in Perl, transfers large-scale genomics data from our LIMS into a PepcDB reportable XML file. This software package has facilitated inter-group communication, improved the quality and accuracy of information in our LIMS, and increased the efficiency and accuracy of our reports to PepcDB.
Aspartoacylase catalyzes hydrolysis of N-acetyl-l-aspartate to aspartate and acetate in the vertebrate brain. Deficiency in this activity leads to spongiform degeneration of the white matter of the brain and is the established cause of Canavan disease, a fatal progressive leukodystrophy affecting young children. We present crystal structures of recombinant human and rat aspartoacylase refined to 2.8- and 1.8-Å resolution, respectively. The structures revealed that the N-terminal domain of aspartoacylase adopts a protein fold similar to that of zinc-dependent hydrolases related to carboxypeptidases A. The catalytic site of aspartoacylase shows close structural similarity to those of carboxypeptidases despite only 10–13% sequence identity between these proteins. About 100 C-terminal residues of aspartoacylase form a globular domain with a two-stranded β-sheet linker that wraps around the N-terminal domain. The long channel leading to the active site is formed by the interface of the N- and C-terminal domains. The C-terminal domain is positioned in a way that prevents productive binding of polypetides in the active site. The structures revealed that residues 158–164 may undergo a conformational change that results in opening and partial closing of the channel entrance. We hypothesize that the catalytic mechanism of aspartoacylase is closely analogous to that of carboxypeptidases. We identify residues involved in zinc coordination, and propose which residues may be involved in substrate binding and catalysis. The structures also provide a structural framework necessary for understanding the deleterious effects of many missense mutations of human aspartoacylase.
The structure of the UDP-glucose pyrophosphorylase encoded by Arabidopsis thaliana gene At3g03250 has been solved to a nominal resolution of 1.86 angstrom. In addition, the structure has been solved in the presence of the substrates/products UTP and UDP-glucose to nominal resolutions of 1.64 angstrom and 1.85 angstrom. The three structures revealed a catalytic domain similar to that of other nucleoticlyl-glucose pyrophosphorylases with a carboxyterminal beta-helix domain in a unique orientation. Conformational changes are observed between the native and substrate-bound complexes. The nucleotide-binding loop and the carboxy-terminal domain, including the suspected catalytically important Lys360, move in and out of the active site in a concerted fashion. TLS refinement was employed initially to model conformational heterogeneity in the UDP-glucose complex followed by the use of multiconformer refinement for the entire molecule. Normal mode analysis generated atomic displacement predictions in good agreement in magnitude and direction with the observed conformational changes and anisotropic displacement parameters generated by TLS refinement. The structures and the observed dynamic changes provide insight into the ordered mechanism of this enzyme and previously described oligomerization effects on catalytic activity. (c) 2006 Elsevier Ltd. All rights reserved.
HsSSAT1 is an important enzyme that converts polyamines such as spermidine or spermine into acetylpolyamines. Because polyamines are major regulators of cell growth and differentiation,1-3 HsSSAT1 has been considered as a target for cancer treatment. Gene locus BC011751 from Homo sapiens has been annotated as HsSSAT24 because of its high sequence identity (46%) and similarity (61%) to HsSSAT1. Recent biochemical studies showed that HsSSAT2 does not transfer acetyl group of AcCoA to polyamines, such as spermidine or spermine, but rather thialysine [2-amino-3-(2-aminoethylsulfanyl)propanoic acid] which is a much better substrate of this enzyme.5 However, the physiological role of HsSSAT2 is still unclear. Thialysine is a structural analog of L-lysine, and has its role as an antimetabolite by competing with L-lysine for the incorporation into polypeptides, which makes proteins inactive.6, 7 Therefore, high thialysine to lysine ratios block the growth of prokaryotic and eukaryotic cells.8-11 In addition to the antimetabolite function, metabolites formed from thialysine have been characterized and identified in mammalian tissues including brain.12, 13 Thialysine can be further converted to the cyclic ketimine, AECK,12 TMA,14 and AECK-DD.15, 16 The functions of AECK, TMA, and AECK-DD in the cell are not clear. However, it has been proposed that AECK and TMA serve neurochemical roles, and AECK-DD has a role in the modulation of oxidative processes in vivo.17 Herein, we report the three-dimensional structure of HsSSAT2 in complex with AcCoA at a resolution of 1.8 Å. SSAT, spermidine/spermine N1-acetyltransferase; AcCoA, acetyl coenzyme A; AECK, S-(aminoethyl)-L-cysteine ketimine; TMA, 4-thiomorpholine-3,5-dicarboxylic acid; AECK-DD, AECK-decarboxylated dimer; SeMet, selenomethionine; TCEP, tris(2-carboxylethyl) phosphine; Bis-Tris, bis(2-hydroxyethyl)imino-tris(hydroxymethyl)methane; MEPEG, methyl ether polyethylene glycol; MOPS, 3-(N-morpholino) propanesulfonic acid; PDB, Protein Data Bank; MAD, multi-wavelength anomalous diffraction; GNAT, GCN5-related N-acetyltransferases; ScHpa2, histone acetyltransferase from Saccharomyces cerevisiae; SeAac(6′)-Iy, aminoglycoside N-acetyltransferase from Salmonella enteritidis; ScGNA1, glucosamine-6-phosphate N-acetyltransferase 1 from Saccharomyces cerevisiae; APS, Advanced Photon Source; ANL, Argonne National Laboratory; DALI, Distance mAtrix aLIgnment; VAST, Vector Alignment Search Tool. The gene encoding the HsSSAT2 was cloned and the SeMet-labeled proteins were expressed and purified following the standard CESG pipeline protocol for cloning,18 protein expression,19 protein purification,20 and overall information management.21 Crystals of HsSSAT2 were grown by the hanging-drop vapor diffusion method from 10 mg/mL protein solution in buffer (50 mM NaCl, 3 mM NaN3, 0.3 mM TCEP, 5 mM Bis-Tris pH 6.0) mixed with an equal amount of reservoir solution containing 16.4% MEPEG 5,000, 100 mM MOPS pH 7.0, and 150 mM potassium glutamate at 277 K. Crystals grew approximately as tetragonal bipyramids with dimensions of approximately 100 × 100 × 50 μm. The selenomethionyl crystals of HsSSAT2 belong to space group P212121, with unit-cell parameters a = 54.6, b = 83.6, c = 87.5 Å. One crystal was transferred to the same reservoir solution where the crystal was grown, followed by stepwise additions of ethylene glycol up to 25%. Subsequently, the crystal was transferred to the final cryoprotectant solution (16.4% MEPEG 5 K, 100 mM MOPS pH 7.0, 150 mM potassium glutamate, and 25% ethylene glycol) for 2 s, and it was frozen by direct immersion in liquid nitrogen at 100 K. X-ray diffraction data were collected at synchrotron beam line 22-ID at the APS of the ANL. The diffraction images were integrated and scaled using HKL2000.22 The selenium substructure of SeMet-labeled HsSSAT2 crystal was determined using Hyss23 and SHELXD.24 The protein structure was phased and the initial phase information was further improved by electron-density modification using two-wavelength MAD data in autoSHARP.25 The automatic tracing procedure of ARP/wARP26 in autoSHARP25 produced an initial model with approximately 86% of residues placed, of which 79% had side-chains assigned. The structure was completed using alternate cycles of manual model building in Coot27 and Xfit,28 and refinement in REFMAC5.29 All steps were monitored using an Rfree value based on 5.0% of the independent reflections. The stereochemical quality of the final model was assessed using PROCHECK30 and MolProbity.31 The structure of HsSSAT2 in complex with AcCoA has been determined to a resolution of 1.8 Å. Data collection, phasing, refinement, and model statistics are summarized in Table I. The final model includes two HsSSAT2 monomers (residues 3–30, 35–60, and 68–169 for monomer A; residues 2–59 and 70–170 for monomer B), one AcCoA molecule, and 305 water molecules. The coordinates and structure factor files have been deposited in the PDB with accession number 2BEI. The three-dimensional structure of HsSSAT2 is a mixed α/β fold containing eight α-helices and seven β-strands [Fig. 1(a)]. The catalytic domain of HsSSAT2 has the common fold of GNAT family which is known as a central, mixed β-sheet that is mainly built up of antiparallel strands34 with one exception between parallel β4 and β5. a: Domain-swapped dimeric structure. The secondary structure elements are numbered in the order of appearance in the primary structure. One monomer with AcCoA is represented in cyan and the other monomer is in green. AcCoA is colored by atom types (carbon: gray; oxygen: red; nitrogen: blue; phosphorous: orange; sulfur: yellow). b: ScHpa2 in complex with AcCoA, the best structural homolog in the PDB as determined by DALI and VAST. ScHpa2 is colored in magenta and HsSSAT2 shows one monomer in cyan as in part (a). c: The electron densities of AcCoA in the 2Fo–Fc map at 1.5 σ level. AcCoA is colored the same as in part (a). Figures were prepared using the PyMol program (http://pymol.sourceforge.net/). We found significant extra electron density from the map calculated from the initial protein model using autoSHARP. The extra electron density looked like an AcCoA in a large hydrophobic cleft located at the site where the two parallel strands, β4 and β5, diverge because of a β-bulge in strand β4 [Fig. 1(c)]. Leu91 and Glu92 in this β-bulge are well conserved among GNAT families including Homo sapiens SSAT1 and other SSAT2 sequences [Fig. 2(b)]. HsSSAT2 acquired the AcCoA from Escherichia coli cells used for overexpression of HsSSAT2. Only one of two molecules in asymmetric unit contained the AcCoA, and this feature is unique among all acetyltransferase structures that contain coenzyme A or AcCoA in the PDB. This naturally acquired AcCoA in only one active site of the dimer suggests that each monomer of HsSSAT2 dimer may participate in "half-the-sites" reactivity, but we have no experimental evidence to support this. The AcCoA interactions with HsSSAT2 and the sequence alignment of SSATs. a: The interactions between HsSSAT2 and AcCoA. AcCoA is colored the same as in part (a) and hydrogen bonds are marked by green arrows. The red dot represents a water. b: The overall sequence alignment results of HsSSAT2, HsSSAT1, and other SSAT2s. Conserved, identical, similar, and different residues are colored in red, blue, cyan, and black, respectively. AcCoA interacting sites are identified by filled circles, and the proposed key thialysine contact residue (Ala128) is identified by a filled triangle. Residues forming the conserved β-bulge are marked with a green triangle, and four GNAT sequence motifs are boxed and labeled. The abbreviations for SSAT species are used: Hs, Homo sapiens (human); Bt, Bos taurus (domestic cow); Cf, Canis familiaris (domestic dog); Ss, Sus scrofa (wild pig); Mm, Mus musculus (house mouse); Rn, Rattus norvegicus (Norway rat). HsSSAT2 shares the {R/Q}-X-X-G-X-G sequence motif (R101-X-X-G104-X-G106 in HsSSAT2) for AcCoA recognition and binding35 with other GNAT superfamily, and AcCoA bends at the pyrophosphate moiety and at the pantetheine moiety. The pyrophosphate moiety of AcCoA mainly interacts with the GNAT sequence motif, and the pantetheine moiety of AcCoA interacts with β4 (I94 and V96) and α6 (N133 and Y140) through hydrogen bonds. The interactions between HsSSAT2 and AcCoA are described in Figure 2(a). Structural homology searches using the DALI server36 and VAST (http://www.ncbi.nlm.nih.gov/Structure/VAST/vast.shtml) frequently produced the structure of the ScHpa2 in complex with AcCoA (PDB ID: 1QSM)37 as the closest structural homolog of HsSSAT2, with Z = 19.5, RMSD = 2.1 Å over 143 aligned residues and 22% sequence identity, and Figure 1(b) shows the structure alignment between HsSSAT2 and ScHpa2. The SeAac(6′)-Iy in complex with coenzyme A and ribostamycin (PDB ID: 1S3Z)38 is the second best structural homolog of HsSSAT2 from DALI, with Z = 18.0, RMSD = 2.3 Å over 141 aligned residues, and 17% sequence identity. After our HsSSAT2 structure was released, HsSSAT1 structures were deposited to the PDB, which have high structural homology to HsSSAT2 (Z = 21.2, RMSD = 1.9 Å over 154 aligned residues, and 47% sequence identity for entry 2B5G, for example). From the structure alignment with the SeAac(6′)-Iy which contains the substrate ribostamycin, we suggest that the backbone carbonyl of Ala128 has an important role in recognizing the nucleophilic amine group of thialysine by a hydrogen bond that enhances the nucleophilic character of the amine. Interestingly, the intermolecular interaction between two monomers of HsSSAT2 shows a pattern of domain swapping39; β7 (residues 154–160) of one molecule inserts between β5 and β6 of the other molecule, and vice versa [Fig. 1(a)]. This domain swapping β-strand exchange between subunits in the dimer is not unusual among GNATs, and has been observed in PDB structures of ScHpa2, SeAac(6′)-Iy, and ScGNA1 (PDB ID: 1I12). The interface of HsSSAT2 dimer is extensive, burying 3,409 Å2 or 30.6% of the total monomeric surface. The dimeric interface is quite nonpolar, with 66.5% of the buried interface area from carbon atoms. The number of residues that are buried by the other molecule by domain swapping comprises 15 amino acids from the N-terminal end of β7 to the C-terminal end of the molecule (residues 154–169). Based on the domain swapping and the surface area calculation results, we conclude that HsSSAT2 is probably a dimer in vivo. The authors acknowledge the Southeast Regional Collaborative Access Team (SER-CAT) for use of beamline 22-ID and 22-BM at the Advanced Photon Source (APS), Argonne National Laboratory (ANL). Supporting institutions may be found at http://www.ser-cat.org/members.html. Use of the APS was supported by the US Department of Energy, Office of Science, Office of Basic Energy Sciences. The authors also acknowledge all CESG members, especially Eduard Bitto, Euiyoung Bae, Dave Aceti, Craig S. Newman, Zhaohui Sun, Russell L. Wrobel, Eric Steffan, Zachary Eggers, Megan Riters, Ronnie O. Frederick, John Kunert, Hassan Sreenath, Brendan T. Burns, Kory D. Seder, Holalkere V. Geetha, Frank C. Vojtik, Won Bae Jeon, Jason M. Ellefson, Andrew C. Olson, Janet E. McCombs, Janelle T. Warick, Bryan Ramirez, Zsolt Zolnai, Peter T. Lee, Mike Runnels, John Cao, Jianhua Zhang, John G. Primm, Donna M. Troestler, Michael R. Sussman, Brian G. Fox, and John L. Markley. A publication on the spermidine/spermine acetyltransferases has recently appeared.40
Eukaryotic pyrimidine 5'-nucleotidase type 1 (P5N-1) catalyzes dephosphorylation of pyrimidine 5'-mononucleotides. Deficiency of P5N-1 activity in red blood cells results in nonspherocytic hemolytic anemia. The enzyme deficiency is either familial or can be acquired through lead poisoning. We present the crystal structure of mouse P5N-1 refined to 2.35 angstrom resolution. The mouse P5N-1 has a 92% sequence identity to its human counterpart. The structure revealed that P5N-1 adopts a fold similar to enzymes of the haloacid dehydrogenase superfamily. The active site of this enzyme is structurally highly similar to those of phosphoserine phosphatases. We propose a catalytic mechanism for P5N-1 that is also similar to that of phosphoserine phosphatases and provide experimental evidence for the mechanism in the form of structures of several reaction cycle states, including: 1) P5N-1 with bound Mg(II) at 2.25 angstrom, 2) phosphoenzyme intermediate analog at 2.30 angstrom, 3) product-transition complex analog at 2.35 A, and 4) product complex at 2.1 angstrom resolution with phosphate bound in the active site. Furthermore the structure of Pb(II)-inhibited P5N-1 (at 2.35 angstrom) revealed that Pb(II) binds within the active site in a way that compromises function of the cationic cavity, which is required for the recognition and binding of the phosphate group of nucleotides.
The structure of the Rieske-type ferredoxin (T4moC) from toluene 4-monooxygenase was determined by X-ray crystallography in the [2Fe-2S](2+) state at a resolution of 1.48 A using single-wavelength anomalous dispersion phasing with the [2Fe-2S] center. The structure consists of ten beta-strands arranged into the three antiparallel beta-sheet topology observed in all Rieske proteins. Trp69 of T4moC is adjacent to the [2Fe-2S] centre, which displaces a loop containing the conserved Pro81 by approximately 8 A away from the [2Fe-2S] cluster compared with the Pro loop in the closest structural and functional homolog, the Rieske-type ferredoxin BphF from biphenyl dioxygenase. In addition, T4moC contains five hydrogen bonds to the [2Fe-2S] cluster compared with three hydrogen bonds in BphF. Moreover, the electrostatic surface of T4moC is distinct from that of BphF. These structural differences are identified as possible contributors to the evolutionary specialization of soluble Rieske-type ferredoxins between the diiron monooxygenases and cis-dihydrodiol-forming dioxygenases.
We describe X-ray crystal and NMR solution structures of the protein coded for by Arabidopsis thaliana gene At1g77540.1 (At1g77540). The crystal structure was determined to 1.15 A with an R factor of 14.9% (Rfree = 17.0%) by multiple-wavelength anomalous diffraction using sodium bromide derivatized crystals. The ensemble of NMR conformers was determined with protein samples labeled with 15N and 13C + 15N. The X-ray structure and NMR ensemble were closely similar with rmsd 1.4 A for residues 8-93. At1g77540 was found to adopt a fold similar to that of GCN5-related N-acetyltransferases. Enzymatic activity assays established that At1g77540 possesses weak acetyltransferase activity against histones H3 and H4. Chemical shift perturbations observed in 15N-HSQC spectra upon the addition of CoA indicated that the cofactor binds and identified its binding site. The molecular details of this interaction were further elucidated by solving the X-ray structure of the At1g77540-CoA complex. This work establishes that the domain family COG2388 represents a novel class of acetyltransferase and provides insight into possible mechanistic roles of the conserved Cys76 and His41 residues of this family.
The Center for Eukaryotic Structural Genomics (CESG) focuses on technology and methodology development for high-throughput X-ray or NMR structure determination of proteins from eukaryotic organisms.1 The goals of this project also include the identification of new or unique protein folds and characterization of proteins of unknown structure or function. Through a process of selecting targets that have no close amino acid sequence relationship to those in Protein Data Bank (PDB),2 CESG selected two open reading frames, At5g11950 and At2g37210, from Arabidopsis thaliana for structural characterization. These two genes encode highly conserved hypothetical proteins with molecular weights of 23.8 and 23.6 kDa, respectively. The biological functions of the At5g11950 and At2g37210 genes in A. thaliana are not yet established. Based on sequence similarities, the protein products of At5g11950 and At2g37210 are annotated as lysine decarboxylase (LDC)-like proteins; however, no indication of the basis for this annotation can be found. In A. thaliana, at least 11 hypothetical proteins are annotated as LDC-like proteins by genome analysis.3 No biochemical evidence supporting this annotation is available. Here, we report the X-ray crystal structures of the proteins from A. thaliana gene loci At5g11950 and At2g37210 and describe the structural context of the characteristic motif of this protein family. The genes were cloned4 and proteins were expressed5 and purified6 by standard CESG protocols. Crystals of Se–Met-labeled At5g11950 were grown by the hanging drop method, from a 10 mg/ml protein solution in Buffer A (5 mM BisTris, 50 mM NaCl, 3.1 mM NaN3, 0.3 mM TCEP, pH 6.0) mixed with an equal volume of well solution containing 13% (w/v) MePEG 2000, 280 mM KNO3, 100 mM MOPS (pH 7.0 at 293 K). Crystals were cryoprotected by placing them serially in well solutions supplemented with increasing concentrations of ethylene glycol, up to a final concentration of 25% (v/v) ethylene glycol. Single-wavelength diffraction data were collected from Se–Met-labeled At5g11950 using an APS 1 detector on beamline 19-BM SBC-CAT at the Advanced Photon Source, Argonne National Laboratory. The data were integrated and scaled using the HKL2000 suite.7 Localization of the Se positions, phasing, and phase improvement were performed with SOLVE8 and RESOLVE9 programs. The initial model was built using the automatic tracing procedure as implemented in ARP/wARP10 and refined to 2.15 Å using Refmac5.11 Crystals of Se–Met-labeled At2g37210 were grown by the hanging drop method, from a 10 mg/ml solution in Buffer A (see above) mixed with an equal volume of well solution containing 22% (w/v) MePEG 2000, 84 mM MgSO4, 100 mM BisTris (pH 6.5 at 296 K). Crystals were cryoprotected by placing them serially in well solutions supplemented with increasing concentrations of ethylene glycol, up to a final concentration of 20% (v/v) ethylene glycol. Single-wavelength diffraction data were collected using a MAR 225 detector on beamline 22-BM SER-CAT at the Advanced Photon Source, Argonne National Laboratory. The data were processed using the HKL2000 suite.7 The structure was solved by molecular replacement using MOLREP12 and the structure of At5g11950 as the phasing model. The structure of At2g37210 was refined to 1.95 Å using Refmac5.11 Table I summarizes data collection, phasing, refinement, and model statistics. Coordinates for the crystal structures and diffraction data have been deposited in the PDB under the accession codes 1YDH and 2A33 for At5g11950 and At2g37210, respectively. The monomeric structure of At5g11950 shows an α/β protein fold comprising eight α-helices and seven β-strands (β1α1β2α2β3α3β4α4β5α5β6α6α 7β7α8) [Fig. 1(A)]. The central feature of this domain is the β-sheet formed by seven parallel β-strands surrounded by the eight α-helices. On one side of the central β-sheet are helices α1, α2, α3, and α8, and on the other side are helices α4, α5, α6, and α7. The tertiary structure of the At2g37210 monomer is almost identical to that of the At5g11950 monomer, with a 0.9 Å root mean square deviation (rmsd) and 71% identity over 167 aligned Cα positions. The only major difference between the structures is an absence of the short α3 helix in the At2g37210 structure. The loop that spans residues 82–89 was highly disordered in the electron density map, and thus, was not built into the final At2g37210 structure. (A) Ribbon diagram of the At5g11950 dimer with each monomer colored separately. (B) An overlay of Cα traces of At5g11950 (red) and At2g37210 (cyan). The residues corresponding to helix α3 and neighboring loops adopt a defined structure within At5g11950 (red). The same region is disordered in the At2g37210 structure (cyan). A fully conserved residue, Arg98, and the PGGxGTxxE motif, which is highly conserved among the LDC like proteins,15 were mapped onto a Cα trace of the At5g11950 structure. (C) Surface view of the two monomers (yellow and cyan) in the At2g37210 dimer. The dark blue area represents the conserved motif PGGxGTxxE and residue Arg98, located at the bottom of a cavity. The red line highlights the residues corresponding to helix α3 of At5g11950. These residues are positioned near the entrance to the cavity and are disordered in the crystal structure of At2g37210, suggesting that they may be involved in controlling access of a substrate to the putative active site. The figures were generated using PyMol.21 Two At5g11950 subunits associate to form a tight dimer in the crystalline asymmetric unit [Fig. 1(A)]. The dimer buries 1806 Å2 surface area of each monomer and includes 12 hydrogen bonds. The interface between the two monomers is mostly hydrophobic (68%) and stabilized by contact of helices α5 and α6. The CASTp server13 was used to search for pockets or cavities on the surface of At5g11950. It found a cleft with a surface area of 623 Å2 and a volume of 322 Å3. The bottom of the cleft is defined by residues from strands β1, β2, and β3 and a loop between β3 and α3, and the wall of the cleft is formed mostly by residues from helices α4 and α5. The coordinates of At5g11950 and At2g37210 were analyzed by the DALI server14 to find structurally similar proteins in the PDB. DALI returned 8 and 19 structural homologs for At5g11590 and At2g37210, respectively, all with a Z score over 7.0. The best match to both At5g11590 and At2g37210 was TT1887 (PDB 1WEH), with Z scores of 19.1 and 20.4, rmsd's of 2.4 and 2.1 Å and 23 and 25% identity, respectively, for the two proteins whose structures we determined. TT1887 is a hypothetical protein from Thermus thermophilus Hb8, which is also annotated as an LDC-like protein.15 However, its true biological activity is still unknown. The second and third top matches were nucleoside 2-deoxyribosyltransferase (PDB 1F8X, Z score 7.5 and 7.8, rmsd 3.9 and 3.4 Å, 12 and 15% identity)16 and UDP–N-acetylglucosamine 2-epimerase (PDB 1F6D, Z score 7.2 and 8.1, rmsd 3.0 and 2.9 Å, 12 and 12% identity),17 respectively. Although they share an apparent structural similarity, both At5g11950 and At2g37210 may have a different function because they do not contain the active site residues of those two enzymes. Interestingly, At5g11950 and At2g37210 are also structurally similar to the negative transcriptional regulator NmrA (PDB 1K6I, Z score 6.7 and 7.6, rmsd 3.5 and 3.3 Å, 8 and 7% identity). NmrA is involved in the signaling pathway of nitrogen metabolite repression in various fungi.18 A VAST search19 found 18 and 14 structural neighbors of At5g11950 and At2g37210, respectively, all with a VAST score over 13.0. Among these structures, four top neighbors with VAST scores greater than 17, rmsd values less than 2.0 Å and a sequence identity of over 23% were annotated as putative LDCs: YvdD (PDB 1T35) from Bacillus subtilis, Tm1055 (PDB 1RCU) from Thermotoga maritima, TT1465 (PDB 1WEK) and TT1887 (PDB 1WEH) from T. thermophilus Hb8.15 All four structures display an α/β protein fold and contain 6, 7, or 8 α-helices flanking a central β sheet in a similar location to the At5g11950 and At2g37210 structures. An FFAS03 search20 confirmed that the four putative LDCs identified by the VAST server share distant sequence homology to At5g11950 and At2g37210, with FFAS03 scores below −49.8 and sequence identity just over 21%. Based upon the crystal structures and sequence homology searches, the protein fold of At5g11950 and At2g37210 was classified as part of the LDC family, pfam03641; however, at present there is no biochemical evidence to support this annotation. Structural analysis with At5g11950 revealed that the consensus motif PGGxGTxxE15 is within helix α5 and constitutes part of a cleft [Fig. 1(B)]. In addition, conserved residues Arg98, Thr118, and Glu121 are positioned at the bottom of the cleft, which is created by the β-sheet and helices α4 and α5 in each monomer [Fig. 1(C)]. With these findings, we speculate that the invariant residues and consensus motif are functionally important for biological activity, perhaps forming part of a catalytic site. Data were collected at Southeast Regional Collaborative Access Team (SER-CAT) 22-BM beamline at the Advanced Photon Source, Argonne National Laboratory. Supporting institutions may be found at www.ser-cat.org/members.html. Use of the Argonne National Laboratory Structural Biology Center beamlines at the Advanced Photon Source, was supported by the U.S. Department of Energy, Office of Energy Research, under Contract No. W-31-109-ENG-38. Special thanks goes to all members of the CESG.
The gene product of At3g22680 from Arabidopsis thaliana codes for a protein of unknown function. The crystal structure of the At3g22680 gene product was determined by multiple-wavelength anomalous diffraction and refined to an R factor of 16.0% (Rfree = 18.4%) at 1.60 A resolution. The refined structure shows one monomer in the asymmetric unit, with one molecule of the non-denaturing detergent CHAPS {3-[(3-cholamidopropyl)dimethylammonio]-1-propane sulfonate} tightly bound. Protein At3g22680 shows no structural homology to any other known proteins and represents a new fold in protein conformation space.