The Protein Data Bank (PDB) is the single, freely available, global archive of structural data for biological macromolecules. It is maintained by the wwPDB consortium consisting of the Research Collaboratory for Structural Bioinformatics (RCSB PDB), the Protein Data Bank in Europe (PDBe), the PDB Japan (PDBj) and the BioMagResBank (BMRB). This chapter describes the organization of the wwPDB, the systems in place for data deposition, annotation and distribution, and a summary of the services provided by the wwPDB member sites. Keywords: Protein Data Bank
The crystal structure of the signal transduction protein TRAP is reported at 1.85 Å resolution. The structure of TRAP consists of a central eight-stranded β-barrel flanked asymmetrically by helices and is monomeric both in solution and in the crystal structure. A formate ion was found bound to TRAP identically in all four molecules in the asymmetric unit.
The EUROCarbDB project is a design study for a technical framework, which provides sophisticated, freely accessible, open-source informatics tools and databases to support glycobiology and glycomic research. EUROCarbDB is a relational database containing glycan structures, their biological context and, when available, primary and interpreted analytical data from high-performance liquid chromatography, mass spectrometry and nuclear magnetic resonance experiments. Database content can be accessed via a web-based user interface. The database is complemented by a suite of glycoinformatics tools, specifically designed to assist the elucidation and submission of glycan structure and experimental data when used in conjunction with contemporary carbohydrate research workflows. All software tools and source code are licensed under the terms of the Lesser General Public License, and publicly contributed structures and data are freely accessible. The public test version of the web interface to the EUROCarbDB can be found at http://www.ebi.ac.uk/eurocarb.
The techniques used in protein production and structural biology have been developing rapidly, but techniques for recording the laboratory information produced have not kept pace. One approach is the development of laboratory information-management systems (LIMS), which typically use a relational database schema to model and store results from a laboratory workflow. The underlying philosophy and implementation of the Protein Information Management System (PiMS), a LIMS development specifically targeted at the flexible and unpredictable workflows of protein-production research laboratories of all scales, is described. PiMS is a web-based Java application that uses either Postgres or Oracle as the underlying relational database-management system. PiMS is available under a free licence to all academic laboratories either for local installation or for use as a managed service.
The Protein Data Bank (PDB) contains a wealth of small molecule - macro molecule complexes the study of which contribute enormously to our understanding of the interactions. However, exploiting and mining this treasure trove of data requires advanced analysis and retrieval methods that take into account both types of molecules. One such method is PDBeMotif, that has been developed by the Protein Data Bank in Europe (PDBe) at EMBL-EBI. Utilizing a relational database model at the back-end, the data structure represents a network of molecule, residue and motif interactions as well as their relative positions in the sequence and in 3D. The loader applies a number of algorithms to analyse PDB and derive necessary information, such as planarity and aromaticity of the chemical compounds, hydrogen-bonds network, coordination geometry, bond types (including pi electron interactions), 3D structural motifs, sequence domains and families. It collects information about sequence features, motifs and catalytic sites from available Distributed Annotation System (DAS) resources. The web application allows for a wide variety of searches and data analysis including protein motifs with chemical fragments association, protein sites characterisation, correlating properties, hits multiple sequence and 3D alignments. The whole system is released under GPL and available with the source code from http://sourceforge.net/projects/pdbsam and on line at http://www.ebi.ac.uk/pdbe-site/PDBeMotif/
We present a suite of software for the complete and easy deposition of NMR data to the PDB and BMRB. This suite uses the CCPN framework and introduces a freely downloadable, graphical desktop application called CcpNmr Entry Completion Interface (ECI) for the secure editing of experimental information and associated datasets through the lifetime of an NMR project. CCPN projects can be created within the CcpNmr Analysis software or by importing existing NMR data files using the CcpNmr FormatConverter. After further data entry and checking with the ECI, the project can then be rapidly deposited to the PDBe using AutoDep, or exported as a complete deposition NMR-STAR file. In full CCPN projects created with ECI, it is straightforward to select chemical shift lists, restraint data sets, structural ensembles and all relevant associated experimental collection details, which all are or will become mandatory when depositing to the PDB. Instructions and download information for the ECI are available from the PDBe web site at http://www.ebi.ac.uk/pdbe/nmr/deposition/eci.html .
Cryo-electron microscopy reconstruction methods are uniquely able to reveal structures of many important macromolecules and macromolecular complexes. EMDataBank.org, a joint effort of the Protein Data Bank in Europe (PDBe), the Research Collaboratory for Structural Bioinformatics (RCSB) and the National Center for Macromolecular Imaging (NCMI), is a global ‘one-stop shop’ resource for deposition and retrieval of cryoEM maps, models and associated metadata. The resource unifies public access to the two major archives containing EM-based structural data: EM Data Bank (EMDB) and Protein Data Bank (PDB), and facilitates use of EM structural data of macromolecules and macromolecular complexes by the wider scientific community.
We present a novel technique for a fast chemical substructure search on a relational database by use of a standard SQL query. The symmetry of a query graph is analyzed to give additional constraints. Our method is based on breadth-first search (BFS) algorithms implementation using Relational Database Management Systems (RDBMS). In addition to the chemical search we apply our technique to the field of intermolecular interactions which involves nonplanar graphs and describe how to achieve linear time performance along with the suggestion on how to sufficiently reduce the linear coefficient. From the algorithms theory perspective these results mean that subgraph isomorphism is a polynomial time problem, hence equal problems have the same complexity. The application to subgraph isomorphism in chemical search is available at http://www.ebi.ac.uk/msd-srv/chemsearch and http://www.ebi.ac.uk/msd-srv/msdmotif/chem. The application to the network of molecule interactions is available at http://www.ebi.ac.uk/msd-srv/msdmotif.
The Protein Data Bank (PDB) is the repository for three-dimensional structures of biological macromolecules, determined by experimental methods. The data in the archive is free and easily available via the Internet from any of the worldwide centers managing this global archive. These data are used by scientists, researchers, bioinformatics specialists, educators, students, and general audiences to understand biological phenomenon at a molecular level. Analysis of this structural data also inspires and facilitates new discoveries in science. This chapter describes the tools and methods currently used for deposition, processing, and release of data in the PDB. References to future enhancements are also included.
The quality of the carbohydrates in PDB entries is rather poor compared with the protein parts [1], with about 30% of the PDB entries with carbohydrates being erroneous [2].The main reasons for this are the complexity of carbohydrates and the lack of check programs for carbohydrate 3D structures.Recently, such check tools were established.Pdb-care (PDB CArbohydrate REsidue check) [www.glycosciences.de/tools/pdb-care/]checks if the carbohydrate residue names used in a PDB file match the monosaccharide units present in the structure [3].Furthermore, the connectivities given in a PDB file are checked.These bond checks can be applied not only to carbohydrates but to any residue.To evaluate the conformation of a carbohydrate chain, plots of the phi / psi angles of the glycosidic linkages can be used, similar to the Ramachandran Plot for proteins [4].However, the preferred glycan torsions depend on the involved residues and the linkage position.Thus, separate plots have to be generated for each type of disaccharide fragment in the structure.This can be done with carp [www.glycosciences.de/tools/carp/].Detected torsions can be compared with either all torsions of the respective type that are present in the PDB or with computed energy maps taken from GlycoMapsDB [www.glycosciences.de/modeling/glycomapsdb/] [5].Use of these tools will help to increase the reliability of carbohydrate 3D structures.
A new scheme has been devised to represent viruses and other biological assemblies with regular noncrystallographic symmetry in the Protein Data Bank (PDB). The scheme describes existing and anticipated PDB entries of this type using generalized descriptions of deposited and experimental coordinate frames, symmetry and frame transformations. A simplified notation has been adopted to express the symmetry generation of assemblies from deposited coordinates and matrix operations describing the required point, helical or crystallographic symmetry. Complete correct information for building full assemblies, subassemblies and crystal asymmetric units of all virus entries is now available in the remediated PDB archive.
We describe the role of the BioMagResBank (BMRB) within the Worldwide Protein Data Bank (wwPDB) and recent policies affecting the deposition of biomolecular NMR data. All PDB depositions of structures based on NMR data must now be accompanied by experimental restraints. A scheme has been devised that allows depositors to specify a representative structure and to define residues within that structure found experimentally to be largely unstructured. The BMRB now accepts coordinate sets representing three-dimensional structural models based on experimental NMR data of molecules of biological interest that fall outside the guidelines of the Protein Data Bank (i.e., the molecule is a peptide with 23 or fewer residues, a polynucleotide with 3 or fewer residues, a polysaccharide with 3 or fewer sugar residues, or a natural product), provided that the coordinates are accompanied by representation of the covalent structure of the molecule (atom connectivity), assigned NMR chemical shifts, and the structural restraints used in generating model. The BMRB now contains an archive of NMR data for metabolites and other small molecules found in biological systems.
BACKGROUND:Protein structures have conserved features - motifs, which have a sufficient influence on the protein function. These motifs can be found in sequence as well as in 3D space. Understanding of these fragments is essential for 3D structure prediction, modelling and drug-design. The Protein Data Bank (PDB) is the source of this information however present search tools have limited 3D options to integrate protein sequence with its 3D structure.RESULTS:We describe here a web application for querying the PDB for ligands, binding sites, small 3D structural and sequence motifs and the underlying database. Novel algorithms for chemical fragments, 3D motifs, phi/psi sequences, super-secondary structure motifs and for small 3D structural motif associations searches are incorporated. The interface provides functionality for visualization, search criteria creation, sequence and 3D multiple alignment options. MSDmotif is an integrated system where a results page is also a search form. A set of motif statistics is available for analysis. This set includes molecule and motif binding statistics, distribution of motif sequences, occurrence of an amino-acid within a motif, correlation of amino-acids side-chain charges within a motif and Ramachandran plots for each residue. The binding statistics are presented in association with properties that include a ligand fragment library. Access is also provided through the distributed Annotation System (DAS) protocol. An additional entry point facilitates XML requests with XML responses.CONCLUSION:MSDmotif is unique by combining chemical, sequence and 3D data in a single search engine with a range of search and visualisation options. It provides multiple views of data found in the PDB archive for exploring protein structures.
The Macromolecular Structure Database (MSD) (http://www.ebi.ac.uk/msd/) [1] group is one of the four partners in the worldwide Protein Data Bank (wwPDB) [2], the consortium entrusted with the collation, maintenance and distribution of the global repository of macromolecular structure data.Structures can be deposited at the MSD using the AutoDep [3] deposition tool or with the other wwPDB partners using the ADIT system.The AutoDep 4.1 is an extension of the previously developed software AutoDep 4.0 [4], [5] which was a complete rewrite of the original AutoDep system.The AutoDep system allows value-added information to be returned in a safe and secure manner into the passwordprotected deposition session only accessible to the depositor, following annotation of the structure by curation staff, within 2 days of deposition.AutoDep 4.1 is also available for download and installation in-house, where a deposition can be completed and validated before uploading the whole deposition session to the MSD site, where submission to the wwPDB can be completed in minutes.The extended version of AutoDep ( 4.1) provides detailed information regarding the ligand binding sites, Uniprot sequence mapping and taxonomy information, applets for viewing the files, in addition to structure factor validation statistics, quaternary structure assessments, and new ligand dictionaries.With structures being determined at an ever increasing rate, it is imperative that deposition tools keep pace with this exponential growth of data.We believe that the latest release of AutoDep has significantly automated the deposition process, reduced the time taken to deposit a structure, and at the same time harnessed services offered by the MSD group in returning useful information to the depositor.
To the Editor:The building of crystallographic models of proteins is guided by understanding of the primary structure of proteins and by well-established and rigorously applied stereochemical principles.However, a cursory survey of Protein Data Bank entries containing oligosaccharides suggests that of the order of one-third of entries contain significant errors in carbohydrate stereochemistry, nomenclature or even consistency with the electron density maps.Many of the stereochemical errors can be detected by reference to conformational studies of glycans 1,2 and to publicly available resources (http://www.glycosciences.de/tools/).However, these errors also indicate that there is a wide discrepancy in the sophistication of building and validation tools available for protein and carbohydrate models.An example of the difficulties that can be encountered when building crystallographic models of glycoproteins is the recent model proposed by Szakonyi et al. 3 of the Epstein-Barr virus major envelope glycoprotein, EBV gp350.EBV gp350 was expressed in Spodoptera frugiperda Sf9 cells, and Szakonyi et al. 3 report that they observed electron density corresponding to the oligosaccharide chains of fourteen N-linked glycosylation sites.The crystallization of such a heavily glycosylated glycoprotein is a notable achievement.However, the proposed model contains not only systematic errors in carbohydrate stereochemistry, presumably resulting from inadequate parameter files, but also hitherto unreported motifs in the primary structures of the glycans.The previously undescribed glycosidic linkages and motifs that Szakonyi et al. 3 propose include Man-(1→3)-GlcNAc and GlcNAc-(1→3)-GlcNAc linkages (of indeterminate anomericity) within the trimannosyl core, hybrid-type glycans containing a terminal Man-(1→3)-GlcNAc linkage on the 3-antennae, and β-galactosyl motifs capping oligomannose-type glycans.We suggest that, in the absence of supporting evidence, electron density at 3.5-Å resolution should not be used to support linkages incompatible
We discuss basic physical-chemical principles underlying the formation of stable macromolecular complexes, which in many cases are likely to be the biological units performing a certain physiological function. We also consider available theoretical approaches to the calculation of macromolecular affinity and entropy of complexation. The latter is shown to play an important role and make a major effect on complex size and symmetry. We develop a new method, based on chemical thermodynamics, for automatic detection of macromolecular assemblies in the Protein Data Bank (PDB) entries that are the results of X-ray diffraction experiments. As found, biological units may be recovered at 80-90% success rate, which makes X-ray crystallography an important source of experimental data on macromolecular complexes and protein-protein interactions. The method is implemented as a public WWW service.
The worldwide Protein Data Bank (wwPDB) is the international collaboration that manages the deposition, processing and distribution of the PDB archive. The online PDB archive is a repository for the coordinates and related information for more than 38 000 structures, including proteins, nucleic acids and large macromolecular complexes that have been determined using X-ray crystallography, NMR and electron microscopy techniques. The founding members of the wwPDB are RCSB PDB (USA), MSD-EBI (Europe) and PDBj (Japan) [H.M. Berman, K. Henrick and H. Nakamura (2003) Nature Struct. Biol., 10, 980]. The BMRB group (USA) joined the wwPDB in 2006. The mission of the wwPDB is to maintain a single archive of macromolecular structural data that are freely and publicly available to the global community. Additionally, the wwPDB provides a variety of services to a broad community of users. The wwPDB website at http://www.wwpdb.org/ provides information about services provided by the individual member organizations and about projects undertaken by the wwPDB.
A major problem faced by structural biology today is the issue of function prediction. With the success of the various Structural genomics initiatives and advances in crystallography, proteomics, and other experimental techniques, there has been an explosion of new protein structures being deposited in the databases. In many cases, however, these proteins have little or no functional annotation. Sequence-based approaches still remain the most effective way to assign function based on homology, but in cases of extreme divergence and analogous proteins these methods can fail. In order to identify these types of relationships, a number of structure-based approaches have been developed, such as the MSDmotif service. No single method is successful in all cases and a more prudent approach involves the utilization of data from a wide range of resources. One such approach is the ProFunc server, developed to help researchers narrow down the number of functional possibilities for experimental validation.
The Worldwide Protein Data Bank (wwPDB; wwpdb.org) is the international collaboration that manages the deposition, processing and distribution of the PDB archive. The online PDB archive at ftp://ftp.wwpdb.org is the repository for the coordinates and related information for more than 47 000 structures, including proteins, nucleic acids and large macromolecular complexes that have been determined using X-ray crystallography, NMR and electron microscopy techniques. The members of the wwPDBRCSB PDB (USA), MSD-EBI (Europe), PDBj (Japan) and BMRB (USA)have remediated this archive to address inconsistencies that have been introduced over the years. The scope and methods used in this project are presented.