This letter announces that PDBx/mmCIF format files will become mandatory for crystallographic depositions to the Protein Data Bank (PDB).
Abstract The Protein Data Bank in Europe (PDBe), a founding member of the Worldwide Protein Data Bank (wwPDB), actively participates in the deposition, curation, validation, archiving and dissemination of macromolecular structure data. PDBe supports diverse research communities in their use of macromolecular structures by enriching the PDB data and by providing advanced tools and services for effective data access, visualization and analysis. This paper details the enrichment of data at PDBe, including mapping of RNA structures to Rfam, and identification of molecules that act as cofactors. PDBe has developed an advanced search facility with ∼100 data categories and sequence searches. New features have been included in the LiteMol viewer at PDBe, with updated visualization of carbohydrates and nucleic acids. Small molecules are now mapped more extensively to external databases and their visual representation has been enhanced. These advances help users to more easily find and interpret macromolecular structure data in order to solve scientific problems.
The Protein Data Bank (PDB) is the single global archive of experimentally determined three-dimensional (3D) structure data of biological macromolecules. Since 2003, the PDB has been managed by the Worldwide Protein Data Bank (wwPDB; wwpdb.org), an international consortium that collaboratively oversees deposition, validation, biocuration, and open access dissemination of 3D macromolecular structure data. The PDB Core Archive houses 3D atomic coordinates of more than 144 000 structural models of proteins, DNA/RNA, and their complexes with metals and small molecules and related experimental data and metadata. Structure and experimental data/metadata are also stored in the PDB Core Archive using the readily extensible wwPDB PDBx/mmCIF master data format, which will continue to evolve as data/metadata from new experimental techniques and structure determination methods are incorporated by the wwPDB. Impacts of the recently developed universal wwPDB OneDep deposition/validation/biocuration system and various methods-specific wwPDB Validation Task Forces on improving the quality of structures and data housed in the PDB Core Archive are described together with current challenges and future plans.
The Protein Data Bank (PDB) is the single global repository for experimentally determined 3D structures of biological macromolecules and their complexes with ligands. The worldwide PDB (wwPDB) is the international collaboration that manages the PDB archive according to the FAIR principles: Findability, Accessibility, Interoperability and Reusability. The wwPDB recently developed OneDep, a unified tool for deposition, validation and biocuration of structures of biological macromolecules. All data deposited to the PDB undergo critical review by wwPDB Biocurators. This article outlines the importance of biocuration for structural biology data deposited to the PDB and describes wwPDB biocuration processes and the role of expert Biocurators in sustaining a highquality archive. Structural data submitted to the PDB are examined for self-consistency, standardized using controlled vocabularies, cross-referenced with other biological data resources and validated for scientific/technical accuracy. We illustrate how biocuration is integral to PDB data archiving, as it facilitates accurate, consistent and comprehensive representation of biological structure data, allowing efficient and effective usage by research scientists, educators, students and the curious public worldwide.
The Protein Data Bank in Europe (PDBe, pdbe.org) is actively engaged in the deposition, annotation, remediation, enrichment and dissemination of macromolecular structure data. This paper describes new developments and improvements at PDBe addressing three challenging areas: data enrichment, data dissemination and functional reusability. New features of the PDBe Web site are discussed, including a context dependent menu providing links to raw experimental data and improved presentation of structures solved by hybrid methods. The paper also summarizes the features of the LiteMol suite, which is a set of services enabling fast and interactive 3D visualization of structures, with associated experimental maps, annotations and quality assessment information. We introduce a library of Web components which can be easily reused to port data and functionality available at PDBe to other services. We also introduce updates to the SIFTS resource which maps PDB data to other bioinformatics resources, and the PDBe REST API.
The Worldwide PDB recently launched a deposition, biocuration, and validation tool: OneDep. At various stages of OneDep data processing, validation reports for three-dimensional structures of biological macromolecules are produced. These reports are based on recommendations of expert task forces representing crystallography, nuclear magnetic resonance, and cryoelectron microscopy communities. The reports provide useful metrics with which depositors can evaluate the quality of the experimental data, the structural model, and the fit between them. The validation module is also available as a stand-alone web server and as a programmatically accessible web service. A growing number of journals require the official wwPDB validation reports (produced at biocuration) to accompany manuscripts describing macromolecular structures. Upon public release of the structure, the validation report becomes part of the public PDB archive. Geometric quality scores for proteins in the PDB archive have improved over the past decade.
Transport proteins can play a pivotal role in the absorption, distribution, metabolism, excretion (ADME) and toxicity of many drug molecules. Understanding the mechanisms and binding characteristics of these membrane proteins can potentially enhance the drug discovery process by reducing the risk of transporter related interactions, which may hinder drug development in its later stages. Data, information and knowledge about transporters can therefore play an important role in the drug discovery and development process. Biological data describing genetic, functional and structural aspects of transport proteins have accumulated in recent decades, and have been followed by an increase in the number of publicly available resources that provide bioactivity and chemical data for small molecules that bind to these transporters. This chapter surveys the publicly available bioinformatics and cheminformatics databases that currently warehouse transporter relevant data and examines emerging computational techniques that make use of these data in order to predict the activity, binding capability and other characteristics of membrane transporters.
Both metabolism and transport are key elements defining the bioavailability and biological activity of molecules, i.e. their adverse and therapeutic effects. Structured and high quality experimental data stored in a suitable container, such as a relational database, facilitates easy computational processing and thus allows for high quality information/knowledge to be efficiently inferred by computational analyses. Our aim was to create a freely accessible database that would provide easy access to data describing interactions between proteins involved in transport and xenobiotic metabolism and their small molecule substrates and modulators. We present Metrabase, an integrated cheminformatics and bioinformatics resource containing curated data related to human transport and metabolism of chemical compounds. Its primary content includes over 11,500 interaction records involving nearly 3,500 small molecule substrates and modulators of transport proteins and, currently to a much smaller extent, cytochrome P450 enzymes. Data was manually extracted from the published literature and supplemented with data integrated from other available resources. Metrabase version 1.0 is freely available under a CC BY-SA 4.0 license at http://www-metrabase.ch.cam.ac.uk.
The Protein Data Bank in Europe (http://pdbe.org) accepts and annotates depositions of macromolecular structure data in the PDB and EMDB archives and enriches, integrates and disseminates structural information in a variety of ways. The PDBe website has been redesigned based on an analysis of user requirements, and now offers intuitive access to improved and value-added macromolecular structure information. Unique value-added information includes lists of reviews and research articles that cite or mention PDB entries as well as access to figures and legends from full-text open-access publications that describe PDB entries. A powerful new query system not only shows all the PDB entries that match a given query, but also shows the 'best structures' for a given macromolecule, ligand complex or sequence family using data-quality information from the wwPDB validation reports. A PDBe RESTful API has been developed to provide unified access to macromolecular structure data available in the PDB and EMDB archives as well as value-added annotations, e.g. regarding structure quality and up-to-date cross-reference information from the SIFTS resource. Taken together, these new developments facilitate unified access to macromolecular structure data in an intuitive way for non-expert users and support expert users in analysing macromolecular structure data.
ChEMBL is an open large-scale bioactivity database (https://www.ebi.ac.uk/chembl), previously described in the 2012 Nucleic Acids Research Database Issue. Since then, a variety of new data sources and improvements in functionality have contributed to the growth and utility of the resource. In particular, more comprehensive tracking of compounds from research stages through clinical development to market is provided through the inclusion of data from United States Adopted Name applications; a new richer data model for representing drug targets has been developed; and a number of methods have been put in place to allow users to more easily identify reliable data. Finally, access to ChEMBL is now available via a new Resource Description Framework format, in addition to the web-based interface, data downloads and web services.
Cancer remains a fundamental burden to public health despite substantial efforts aimed at developing effective chemotherapeutics and significant advances in chemotherapeutic regimens. The major challenge in anti-cancer drug design is to selectively target cancer cells with high specificity. Research into treating malignancies by targeting altered metabolism in cancer cells is supported by computational approaches, which can take a leading role in identifying candidate targets for anti-cancer therapy as well as assist in the discovery and optimisation of anti-cancer agents. Natural products appear to have privileged structures for anti-cancer drug development and the bulk of this particularly valuable chemical space still remains to be explored. In this review we aim to provide a comprehensive overview of current strategies for computer-guided anti-cancer drug development. We start with a discussion of state-of-the art bioinformatics methods applied to the identification of novel anti-cancer targets, including machine learning techniques, the Connectivity Map and biological network analysis. This is followed by an extensive survey of molecular modelling and cheminformatics techniques employed to develop agents targeting proteins involved in the glycolytic, lipid, NAD+, mitochondrial (TCA cycle), amino acid and nucleic acid metabolism of cancer cells. A dedicated section highlights the most promising strategies to develop anti-cancer therapeutics from natural products and the role of metabolism and some of the many targets which are under investigation are reviewed. Recent success stories are reported for all the areas covered in this review. We conclude with a brief summary of the most interesting strategies identified and with an outlook on future directions in anti-cancer drug development.
We present a novel approach to crystallographic ligand density interpretation based on Zernike shape descriptors. Electron density for a bound ligand is expanded in an orthogonal polynomial series (3D Zernike polynomials) and the coefficients from this expansion are employed to construct rotation-invariant descriptors. These descriptors can be compared highly efficiently against large databases of descriptors computed from other molecules. In this manuscript we describe this process and show initial results from an electron density interpretation study on a dataset containing over a hundred OMIT maps. We could identify the correct ligand as the first hit in about 30 % of the cases, within the top five in a further 30 % of the cases, and giving rise to an 80 % probability of getting the correct ligand within the top ten matches. In all but a few examples, the top hit was highly similar to the correct ligand in both shape and chemistry. Further extensions and intrinsic limitations of the method are discussed.
The use of spherical harmonics in the molecular sciences is widespread. They have been employed with success in, for instance, the crystallographic fast rotation function, small-angle scattering particle reconstruction, molecular surface visualisation, protein-protein docking, active site analysis and protein function prediction. An extension of the spherical harmonic expansion method is presented here that enables regions (bodies) rather than contours (surfaces) to be described and which lends itself favourably to the construction of rotationally invariant shape descriptors. This method introduces a radial term that extends the spherical harmonics to 3D polynomials. These polynomials maintain the advantages of the spherical harmonics (orthonormality, completeness, uniqueness and fast computation) but correct the drawbacks (contour based shape description and star-shape objects) and give rise to powerful invariant descriptors. We provide proof-of-principle examples illustrating the potential of this method for accurate object representation, an analysis of the descriptor classification power, and comparisons to other methods.
Root nodule extensins (RNEs) are highly glycosylated plant glycoproteins localized in the extracellular matrix of legume tissues and in the lumen of Rhizobium-induced infection threads. In pea and other legumes, a family of genes encode glycoproteins of different overall length but with the same basic composition. The predicted polypeptide sequence reveals repeating and alternating motifs characteristic of extensins and arabinogalactan proteins. In order to monitor the behavior of individual RNE gene products in the plant extracellular matrix, the coding sequence of PsRNE1 from Pisum sativum was expressed in insect cells and in tobacco leaves. RNE products extracted from tobacco tissues were of high molecular weight (in excess of 80 kDa), indicating extensive glycosylation similar to that in pea tissues. Epitope-tagged derivatives of PsRNE1 could be localized in cell walls. However, the introduction of epitope tags at the C-terminus of RNE altered the behavior of RNE in the extracellular matrix, apparently preventing intermolecular crosslinking of RNE molecules and their covalent association with other cell wall components. These observations are discussed in the light of a computational model for the RNE glycoprotein that is consistent with an extended rod-like structure. It is proposed that RNE can undergo three classes of tyrosine-based crosslinking. Intramolecular crosslinking of vicinal Tyr residues is rod stiffening, end-to-end linkage is rod lengthening, and side-to-side intermolecular crosslinking is rod bundling. The control of these interconversions could have important implications for the biomechanics of infection thread growth.