This chapter contains sections titled: Database Design Process Loading and Exporting Data in mmCIF Exporting mmCIFs or XML Files from the Deposition Database Subtypes and 'Leaf Views' Maintenance Aspects Data Clean-up The Search Database Transformation Incremental Transformation Replication Oracle Cartridge Applications Related Data Warehouse Acknowledgements References
The E-MSD macromolecular structure relational database (http://www.ebi.ac.uk/msd) is designed to be a single access point for protein and nucleic acid structures and related information. The database is derived from Protein Data Bank (PDB) entries. Relational database technologies are used in a comprehensive cleaning procedure to ensure data uniformity across the whole archive. The search database contains an extensive set of derived properties, goodness-of-fit indicators, and links to other EBI databases including InterPro, GO, and SWISS-PROT, together with links to SCOP, CATH, PFAM and PROSITE. A generic search interface is available, coupled with a fast secondary structure domain search tool.
High throughput structural genomics projects are now underway.These projects will collect comprehensive data on protein structure.The e-msd has contributed to the detailed data representation model(s) and exchange data formats and mechanisms between each step in a structure determination.This model, incorporates the means for making the information reliable, accurate and up to date and to include indicators of reliability.The e-msd recognizes that in order for data to be used efficiently for searches within the database, and have ensured data uniformity in that the meta-description of the data is consistent across all entries.The e-msd not only allows for data harvesting and archival of clean cross referenced structural data, search mechanisms are being developed to allow research workers to for example automatically annotate structure motifs and binding site properties.