The Protein Data Bank (PDB) is the repository for three-dimensional structures of biological macromolecules, determined by experimental methods. The data in the archive is free and easily available via the Internet from any of the worldwide centers managing this global archive. These data are used by scientists, researchers, bioinformatics specialists, educators, students, and general audiences to understand biological phenomenon at a molecular level. Analysis of this structural data also inspires and facilitates new discoveries in science. This chapter describes the tools and methods currently used for deposition, processing, and release of data in the PDB. References to future enhancements are also included.
ike most scientists, annotators at the Research Collaboratory for Structural Bioinformatics (RCSB) (http://www.pdb.org)dread the immortal cocktail party question ''So, what do you do?''Unlike for some jobs, however, their answer can leave other scientists at the party with no response.Even within the structural biology community, our job is not well-understood.Throughout this perspective, we will shed light on the daily challenges faced by annotators at the RCSB and give the reader a glimpse at the juggling act that defines the job of a biocurator.
The Protein Data Bank (PDB; http://www.pdb.org ) is the world-wide repository for three-dimensional structural data determined using various experimental methods. The options and procedures for searching and downloading structural data from the PDB, which is maintained by the Research Collaboratory for Structural Bioinformatics (RCSB), are described here along with tools for depositing and assessing the quality of structures. Several types of information are associated with each structure deposition including atomic coordinates of the structure, experimental data used to solve the structure, sequences of all macromolecules that constitute the structures, details about the structure solution method, images showing different views of the structure, derived geometric data, and a variety of links to other resources. These data and resources may be used for understanding and designing biochemical, genetic, or other experiments to study the stability or function of the molecule. They can also be used for molecular modeling and drug design.
This chapter describes the structure-validation procedures that are used by the Protein Data Bank (PDB) to maintain data quality. It also presents an analysis of the way these procedures are used to process new entries and postprocess or revisit the legacy data. The PDB was established in 1971 by Walter Hamilton at Brookhaven National Laboratory in response to community requirements for a central repository for information about biological macromolecular structures. The PDB collects the results of structure determination experiments, organizes the data, and makes it available to an increasingly broad community of users. As part of this process, it is the responsibility of the PDB to ensure that the data released from the PDB are represented accurately and that diagnostics are provided to assist depositors in correcting errors in the deposited structures. Combined with similar advances in the field of structural biology, there have been dramatic increases in the size of the archive and the diversity of structure data, and the community of PDB users has broadened significantly.
The Protein Data Bank [PDB; Berman, Westbrook et al. (2000), Nucleic Acids Res. 28 , 235–242; http://www.pdb.org/ ] is the single worldwide archive of primary structural data of biological macromolecules. Many secondary sources of information are derived from PDB data. It is the starting point for studies in structural bioinformatics. This article describes the goals of the PDB, the systems in place for data deposition and access, how to obtain further information and plans for the future development of the resource. The reader should come away with an understanding of the scope of the PDB and what is provided by the resource.
Areas of Specialization struCtural Biology & BioinformatiCs Plant genetiCs Protein BioChemistry DeveloPmental Biology ComPutational BioinformatiCs laBoratory moleCular genetiCs of leaf DeveloPment PlastiD moleCular genetiCs moleCular genetiCs of meiotiC reComBination & Chromosome segregation moleCular Biology of Plant DeveloPment DeveloPmental & moleCular genetiCs segmentation During animal DeveloPment synaPse formation & the Central nervous system meChanisms of transCriPtion in miCroorganisms reProDuCtive Biology, Cell-Cell interaCtions emBryoniC Patterning, hematoPoiesis & ePigenetiC Control of gene exPression in DrosoPhilia outreaCh aCtivities Cell anD Cell ProDuCts fermentation faCility transCriPtional regulation in yeast BeneDiCt miChael fellow Charles anD Johanna BusCh fellow list of PuBliCations Waksman institute faculty DiRectoRy Mission Statement The Waksman Institute's mission is to conduct research in microbial molecular genetics, developmental molecular genetics, plant molecular genetics, and structural and computational biology. The Institute also provides a catalyst for general university initiatives, a life science infrastructure, undergraduate and graduate education, and a public service function for the state. Background The principal mission of the Waksman Institute is research. While the initial emphasis of the institute at its founding was microbiology, its focus soon turned toward molecular genetics, and was later broadened to include organisms other than viruses, bacteria, and fungi. As a reflection of this new, broadened, vision of research at the Institute, under previous directors fruit flies and plants were also studied. Since my assumption of the Institute's directorship, I have strived to expand its investigative horizons to include computational and structural biology, and a further emphasis on the molecular genetics of regulation of gene expression and biomolecular interactions. This new expansion of the Waksman Institute's investigative goals has stimulated the introduction of interdisciplinary programs with chemistry, computer science, and plant science. Indeed, the Institute's research mission has evolved from a diversity of disciplines centered on antibiotics to a unified discipline of molecular genetics with a more diverse set of biological problems. The Institute today employs faculty teams that concentrate on certain classes of organisms amenable to genetic analysis such as bacteria and fungi (Escherichia coli and yeast), animal systems (e.g., Drosophila and C. elegans), and plants (Arabidopsis, tobacco, and maize). Although the Institute focuses on basic academic questions in microbial, animal, and plant research, it continues to seek practical and commercially viable applications of its discoveries. Historically, in fact, the Institute owes its existence to the symbiotic relationship that exists between academic research institutions and the private sector. In 1939 Dr. Selman Waksman, the Institute's founder …