The cryogenic sample-electron microscopy (cryoEM) field has generated significant amounts of 3D Electron Microscopy (3DEM) volumetric data and associated metadata, now comprehensively archived in the Electron Microscopy Data Bank (EMDB - www.emdatabank.org ) and the Electron Microscopy Public Image Archive (EMPIAR - www.empiar.org ). Harnessing the full potential of these resources requires robust, flexible, and publicly accessible tools for data exploration, analysis and retrieval. Here, we present Chart Builder, an interactive web-based platform that enables researchers to create customizable, publication-quality visualizations directly from archival metadata, validation assessments, and cross-reference annotations. Chart Builder integrates the same query-driven and flexible Solr search system as EMDB search, into a user interface with tools to assist users to filter, group, and compare data without programming expertise. It supports multiple chart types (including line, bar, area, scatter (2D and 3D), histogram, bubble, pie, geographic and Venn diagrams) with customizable axes, data series, and statistical operators. Users can apply global filters, define temporal, categorical, or custom query-based axes, and explore multi-dimensional relationships interactively. Chart data-points are linked to their underlying datasets, such that visualisation interaction opens entry-level or archive-level search results for inspection and datasets from charts may be exported in several ways. Findability, accessibility, interoperability and reusability of data are facilitated by these direct access and export mechanisms, including HTML embedding, persistent URL sharing and chart/data download options. By combining interactivity and ease of use with up-to-date access to the EMDB and EMPIAR archive metadata, both computational and experimental communities may explore and visualize current metadata and export to formats for further analysis or as publication-ready figures. Chart Builder promotes community-driven data analysis and empowers users to evaluate trends in the biological 3DEM field. Chart Builder is freely accessible and fully integrated into the EMDB website at https://www.ebi.ac.uk/emdb/statistics/builder/ .
Motivation:The electron microscopy data bank (EMDB) is a key repository for three-dimensional electron microscopy (3DEM) data but lacks comprehensive annotations and connections to many related biological, functional, and structural data resources. This limitation arises from the optional nature of such information to reduce depositor burden and the complexity of maintaining up-to-date external references, often requiring depositor consent. To address these challenges, we developed EMDB Integration with Complexes, Structures, and Sequences (EMICSS), an independent system that automatically updates cross-references with over 20 external resources, including UniProt, AlphaFold DB, PubMed, Complex Portal, and Gene Ontology. Results:EMICSS (https://www.ebi.ac.uk/emdb/emicss) annotations are accessible in multiple formats for every EMDB entry and its linked resources, and programmatically via the EMDB application programming interface. EMICSS plays a crucial role supporting the EMDB website, with annotations being used on entry pages, statistics, and in the search system. Availability and implementation:EMICSS is implemented in Python and it is an open-source, distributed under the Apache license version 2.0, with core code available on GitHub (https://github.com/emdb-empiar/added_annotations).
Molecular structure determination using electron cryomicroscopy (cryoEM) is poised in early 2025 to surpass X-ray crystallography as the most used method for experimentally determining new structures. But the technique has not reached the physical limits set by radiation damage and the signal-to-noise ratio in individual images of molecules. By examining these limits and comparing the number and resolution of structures determined versus molecular weight, we identify opportunities for extending the application of single-particle cryoEM. This will help guide technology development to continue the exponential growth of structural biology.
The Protein Data Bank (PDB) is the global repository for public-domain experimentally determined 3D biomolecular structural information. The archival nature of the PDB presents certain challenges pertaining to updating or adding associated annotations from trusted external biodata resources. While each Worldwide PDB (wwPDB) partner has made best efforts to provide up-to-date external annotations, accessing and integrating information from disparate wwPDB data centers can be an involved process. To address this issue, the wwPDB has established the PDB Next Generation (or NextGen) Archive, developed to centralize and streamline access to enriched structural annotations from wwPDB partners and trusted external sources. At present, the NextGen Archive provides mappings between experimentally determined 3D structures of proteins and UniProt amino acid sequences, domain annotations from Pfam, SCOP2 and CATH databases and intra-molecular connectivity information. Since launch, the PDB NextGen Archive has seen substantial user engagement with over 3.5 million data file downloads, ensuring researchers have access to accurate, up-to-date and easily accessible structural annotations. Database URL: http://www.wwpdb.org/ftp/pdb-nextgen-archive-site.
Organised data is easy to use but the rapid developments in the field of bioimaging, with improvements in instrumentation, detectors, software and experimental techniques, have resulted in an explosion of the volumes of data being generated, making well-organised data an elusive goal. This guide offers a handful of recommendations for bioimage depositors, analysts and microscope and software developers, whose implementation would contribute towards better organised data in preparation for archival. Based on our experience archiving large image datasets in EMPIAR, the BioImage Archive and BioStudies, we propose a number of strategies that we believe would improve the usability (clarity, orderliness, learnability, navigability, self-documentation, coherence and consistency of identifiers, accessibility, succinctness) of future data depositions more useful to the bioimaging community (data authors and analysts, researchers, clinicians, funders, collaborators, industry partners, hardware/software producers, journals, archive developers as well as interested but non-specialist users of bioimaging data). The recommendations that may also find use in other data-intensive disciplines. To facilitate the process of analysing data organisation, we present bandbox, a Python package that provides users with an assessment of their data by flagging potential issues, such as redundant directories or invalid characters in file or folder names, that should be addressed before archival. We offer these recommendations as a starting point and hope to engender more substantial conversations across and between the various data-rich communities.
In January 2020, a workshop was held at EMBL-EBI (Hinxton, UK) to discuss data requirements for the deposition and validation of cryoEM structures, with a focus on single-particle analysis. The meeting was attended by 47 experts in data processing, model building and refinement, validation, and archiving of such structures. This report describes the workshop's motivation and history, the topics discussed, and the resulting consensus recommendations. Some challenges for future methods-development efforts in this area are also highlighted, as is the implementation to date of some of the recommendations.
Organised data is easy to use but the rapid developments in the field of bioimaging, with improvements in instrumentation, detectors, software and experimental techniques, have resulted in an explosion of the volumes of data being generated, making well-organised data an elusive goal. This guide offers a handful of recommendations for bioimage depositors, analysts and microscope and software developers, whose implementation would contribute towards better organised data in preparation for archival. Based on our experience archiving large image datasets in EMPIAR, the BioImage Archive and BioStudies, we propose a number of strategies that we believe would improve the usability (clarity, orderliness, learnability, navigability, self-documentation, coherence and consistency of identifiers, accessibility, succinctness) of future data depositions more useful to the bioimaging community (data authors and analysts, researchers, clinicians, funders, collaborators, industry partners, hardware/software producers, journals, archive developers as well as interested but non-specialist users of bioimaging data). The recommendations that may also find use in other data-intensive disciplines. To facilitate the process of analysing data organisation, we present bandbox, a Python package that provides users with an assessment of their data by flagging potential issues, such as redundant directories or invalid characters in file or folder names, that should be addressed before archival. We offer these recommendations as a starting point and hope to engender more substantial conversations across and between the various data-rich communities.
IHMCIF (github.com/ihmwg/IHMCIF) is a data information framework that supports archiving and disseminating macromolecular structures determined by integrative or hybrid modeling (IHM), and making them Findable, Accessible, Interoperable, and Reusable (FAIR). IHMCIF is an extension of the Protein Data Bank Exchange/macromolecular Crystallographic Information Framework (PDBx/mmCIF) that serves as the framework for the Protein Data Bank (PDB) to archive experimentally determined atomic structures of biological macromolecules and their complexes with one another and small molecule ligands (e.g., enzyme cofactors and drugs). IHMCIF serves as the foundational data standard for the PDB-Dev prototype system, developed for archiving and disseminating integrative structures. It utilizes a flexible data representation to describe integrative structures that span multiple spatiotemporal scales and structural states with definitions for restraints from a variety of experimental methods contributing to integrative structural biology. The IHMCIF extension was created with the benefit of considerable community input and recommendations gathered by the Worldwide Protein Data Bank (wwPDB) Task Force for Integrative or Hybrid Methods (wwpdb.org/task/hybrid). Herein, we describe the development of IHMCIF to support evolving methodologies and ongoing advancements in integrative structural biology. Ultimately, IHMCIF will facilitate the unification of PDB-Dev data and tools with the PDB archive so that integrative structures can be archived and disseminated through PDB.
Biomolecular structure analysis from experimental NMR studies generally relies on restraints derived from a combination of experimental and knowledge-based data. A challenge for the structural biology community has been a lack of standards for representing these restraints, preventing the establishment of uniform methods of model-vs-data structure validation against restraints and limiting interoperability between restraint-based structure modeling programs. The NEF and NMR-STAR formats provide a standardized approach for representing commonly used NMR restraints. Using these restraint formats, a standardized validation system for assessing structural models of biopolymers against restraints has been developed and implemented in the wwPDB OneDep data deposition-validation-biocuration system. The resulting wwPDB restraint violation report provides a model vs. data assessment of biomolecule structures determined using distance and dihedral restraints, with extensions to other restraint types currently being implemented. These tools are useful for assessing NMR models, as well as for assessing biomolecular structure predictions based on distance restraints.
Public archiving in structural biology is well established with the Protein Data Bank (PDB; wwPDB.org) catering for atomic models and the Electron Microscopy Data Bank (EMDB; emdb-empiar.org) for 3D reconstructions from cryo-EM experiments. Even before the recent rapid growth in cryo-EM, there was an expressed community need for a public archive of image data from cryo-EM experiments for validation, software development, testing and training. Concomitantly, the proliferation of 3D imaging techniques for cells, tissues and organisms using volume EM (vEM) and X-ray tomography (XT) led to calls from these communities to publicly archive such data as well. EMPIAR (empiar.org) was developed as a public archive for raw cryo-EM image data and for 3D reconstructions from vEM and XT experiments and now comprises over a thousand entries totalling over 2 petabytes of data. EMPIAR resources include a deposition system, entry pages, facilities to search, visualize and download datasets, and a REST API for programmatic access to entry metadata. The success of EMPIAR also poses significant challenges for the future in dealing with the very fast growth in the volume of data and in enhancing its reusability.
The Protein Data Bank (PDB) is the single global archive of atomic-level, three-dimensional structures of biological macromolecules experimentally determined by macromolecular crystallography, nuclear magnetic resonance spectroscopy or three-dimensional cryo-electron microscopy. The PDB is growing continuously, with a recent rapid increase in new structure depositions from Asia. In 2022, the Worldwide Protein Data Bank (wwPDB; https://www.wwpdb.org/) partners welcomed Protein Data Bank China (PDBc; https://www.pdbc.org.cn) to the organization as an Associate Member. PDBc is based in the National Facility for Protein Science in Shanghai which is associated with the Shanghai Advanced Research Institute of Chinese Academy of Sciences, the Shanghai Institute for Advanced Immunochemical Studies and the iHuman Institute of ShanghaiTech University. This letter describes the history of the wwPDB, recently established mechanisms for adding new wwPDB data centers and the processes developed to bring PDBc into the partnership.
Volume electron microscopy (vEM) techniques produce scientifically important datasets which are time and resource intensive to generate (Peddie et al., 2022). Public archival of such datasets, usually described in the literature, provides many benefits to the data depositors, to those making use of research results based on the datasets, and to the vEM community at large, both now and in the future. In this chapter we discuss these benefits, explain how EMBL-EBI's image data services support archival of both vEM and correlative imaging data, and discuss how future developments will unlock more value from these vEM datasets.
Volume electron microscopy (vEM) is a group of techniques that reveal the 3D ultrastructure of cells and tissues through continuous depths of at least 1 micrometer. A burgeoning grassroots community effort is fast building the profile and revealing the impact of vEM technology in the life sciences and clinical research.
ModelCIF (github.com/ihmwg/ModelCIF) is a data information framework developed for and by computational structural biologists to enable delivery of Findable, Accessible, Interoperable, and Reusable (FAIR) data to users worldwide. ModelCIF describes the specific set of attributes and metadata associated with macromolecular structures modeled by solely computational methods and provides an extensible data representation for deposition, archiving, and public dissemination of predicted three-dimensional (3D) models of macromolecules. It is an extension of the Protein Data Bank Exchange / macromolecular Crystallographic Information Framework (PDBx/mmCIF), which is the global data standard for representing experimentally-determined 3D structures of macromolecules and associated metadata. The PDBx/mmCIF framework and its extensions (e.g., ModelCIF) are managed by the Worldwide Protein Data Bank partnership (wwPDB, wwpdb.org) in collaboration with relevant community stakeholders such as the wwPDB ModelCIF Working Group (wwpdb.org/task/modelcif). This semantically rich and extensible data framework for representing computed structure models (CSMs) accelerates the pace of scientific discovery. Herein, we describe the architecture, contents, and governance of ModelCIF, and tools and processes for maintaining and extending the data standard. Community tools and software libraries that support ModelCIF are also described.
The Electron Microscopy Data Bank (EMDB) is the central archive of the electron cryo-microscopy (cryo-EM) community for storing and disseminating volume maps and tomograms. With input from the community, EMDB has developed new resources for the validation of cryo-EM structures, focusing on the quality of the volume data alone and that of the fit of any models, themselves archived in the Protein Data Bank (PDB), to the volume data. Based on recommendations from community experts, the validation resources are developed in a three-tiered system. Tier 1 covers an extensive and evolving set of validation metrics, including tried and tested metrics as well as more experimental ones, which are calculated for all EMDB entries and presented in the Validation Analysis (VA) web resource. This system is particularly useful for cryo-EM experts, both to validate individual structures and to assess the utility of new validation metrics. Tier 2 comprises a subset of the validation metrics covered by the VA resource that have been subjected to extensive testing and are considered to be useful for specialists as well as nonspecialists. These metrics are presented on the entry-specific web pages for the entire archive on the EMDB website. As more experience is gained with the metrics included in the VA resource, it is expected that consensus will emerge in the community regarding a subset that is suitable for inclusion in the tier 2 system. Tier 3, finally, consists of the validation reports and servers that are produced by the Worldwide Protein Data Bank (wwPDB) Consortium. Successful metrics from tier 2 will be proposed for inclusion in the wwPDB validation pipeline and reports. The details of the new resource are described, with an emphasis on the tier 1 system. The output of all three tiers is publicly available, either through the EMDB website (tiers 1 and 2) or through the wwPDB ftp sites (tier 3), although the content of all three will evolve over time (fastest for tier 1 and slowest for tier 3). It is our hope that these validation resources will help the cryo-EM community to obtain a better understanding of the quality and of the best ways to assess the quality of cryo-EM structures in EMDB and PDB.
Despite the huge impact of data resources in genomics and structural biology, until now there has been no central archive for biological data for all imaging modalities. The BioImage Archive is a new data resource at the European Bioinformatics Institute (EMBL-EBI) designed to fill this gap. In its initial development BioImage Archive accepts bioimaging data associated with publications, in any format, from any imaging modality from the molecular to the organism scale, excluding medical imaging. The BioImage Archive will ensure reproducibility of published studies that derive results from image data and reduce duplication of effort. Most importantly, the BioImage Archive will help scientists to generate new insights through reuse of existing data to answer new biological questions, and provision of training, testing and benchmarking data for development of tools for image analysis. The archive is available at https://www.ebi.ac.uk/bioimage-archive/ . Highlights The BioImage Archive is a new archival data resource at the European Bioinformatics Institute (EMBL-EBI). The BioImage Archive aims to accept all biological imaging data associated with peer-reviewed publications using approaches that probe biological structure, mechanism and dynamics, as well as other important datasets that can serve as reference examples for particular biological or technical domains. The BioImage Archive aims to encourage the use of valuable imaging data, to improve reproducibility of published results that rely on image data, and to facilitate extraction of novel biological insights from existing data and development of new image analysis methods. The BioImage Archive forms the foundation for an ecosystem of related databases, supporting those resources with storage infrastructure and indexing across databases. Across this ecosystem, the BioImage Archive already stores and provides access to over 1.5 petabytes of image data from many different imaging modalities and biological domains. Future development of the BioImage Archive will support the fast-emerging next generation file formats (NGFFs) for bioimaging data, providing access mechanisms tailored toward modern visualisation and data exploration tools, as well as unlocking the power of modern AI-based image-analysis approaches.
Recent technological advances in electron cryo-microscopy (cryo-EM) have led to significant improvements in the resolution of many single-particle reconstructions and a sharp increase in the number of entries released in the Electron Microscopy Data Bank (EMDB) every year, which in turn has opened new possibilities for data mining. Here we present a resolution-dependent library of rotamer-specific amino-acid map motifs mined from entries in the EMDB archive with reported resolution between 2.0 and 4.0Å. We further describe 3D-Strudel, a method for map/model validation based on these libraries. 3D-Strudel calculates linear correlation coefficients between the map values of a map-motif from the library and the experimental map values around a target residue. We also present “Strudel Score”, a plug-in for ChimeraX, as a user-friendly tool for visualisation of 3D-Strudel validation results.
This protocol illustrates the steps necessary to deposit correlated 3D cryo-imaging data from cryo-structured illumination microscopy and cryo-soft X-ray tomography with the BioStudies and EMPIAR deposition databases of the European Bioinformatics Institute. There is currently a real need for a robust method of data deposition to ensure unhindered access to and independent validation of correlative light and X-ray microscopy data to allow use in further comparative studies, educational activities, and data mining. For complete details on the use and execution of this protocol, please refer to Kounatidis et al. (2020).