The 2024 Nobel Prize in Chemistry was awarded in part for de novo protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and three-dimensional structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments. Herein, we provide an overview of the deterministic relationship between sequence and structure, which was hypothesized over half a century ago with profound implications for the biological sciences ever since. We postulate that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting protein conformational ensembles. Accordingly, we describe a conceptual model architecture, which may be trained on sequence data in combination with conformationally sensitive structural information, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Notwithstanding certain limitations in this context, NMR offers abundant structural heterogeneity conducive to conformational ensemble prediction. As NMR and other data continue to accumulate, sequence-informed prediction of protein structural dynamics with AI/ML has the potential to emerge as a transformative capability across the biological sciences.
Protein structure prediction via artificial intelligence/machine learning (AI/ML) approaches has sparked substantial research interest in structural biology and adjacent disciplines. More recently, AlphaFold2 (AF2) has been adapted for the prediction of multiple structural conformations-beyond the original scope of predicting single-state structures. This is accomplished by using multiple random seeds and subsampling the multiple sequence alignment (MSA). Research using this novel approach has focused on proteins (typically 50 residues in length or greater), while multi-conformation prediction of shorter peptides has not yet been explored in this context. Here, we report AF2-based structural conformation prediction of a total of 557 peptides (ranging in length from 10 to 40 residues) for a benchmark dataset with corresponding nuclear magnetic resonance (NMR)-determined conformational ensembles. De novo structure predictions were accompanied by structural comparison analyses to assess prediction accuracy. We found that the prediction of conformational ensembles of peptides with AF2 varied in accuracy versus NMR data, with average root-mean-square deviation (RMSD) among structured regions under 2.5 Å and average root-mean-square fluctuation (RMSF) differences under 1.5 Å for the entire set of 557 peptides. Our results reveal notable capabilities of AF2-based structural conformation prediction for peptides but also highlight considerable limitations, underscoring the necessity for interpretation discretion and the need for improved conformational ensemble prediction approaches.
Data visualization is a pivotal component of a structural biologist's arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.
Structures of many large biomolecular assemblies are now being determined using integrative approaches. In these approaches, information derived from multiple experimental and computational methods is combined to compute three-dimensional structures of multi-protein complexes and other macromolecular machines. A standalone prototype data resource for integrative structures called PDB-Dev was built, based on recommendations of the Integrative and Hybrid Methods (IHM) Task Force of the Worldwide Protein Data Bank (wwPDB). This effort included developing data standards and software tools for collecting, curating, validating, visualizing, archiving, and disseminating integrative structures that span diverse spatiotemporal scales and conformational states. Mechanisms have been created to validate integrative structures based on the experimental data underpinning them. Building upon this foundational framework, PDB-Dev has been further expanded to handle large dynamic macromolecular systems and integrative structures that combine, for example, experimental restraints with atomic coordinates computed by machine learning algorithms. Data standards and supporting tools have also been extended to capture information about biomolecular dynamics, such as conformational transitions and related kinetic data derived from biophysical methods. Recently, PDB-Dev was unified with the PDB archive and rebranded as PDB-IHM (pdb-ihm.org), further promoting FAIR (Findable, Accessible, Interoperable, and Reusable) principles of data stewardship for integrative structural biology.
The 2024 Nobel Prize in Chemistry was awarded in part for protein structure prediction using AlphaFold2, an artificial intelligence/machine learning (AI/ML) model trained on vast amounts of sequence and 3D structure data. AlphaFold2 and related models, including RoseTTAFold and ESMFold, employ specialized neural network architectures driven by attention mechanisms to infer relationships between sequence and structure. At a fundamental level, these AI/ML models operate on the long-standing hypothesis that the structure of a protein is determined by its amino acid sequence. More recently, AlphaFold2 has been adapted for the prediction of multiple protein conformations by subsampling multiple sequence alignments (MSAs). The deterministic relationship between sequence and structure was hypothesized over half a century ago with profound implications for the biological sciences ever since. Based on this relationship, we hypothesize that protein conformational dynamics are also determined, at least in part, by amino acid sequence and that this relationship may be leveraged for construction of AI/ML models dedicated to predicting ensembles of protein structures (i.e., distinct conformations). Accordingly, we conceptualized an AI/ML model architecture which may be trained on sequence data in combination with conformationally-sensitive structure data, coming primarily from nuclear magnetic resonance (NMR) spectroscopy. Sequence-informed prediction of protein structural dynamics has the potential to emerge as a transformative capability across the biological sciences, and its implementation could very well be on the horizon.
Atomic coordinate models are important for the interpretation of 3D maps produced with cryoEM and cryoET (3D electron microscopy; 3DEM). In addition to visual inspection of such maps and models, quantitative metrics can inform about the reliability of the atomic coordinates, in particular how well the model is supported by the experimentally determined 3DEM map. A recently introduced metric, Q-score, was shown to correlate well with the reported resolution of the map for well fitted models. Here, we present new statistical analyses of Q-score based on its application to ∼10 000 maps and models archived in the EMDB (Electron Microscopy Data Bank) and PDB (Protein Data Bank). Further, we introduce two new metrics based on Q-score to represent each map and model relative to all entries in the EMDB and those with similar resolution. We explore through illustrative examples of proteins, nucleic acids and small molecules how Q-scores can indicate whether the atomic coordinates are well fitted to 3DEM maps and also whether some parts of a map may be poorly resolved due to factors such as molecular flexibility, radiation damage and/or conformational heterogeneity. These examples and statistical analyses provide a basis for how Q-scores can be interpreted effectively in order to evaluate 3DEM maps and atomic coordinate models prior to publication and archiving.
The easiest and often most useful way to work with experimentally determined or computationally predicted structures of biomolecules is by viewing their three‐dimensional (3D) shapes using a molecular visualization tool. Mol* was collaboratively developed by RCSB Protein Data Bank (RCSB PDB, RCSB.org ) and Protein Data Bank in Europe (PDBe, PDBe.org ) as an open‐source, web‐based, 3D visualization software suite for examination and analyses of biostructures. It is capable of displaying atomic coordinates and related experimental data of biomolecular structures together with a variety of annotations, facilitating basic and applied research, training, education, and information dissemination. Across RCSB.org , the RCSB PDB research‐focused web portal, Mol* has been implemented to support single‐mouse‐click atomic‐level visualization of biomolecules (e.g., proteins, nucleic acids, carbohydrates) with bound cofactors, small‐molecule ligands, ions, water molecules, or other macromolecules. RCSB.org Mol* can seamlessly display 3D structures from various sources, allowing structure interrogation, superimposition, and comparison. Using influenza A H5N1 virus as a topical case study of an important pathogen, we exemplify how Mol* has been embedded within various RCSB.org tools—allowing users to view polymer sequence and structure‐based annotations integrated from trusted bioinformatics data resources, assess patterns and trends in groups of structures, and view structures of any size and compositional complexity. In addition to being linked to every experimentally determined biostructure and Computed Structure Model made available at RCSB.org , Standalone Mol* is freely available for visualizing any atomic‐level or multi‐scale biostructure at rcsb.org/3d-view .
This review article describes the co-evolution of structural biology as a discipline and the Protein Data Bank (PDB), established in 1971 as the first open-access data resource in biology by like-minded structural scientists. As the PDB archive grew in size and scope to encompass macromolecular crystallography, NMR spectroscopy, and cryo-electron microscopy, new technologies were developed to ingest, validate, curate, store, and distribute the information. Community engagement ensured that the needs of structural biologists (data depositors) and data consumers were met. Today, the archive houses more than 230,000 experimentally determined structures of proteins, nucleic acids, and macromolecular machines and their complexes with one another and small-molecule ligands. Aggregate costs of PDB data preservation are ~1% of the cost of structure determination. The enormous impact of PDB data on basic and applied research and education across the natural and medical sciences is presented and highlighted with illustrative examples. Enablement of de novo protein structure prediction (AlphaFold2, RoseTTAfold, OpenFold, etc.) is the most widely appreciated benefit of having a corpus of rigorously validated, expertly curated 3D biostructure data.
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry methods development. This update reports significant growth and enhancements since our last review in 2016. Of note, the database now contains 2.9 million binding measurements spanning 1.3 million compounds and thousands of protein targets. This growth is largely attributable to our unique focus on curating data from US patents, which has yielded a substantial influx of novel binding data. Recent improvements include a remake of the website following responsive web design principles, enhanced search and filtering capabilities, new data download options and webservices, and establishment of a long-term data archive replicated across dispersed sites. We also discuss BindingDB's positioning relative to related resources, its open data sharing policies, insights gleaned from the dataset, and plans for future growth and development.
The Protein Data Bank (PDB) was founded in 1971 as the first open-access digital data resource in biology to serve as the single global archive for three-dimensional (3D) macromolecular structure data. Current PDB holdings exceed 230,000 experimentally determined structures of proteins, nucleic acids, viruses, and macromolecular machines. The RCSB Protein Data Bank RCSB.org research-focused web portal facilitates search, analyses, and visualization of every PDB structure along with more than one million Computed Structure Models from AlphaFold DB and the ModelArchive. It is powered by a set of publicly available Application Programming Interfaces (APIs) that both support RCSB.org users and provide programmatic access to PDB data. Given the breadth and levels of granularity encompassed in this rich data collection, efficiently accessing the information programmatically may be challenging for new users. RCSB PDB has developed a Python software package, rcsb-api, that facilitates easy and efficient use of RCSB PDB APIs within a Python environment. This software tool is designed to streamline access to the extensive corpus of data housed within the PDB, enabling researchers to search, retrieve, and analyze 3D biostructure data seamlessly. Its use will accelerate research in structural biology, molecular biology and biochemistry, drug discovery, and bioinformatics by providing more efficient tools for data integration and analysis. The new toolkit is available on GitHub (github.com/rcsb/py-rcsb-api) and published to the public Python package repository (PyPI) to foster wider usage and support basic and applied research in fundamental biology, biomedicine, and the energy sciences.
In humans, protein-protein interactions mediate numerous biological processes and are central to both normal physiology and disease. Extensive research efforts have aimed to elucidate the human protein interactome, and comprehensive databases now catalog these interactions at scale. However, structural coverage of the human protein interactome is limited and remains challenging to resolve through experimental methodology alone. Recent advances in artificial intelligence/machine learning (AI/ML)-based approaches for protein interaction structure prediction present opportunities for large-scale structural characterization of the human interactome. One such model, Boltz-2, which is capable of predicting the structures of protein complexes, may serve this objective. Here, we present de novo computed models of 1,394 binary human protein interaction structures predicted using Boltz-2 based on biochemically determined interaction data sourced from the IntAct database. We assessed the predicted interaction structures through different confidence metrics, which consider both overall structure and the interaction interface. These analyses indicated that prediction confidence tended to be greater for smaller complexes, while increased multiple sequence alignment (MSA) depth tended to improve prediction confidence. Additionally, we examined annotated protein domains and found that 679 of the predicted structural complexes contained a variety of domains with putative interaction involvement on the basis of interaction interface proximity. Furthermore, our analyses revealed intricate interaction networks within the context of biological function and cancer. This work demonstrates the utility of Boltz-2 for in silico structural modeling of the human protein interactome, highlighting both strengths and limitations, while also providing a novel view of broad functional contextualization. Ultimately, such modeling is expected to yield broad structural insights with relevance across multiple domains of biomedical research.
BACKGROUND:Trauma is a leading cause of mortality, but injury-specific molecular targets remain largely unknown. We hypothesized that distinctive yet unrecognized tissue targets accessible to circulating ligands might emerge during trauma, thereby underscoring a trauma-related proteome. METHODS:We screened a peptide library to discover targets in a porcine model of major trauma: compound femur fracture with hemorrhagic shock. Bioinformatics yielded conserved motifs, and candidate receptors were affinity purified. In silico and in vitro approaches served to investigate possible associations between candidate receptors and calcium, a major component of skeletal muscle and bone. In vivo homing and molecular imaging (PET/MRI and SPECT/CT) studies of the most promising ligand peptide candidate were performed in the porcine model and were also confirmed in a corresponding rat model of major trauma. Optical methodologies and molecular dynamics simulations served to explore the molecular attributes of the ligand-receptor binding. FINDINGS:Nearly all molecular targets of the selected ligand peptides were calcium-dependent proteins, which become accessible upon trauma. We validated specific binding of homing peptides to these receptors in injured tissues, including CLRGFPALVC:CASQ1, CSEIGVRAC:HSP27, and CRQRPASGC:CALR. Notably, we determined that ligand peptide CRQRPASGC targets an injury-specific calcium-facilitated conformation of calreticulin, enabling specific molecular imaging of trauma. CONCLUSIONS:We conceptually propose the term "traumome" for the functional receptor repertoire that becomes readily amenable for ligand-directed targeting upon major trauma. These preclinical findings pave the way toward clinic-ready targeted theragnostic approaches in the setting of trauma. FUNDING:Major funding was provided by the Defense Advanced Research Projects Agency (DARPA).
The Protein Data Bank (PDB) was established in 1971 as the first open-access digital data resource in biology with just seven X-ray crystallographic structures of proteins. Today, the single global PDB archive houses more than 215,000 experimentally-determined, atomic-level three-dimensional (3D) structures of biological macromolecules that are made freely available to many millions of users worldwide with no limitations on data usage. 3D biostructure information facilitates basic and applied research and education across the sciences, impacting fundamental biology, biomedicine, biotechnology, and energy sciences. The Worldwide Protein Data Bank partnership (wwPDB, wwpdb.org) currently includes five Full Members (RCSB PDB, PDBe, PDBj, BMRB, and EMDB) and one Associate Member (PDBc), which together manage the PDB, EMDB, and BMRB Core Archives. wwPDB Members are committed to ensuring that structural biology data are Findable, Accessible, Interoperable, and Reusable (FAIR). With accelerating growth of 3D structure depositions to the PDB, the wwPDB data centers anticipate that all possible four-character PDB accession codes (PDB IDs) will be consumed by 2029. In preparation for this milestone, the wwPDB has revised the PDB accession code format by extending its length to 12 characters. The format of extended PDB ID is prefix “pdb_” followed by eight alphanumeric characters, e.g., pdb_10021abc. This process will facilitate more robust text mining detection of PDB IDs across the published literature and allow for more informative and transparent delivery of revised data files. Once all available four-character PDB IDs have been assigned, newly deposited PDB entries will be issued with extended PDB ID codes. Such entries will not be compatible with the legacy PDB format and will be distributed only in Protein Data Bank exchange/macromolecular CIF (PDBx/mmCIF) format. All existing PDB entries bearing legacy (four-character) IDs will also be persistently identified with extended PDB IDs stored in the PDBx/mmCIF atomic coordinate file (e.g., PDB ID “1abc” will have extended ID format as “pdb_00001abc”). Herein, we describe the five-year transition plan to help PDB depositors, users, and scientific journals embrace the extended PDB ID and exclusively rely on PDBx/mmCIF format data files. Resources for supporting the extended PDB ID format will be provided to PDB users in the form of FAQ on PDB ID extension and learning materials/links on PDBx/mmCIF format. Additionally, a PDB “beta” archive will be provided starting from 2026, that will have a file directory structure consistent with the current PDB archive. The PDB “beta” archive will become the PDB main archive once all the four-character PDB accession codes are exhausted, and the current PDB archive will be removed. As the transition phase progresses, additional training resources will be provided. Community-wide adoption of PDBx/mmCIF format is the goal of the wwPDB. Since 2014, this format has been the official master format underpinning both the PDB Core Archive and the wwPDB global OneDep software system for complete deposition, rigorous validation, and expert biocuration of incoming 3D biostructure data. PDBx/mmCIF has many advantages versus the legacy PDB format. It is flexible, fully extensible, both human- and machine-readable, and can accommodate 3D biostructures of any size and composition. This presentation will cover (a) a guide to extended PDB ID and PDBx/mmCIF format for PDB users and scientific journals editors and expert referees; (b) how the PDBx/mmCIF format accommodates extended IDs (PDB IDs, ligand codes), and where and how extended PDB IDs are stored in PDBx/mmCIF data files; (c) how to derive the new extended PDB ID from an old PDB ID for more robust identification of all PDB structures in the scientific literature; (d) the PDB “beta” archive that will be provided in 2026 to support this persistent identifier transition, and concomitant changes in file directory architecture; and (e) example files containing extended PDB IDs that can be accessed during testing of software changes designed to enable their adoption and management. RCSB PDB is funded by the National Science Foundation (DBI-1832184), the US Department of Energy (DE-SC0019749), and the National Cancer Institute, National Institute of Allergy and Infectious Diseases, and National Institute of General Medical Sciences of the National Institutes of Health under grant R01GM133198.
In addition to providing open access to >220,000 experimentally-determined three-dimensional (3D) structures of biological macromolecules stored in the Protein Data Bank archive (PDB), the Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) research-focused web portal (RCSB.org) serves >1 million Computed Structure Models (CSMs, predicted protein structures) generated using AlphaFold2 (from AlphaFold DB) or RoseTTAFold2/AlphaFold2 (from ModelArchive). All PDB structures and CSMs made freely available from RCSB.org are integrated weekly with related functional annotations from trusted external resources, providing up-to-date information for each 3D biostructure. Both PDB structures and CSMs can be searched, analyzed, and visualized with an extensive array of user-friendly RCSB PDB online resources. Provenance and reliability of PDB data and CSMs are clearly identified at RCSB.org. For PDB structures, users can access primary PDB structure quality metrics, including those summarized in the validation slider graphic and explore how structure quality varies within a PDB structure (using Mol* 3D visualization) and across the entire archive. These RCSB.org tools can help identify experimentally-determined structure(s) that are most appropriate for developing testable hypotheses and supporting basic and applied research across the biological and biomedical sciences. CSMs offer great alternatives and/or starting models for data analysis and hypothesis generation when a relevant experimentally-determined structure is not available at the PDB. Simultaneous delivery of PDB data and CSMs provides access to 3D structure information across the proteomes of H. sapiens, various model organisms (i.e., mouse, worm, fly, etc.), select human pathogens, and organisms important in the fight against climate change (e.g., A. thaliana, S. divinum). CSMs do, however, have their limitations. Typically, they are comparable in accuracy to lower-resolution experimental structures and should not be relied on when a corresponding PDB structure(s) is available. Similar to the validation slider for experimentally- determined structures archived in the PDB, all CSMs delivered via RCSB.org report global and local prediction confidence as pLDDT scores. CSM prediction confidence scoring is also displayed in 3D using Mol*. RCSB PDB Core Operations are funded by the National Science Foundation (DBI-2321666), the US Department of Energy (DE- SC0019749), and the National Cancer Institute, National Institute of Allergy and Infectious Diseases, and National Institute of General Medical Sciences of the National Institutes of Health under grant R01GM133198.
The online Molecule of the Month series authored by David S. Goodsell and published by the Research Collaboratory for Structural Biology Protein Data Bank at PDB101.RCSB.org has highlighted stories about the biomolecular structures driving fundamental biology, biomedicine, bioenergy, and biotechnology since January 2000. A new chapter begins in 2025: Janet Iwasa has taken over as the series creator of stories about critically important biological macromolecules in a rapidly changing world.
The Protein Data Bank (PDB) is the global repository for public-domain experimentally determined 3D biomolecular structural information. The archival nature of the PDB presents certain challenges pertaining to updating or adding associated annotations from trusted external biodata resources. While each Worldwide PDB (wwPDB) partner has made best efforts to provide up-to-date external annotations, accessing and integrating information from disparate wwPDB data centers can be an involved process. To address this issue, the wwPDB has established the PDB Next Generation (or NextGen) Archive, developed to centralize and streamline access to enriched structural annotations from wwPDB partners and trusted external sources. At present, the NextGen Archive provides mappings between experimentally determined 3D structures of proteins and UniProt amino acid sequences, domain annotations from Pfam, SCOP2 and CATH databases and intra-molecular connectivity information. Since launch, the PDB NextGen Archive has seen substantial user engagement with over 3.5 million data file downloads, ensuring researchers have access to accurate, up-to-date and easily accessible structural annotations. Database URL: http://www.wwpdb.org/ftp/pdb-nextgen-archive-site.
With the ever‐expanding toolkit of molecular viewers, the ability to visualize macromolecular structures has never been more accessible. Yet, the idiosyncratic technical intricacies across tools and the integration complexities associated with handling structure annotation data present significant barriers to seamless interoperability and steep learning curves for many users. The necessity for reproducible data visualizations is at the forefront of the current challenges. Recently, we introduced MolViewSpec (homepage: https://molstar.org/mol‐view‐spec/, GitHub project: https://github.com/molstar/mol‐view‐spec), a specification approach that defines molecular visualizations, decoupling them from the varying implementation details of different molecular viewers. Through the protocols presented herein, we demonstrate how to use MolViewSpec and its 3D view–building Python library for creating sophisticated, customized 3D views covering all standard molecular visualizations. MolViewSpec supports representations like cartoon and ball‐and‐stick with coloring, labeling, and applying complex transformations such as superposition to any macromolecular structure file in mmCIF, BinaryCIF, and PDB formats. These examples showcase progress towards reusability and interoperability of molecular 3D visualization in an era when handling molecular structures at scale is a timely and pressing matter in structural bioinformatics as well as research and education across the life sciences. © 2024 The Authors. Current Protocols published by Wiley Periodicals LLC. Basic Protocol 1: Creating a MolViewSpec view using the MolViewSpec Python package Basic Protocol 2: Creating a MolViewSpec view with reference to MolViewSpec annotation files Basic Protocol 3: Creating a MolViewSpec view with labels and other advanced features Support Protocol 1: Computing rotation and translation vectors Support Protocol 2: Creating a MolViewSpec annotation file
Molecular origami offers an offline way to explore the 3D structures of biology. Visit PDB101.rcsb.org to download free paper models of DNA, green fluorescent protein, viruses, and more.