The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB, RCSB.org), the US Worldwide Protein Data Bank (wwPDB, wwPDB.org) data center for the global PDB archive, provides access to the PDB data via its RCSB.org research-focused web portal. We report substantial additions to the tools and visualization features available at RCSB.org, which now delivers more than 227000 experimentally determined atomic-level three-dimensional (3D) biostructures stored in the global PDB archive alongside more than 1 million Computed Structure Models (CSMs) of proteins (including models for human, model organisms, select human pathogens, crop plants and organisms important for addressing climate change). In addition to providing support for 3D structure motif searches with user-provided coordinates, new features highlighted herein include query results organized by redundancy-reduced Groups and summary pages that facilitate exploration of groups of similar proteins. Newly released programmatic tools are also described, as are enhanced training opportunities.
Abstract The Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB), founding member of the Worldwide Protein Data Bank (wwPDB), is the US data center for the open-access PDB archive. As wwPDB-designated Archive Keeper, RCSB PDB is also responsible for PDB data security. Annually, RCSB PDB serves >10 000 depositors of three-dimensional (3D) biostructures working on all permanently inhabited continents. RCSB PDB delivers data from its research-focused RCSB.org web portal to many millions of PDB data consumers based in virtually every United Nations-recognized country, territory, etc. This Database Issue contribution describes upgrades to the research-focused RCSB.org web portal that created a one-stop-shop for open access to ∼200 000 experimentally-determined PDB structures of biological macromolecules alongside >1 000 000 incorporated Computed Structure Models (CSMs) predicted using artificial intelligence/machine learning methods. RCSB.org is a ‘living data resource.’ Every PDB structure and CSM is integrated weekly with related functional annotations from external biodata resources, providing up-to-date information for the entire corpus of 3D biostructure data freely available from RCSB.org with no usage limitations. Within RCSB.org, PDB structures and the CSMs are clearly identified as to their provenance and reliability. Both are fully searchable, and can be analyzed and visualized using the full complement of RCSB.org web portal capabilities.
Biology is advanced by producing structural models of biological systems, such as protein complexes. Some systems are recalcitrant to traditional structure determination methods. In such cases, it may still be possible to produce useful models by integrative structure determination that depends on simultaneous use of multiple types of data. An ensemble of models that are sufficiently consistent with the data is produced by a structural sampling method guided by a data-dependent scoring function. The variation in the ensemble of models quantified the uncertainty of the structure, generally resulting from the uncertainty in the input information and actual structural heterogeneity in the samples used to produce the data. Here, we describe how to generate, assess, and interpret ensembles of integrative structural models using our open source Integrative Modeling Platform program (https://integrativemodeling.org).
This repository contains the input experimental data used in a tutorial on modeling of RNA Polymerase III and the largest cluster of output models.
Integrative structure modeling provides 3D models of macromolecular systems that are based on information from multiple types of experiments, physical principles, statistical inferences, and prior structural models. Here, we provide a hands-on realistic example of integrative structure modeling of the quaternary structure of the actin, tropomyosin, and gelsolin protein assembly based on electron microscopy, solution X-ray scattering, and chemical crosslinking data for the complex as well as excluded volume, sequence connectivity, and rigid atomic X-ray structures of the individual subunits. We follow the general four-stage process for integrative modeling, including gathering the input information, converting the input information into a representation of the system and a scoring function, sampling alternative model configurations guided by the scoring function, and analyzing the results. The computational aspects of this approach are implemented in our open-source Integrative Modeling Platform (IMP), a comprehensive and extensible software package for integrative modeling ( https://integrativemodeling.org ). In particular, we rely on the Python Modeling Interface (PMI) module of IMP that provides facile mixing and matching of macromolecular representations, restraints based on different types of information, sampling algorithms, and analysis including validations of the input data and output models. Finally, we also outline how to deposit an integrative structure and corresponding experimental data into PDB-Dev, the nascent worldwide Protein Data Bank (wwPDB) resource for archiving and disseminating integrative structures ( https://pdb-dev.wwpdb.org ). The example application provides a starting point for a user interested in using IMP for integrative modeling of other biomolecular systems.
Small-angle X-ray scattering (SAXS) is an experimental technique that allows structural information on biomolecules in solution to be gathered. High-quality SAXS profiles have typically been obtained by manual merging of scattering profiles from different concentrations and exposure times. This procedure is very subjective and results vary from user to user. Up to now, no robust automatic procedure has been published to perform this step, preventing the application of SAXS to high-throughput projects. Here, SAXS Merge, a fully automated statistical method for merging SAXS profiles using Gaussian processes, is presented. This method requires only the buffer-subtracted SAXS profiles in a specific order. At the heart of its formulation is non-linear interpolation using Gaussian processes, which provides a statement of the problem that accounts for correlation in the data.
MOTIVATION:Statistical potentials have been widely used for modeling whole proteins and their parts (e.g. sidechains and loops) as well as interactions between proteins, nucleic acids and small molecules. Here, we formulate the statistical potentials entirely within a statistical framework, avoiding questionable statistical mechanical assumptions and approximations, including a definition of the reference state.RESULTS:We derive a general Bayesian framework for inferring statistically optimized atomic potentials (SOAP) in which the reference state is replaced with data-driven 'recovery' functions. Moreover, we restrain the relative orientation between two covalent bonds instead of a simple distance between two atoms, in an effort to capture orientation-dependent interactions such as hydrogen bonds. To demonstrate this general approach, we computed statistical potentials for protein-protein docking (SOAP-PP) and loop modeling (SOAP-Loop). For docking, a near-native model is within the top 10 scoring models in 40% of the PatchDock benchmark cases, compared with 23 and 27% for the state-of-the-art ZDOCK and FireDock scoring functions, respectively. Similarly, for modeling 12-residue loops in the PLOP benchmark, the average main-chain root mean square deviation of the best scored conformations by SOAP-Loop is 1.5 Å, close to the average root mean square deviation of the best sampled conformations (1.2 Å) and significantly better than that selected by Rosetta (2.1 Å), DFIRE (2.3 Å), DOPE (2.5 Å) and PLOP scoring functions (3.0 Å). Our Bayesian framework may also result in more accurate statistical potentials for additional modeling applications, thus affording better leverage of the experimentally determined protein structures.AVAILABILITY AND IMPLEMENTATION:SOAP-PP and SOAP-Loop are available as part of MODELLER (http://salilab.org/modeller).
Protein structures evolved through a complex interplay of cooperative interactions, and it is still very challenging to design new protein folds de novo. Here we present a strategy to design self-assembling polypeptide nanostructured polyhedra based on modularization using orthogonal dimerizing segments. We designed and experimentally demonstrated the formation of the tetrahedron that self-assembles from a single polypeptide chain comprising 12 concatenated coiled coil-forming segments separated by flexible peptide hinges. The path of the polypeptide chain is guided by a defined order of segments that traverse each of the six edges of the tetrahedron exactly twice, forming coiled-coil dimers with their corresponding partners. The coincidence of the polypeptide termini in the same vertex is demonstrated by reconstituting a split fluorescent protein in the polypeptide with the correct tetrahedral topology. Polypeptides with a deleted or scrambled segment order fail to self-assemble correctly. This design platform provides a foundation for constructing new topological polypeptide folds based on the set of orthogonal interacting polypeptide segments.
A set of software tools for building and distributing models of macromolecular assemblies uses an integrative structure modeling approach, which casts the building of models as a computational optimization problem where information is encoded into a scoring function used to evaluate candidate models.
A set of software tools for building and distributing models of macromolecular assemblies uses an integrative structure modeling approach, which casts the building of models as a computational optimization problem where information is encoded into a scoring function used to evaluate candidate models.
Advances in electron microscopy (EM) allow for structure determination of large biological assemblies at increasingly higher resolutions. A key step in this process is fitting multiple component structures into an EM-derived density map of their assembly. Here, we describe a web server for this task. The server takes as input a set of protein structures in the PDB format and an EM density map in the MRC format. The output is an ensemble of models ranked by their quality of fit to the density map. The models can be viewed online or downloaded from the website. The service is available at; http://salilab.org/multifit/ and http://bioinfo3d.cs.tau.ac.il/.
Structural modeling of macromolecular complexes greatly benefits from interactive visualization capabilities. Here we present the integration of several modeling tools into UCSF Chimera. These include comparative modeling by MODELLER, simultaneous fitting of multiple components into electron microscopy density maps by IMP MultiFit, computing of small-angle X-ray scattering profiles and fitting of the corresponding experimental profile by IMP FoXS, and assessment of amino acid sidechain conformations based on rotamer probabilities and local interactions by Chimera.
Proteomics techniques have been used to generate comprehensive lists of protein interactions in a number of species. However, relatively little is known about how these interactions result in functional multiprotein complexes. This gap can be bridged by combining data from proteomics experiments with data from established structure determination techniques. Correspondingly, integrative computational methods are being developed to provide descriptions of protein complexes at varying levels of accuracy and resolution, ranging from complex compositions to detailed atomic structures.
For many macromolecular assemblies, both a cryo-electron microscopy map and atomic structures of its component proteins are available. Here we describe a method for fitting and refining a component structure within its map at intermediate resolution (<15 Å). The atomic positions are optimized with respect to a scoring function that includes the crosscorrelation coefficient between the structure and the map as well as stereochemical and nonbonded interaction terms. A heuristic optimization that relies on a Monte Carlo search, a conjugate-gradients minimization, and simulated annealing molecular dynamics is applied to a series of subdivisions of the structure into progressively smaller rigid bodies. The method was tested on 15 proteins of known structure with 13 simulated maps and 3 experimentally determined maps. At ∼10 Å resolution, Cα rmsd between the initial and final structures was reduced on average by ∼53%. The method is automated and can refine both experimental and predicted atomic structures.
Functional characterization of a protein sequence is one of the most frequent problems in biology. This task is usually facilitated by accurate three-dimensional (3-D) structure of the studied protein. In the absence of an experimentally determined structure, comparative or homology modeling can sometimes provide a useful 3-D model for a protein that is related to at least one known protein structure. Comparative modeling predicts the 3-D structure of a given protein sequence (target) based primarily on its alignment to one or more proteins of known structure (templates). The prediction process consists of fold assignment, target-template alignment, model building, and model evaluation. This unit describes how to calculate comparative models using the program MODELLER and discusses all four steps of comparative modeling, frequently observed errors, and some applications. Modeling lactate dehydrogenase from Trichomonas vaginalis (TvLDH) is described as an example. The download and installation of the MODELLER software is also described.
QM/MM methods were used to study the isomerization step from (2R)-methylmalonyl-CoA to succinyl-CoA. A pathway via a "fragmentation-recombination" mechanism is ruled out on energetic grounds. For the other radicalic pathway, involving an addition recombination step, geometries and vibrational contributions have been determined, and a barrier height of 11.70 kcal/mol was found. The effect of adjacent hydrogen-donating groups was found to reduce the energy barrier by 1-2 kcal/mol each and thus to provide a significant catalytic effect for this reaction. By means of molecular dynamics studies, the stereochemistry of the methylmalonyl-CoA mutase catalyzed reaction was examined. It is shown that TYR89 is essential for maintaining stereoselectivity of the abstraction of a hydrogen in the backreaction. The subsequent selective formation of one isomer of methylmalonyl-CoA is probably due to the presence of a bulky side chain.