We address the problem of predicting high-detail RNA structure geometry from the information available in low-detail experimental maps. Here, low-detail refers to resolutions ≈ 2.5-3.5Å, where the location of the phosphate groups and the glycosidic bonds can be determined from experimental maps but all other backbone atom positions cannot. In contrast, higher-resolution maps allow high-detail determinations of all backbone atomic positions. To this end, we first create a gold standard dataset of highly curated, experimentally supported RNA suites. Second, we develop and employ a modified version of the previously devised algorithm MINT-AGE to learn clusters that are in high correspondence with the gold standard's conformational classes of suites based on 3D RNA structure. Since some of the gold standard classes are of very small size, a new modified version of MINT-AGE is able to also identify very small clusters. Third, we create a new conformer prediction algorithm, RNAprecis, which assigns low-detail structures to newly designed 3D shape coordinates. Our improvements include: (i) learned classes augmented to cover also very low sample sizes and (ii) replacing distances from clusters by Bayesian posterior probabilities. On test data containing suites modeled as conformational outliers, RNAprecis shows good results suggesting that our learning method generalizes well. In particular, we show that the modified MINT-AGE clustering can more finely delineate between previously unseen suite conformer separations. For example, the 0a conformer has been separated into two clusters seen in different structural contexts. Such new distinctions can have implications for biochemical interpretation of RNA structure.
AlphaFold2 protein structure predictions are widely available for structural biology uses. These predictions, especially for eukaryotic proteins, frequently contain extensive regions predicted below the pLDDT 70 level, the rule-of-thumb cutoff for high confidence. This work identifies major modes of behavior within low-pLDDT regions through a survey of human proteome predictions provided by the AlphaFold Protein Structure Database. The near-predictive mode resembles folded protein and can be a nearly accurate prediction. Barbed wire is extremely unproteinlike, being recognized by wide looping coils, an absence of packing contacts, and numerous signature validation outliers, and it likely represents a nonpredicted region. Pseudostructure presents an intermediate behavior with a misleading appearance of isolated and badly formed secondary structure-like elements. These prediction modes are compared with annotations of disorder from MobiDB, showing general correlation between barbed wire/pseudostructure and many measures of disorder, an association between pseudostructure and signal peptides, and an association between near-predictive and regions of conditional folding. To enable users to identify these regions within a prediction, a new Phenix tool is developed encompassing the results of this work, including prediction annotation, visual markup, and residue selection based on these prediction modes. This tool will help users develop expertise in interpreting difficult AlphaFold predictions and identify the near-predictive regions that can aid in molecular replacement when a prediction does not contain enough high-pLDDT regions.
Cis -nonPro peptides, a very rare feature in protein structures, are of considerable importance for two opposite reasons. On one hand, their genuine occurrences are mostly found at sites critical to biological function, from the active sites of carbohydrate enzymes to rare adjacent-residue disulfide bonds. On the other hand, a cis -nonPro can easily be misfit into weak or ambiguous electron density, which has led to a high incidence of unjustified cis -nonPro over the last decade. This paper uses the greatly expanded crystallographic data and newly stringent quality-filtering to identify the genuine occurrences and survey both individual examples and broad patterns of their functionality. The accompanying paper describes the problem of cis -nonPro over-use, including its causes, validation, and correction.We explain the procedure developed to identify genuine cis -nonPro examples with almost no false positives, including the new observation that peptides with a glycine on one side or the other need extra care to avoid mis-assignment as cis -nonPro. We then survey a sample of the varied functional roles and structural contexts of cis -nonPro, emphasizing aspects not previously covered systematically: the preferred occurrence at β-strand ends in TIM barrel structures, the concentration of occurrence in proteins that process, bind, or contain carbohydrates, and the resulting complications in defining a simple occurrence frequency.
AlphaFold2 protein structure predictions are widely available for structural biology uses. These predictions, especially for eukaryotic proteins, frequently contain extensive regions predicted below the pLDDT = 70 level, the rule-of-thumb cutoff for high confidence. This work identifies major modes of behavior within low-pLDDT regions through a survey of human proteome predictions provided by the AlphaFold Protein Structure Database. The near-predictive mode resembles folded protein and can be a nearly accurate prediction. Barbed wire is extremely unprotein-like, being recognized by wide looping coils, an absence of packing contacts and numerous signature validation outliers, and it represents a region where the conformation has no predictive value. Pseudostructure presents an intermediate behavior with a misleading appearance of isolated and badly formed secondary-structure-like elements. These prediction modes are compared with annotations of disorder from MobiDB, showing general correlation between barbed wire/pseudostructure and many measures of disorder, an association between pseudostructure and signal peptides, and an association between near-predictive and regions of conditional folding. To enable users to identify these regions within a prediction, a new Phenix tool is developed encompassing the results of this work, including prediction annotation, visual markup and residue selection based on these prediction modes. This tool will help users develop expertise in interpreting difficult AlphaFold predictions and identify the near-predictive regions that can aid in molecular replacement when a prediction does not contain enough high-pLDDT regions.
It is important to remember that structural biology generates static “models” of protein and nucleic acid structures. These models, often generated by the best of our science, technology, and efforts, can be a very good estimation of the actual structures, but they are not the real structures. In the worst case, they can be quite unlike reality. Over decades of developing software tools for validation of these models, the Richardson lab has encountered countless examples where these models contain local errors, such as unlikely covalent bond lengths and/or angles, non-bonded atoms modelled too closely together resulting in steric clashes, or arrangements of backbone or sidechain atoms in unusual dihedral angle configurations. Many of these examples are “mostly harmless”, while others are decidedly not. This talk will explore examples of common modeling errors, many of which were found after publication and are publicly available in the PDB, meaning they made it through the review process, demonstrating a need for reviewers to have adequate access to data and tools to suitably evaluate whether a structure should be published. This talk also considers a couple of case studies to illustrate larger issues that reviewers of macromolecular structures should watch out for. The first case, described as “a scientific scandal of epic proportions [that] shook macromolecular crystallography to its core”[1], involving a number of structures based on fabricated data, prompted the creation of several validation task forces and the addition of validation reports to the PDB repository. In the second case, on a more positive note, our tools were critical to finding and fixing areas with errors within SARS-CoV-2 protein structures that were publicly released early, allowing peer review even prior to publication[2]. Although improvements in refinement tools and especially the rise of AlphaFold and other AI tools have improved the quality of structures overall, we see cases of overfitting to the validation metrics, meaning there is still a corresponding need for improvement and adoption of new validation tools. Unfortunately, there is no simple answer for these issues; reviewers must rely on their best judgement when evaluating structures, and must therefore remain well-educated on current methods and vigilant toward new challenges.
At last year’s ACA meeting, we presented a high-quality, residue-filtered dataset of RNA residues. This RNA2023 dataset is useful for defining the most common modes of RNA conformations, and for assembling libraries of expected behaviors. However, structural biology does not depend solely on common behaviors. Rare conformations and their justifications are fascinating and crucial to understanding structure and function. Here we present unusual RNA features gleaned from our high-quality RNA2023 dataset. These include: Previously “wannabe” backbone conformations - some of which are supported by the new data and some of which are not, non-standard ribose sugar puckers such as C1’-exo, and residues with unusual, non-rotameric values for the gamma backbone dihedral. These rare features should be identified, appreciated, and not lost to over-regularization when supported by experimental data.
In January 2020, a workshop was held at EMBL-EBI (Hinxton, UK) to discuss data requirements for the deposition and validation of cryoEM structures, with a focus on single-particle analysis. The meeting was attended by 47 experts in data processing, model building and refinement, validation, and archiving of such structures. This report describes the workshop's motivation and history, the topics discussed, and the resulting consensus recommendations. Some challenges for future methods-development efforts in this area are also highlighted, as is the implementation to date of some of the recommendations.
Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
Introduction -------------------------------------------------------------------------------- This is the RNA2023 dataset by the Richardson Lab at Duke University These are high-quality residues from high-quality, low-redundancy RNA chains in the PDB. For a similar set of quality-filtered protein residues, see the top2018 datasets at: https://doi.org/10.5281/zenodo.4626149 https://doi.org/10.5281/zenodo.5115232 Corresponding authors -------------------------------------------------------------------------------- dcrjsr at kinemage.biochem.duke.edu christopher.sci.williams at gmail.com Usage recommendations -------------------------------------------------------------------------------- RNA residues that fail the filtering criteria described below have been removed from the files. As a result, these files can be considered pre-filtered and will return only results for residues of good model quality with supporting experimental data. Files already contain hydrogens added by Reduce in the context of the original full models. Two datasets are provided. The standard dataset is rna2023_pruned. We recommend this version as the default. The RNA backbone conformational space is highly diverse, and some real conformations fall below the statistical threshold for recognition as suites. Therefore we do not recommend excluding suite outliers from the dataset except in specialty cases. We also provide a rna2023_nosuiteout dataset. In this case, no residues having "!!" outlier suite identifications are permitted. This set may be useful in specialist cases where suite outliers are undesireable or where losing some real conformations is an acceptable sacrifice for maximal filtering. Each dataset also has a mmCIF version. Note: Chains are named based on author chain ids, except for 8b0x, chain a. To avoid conflicts with 8b0x chain A in file systems that do not support case-sensitive file names, 8b0x chain a has been renamed to chain AB, matching its PDB/mmCIF designation. Additional files -------------------------------------------------------------------------------- rna2023_pdbmetadata.csv contains information on release date, resolution, title, authors, etc for each source pdb. rna2023_chain_list contains a list of all included chains, plus statistics on the number residues from the original chain passed the quality filters. rna2023_suitename_table.csv and rna2023_suitename_table_nosuiteout.csv contain suitename identifications of rotameric RNA backbone conformations (1a, 1c, 2u, 6d, etc) precomputed for convenience. Filtering criteria: Chain level -------------------------------------------------------------------------------- The chain list was derived from http://rna.bgsu.edu/rna3dhub/nrlist, version 3.150 as of 2020/10/28, with a 1.9Å resolution cutoff. We added 6ugg chain A and two recent EM ribosome structures: 8a3d and 8b0x After residue-level filtering, chains with no complete suites were removed. Filtering criteria: Residue level -------------------------------------------------------------------------------- Even excellent structures usually contain some poorly-resolved regions. Residue-level filtering helps avoid including these regions in otherwise high-quality data Residues are required to meet the following validation quality contain: No sugar pucker outliers No steric overlaps or "clashes", as per Probe >= 0.5Å No covalent bond or angle geometry outliers Optionally, no !! suite outliers Residues from xray structures are required for meet the following fit-to-map criteria: Average of worst 2 atoms' 2Fo-Fc map values >= 1.2 Average of worst 2 atoms' RSCC scores >= 0.7 No atoms modeled at partial occupancy Residues from em structures are required for meet the following fit-to-map criteria: RSCC >= 0.7 Residue inclusion fraction = 1.0 or >= 0.95, depending on structure No atoms modeled at partial occupancy Filtering is documented in each pruned file. See USER DOC lines in .pdb and data_rna2023_dataset loops in .cif Version history -------------------------------------------------------------------------------- Version 1.0 Jun 30, 2023 Initial version
Experimental structure determination can be accelerated with AI-based structure prediction methods such as AlphaFold. Here we present an automatic procedure requiring only sequence information and crystallographic data that uses AlphaFold predictions to produce an electron density map and a structural model. Iterating through cycles of structure prediction is a key element of our procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. We applied this procedure to X-ray data for 215 structures released by the Protein Data Bank in a recent 6-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2Å. Predictions from our iterative template-guided prediction procedure were more accurate than those obtained without templates. We suggest a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization.
Model building and refinement, and the validation of their correctness, are very effective and reliable at local resolutions better than about 2.5 Å for both crystallography and cryo-EM. However, at local resolutions worse than 2.5 Å both the procedures and their validation break down and do not ensure reliably correct models. This is because in the broad density at lower resolution, critical features such as protein backbone carbonyl O atoms are not just less accurate but are not seen at all, and so peptide orientations are frequently wrongly fitted by 90-180°. This puts both backbone and side chains into the wrong local energy minimum, and they are then worsened rather than improved by further refinement into a valid but incorrect rotamer or Ramachandran region. On the positive side, new tools are being developed to locate this type of pernicious error in PDB depositions, such as CaBLAM, EMRinger, Pperp diagnosis of ribose puckers, and peptide flips in PDB-REDO, while interactive modeling in Coot or ISOLDE can help to fix many of them. Another positive trend is that artificial intelligence predictions such as those made by AlphaFold2 contribute additional evidence from large multiple sequence alignments, and in high-confidence parts they provide quite good starting models for loops, termini or whole domains with otherwise ambiguous density.
Machine-learning prediction algorithms such as AlphaFold and RoseTTAFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including new experimental information such as a density map, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt on the basis of experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We show that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for interpretation of crystallographic and electron cryo-microscopy maps.
Machine learning prediction algorithms such as AlphaFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including experimental information, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt based on experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We find that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for crystallographic and electron cryo-microscopy map interpretation.AlphaFold modeling can be improved synergistically by including information from experimental density maps.
This introduction to the session will be a reminder of some of the changes with resolution that we probably all know about. At what resolution can we see holes in rings? When do the carbonyl groups fade into the backbone? When do nucleic acids shift from connected base pairs to connected density along the base-stacking direction? What feature visibilities differ between crystallography and cryo-EM? And a reminder that it's the local resolution/disorder that matters, not the overall resolution. The image shows two specific regions in T4 lysozyme, the only protein in the PDB with a deposited structure at 1A, 2Å, 3Å, and 4Å.
We have curated a high-quality, "best-parts" reference dataset of about 3 million protein residues in about 15,000 PDB-format coordinate files, each containing only residues with good electron density support for a physically acceptable model conformation. The resulting prefiltered data typically contain the entire core of each chain, in quite long continuous fragments. Each reference file is a single protein chain, and the total set of files were selected for low redundancy, high resolution, good MolProbity score, and other chain-level criteria. Then each residue was critically tested for adequate local map quality to firmly support its conformation, which must also be free of serious clashes or covalent-geometry outliers. The resulting Top2018 prefiltered datasets have been released on the Zenodo online web service and are freely available for all uses under a Creative Commons license. Currently, one dataset is residue filtered on main chain plus Cβ atoms, and a second dataset is full-residue filtered; each is available at four different sequence-identity levels. Here, we illustrate both statistics and examples that show the beneficial consequences of residue-level filtering. That process is necessary because even the best of structures contain a few highly disordered local regions with poor density and low-confidence conformations that should not be included in reference data. Therefore, the open distribution of these very large, prefiltered reference datasets constitutes a notable advance for structural bioinformatics and the fields that depend upon it.
This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.
Ever since the first structures of proteins were determined in the 1960s, structural biologists have required methods to visualize biomolecular structures, both as an essential tool for their research and also to promote 3D comprehension of structural results by a wide audience of researchers, students, and the general public. In this review to celebrate the 50th anniversary of the Protein Data Bank, we present our own experiences in developing and applying methods of visualization and analysis to the ever-expanding archive of protein and nucleic acid structures in the worldwide Protein Data Bank. Across that timespan, Jane and David Richardson have concentrated on the organization inside and between the macromolecules, with ribbons to show the overall backbone "fold" and contact dots to show how the all-atom details fit together locally. David Goodsell has explored surface-based representations to present and explore biological subjects that range from molecules to cells. This review concludes with some ideas about the current challenges being addressed by the field of biomolecular visualization.
This work builds upon the record-breaking speed and generous immediate release of new experimental three-dimensional structures of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) proteins and complexes, which are crucial to downstream vaccine and drug development. We have surveyed those structures to catch the occasional errors that could be significant for those important uses and for which we were able to provide demonstrably higher-accuracy corrections. This process relied on new validation and correction methods such as CaBLAM and ISOLDE, which are not yet in routine use. We found such important and correctable problems in seven early SARS-CoV-2 structures. Two of the structures were soon superseded by new higher-resolution data, confirming our proposed changes. For the other five, we emailed the depositors a documented and illustrated report and encouraged them to make the model corrections themselves and use the new option at the worldwide Protein Data Bank for depositors to re-version their coordinates without changing the Protein Data Bank code. This quickly and easily makes the better-accuracy coordinates available to anyone who examines or downloads their structure, even before formal publication. The changes have involved sequence misalignments, incorrect RNA conformations near a bound inhibitor, incorrect metal ligands, and cis-trans or peptide flips that prevent good contact at interaction sites. These improvements have propagated into nearly all related structures done afterward. This process constitutes a new form of highly rigorous peer review, which is actually faster and more strict than standard publication review because it has access to coordinates and maps; journal peer review would also be strengthened by such access.
The refinement of biomolecular crystallographic models relies on geometric restraints to help to address the paucity of experimental data typical in these experiments. Limitations in these restraints can degrade the quality of the resulting atomic models. Here, an integration of the full all-atom Amber molecular-dynamics force field into Phenix crystallographic refinement is presented, which enables more complete modeling of biomolecular chemistry. The advantages of the force field include a carefully derived set of torsion-angle potentials, an extensive and flexible set of atom types, Lennard-Jones treatment of nonbonded interactions and a full treatment of crystalline electrostatics. The new combined method was tested against conventional geometry restraints for over 22 000 protein structures. Structures refined with the new method show substantially improved model quality. On average, Ramachandran and rotamer scores are somewhat better, clashscores and MolProbity scores are significantly improved, and the modeling of electrostatics leads to structures that exhibit more, and more correct, hydrogen bonds than those refined using traditional geometry restraints. In general it is found that model improvements are greatest at lower resolutions, prompting plans to add the Amber target function to real-space refinement for use in electron cryo-microscopy. This work opens the door to the future development of more advanced applications such as Amber-based ensemble refinement, quantum-mechanical representation of active sites and improved geometric restraints for simulated annealing.