AlphaFold2 protein structure predictions are widely available for structural biology uses. These predictions, especially for eukaryotic proteins, frequently contain extensive regions predicted below the pLDDT 70 level, the rule-of-thumb cutoff for high confidence. This work identifies major modes of behavior within low-pLDDT regions through a survey of human proteome predictions provided by the AlphaFold Protein Structure Database. The near-predictive mode resembles folded protein and can be a nearly accurate prediction. Barbed wire is extremely unproteinlike, being recognized by wide looping coils, an absence of packing contacts, and numerous signature validation outliers, and it likely represents a nonpredicted region. Pseudostructure presents an intermediate behavior with a misleading appearance of isolated and badly formed secondary structure-like elements. These prediction modes are compared with annotations of disorder from MobiDB, showing general correlation between barbed wire/pseudostructure and many measures of disorder, an association between pseudostructure and signal peptides, and an association between near-predictive and regions of conditional folding. To enable users to identify these regions within a prediction, a new Phenix tool is developed encompassing the results of this work, including prediction annotation, visual markup, and residue selection based on these prediction modes. This tool will help users develop expertise in interpreting difficult AlphaFold predictions and identify the near-predictive regions that can aid in molecular replacement when a prediction does not contain enough high-pLDDT regions.
Cis -nonPro peptides, a very rare feature in protein structures, are of considerable importance for two opposite reasons. On one hand, their genuine occurrences are mostly found at sites critical to biological function, from the active sites of carbohydrate enzymes to rare adjacent-residue disulfide bonds. On the other hand, a cis -nonPro can easily be misfit into weak or ambiguous electron density, which has led to a high incidence of unjustified cis -nonPro over the last decade. This paper uses the greatly expanded crystallographic data and newly stringent quality-filtering to identify the genuine occurrences and survey both individual examples and broad patterns of their functionality. The accompanying paper describes the problem of cis -nonPro over-use, including its causes, validation, and correction.We explain the procedure developed to identify genuine cis -nonPro examples with almost no false positives, including the new observation that peptides with a glycine on one side or the other need extra care to avoid mis-assignment as cis -nonPro. We then survey a sample of the varied functional roles and structural contexts of cis -nonPro, emphasizing aspects not previously covered systematically: the preferred occurrence at β-strand ends in TIM barrel structures, the concentration of occurrence in proteins that process, bind, or contain carbohydrates, and the resulting complications in defining a simple occurrence frequency.
AlphaFold2 protein structure predictions are widely available for structural biology uses. These predictions, especially for eukaryotic proteins, frequently contain extensive regions predicted below the pLDDT = 70 level, the rule-of-thumb cutoff for high confidence. This work identifies major modes of behavior within low-pLDDT regions through a survey of human proteome predictions provided by the AlphaFold Protein Structure Database. The near-predictive mode resembles folded protein and can be a nearly accurate prediction. Barbed wire is extremely unprotein-like, being recognized by wide looping coils, an absence of packing contacts and numerous signature validation outliers, and it represents a region where the conformation has no predictive value. Pseudostructure presents an intermediate behavior with a misleading appearance of isolated and badly formed secondary-structure-like elements. These prediction modes are compared with annotations of disorder from MobiDB, showing general correlation between barbed wire/pseudostructure and many measures of disorder, an association between pseudostructure and signal peptides, and an association between near-predictive and regions of conditional folding. To enable users to identify these regions within a prediction, a new Phenix tool is developed encompassing the results of this work, including prediction annotation, visual markup and residue selection based on these prediction modes. This tool will help users develop expertise in interpreting difficult AlphaFold predictions and identify the near-predictive regions that can aid in molecular replacement when a prediction does not contain enough high-pLDDT regions.
Model building and refinement, and the validation of their correctness, are very effective and reliable at local resolutions better than about 2.5 Å for both crystallography and cryo-EM. However, at local resolutions worse than 2.5 Å both the procedures and their validation break down and do not ensure reliably correct models. This is because in the broad density at lower resolution, critical features such as protein backbone carbonyl O atoms are not just less accurate but are not seen at all, and so peptide orientations are frequently wrongly fitted by 90-180°. This puts both backbone and side chains into the wrong local energy minimum, and they are then worsened rather than improved by further refinement into a valid but incorrect rotamer or Ramachandran region. On the positive side, new tools are being developed to locate this type of pernicious error in PDB depositions, such as CaBLAM, EMRinger, Pperp diagnosis of ribose puckers, and peptide flips in PDB-REDO, while interactive modeling in Coot or ISOLDE can help to fix many of them. Another positive trend is that artificial intelligence predictions such as those made by AlphaFold2 contribute additional evidence from large multiple sequence alignments, and in high-confidence parts they provide quite good starting models for loops, termini or whole domains with otherwise ambiguous density.
This introduction to the session will be a reminder of some of the changes with resolution that we probably all know about. At what resolution can we see holes in rings? When do the carbonyl groups fade into the backbone? When do nucleic acids shift from connected base pairs to connected density along the base-stacking direction? What feature visibilities differ between crystallography and cryo-EM? And a reminder that it's the local resolution/disorder that matters, not the overall resolution. The image shows two specific regions in T4 lysozyme, the only protein in the PDB with a deposited structure at 1A, 2Å, 3Å, and 4Å.
We have curated a high-quality, "best-parts" reference dataset of about 3 million protein residues in about 15,000 PDB-format coordinate files, each containing only residues with good electron density support for a physically acceptable model conformation. The resulting prefiltered data typically contain the entire core of each chain, in quite long continuous fragments. Each reference file is a single protein chain, and the total set of files were selected for low redundancy, high resolution, good MolProbity score, and other chain-level criteria. Then each residue was critically tested for adequate local map quality to firmly support its conformation, which must also be free of serious clashes or covalent-geometry outliers. The resulting Top2018 prefiltered datasets have been released on the Zenodo online web service and are freely available for all uses under a Creative Commons license. Currently, one dataset is residue filtered on main chain plus Cβ atoms, and a second dataset is full-residue filtered; each is available at four different sequence-identity levels. Here, we illustrate both statistics and examples that show the beneficial consequences of residue-level filtering. That process is necessary because even the best of structures contain a few highly disordered local regions with poor density and low-confidence conformations that should not be included in reference data. Therefore, the open distribution of these very large, prefiltered reference datasets constitutes a notable advance for structural bioinformatics and the fields that depend upon it.
Ever since the first structures of proteins were determined in the 1960s, structural biologists have required methods to visualize biomolecular structures, both as an essential tool for their research and also to promote 3D comprehension of structural results by a wide audience of researchers, students, and the general public. In this review to celebrate the 50th anniversary of the Protein Data Bank, we present our own experiences in developing and applying methods of visualization and analysis to the ever-expanding archive of protein and nucleic acid structures in the worldwide Protein Data Bank. Across that timespan, Jane and David Richardson have concentrated on the organization inside and between the macromolecules, with ribbons to show the overall backbone "fold" and contact dots to show how the all-atom details fit together locally. David Goodsell has explored surface-based representations to present and explore biological subjects that range from molecules to cells. This review concludes with some ideas about the current challenges being addressed by the field of biomolecular visualization.
This work builds upon the record-breaking speed and generous immediate release of new experimental three-dimensional structures of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) proteins and complexes, which are crucial to downstream vaccine and drug development. We have surveyed those structures to catch the occasional errors that could be significant for those important uses and for which we were able to provide demonstrably higher-accuracy corrections. This process relied on new validation and correction methods such as CaBLAM and ISOLDE, which are not yet in routine use. We found such important and correctable problems in seven early SARS-CoV-2 structures. Two of the structures were soon superseded by new higher-resolution data, confirming our proposed changes. For the other five, we emailed the depositors a documented and illustrated report and encouraged them to make the model corrections themselves and use the new option at the worldwide Protein Data Bank for depositors to re-version their coordinates without changing the Protein Data Bank code. This quickly and easily makes the better-accuracy coordinates available to anyone who examines or downloads their structure, even before formal publication. The changes have involved sequence misalignments, incorrect RNA conformations near a bound inhibitor, incorrect metal ligands, and cis-trans or peptide flips that prevent good contact at interaction sites. These improvements have propagated into nearly all related structures done afterward. This process constitutes a new form of highly rigorous peer review, which is actually faster and more strict than standard publication review because it has access to coordinates and maps; journal peer review would also be strengthened by such access.
The MolProbity web service provides macromolecular model validation to help correct local errors, for the structural biology community worldwide. Here we highlight new validation features, and also describe how we are fighting back against outside developments which compromise that mission. Our new tool called UnDowser analyzes the properties and context of clashing HOH “waters” to diagnose what they might actually represent; a dozen distinct scenarios are illustrated and described. We now treat alternate conformations more thoroughly, and switching to the Neo4j database (graphical rather than relational) enables cleaner, more comprehensive, and much larger reference datasets. A problematic outside change is that refinement software now increasingly restrains traditional validation criteria (geometry, clashes, rotamers, and even Ramachandran) in order to supplement the sparser experimental data at 3–4 Å resolutions typical of modern cryoEM. But unfortunately the broad density allows model optimization without fixing underlying problems, which means these structures often score much better on validation than they really are. CaBLAM, our tool designed for evaluating peptide orientations at lower resolutions, was described in the previous Tools issue, and here we demonstrate its effectiveness in diagnosing local errors even when other validation outliers have been artificially removed. Sophisticated hacking of the MolProbity server has required continual monitoring and various security measures short of restricting user access. The deprecation of Java applets now prevents KiNG interactive online display of outliers on the 3D model during a MolProbity run, but that important functionality has now been recaptured with a modified version of the Javascript NGL Viewer.
Diffraction (X-ray, neutron and electron) and electron cryo-microscopy are powerful methods to determine three-dimensional macromolecular structures, which are required to understand biological processes and to develop new therapeutics against diseases. The overall structure-solution workflow is similar for these techniques, but nuances exist because the properties of the reduced experimental data are different. Software tools for structure determination should therefore be tailored for each method. Phenix is a comprehensive software package for macromolecular structure determination that handles data from any of these techniques. Tasks performed with Phenix include data-quality assessment, map improvement, model building, the validation/rebuilding/refinement cycle and deposition. Each tool caters to the type of experimental data. The design of Phenix emphasizes the automation of procedures, where possible, to minimize repetitive and time-consuming manual tasks, while default parameters are chosen to encourage best practice. A graphical user interface provides access to many command-line features of Phenix and streamlines the transition between programs, project tracking and re-running of previous tasks.
We find that the overall quite good methods used in the CryoEM Model Challenge could still benefit greatly from several strategies for improving local conformations. Our assessments primarily use validation criteria from the MolProbity web service. Those criteria include MolProbity's all-atom contact analysis, updated versions of standard conformational validations for protein and RNA, plus two recent additions: first, flags for cis-nonPro and twisted peptides, and second, the CaBLAM system for diagnosing secondary structure, validating Ca backbone, and validating adjacent peptide CO orientations in the context of the Ca trace. In general, automated ab initio building of starting models is quite good at backbone connectivity but often fails at local conformation or sequence register, especially at poorer than 3.5 angstrom resolution. However, we show that even if criteria (such as Ramachandran or rotamer) are explicitly restrained to improve refinement behavior and overall validation scores, automated optimization of a deposited structure seldom corrects specific misfittings that start in the wrong local minimum, but just hides them. Therefore, local problems should be identified, and as many as possible corrected, before starting refinement. Secondary structures are confusing at 3-4 angstrom but can be better recognized at 6-8 angstrom. In future model challenges, specific steps being tested (such as segmentation) and the required documentation (such as PDB code of starting model) should each be explicitly defined, so competing methods on a given task can be meaningfully compared. Individual local examples are presented here, to understand what local mistakes and corrections look like in 3D, how they probably arise, and what possible improvements to methodology might help avoid them. At these resolutions, both structural biologists and end-users need meaningful estimates of local uncertainty, perhaps through explicit ensembles. Fitting problems can best be diagnosed by validation that spans multiple residues; CaBLAM is such a multi-residue tool, and its effectiveness is demonstrated.
Traditionally, validation was considered to be a final gatekeeping function, but refinement is smoother and results are better if model validation actively guides corrections throughout structure solution. This shifts emphasis from global to local measures: primarily geometry, conformations and sterics. A fit into the wrong local minimum conformation usually produces outliers in multiple measures. Moving to the right local minimum should be prioritized, rather than small shifts across arbitrary borderlines. Steric criteria work best with all explicit H atoms. `Backrub' motions should be used for side chains and `P-perp' diagnostics to correct ribose puckers. A `water' may actually be an ion, a relic of misfitting or an unmodeled alternate. Beware of wishful thinking in modeling ligands. At high resolution, internally consistent alternate conformations should be modeled and geometry in poor density should not be downweighted. At low resolution, CaBLAM should be used to diagnose protein secondary structure and ERRASER to correct RNA backbone. All atoms should not be forced inside density, beware of sequence misalignment, and very rare conformations such as cis-non-Pro peptides should be avoided. Automation continues to improve, but the crystallographer still must look at each outlier, in the context of density, and correct most of them. For the valid few with unambiguous density and something that is holding them in place, a functional reason should be sought. The expectation is a few outliers, not zero.
Hoogsteen base pairs are seen in DNA crystal structures, but only rarely. This study tests whether Hoogsteens or other syn purines are either under-modeled or over-modeled, which are known problems for rare conformations. Candidate purines needing a syn/anti 180° flip were identified by diagnostic patterns of difference electron-density peaks. Manual inspection narrowed 105 flip candidates to 20 convincing cases, all at ≤2.7 Å resolution. Rebuilding and refinement confirmed that 14 of these were authentic purine flips. Seven examples are modeled as Watson-Crick base pairs but should be Hoogsteens (commonest at duplex termini), and three had the opposite issue. Syn/anti flips were also needed for some single-stranded purines. Five of the 20 convincing cases arose from an unmodeled alternate duplex running in the opposite direction. These are in semi-palindromic DNA sequences bound by a homodimeric protein and show flipped-purine-like difference peaks at residues where the palindrome is imperfect. This study documents types of incorrect modeling which are worth avoiding. However, the primary conclusions are that such mistakes are infrequent, the bias towards fitting anti purines is very slight, and the occurrence rate of Hoogsteen base pairs in DNA crystal structures remains unchanged from earlier estimates at ∼0.3%.
Vicinal disulfides between sequence-adjacent cysteine residues are very rare and rather startling structural features which play a variety of functional roles. Typically discussed as an isolated curiosity, they have never received a general treatment covering both cis and trans forms. Enabled by the growing database of high-resolution structures, required deposition of diffraction data, and improved methods for discriminating reliable from dubious cases, we identify and describe distinct protein families with reliably genuine examples of cis or trans vicinal disulfides and discuss their conformations, conservation, and functions. No cis-trans interconversions and only one case of catalytic redox function are seen. Some vicinal disulfides are essential to large, functionally coupled motions, whereas most form the centers of tightly packed internal regions. Their most widespread biological role is providing a rigid hydrophobic contact surface under the undecorated side of a sugar or multiring ligand, contributing an important aspect of binding specificity.
This paper describes the current update on macromolecular model validation services that are provided at the MolProbity website, emphasizing changes and additions since the previous review in 2010. There have been many infrastructure improvements, including rewrite of previous Java utilities to now use existing or newly written Python utilities in the open-source CCTBX portion of the Phenix software system. This improves long-term maintainability and enhances the thorough integration of MolProbity-style validation within Phenix. There is now a complete MolProbity mirror site at http://molprobity.manchester.ac.uk. GitHub serves our open-source code, reference datasets, and the resulting multi-dimensional distributions that define most validation criteria. Coordinate output after Asn/Gln/His "flip" correction is now more idealized, since the post-refinement step has apparently often been skipped in the past. Two distinct sets of heavy-atom-to-hydrogen distances and accompanying van der Waals radii have been researched and improved in accuracy, one for the electron-cloud-center positions suitable for X-ray crystallography and one for nuclear positions. New validations include messages at input about problem-causing format irregularities, updates of Ramachandran and rotamer criteria from the million quality-filtered residues in a new reference dataset, the CaBLAM Cα-CO virtual-angle analysis of backbone and secondary structure for cryoEM or low-resolution X-ray, and flagging of the very rare cis-nonProline and twisted peptides which have recently been greatly overused. Due to wide application of MolProbity validation and corrections by the research community, in Phenix, and at the worldwide Protein Data Bank, newly deposited structures have continued to improve greatly as measured by MolProbity's unique all-atom clashscore.
Here we describe the updated MolProbity rotamer-library distributions derived from an order-of-magnitude larger and more stringently quality-filtered dataset of about 8000 (vs. 500) protein chains, and we explain the resulting changes and improvements to model validation as seen by users. To include only side-chains with satisfactory justification for their given conformation, we added residue-specific filters for electron-density value and model-to-density fit. The combined new protocol retains a million residues of data, while cleaning up false-positive noise in the multi- datapoint distributions. It enables unambiguous characterization of conformational clusters nearly 1000-fold less frequent than the most common ones. We describe examples of local interactions that favor these rare conformations, including the role of authentic covalent bond-angle deviations in enabling presumably strained side-chain conformations. Further, along with favored and outlier, an allowed category (0.3–2.0% occurrence in reference data) has been added, analogous to Ramachandran validation categories. The new rotamer distributions are used for current rotamer validation in MolProbity and PHENIX, and for rotamer choice in PHENIX model-building and refinement. The multi-dimensional distributions and Top8000 reference dataset are freely available on GitHub. These rotamers are termed “ultimate” because data sampling and quality are now fully adequate for this task, and also because we believe the future of conformational validation should integrate side-chain with backbone criteria. Proteins 2016; 84:1177–1189. © 2016 Wiley Periodicals, Inc.
With increasing recognition of the roles RNA molecules and RNA/protein complexes play in an unexpected variety of biological processes, understanding of RNA structure-function relationships is of high current importance. To make clean biological interpretations from three-dimensional structures, it is imperative to have high-quality, accurate RNA crystal structures available, and the community has thoroughly embraced that goal. However, due to the many degrees of freedom inherent in RNA structure (especially for the backbone), it is a significant challenge to succeed in building accurate experimental models for RNA structures. This chapter describes the tools and techniques our research group and our collaborators have developed over the years to help RNA structural biologists both evaluate and achieve better accuracy. Expert analysis of large, high-resolution, quality-conscious RNA datasets provides the fundamental information that enables automated methods for robust and efficient error diagnosis in validating RNA structures at all resolutions. The even more crucial goal of correcting the diagnosed outliers has steadily developed toward highly effective, computationally based techniques. Automation enables solving complex issues in large RNA structures, but cannot circumvent the need for thoughtful examination of local details, and so we also provide some guidance for interpreting and acting on the results of current structure validation for RNA.