Structural predictions have reached unprecedented accuracy. They leverage sequence-specific data to capture all potential interactions a sequence has evolved to fulfill. AlphaFold derives information from three sources: learned parameters capturing intrinsic amino acid secondary structure and environment propensity; models of related proteins providing structural templates; and aligned sequences encoding profiles and concerted evolutionary changes of residues involved in contacts. However, function demands dynamic changes; hence not all possible interactions can coexist simultaneously. Comprehensive information entails contradictions, which resolved in favor of the better-informed structure will silence less stable states and associations. Here, we introduce a method using all three channels to include prior knowledge: site-specific variants, predefined alignments and templates. Selecting information relevant to a particular state delimits the functional context of a prediction. Our program VAIRO allows us to rescue asymmetric and weaker interactions to complete the view of molecular assemblies in the architecture of a bacterial surface layer, and reveals otherwise inaccessible dynamic states in a pneumococcal multimeric membrane protein complex. VAIRO is distributed via the python package index (PyPI) (https://pypi.org/project/vairo) and the code is also available on Github (https://github.com/arcimboldo-team/vairo).
Crystallography at low resolution must determine the atomic model from less experimental observations, which is challenging in the absence of a model. In addition, model bias is more severe when independent experimental data are scarce. Our methods solve the phase problem by combining the location of accurate model fragments using Phaser with density modification and interpretation of the resulting maps using SHELXE. From a partial, correct structure, the density modification process and the stereochemical constraints draw the rest of the structure, validating the result. This same principle is now exploited at low resolution. Coiled coils are important, ubiquitous structures but notoriously difficult to phase and to predict. Both correct solutions and incorrect ones are poorly discriminated by the crystallographic figures of merit as long as helices are correctly oriented. We incorporate coiled-coil verification, designed to set up competing, incompatible structural hypotheses to probe both the results and establish the power of the data to discriminate them. Efficiency of coiled-coil phasing and validation in test cases from 3 to 4 Å is demonstrated in ARCIMBOLDO_LITE, placing single helices, and in ARCIMBOLDO_SHREDDER, with fragments derived from AlphaFold models. SHELXE tracing at low resolution has been enhanced, maintaining its local character but extending the environment assessment. For non-helical structures, verification is demonstrated in the fragment location process. Its use is exemplified with the solution of the VSR1 structure at 3.5 Å, depending on LLG optimization and the emergence of new features in the electron density. Relying on verification, we have extended the use of the ARCIMBOLDO software to low resolution.
The Collaborative Computational Project No. 4 (CCP4) is a UK-led international collective with a mission to develop, test, distribute and promote software for macromolecular crystallography. The CCP4 suite is a multiplatform collection of programs brought together by familiar execution routines, a set of common libraries and graphical interfaces. The CCP4 suite has experienced several considerable changes since its last reference article, involving new infrastructure, original programs and graphical interfaces. This article, which is intended as a general literature citation for the use of the CCP4 software suite in structure determination, will guide the reader through such transformations, offering a general overview of the new features and outlining future developments. As such, it aims to highlight the individual programs that comprise the suite and to provide the latest references to them for perusal by crystallographers around the world.
In late 2020, the results of CASP14, the 14th event in a series of competitions to assess the latest developments in computational protein structure-prediction methodology, revealed the giant leap forward that had been made by Google's Deepmind in tackling the prediction problem. The level of accuracy in their predictions was the first instance of a competitor achieving a global distance test score of better than 90 across all categories of difficulty. This achievement represents both a challenge and an opportunity for the field of experimental structural biology. For structure determination by macromolecular X-ray crystallography, access to highly accurate structure predictions is of great benefit, particularly when it comes to solving the phase problem. Here, details of new utilities and enhanced applications in the CCP4 suite, designed to allow users to exploit predicted models in determining macromolecular structures from X-ray diffraction data, are presented. The focus is mainly on applications that can be used to solve the phase problem through molecular replacement.
Structure predictions have matched the accuracy of experimental structures from close homologues, providing suitable models for molecular replacement phasing. Even in predictions that present large differences due to the relative movement of domains or poorly predicted areas, very accurate regions tend to be present. These are suitable for successful fragment-based phasing as implemented in ARCIMBOLDO. The particularities of predicted models are inherently addressed in the new predicted_model mode, rendering preliminary treatment superfluous but also harmless. B-value conversion from predicted LDDT or error estimates, the removal of unstructured polypeptide, hierarchical decomposition of structural units from domains to local folds and systematically probing the model against the experimental data will ensure the optimal use of the model in phasing. Concomitantly, the exhaustive use of models and stereochemistry in phasing, refinement and validation raises the concern of crystallographic model bias and the need to critically establish the information contributed by the experiment. Therefore, in its predicted_model mode ARCIMBOLDO_SHREDDER will first determine whether the input model already constitutes a solution or provides a straightforward solution with Phaser. If not, extracted fragments will be located. If the landscape of solutions reveals numerous, clearly discriminated and consistent probes or if the input model already constitutes a solution, model-free verification will be activated. Expansions with SHELXE will omit the partial solution seeding phases and all traces outside their respective masks will be combined in ALIXE, as far as consistent. This procedure completely eliminates the molecular replacement search model in favour of the inferences derived from this model. In the case of fragments, an incorrect starting hypothesis impedes expansion. The predicted_model mode has been tested in different scenarios.
Detection of translational noncrystallographic symmetry (TNCS) can be critical for success in crystallographic phasing, particularly when molecular-replacement models are poor or anomalous phasing information is weak. If the correct TNCS is detected then expected intensity factors for each reflection can be refined, so that the maximum-likelihood functions underlying molecular replacement and single-wavelength anomalous dispersion use appropriate structure-factor normalization and variance terms. Here, an analysis of a curated database of protein structures from the Protein Data Bank to investigate how TNCS manifests in the Patterson function is described. These studies informed an algorithm for the detection of TNCS, which includes a method for detecting the number of vectors involved in any commensurate modulation (the TNCS order). The algorithm generates a ranked list of possible TNCS associations in the asymmetric unit for exploration during structure solution.
The meiotic chromosome axis plays key roles in meiotic chromosome organization and recombination, yet the underlying protein components of this structure are highly diverged. Here, we show that 'axis core proteins' from budding yeast (Red1), mammals (SYCP2/SYCP3), and plants (ASY3/ASY4) are evolutionarily related and play equivalent roles in chromosome axis assembly. We first identify 'closure motifs' in each complex that recruit meiotic HORMADs, the master regulators of meiotic recombination. We next find that axis core proteins form homotetrameric (Red1) or heterotetrameric (SYCP2:SYCP3 and ASY3:ASY4) coiled-coil assemblies that further oligomerize into micron-length filaments. Thus, the meiotic chromosome axis core in fungi, mammals, and plants shares a common molecular architecture, and likely also plays conserved roles in meiotic chromosome axis assembly and recombination control.
TRANSLATIONAL NON-CRYSTALLOGRAPHIC SYMMETRY Caballero, Iracema (SBU IBMB CSIC, Barcelona, ESP); Sammito, Massimo (Department of Haematology, Cambridge Institute for Medical Research, University of Cambridge, Cambridge, GBR); Usón, Isabel (SBU IBMB CSIC, Barcelona, ESP); Read, Randy (Department of Haematology, Cambridge Institute for Medical Research, University of Cambridge, Cambridge, GBR); McCoy, Airlie J (University of Cambridge, Cambridge, GBR)
The ARCIMBOLDO method of phasing through the location of small fragments combined with density modification and autotracing is particularly suited to helical structures, but coiled coils remain challenging. Features designed for solving coiled coils at resolutions of up to 3 Å were tested on a pool of 150 structures.
The meiotic chromosome plays key roles in meiotic chromosome organization and recombination, yet the underlying protein components of this structure are highly diverged. Here, we show that axis core proteins from budding yeast (Red1), mammals (SYCP2/SYCP3), and plants (ASY3/ASY4) are evolutionarily related and play equivalent roles in chromosome assembly. We first identify motifs in each complex that recruit meiotic HORMADs, the master regulators of meiotic recombination. We next find that core complexes form homotetrameric (Red1) or heterotetrameric (SYCP2:SYCP3 and ASY3:ASY4) coiled-coil assemblies that further oligomerize into micron-length filaments. Thus, the meiotic chromosome core in fungi, mammals, and plants shares a common molecular architecture and role in assembly and recombination control. We propose that the meiotic chromosome self-assembles through cooperative interactions between dynamic DNA loop-extruding cohesin complexes and the filamentous core, then serves as a platform for chromosome organization, recombination, and synaptonemal complex assembly.
ARCIMBOLDO (1) combines the search of small and accurate fragments with PHASER (2), with their expansion to a full structure solution through density modification and autotracing with SHELXE (3).Fragments can be as small and general as a single ideal polyalanine helix, libraries of tertiary structure local folds such as beta sheets, or even fragments extracted from a distant homolog.These three approaches are accessible through our programs ARCIMBOLDO_LITE, ARCIMBOLDO_BORGES, and ARCIMBOLDO_SHREDDER.Our recent implementations include strategies tailored to deal with challenging cases, tackling lower resolution, larger size, data pathologies or the ambiguity in the crystal contents.This communication will focus on the latest developments in our group, illustrated through examples of their application to previously unknown structures.Coiled coil structures required to develop new features in order to tackle their inherent phasing difficulties.PHASER's new packing constraints can now be used at the translation search, and automatic handling of the anisotropy and translational non crystallographic symmetry corrections have been included.Both helix directions are tested for all fragments in partial solutions to increase the search base at low resolution.Moreover, we now make use of a different autotracing algorithm from the SHELXE beta version.ARCIMBOLDO_SHREDDER derives and improves model fragments starting from a distant homolog template, and implements new approaches to refine them in order to reduce the deviations within an overall correct fold.Finally, we also have the possibility of using phase combination of partial solutions in order to increase their information content of the starting map to be expanded into a full solution.[1]