Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method - AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements, that are beyond the scope of current state-of-the-art methods and, in some instances, even manual human modeling. The ability to efficiently sample these barrier-crossing rearrangements unlocks a new horizon for scalable and automated model building. Crucially, ROCKET does not require retraining of AlphaFold2 and is readily adaptable to multimers, ligand-cofolding, and other data modalities. Conversely, our differentiable crystallographic and cryo-EM target functions are capable of augmenting other structure prediction methods. ROCKET thus provides an extensible framework for the integration of experimental observables with biomolecular machine learning.
The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.
Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.
Experimental structure determination can be accelerated with AI-based structure prediction methods such as AlphaFold. Here we present an automatic procedure requiring only sequence information and crystallographic data that uses AlphaFold predictions to produce an electron density map and a structural model. Iterating through cycles of structure prediction is a key element of our procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. We applied this procedure to X-ray data for 215 structures released by the Protein Data Bank in a recent 6-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2Å. Predictions from our iterative template-guided prediction procedure were more accurate than those obtained without templates. We suggest a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization.
The Collaborative Computational Project No. 4 (CCP4) is a UK-led international collective with a mission to develop, test, distribute and promote software for macromolecular crystallography. The CCP4 suite is a multiplatform collection of programs brought together by familiar execution routines, a set of common libraries and graphical interfaces. The CCP4 suite has experienced several considerable changes since its last reference article, involving new infrastructure, original programs and graphical interfaces. This article, which is intended as a general literature citation for the use of the CCP4 software suite in structure determination, will guide the reader through such transformations, offering a general overview of the new features and outlining future developments. As such, it aims to highlight the individual programs that comprise the suite and to provide the latest references to them for perusal by crystallographers around the world.
Fast, reliable docking of models into cryo-EM maps requires understanding of the errors in the maps and the models. Likelihood-based approaches to errors have proven to be powerful and adaptable in experimental structural biology, finding applications in both crystallography and cryo-EM. Indeed, previous crystallographic work on the errors in structural models is directly applicable to likelihood targets in cryo-EM. Likelihood targets in Fourier space are derived here to characterize, based on the comparison of half-maps, the direction- and resolution-dependent variation in the strength of both signal and noise in the data. Because the signal depends on local features, the signal and noise are analysed in local regions of the cryo-EM reconstruction. The likelihood analysis extends to prediction of the signal that will be achieved in any docking calculation for a model of specified quality and completeness. A related calculation generalizes a previous measure of the information gained by making the cryo-EM reconstruction.
The AlphaFold2 results in the 14th edition of Critical Assessment of Structure Prediction (CASP14) showed that accurate (low root-mean-square deviation) in silico models of protein structure domains are on the horizon, whether or not the protein is related to known structures through high-coverage sequence similarity. As highly accurate models become available, generated by harnessing the power of correlated mutations and deep learning, one of the aspects of structural biology to be impacted will be methods of phasing in crystallography. Here, the data from CASP14 are used to explore the prospects for changes in phasing methods, and in particular to explore the prospects for molecular-replacement phasing using in silico models.
The introduction of disulfide bonds into periplasmic proteins is a critical process in many Gram-negative bacteria. The formation and regulation of protein disulfide bonds have been linked to the production of virulence factors. Understanding the different pathways involved in this process is important in the development of strategies to disarm pathogenic bacteria. The well characterized disulfide bond-forming (DSB) proteins play a key role by introducing or isomerizing disulfide bonds between cysteines in substrate proteins. Curiously, the suppressor of copper sensitivity C proteins (ScsCs), which are part of the bacterial copper-resistance response, share structural and functional similarities with DSB oxidase and isomerase proteins, including the presence of a catalytic thioredoxin domain. However, the oxidoreductase activity of ScsC varies with its oligomerization state, which depends on a poorly conserved N-terminal domain. Here, the structure and function of Caulobacter crescentus ScsC (CcScsC) have been characterized. It is shown that CcScsC binds copper in the copper(I) form with subpicomolar affinity and that its isomerase activity is comparable to that of Escherichia coli DsbC, the prototypical dimeric bacterial isomerase. It is also reported that CcScsC functionally complements trimeric Proteus mirabilis ScsC (PmScsC) in vivo, enabling the swarming of P. mirabilis in the presence of copper. Using mass photometry and small-angle X-ray scattering (SAXS) the protein is demonstrated to be trimeric in solution, like PmScsC, and not dimeric like EcDsbC. The crystal structure of CcScsC was also determined at a resolution of 2.6 Å, confirming the trimeric state and indicating that the trimerization results from interactions between the N-terminal α-helical domains of three CcScsC protomers. The SAXS data analysis suggested that the protomers are dynamic, like those of PmScsC, and are able to sample different conformations in solution.
It is remarkable that the first review of the 'molecular replacement' phasing method for macromolecules was published 50 years ago and just a year after the Protein Data Bank was established [Crystallography: Protein Data Bank.Nature New Biology 233, 223 (1971); Rossmann, The Molecular Replacement Method.New York: Gordon & Breach (1972)].What began as a niche method for exploiting the phase information inherent to noncrystallographic symmetry and/or different crystal forms has evolved into the predominant phasing method in macromolecular crystallography, regardless of crystal form or whether similar structures are present in the Protein Data Bank.The evolution of molecular replacement has seen advances in its speed and sophistication, driven by the passion of crystallographers to answer ever more elaborate biological questions as quickly as possible, and has been facilitated by computational progress that would delight the pioneers.In this personal view, I will explore how molecular replacement became the most versatile of macromolecular phasing methods, the current best practice, and make some predictions for its future.
The assessment of CASP models for utility in molecular replacement is a measure of their use in a valuable real-world application. In CASP7, the metric for molecular replacement assessment involved full likelihood-based molecular replacement searches; however, this restricted the assessable targets to crystal structures with only one copy of the target in the asymmetric unit, and to those where the search found the correct pose. In CASP10, full molecular replacement searches were replaced by likelihood-based rigid-body refinement of models superimposed on the target using the LGA algorithm, with the metric being the refined likelihood (LLG) score. This enabled multi-copy targets and very poor models to be evaluated, but a significant further issue remained: the requirement of diffraction data for assessment. We introduce here the relative-expected-LLG (reLLG), which is independent of diffraction data. This reLLG is also independent of any crystal form, and can be calculated regardless of the source of the target, be it X-ray, NMR or cryo-EM. We calibrate the reLLG against the LLG for targets in CASP14, showing that it is a robust measure of both model and group ranking. Like the LLG, the reLLG shows that accurate coordinate error estimates add substantial value to predicted models. We find that refinement by CASP groups can often convert an inadequate initial model into a successful MR search model. Consistent with findings from others, we show that the AlphaFold2 models are sufficiently good, and reliably so, to surpass other current model generation strategies for attempting molecular replacement phasing.
Crystallographic phasing strategies increasingly require the exploration and ranking of many hypotheses about the number, types and positions of atoms, molecules and/or molecular fragments in the unit cell, each with only a small chance of being correct. Accelerating this move has been improvements in phasing methods, which are now able to extract phase information from the placement of very small fragments of structure, from weak experimental phasing signal or from combinations of molecular replacement and experimental phasing information. Describing phasing in terms of a directed acyclic graph allows graph-management software to track and manage the path to structure solution. The crystallographic software supporting the graph data structure must be strictly modular so that nodes in the graph are efficiently generated by the encapsulated functionality. To this end, the development of new software, Phasertng, which uses directed acyclic graphs natively for input/output, has been initiated. In Phasertng, the codebase of Phaser has been rebuilt, with an emphasis on modularity, on scripting, on speed and on continuing algorithm development. As a first application of phasertng, its advantages are demonstrated in the context of phasertng.xtricorder, a tool to analyse and triage merged data in preparation for molecular replacement or experimental phasing. The description of the phasing strategy with directed acyclic graphs is a generalization that extends beyond the functionality of Phasertng, as it can incorporate results from bioinformatics and other crystallographic tools, and will facilitate multifaceted search strategies, dynamic ranking of alternative search pathways and the exploitation of machine learning to further improve phasing strategies.
Detection of translational noncrystallographic symmetry (TNCS) can be critical for success in crystallographic phasing, particularly when molecular-replacement models are poor or anomalous phasing information is weak. If the correct TNCS is detected then expected intensity factors for each reflection can be refined, so that the maximum-likelihood functions underlying molecular replacement and single-wavelength anomalous dispersion use appropriate structure-factor normalization and variance terms. Here, an analysis of a curated database of protein structures from the Protein Data Bank to investigate how TNCS manifests in the Patterson function is described. These studies informed an algorithm for the detection of TNCS, which includes a method for detecting the number of vectors involved in any commensurate modulation (the TNCS order). The algorithm generates a ranked list of possible TNCS associations in the asymmetric unit for exploration during structure solution.
The cation-independent mannose 6-phosphate (M6P)/Insulin-like growth factor-2 receptor (CI-MPR/IGF2R) is an ∼300 kDa transmembrane protein responsible for trafficking M6P-tagged lysosomal hydrolases and internalizing IGF2. The extracellular region of the CI-MPR has 15 homologous domains, including M6P-binding domains (D) 3, 5, 9, and 15 and IGF2-binding domain 11. We have focused on solving the first structures of human D7-10 within two multi-domain constructs, D9-10 and D7-11, and provide the first high-resolution description of the high-affinity M6P-binding D9. Moreover, D9 stabilizes a well-defined hub formed by D7-11 whereby two penta-domains intertwine to form a dimeric helical-type coil via an N-glycan bridge on D9. Remarkably the D7-11 structure matches an IGF2-bound state of the receptor, suggesting this may be an intrinsically stable conformation at neutral pH. Interdomain clusters of histidine and proline residues may impart receptor rigidity and play a role in structural transitions at low pH.
Good prior estimates of the effective root-mean-square deviation (r.m.s.d.) between the atomic coordinates of the model and the target optimize the signal in molecular replacement, thereby increasing the success rate in difficult cases. Previous studies using protein structures solved by X-ray crystallography as models showed that optimal error estimates (refined after structure solution) were correlated with the sequence identity between the model and target, and with the number of residues in the model. Here, this work has been extended to find additional correlations between parameters of the model and the target and hence improved prior estimates of the coordinate error. Using a graph database, a curated set of 6030 molecular-replacement calculations using models that had been solved by X-ray crystallography was analysed to consider about 120 model and target parameters. Improved estimates were achieved by replacing the sequence identity with the Gonnet score for sequence similarity, as well as by considering the resolution of the target structure and the MolProbity score of the model. This approach was extended by analysing 12 610 additional molecular-replacement calculations where the model was determined by NMR. The median r.m.s.d. between pairs of models in an ensemble was found to be correlated with the estimated r.m.s.d. to the target. For models solved by NMR, the overall coordinate error estimates were larger than for structures determined by X-ray crystallography, and were more highly correlated with the number of residues.
The information gained by making a measurement, termed the Kullback-Leibler divergence, assesses how much more precisely the true quantity is known after the measurement was made (the posterior probability distribution) than before (the prior probability distribution). It provides an upper bound for the contribution that an observation can make to the total likelihood score in likelihood-based crystallographic algorithms. This makes information gain a natural criterion for deciding which data can legitimately be omitted from likelihood calculations. Many existing methods use an approximation for the effects of measurement error that breaks down for very weak and poorly measured data. For such methods a different (higher) information threshold is appropriate compared with methods that account well for even large measurement errors. Concerns are raised about a current trend to deposit data that have been corrected for anisotropy, sharpened and pruned without including the original unaltered measurements. If not checked, this trend will have serious consequences for the reuse of deposited data by those who hope to repeat calculations using improved new methods.
(Developmental Cell 50, 494–508.e1–e11; August 19, 2019) As a result of an author oversight in the originally published version of this article, the author Susanne Salomon was omitted from the author list. This error has now been corrected in the article online. The authors apologize for the error and any inconvenience that may have resulted. Temporal Ordering in Endocytic Clathrin-Coated Vesicle Formation via AP2 PhosphorylationWrobel et al.Developmental CellAugust 19, 2019In BriefWrobel et al. show that phosphorylation of the mammalian endocytic AP2 adaptor causes a conformational change that allows it to efficiently bind the protein NECAP. This, in turn, recruits membrane-remodeling proteins into clathrin-coated pits, which drive their formation toward final scission from the parent membrane. Full-Text PDF Open Access
The 3D-Reflection data viewer in Phenix [1] is based on the OpenGL library which is deprecated on MacOS.A replacement based on supported libraries, was urgently required.The NGL-HKL-viewer has been developed in a way that leverages the codebase of the original viewer while extending the functionality.Protein crystallography involves processing raw X-ray images to indexed intensities.Subsequent steps may associate additional data with these indices.As part of the changes to the 3D-Reflection data viewer the types of reflection data that can be displayed were expanded.Real valued data are displayed as spheres scaled according to their magnitudes.Colour coding through both hue and saturation can be applied as well.The viewer allows associating one data parameter with the sizes of the spheres but colour them according to another data parameter.For example, phased structure factors can be displayed as spheres scaled by amplitudes, coloured according to phases and colour saturation used to represent figures of merit.Data can be sorted into bins and be shown or hidden.NGL-HKL-viewer is scriptable from Python and part of CCTBX [2].It is based on NGL [3] and therefore portable to all computing platforms with a modern web browser.The viewer can be embedded in graphical user interfaces such as Qt5 or PySide.It is part of Phaser.Voyager [4].
Diffraction (X-ray, neutron and electron) and electron cryo-microscopy are powerful methods to determine three-dimensional macromolecular structures, which are required to understand biological processes and to develop new therapeutics against diseases. The overall structure-solution workflow is similar for these techniques, but nuances exist because the properties of the reduced experimental data are different. Software tools for structure determination should therefore be tailored for each method. Phenix is a comprehensive software package for macromolecular structure determination that handles data from any of these techniques. Tasks performed with Phenix include data-quality assessment, map improvement, model building, the validation/rebuilding/refinement cycle and deposition. Each tool caters to the type of experimental data. The design of Phenix emphasizes the automation of procedures, where possible, to minimize repetitive and time-consuming manual tasks, while default parameters are chosen to encourage best practice. A graphical user interface provides access to many command-line features of Phenix and streamlines the transition between programs, project tracking and re-running of previous tasks.