The application Scotty, implemented within the Phasertng codebase, was used to perform a PDB-wide analysis of lattice coincidences. Using a broad definition of lattice coincidence, the number of distinct lattice clusters is approximately half the total number of crystallographic PDB entries. In over one thousand lattice clusters entries are reported in different space groups, consistent with pseudo-symmetric variation within a common lattice framework. Space-group frequencies computed at the lattice-coincidence level update those obtained by entry-based counting, and more accurately reflect priors for novel crystal forms. Combining lattice clustering with sequence identity and deposited oligomeric annotations reveals multiple cases of inconsistent biological assembly assignments among structures sharing near-identical lattices, suggesting an opportunity to improve annotations. The survey also identifies a range of protein systems forming extended in cellulo paracrystalline arrays, including storage, sequestration, toxin and membrane-associated proteins, an understudied area of structural biology. Overall, the results demonstrate that lattice-level analysis provides a valuable perspective on macromolecular self-association.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, key challenges persist in modeling side-chain packing, condition-dependent conformational changes and biomolecular interactions, largely because of limited high-quality training data. At the same time, emerging experimental techniques such as cryo-electron microscopy (cryo-EM), cryo-electron tomography (cryo-ET) and high-throughput crystallography are generating vast amounts of structural information but converting these data into mechanistically interpretable atomic models often remains difficult. Here we show that integrating experimental measurements directly into protein structure prediction can overcome these limitations. We introduce ROCKET, an augmentation of AlphaFold2 that refines predicted structures using cryo-EM, cryo-ET and X-ray crystallography data. By optimizing structures in the space of coevolutionary embeddings rather than Cartesian coordinates, ROCKET captures biologically meaningful structural variation that is inaccessible to AlphaFold2 alone and to existing automated modeling approaches, especially when the signal-to-noise ratio is low. ROCKET enables scalable, automated model building without retraining and provides a general framework for integrating experimental observables with biomolecular machine learning.
When a macromolecular crystal lattice targeted for study is closely related to a previously studied crystal lattice, phasing the target crystal by difference (fast-)Fourier transform (DFFT) methods is preferable to performing molecular replacement. The application Scotty, within the Phasertng codebase, is software for the identification of coincident lattices and downstream processing. All crystallographic PDB entries are organized into a scikit-learn `BallTree' index under the Niggli cell distance metric `NCDist'. Nearest neighbours under the metric are progressed to test structure-factor intensity correlation. The structure from the lattice with the highest correlation with the target is used to phase the target lattice, and the coordinates are taken forward to coordinate refinement using a wide convergence radius protocol. The method can identify lattice coincidences accounting for very significant non-isomorphism. The nearest-neighbour search is space-group agnostic, so that coincident lattices are identified even when the nominal space groups are different, the symmetry of one lattice being described as a subgroup of the other or a higher metric symmetry.
Glycine N-acyltransferase (GLYAT; EC 2.3.1.13, Accession ID: AAI12537) is a key enzyme in mammalian homeostasis that has been linked to several pathologies in humans, including cancer. Here we report the first crystal structure of a member of the GLYAT family, both in the apo form as well as bound to benzoyl-CoA. Binding of glycine could be inferred from an acetate molecule from the crystallization solution. A detailed analysis of its structure and the effects of mutations of key residues helped elucidate the catalytic mechanism, showing a general base-catalyzed reaction driven by a potential low-barrier hydrogen bond (LBHB) formed between the catalytic Glu-His dyad. This work will aid further studies of GLYAT and other members of the family.
The CASP16 evaluation of model accuracy (EMA) experiment assessed the ability of predictors to estimate the accuracy of predicted models, with a particular emphasis on multimeric assemblies. Expanding on the CASP15 framework, CASP16 introduced a new evaluation mode (QMODE3) focused on selecting high-quality models from large-scale AlphaFold2-derived model pools generated by MassiveFold. Three primary evaluation tasks were therefore conducted: QMODE1 assessed global structure accuracy, QMODE2 focused on the accuracy of interface residues, and QMODE3 tested model selection performance. Predictors were evaluated using a diverse set of OpenStructure-based metrics, and a novel penalty-based ranking scheme was developed for QMODE3 to handle score interdependence and varying prediction quality distributions. Additionally, we explored the accuracy and utility of predicted local confidence measures now made available on a per-atom basis by methods that invoke AlphaFold3. Results showed that methods incorporating AlphaFold3-derived features-particularly per-atom pLDDT-performed best in estimating local accuracy and in utility for experimental structure solution. For QMODE3, performance varied significantly across monomeric, homomeric, and heteromeric target categories and underscored the ongoing challenge of evaluating complex assemblies.
Model quality assessment (MQA) remains a critical component of structural bioinformatics for both structure predictors and experimentalists seeking to use predictions for downstream applications. In CASP16, the Evaluation of Model Accuracy (EMA) category featured both global and local quality estimation for multimeric assemblies (QMODE1 and QMODE2), as well as a novel QMODE3 challenge-requiring predictors to identify the best five models from thousands generated by MassiveFold. This paper presents detailed results from several leading CASP16 EMA methods, highlighting the strengths and limitations of the approaches.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method - AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements, that are beyond the scope of current state-of-the-art methods and, in some instances, even manual human modeling. The ability to efficiently sample these barrier-crossing rearrangements unlocks a new horizon for scalable and automated model building. Crucially, ROCKET does not require retraining of AlphaFold2 and is readily adaptable to multimers, ligand-cofolding, and other data modalities. Conversely, our differentiable crystallographic and cryo-EM target functions are capable of augmenting other structure prediction methods. ROCKET thus provides an extensible framework for the integration of experimental observables with biomolecular machine learning.
Advances in machine learning have transformed structural biology, enabling swift and accurate prediction of protein structure from sequence. However, challenges persist in capturing sidechain packing, condition-dependent conformational dynamics, and biomolecular interactions, primarily due to scarcity of high-quality training data. Emerging techniques, including cryo-electron tomography (cryo-ET) and high-throughput crystallography, promise vast new sources of structural data, but translating experimental observations into mechanistically interpretable atomic models remains a key bottleneck. Here, we address these challenges by improving the efficiency of structural analysis through combining experimental measurements with a landmark protein structure prediction method – AlphaFold2. We present an augmentation of AlphaFold2, ROCKET, that refines its predictions using cryo-EM, cryo-ET, and X-ray crystallography data, and demonstrate that this approach captures biologically important structural variation that AlphaFold2 does not. By performing structure optimization in the space of coevolutionary embeddings, rather than Cartesian coordinates, ROCKET automates difficult modeling tasks, such as flips of functional loops and domain rearrangements at low resolution. ROCKET does not require retraining of AlphaFold2 and is readily adaptable to other data modalities. This new type of structure refinement that optimizes latent representations in evolutionary space could unlock possibilities for high-throughput ligand screening, assemblies solved at low resolution, and conformational landscapes.
Analysis of crystallographic diffraction data before phasing gives the crystallographer a ‘first look’ at the nature of the problem and the context in which the structure determination will be performed. We here report the development of Xtricorder , an application that targets analysis of crystallographic data specifically for likelihood-based phasing. As well as porting many of the analyses previously available but relatively inaccessible in our Phaser codebase, Xtricorder offers a likelihood-enhanced self-rotation function. A novel and intuitive graphical representation of the self-rotation function presents the results for user inspection, and has the added advantage that, in an adapted form, is appropriate for training a convolutional neural network to enhance the standard Matthews analysis and more accurately predict the number of copies in the asymmetric unit. We investigate the usefulness of the likelihood-enhanced self-rotation function in ‘first look’ analyses, exploring the circumstances under which the self-rotation function results are useful, and discuss the application to AI-generated structure prediction. Synopsis Xtricorder is a new tool for analysing crystallographic data prior to phasing, featuring a likelihood-enhanced self-rotation function and graphical output that aids both user interpretation and machine learning-based prediction of asymmetric unit content. ### Competing Interest Statement The authors have declared no competing interest. Biotechnology and Biological Sciences Research Council, https://ror.org/00cwqg982, BB/Y009398/1
Analysis of crystallographic diffraction data after collection and integration but before phasing gives the crystallographer a `first-look' assessment of data quality and flags potential challenges in subsequent structure determination. We here report the development of Xtricorder, a `first-look' application specifically targeted at likelihood-based phasing. Xtricorder incorporates the full array of analyses previously available in the Phaser codebase, with some enhancements and updates, in a more streamlined and accessible implementation. In addition, Xtricorder offers a likelihood-enhanced self-rotation function. A novel graphical representation of the self-rotation function, the `composite-section diagram', presents the results for user inspection and has the added advantage that, in an adapted form, it is appropriate for training a convolutional neural network to enhance the standard Matthews analysis and double the accuracy of asymmetric unit copy-number prediction. We investigate the usefulness of the likelihood-enhanced self-rotation function in `first-look' analyses, exploring the circumstances under which the self-rotation function results are useful, and discuss the application to AI-generated structure prediction.
The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.
Five new Co-editors are appointed to the Editorial Board of Acta Cryst. D - Structural Biology.
Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.
A defining pathological feature of most neurodegenerative diseases is the assembly of proteins into amyloid that form disease-specific structures1. In Alzheimer's disease, this is characterized by the deposition of β-amyloid and tau with disease-specific conformations. The in situ structure of amyloid in the human brain is unknown. Here, using cryo-fluorescence microscopy-targeted cryo-sectioning, cryo-focused ion beam-scanning electron microscopy lift-out and cryo-electron tomography, we determined in-tissue architectures of β-amyloid and tau pathology in a postmortem Alzheimer's disease donor brain. β-amyloid plaques contained a mixture of fibrils, some of which were branched, and protofilaments, arranged in parallel arrays and lattice-like structures. Extracellular vesicles and cuboidal particles defined the non-amyloid constituents of β-amyloid plaques. By contrast, tau inclusions formed parallel clusters of unbranched filaments. Subtomogram averaging a cluster of 136 tau filaments in a single tomogram revealed the polypeptide backbone conformation and filament polarity orientation of paired helical filaments within tissue. Filaments within most clusters were similar to each other, but were different between clusters, showing amyloid heterogeneity that is spatially organized by subcellular location. The in situ structural approaches outlined here for human donor tissues have applications to a broad range of neurodegenerative diseases.
Two new Co-editors are welcomed to Acta Cryst. D - Structural Biology.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
Cas12a is a programmable nuclease for adaptive immunity against invading nucleic acids in CRISPR-Cas systems. Here, we report the crystal structures of apo Cas12a from Lachnospiraceae bacterium MA2020 (Lb2) and the Lb2Cas12a+crRNA complex, as well as the cryo-EM structure and functional studies of the Lb2Cas12a+crRNA+DNA complex. We demonstrate that apo Lb2Cas12a assumes a unique, elongated conformation, whereas the Lb2Cas12a+crRNA binary complex exhibits a compact conformation that subsequently rearranges to a semi-open conformation in the Lb2Cas12a+crRNA+DNA ternary complex. Notably, in solution, apo Lb2Cas12a is dynamic and can exist in both elongated and compact forms. Residues from Met493 to Leu523 of the WED domain undergo major conformational changes to facilitate the required structural rearrangements. The REC lobe of Lb2Cas12a rotates 103° concomitant with rearrangement of the hinge region close to the WED and RuvC II domains to position the RNA-DNA duplex near the catalytic site. Our findings provide insight into crRNA recognition and the mechanism of target DNA cleavage.
Experimental structure determination can be accelerated with AI-based structure prediction methods such as AlphaFold. Here we present an automatic procedure requiring only sequence information and crystallographic data that uses AlphaFold predictions to produce an electron density map and a structural model. Iterating through cycles of structure prediction is a key element of our procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. We applied this procedure to X-ray data for 215 structures released by the Protein Data Bank in a recent 6-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2Å. Predictions from our iterative template-guided prediction procedure were more accurate than those obtained without templates. We suggest a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization.