Multipolar scattering models, such as the transferable aspherical atom model, account for atomic chemical interactions and provide a more accurate representation of experimental data. However, the simpler independent atom model (IAM), which assumes non-interacting atoms, is the only model available in the most widely used macromolecular refinement programs. This is primarily because IAM offers a hard-to-beat combination of computational efficiency and modelling power at typical macromolecular resolutions. By contrast, more accurate multipolar modelling has historically been limited due to its computational cost and the absence of an interface between software capable of calculating structure factors and gradients based on multipolar models and software designed for macromolecular refinement. This work introduces pyDiSCaMB, a Python software package designed to integrate between the computational crystallography toolbox (cctbx) and the quantum crystallography library DiSCaMB (Densities in Structural Chemistry and Molecular Biology), thus enabling multipolar scattering models in Phenix's toolkit. The implementation, features and capabilities of pyDiSCaMB are presented, the runtimes for the calculation of structure factor and target gradients with respect to atomic parameters are explored, and Fourier images of electrostatic potential, electron density and deformation maps are computed as illustrative examples. The pyDiSCaMB library will make multipolar modelling widely available to the structural biology community, potentially transforming refinement and model-building for both crystallography and cryogenic electron microscopy (cryoEM).
Photosynthesis provides most of the bio-available energy and oxygen to our planet. Yet, while several other mechanisms in oxygenic and anoxygenic photosynthesis have been thoroughly investigated, the core water-splitting reaction of photosystem II (PSII) remains largely uncharacterized. The recent advent of X-ray free-electron lasers has provided us with a tool to probe the structure of PSII's oxygen-evolving complex under ambient conditions and with microsecond time resolution. Still, the quality of diffraction data offered by these experiments is insufficient to reliably observe one-electron differences between individual time points.[1] Instead of probing the imprecise scatterer distribution, oxidation states of individual metal atoms can be assigned by investigating their X-ray absorption edges. Information from classical spectroscopy can not be matched to individual atoms; however, by performing serial diffraction experiments with a pink beam tuned to the metal absorption edge, the anomalous dispersion of each atom becomes embedded in the diffraction image. A careful analysis of Bragg reflection profiles can be thus applied to retrieve the atomic form factors as a function of energy. Refined absorption curves can be then used to characterize the electronic structure of each atom. This spatially resolved anomalous dispersion (SPREAD) technique has been previously successfully applied to data simulated for ferredoxin: a 25 kDa protein containing two differently charged iron centers.[2] The present work describes our recent advances in scaling the pipeline for a 750 kDa PSII with a four-manganese cluster and adapting it to experimental data. In particular, we describe the first working refinement of experimental data, as well as issues encountered with reliability, restraints, mosaicity, and memory use.
With Phenix 1.21, we have incorporated the use of AlphaFold structure prediction into our automated model building pipelines for crystallography and cryo-EM. To help the structural biology community efficiently run these pipelines without having to manage their own AlphaFold installations, we have offloaded that part of the pipeline onto a remote server. In this talk, we will describe how to use these pipelines in Phenix, some example results that can be achieved, other updates in Phenix, and some future developments.
The cctbx.xfel suite of processing programs and tools allows fast, visual analysis of serial diffraction images from synchrotrons and XFELs. Built on DIALS and cctbx, cctbx.xfel is designed for real-time and post-experiment processing with a fully featured graphical user interface. Users can quickly identify hitrates, view diffraction patterns, analyze unit-cell isomorphism using clustering, and merge data using a metadata tagging approach that allows on-the-fly organization and visualization of processing results. This paper describes the fundamental algorithms and command-line programs used by cctbx.xfel, including the two main program dials.stills_process, which performs spot-finding, indexing, geometric refinement, and integration, and cctbx.xfel.merge, which performs scaling, post-refinement, and merging. A discussion of merging statistics is presented and newer features are described, including random sub-sampling for indexing multi-lattice hits and ΔCC 1/2 filtering to remove outliers. Finally we show a complex, heterogeneous sample containing hexagonal and monoclinic isoforms in P63 and P21. The isoforms are separated by unit cell clustering, and for each isoform we resolve a (pseudo-)merohedral indexing ambiguity.
In this presentation, we will highlight the newest developments in Phenix. Our newest release, Phenix 2.0, includes the following new features: more accurate restraints for ligands in situ using quantum mechanical calculations, quantum refinement for proteins using machine learning, improved tools for accessing RCSB web services (such as the PDB validation report), and simplified usage of reference models for reciprocal space refinement which is especially useful at low resolution. We will describe how to use these new tools, discuss other updates in Phenix, and explore future developments.
Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.
The interpretation of cryo-EM maps often includes the docking of known or predicted structures of the components, which is particularly useful when the map resolution is worse than 4 Å. Although it can be effective to search the entire map to find the best placement of a component, the process can be slow when the maps are large. However, frequently there is a well-founded hypothesis about where particular components are located. In such cases, a local search using a map subvolume will be much faster because the search volume is smaller, and more sensitive because optimizing the search volume for the rotation-search step enhances the signal to noise. A Fourier-space likelihood-based local search approach, based on the previously published em_placement software, has been implemented in the new emplace_local program. Tests confirm that the local search approach enhances the speed and sensitivity of the computations. An interactive graphical interface in the ChimeraX molecular-graphics program provides a convenient way to set up and evaluate docking calculations, particularly in defining the part of the map into which the components should be placed.
Advances in machine learning have enabled sufficiently accurate predictions of protein structure to be used in macromolecular structure determination with crystallography and cryo-electron microscopy data. The Phenix software suite has AlphaFold predictions integrated into an automated pipeline that can start with an amino acid sequence and data, and automatically perform model-building and refinement to return a protein model fitted into the data. Due to the steep technical requirements of running AlphaFold efficiently, we have implemented a Phenix-AlphaFold webservice that enables all Phenix users to run AlphaFold predictions remotely from the Phenix GUI starting with the official 1.21 release. This webservice will be improved based on how it is used by the research community and the future research directions for Phenix.
SummaryThe upcoming exascale computing systems Frontier and Aurora will draw much of their computing power from GPU accelerators. The hardware for these systems will be provided by AMD and Intel, respectively, each supporting their own GPU programming model. The challenge for applications that harness one of these exascale systems will be to avoid lock‐in and to preserve performance portability. We report here on our results of using Kokkos to accelerate a real‐world application on NERSC's Perlmutter Phase 1 (using NVIDIA A100 accelerators) and Crusher, the testbed system for OLCF's Frontier (using AMD MI250X). By porting to Kokkos, we successfully ran the same X‐ray tracing code on both systems and achieved speed‐ups between 13 % and 66 % compared to the original CUDA code. These results are a highly encouraging demonstration of using Kokkos to accelerate production science code.
ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.
Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.
Neutron diffraction is one of the three crystallographic techniques (X-ray, neutron and electron diffraction) used to determine the atomic structures of molecules. Its particular strengths derive from the fact that H (and D) atoms are strong neutron scatterers, meaning that their positions, and thus protonation states, can be derived from crystallographic maps. However, because of technical limitations and experimental obstacles, the quality of neutron diffraction data is typically much poorer (completeness, resolution and signal to noise) than that of X-ray diffraction data for the same sample. Further, refinement is more complex as it usually requires additional parameters to describe the H (and D) atoms. The increase in the number of parameters may be mitigated by using the `riding hydrogen' refinement strategy, in which the positions of H atoms without a rotational degree of freedom are inferred from their neighboring heavy atoms. However, this does not address the issues related to poor data quality. Therefore, neutron structure determination often relies on the presence of an X-ray data set for joint X-ray and neutron (XN) refinement. In this approach, the X-ray data serve to compensate for the deficiencies of the neutron diffraction data by refining one model simultaneously against the X-ray and neutron data sets. To be applicable, it is assumed that both data sets are highly isomorphous, and preferably collected from the same crystals and at the same temperature. However, the approach has a number of limitations that are discussed in this work by comparing four separately re-refined neutron models. To address the limitations, a new method for joint XN refinement is introduced that optimizes two different models against the different data sets. This approach is tested using neutron models and data deposited in the Protein Data Bank. The efficacy of refining models with H atoms as riding or as individual atoms is also investigated.
This chapter discusses the use of diffraction simulators to improve experimental outcomes in macromolecular crystallography, in particular for future experiments aimed at diffuse scattering. Consequential decisions for upcoming data collection include the selection of either a synchrotron or free electron laser X-ray source, rotation geometry or serial crystallography, and fiber-coupled area detector technology vs. pixel-array detectors. The hope is that simulators will provide insights to make these choices with greater confidence. Simulation software, especially those packages focused on physics-based calculation of the diffraction, can help to predict the location, size, shape, and profile of Bragg spots and diffuse patterns in terms of an underlying physical model, including assumptions about the crystal's mosaic structure, and therefore can point to potential issues with data analysis in the early planning stages. Also, once the data are collected, simulation may offer a pathway to improve the measurement of diffraction, especially with weak data, and might help to treat problematic cases such as overlapping patterns.
Experimental structure determination can be accelerated with AI-based structure prediction methods such as AlphaFold. Here we present an automatic procedure requiring only sequence information and crystallographic data that uses AlphaFold predictions to produce an electron density map and a structural model. Iterating through cycles of structure prediction is a key element of our procedure: a predicted model rebuilt in one cycle is used as a template for prediction in the next cycle. We applied this procedure to X-ray data for 215 structures released by the Protein Data Bank in a recent 6-month period. In 87% of cases our procedure yielded a model with at least 50% of C α atoms matching those in the deposited models within 2Å. Predictions from our iterative template-guided prediction procedure were more accurate than those obtained without templates. We suggest a general strategy for macromolecular structure determination that includes AI-based prediction both as a starting point and as a method of model optimization.
Cryo-EM observation of biological samples enables visualization of sample heterogeneity, in the form of discrete states that are separable, or continuous heterogeneity as a result of local protein motion before flash freezing. Variability analysis of this continuous heterogeneity describes the variance between a particle stack and a volume, and results in a map series describing the various steps undertaken by the sample in the particle stack. While this observation is absolutely stunning, it is very hard to pinpoint structural details to elements of the maps. In order to bridge the gap between observation and explanation, we designed a tool that refines an ensemble of structures into all the maps from variability analysis. Using this bundle of structures, it is easy to spot variable parts of the structure, as well as the parts that are not moving. Comparison with molecular dynamics simulations highlights the fact that the movements follow the same directions, albeit with different amplitudes. Ligand can also be investigated using this method. Variability refinement is available in the Phenix software suite, accessible under the program name phenix.varref.
In macromolecular crystallographic structure refinement, ligands present challenges for the generation of geometric restraints due to their large chemical variability, their possible novel nature and their specific interaction with the binding pocket of the protein. Quantum-mechanical approaches are useful for providing accurate ligand geometries, but can be plagued by the number of minima in flexible molecules. In an effort to avoid these issues, the Quantum Mechanical Restraints (QMR) procedure optimizes the ligand geometry in situ, thus accounting for the influence of the macromolecule on the local energy minima of the ligand. The optimized ligand geometry is used to generate target values for geometric restraints during the crystallographic refinement. As demonstrated using a sample of >2330 ligand instances in >1700 protein–ligand models, QMR restraints generally result in lower deviations from the target stereochemistry compared with conventionally generated restraints. In particular, the QMR approach provides accurate torsion restraints for ligands and other entities.
Machine-learning prediction algorithms such as AlphaFold and RoseTTAFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including new experimental information such as a density map, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt on the basis of experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We show that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for interpretation of crystallographic and electron cryo-microscopy maps.
Advancements in x-ray free-electron lasers on producing ultrashort, ultrabright, and coherent x-ray pulses enable single-shot imaging of fragile nanostructures, such as superfluid helium droplets. This imaging technique gives unique access to the sizes and shapes of individual droplets. In the past, such droplet characteristics have only been indirectly inferred by ensemble averaging techniques. Here, we report on the size distributions of both pure and doped droplets collected from single-shot x-ray imaging and produced from the free-jet expansion of helium through a 5 μm diameter nozzle at 20 bars and nozzle temperatures ranging from 4.2 to 9 K. This work extends the measurement of large helium nanodroplets containing 109-1011 atoms, which are shown to follow an exponential size distribution. Additionally, we demonstrate that the size distributions of the doped droplets follow those of the pure droplets at the same stagnation condition but with smaller average sizes.
AlphaFold has recently become an important tool in providing models for experimental structure determination by X-ray crystallography and cryo-EM. Large parts of the predicted models typically approach the accuracy of experimentally determined structures, although there are frequently local errors and errors in the relative orientations of domains. Importantly, residues in the model of a protein predicted by AlphaFold are tagged with a predicted local distance difference test score, informing users about which regions of the structure are predicted with less confidence. AlphaFold also produces a predicted aligned error matrix indicating its confidence in the relative positions of each pair of residues in the predicted model. The phenix.process_predicted_ model tool downweights or removes low-confidence residues and can break a model into confidently predicted domains in preparation for molecular replacement or cryo-EM docking. These confidence metrics are further used in ISOLDE to weight torsion and atom-atom distance restraints, allowing the complete AlphaFold model to be interactively rearranged to match the docked fragments and reducing the need for the rebuilding of connecting regions.
Machine learning prediction algorithms such as AlphaFold can create remarkably accurate protein models, but these models usually have some regions that are predicted with low confidence or poor accuracy. We hypothesized that by implicitly including experimental information, a greater portion of a model could be predicted accurately, and that this might synergistically improve parts of the model that were not fully addressed by either machine learning or experiment alone. An iterative procedure was developed in which AlphaFold models are automatically rebuilt based on experimental density maps and the rebuilt models are used as templates in new AlphaFold predictions. We find that including experimental information improves prediction beyond the improvement obtained with simple rebuilding guided by the experimental data. This procedure for AlphaFold modeling with density has been incorporated into an automated procedure for crystallographic and electron cryo-microscopy map interpretation.AlphaFold modeling can be improved synergistically by including information from experimental density maps.