Determination of specimen structure from cryoelectron microscopy (cryo-EM) experiments relies on an accurate model of the electrostatic potential of the specimen. For biological macromolecules, the potential is strongly influenced by the presence of chemical bonds between atoms, a fact unaccounted for by models of electron scattering that are currently standard in the field. We propose a Bayesian approach to the estimation of atomic scattering factors which incorporates the effect of the molecular environment while remaining fast, interpretable, and transferable between molecules. Our algorithm infers atomic scattering factors directly from maps of the electrostatic potential determined by cryo-EM single particle analysis, bypassing the need for computationally intensive theoretical calculations. The algorithm is used to infer empirical scattering factors from high-resolution reconstructions of catalase enzymes. To illustrate its broad applicability, the algorithm is also applied to a training set of publicly available cryo-EM data. The empirical scattering factors show improved agreement with a test set of cryo-EM reconstructions, decreasing the variance of unmodeled signal in the data by up to a factor of three in the resolution range 1/15 to 1/3 Å-1. The predictions are further validated by comparison with magnetic susceptibility values of organic compounds, as well as by application to the refinement of atomic models.
The Collaborative Computational Project No. 4 (CCP4) is a UK-led international collective with a mission to develop, test, distribute and promote software for macromolecular crystallography. The CCP4 suite is a multiplatform collection of programs brought together by familiar execution routines, a set of common libraries and graphical interfaces. The CCP4 suite has experienced several considerable changes since its last reference article, involving new infrastructure, original programs and graphical interfaces. This article, which is intended as a general literature citation for the use of the CCP4 software suite in structure determination, will guide the reader through such transformations, offering a general overview of the new features and outlining future developments. As such, it aims to highlight the individual programs that comprise the suite and to provide the latest references to them for perusal by crystallographers around the world.
Macromolecular refinement uses experimental data together with prior chemical knowledge (usually digested into geometrical restraints) to optimally fit an atomic structural model into experimental data, while ensuring that the model is chemically plausible. In the CCP4 suite this chemical knowledge is stored in a Monomer Library, which comprises a set of restraint dictionaries. To use restraints in refinement, the model is analysed and template restraints from the dictionary are used to infer (i) restraints between concrete atoms and (ii) the positions of riding hydrogen atoms. Recently, this mundane process has been overhauled. This was also an opportunity to enhance the Monomer Library with new features, resulting in a small improvement in REFMAC5 refinement. Importantly, the overhaul of this part of CCP4 has increased flexibility and eased experimentation, opening up new possibilities.
The third area is working with electron density maps – real or complex values on a 3D grid. Electron density can be calculated from both the structural model and experimental data. The functionality here includes reading and writing files in the MRC/CCP4 map format, analysing and modifying the density, and using the fast Fourier transform to switch between the so-called direct space and the reciprocal space.
Nowadays, progress in the determination of three-dimensional macromolecular structures from diffraction images is achieved partly at the cost of increasing data volumes. This is due to the deployment of modern high-speed, high-resolution detectors, the increased complexity and variety of crystallographic software, the use of extensive databases and high-performance computing. This limits what can be accomplished with personal, offline, computing equipment in terms of both productivity and maintainability. There is also an issue of long-term data maintenance and availability of structure-solution projects as the links between experimental observations and the final results deposited in the PDB. In this article, CCP4 Cloud, a new front-end of the CCP4 software suite, is presented which mitigates these effects by providing an online, cloud-based environment for crystallographic computation. CCP4 Cloud was developed for the efficient delivery of computing power, database services and seamless integration with web resources. It provides a rich graphical user interface that allows project sharing and long-term storage for structure-solution projects, and can be linked to data-producing facilities. The system is distributed with the CCP4 software suite version 7.1 and higher, and an online publicly available instance of CCP4 Cloud is provided by CCP4.
Covalent linkages between constituent blocks of macromolecules and ligands have been subject to inconsistent treatment during the model-building, refinement and deposition process. This may stem from a number of sources, including difficulties with initially detecting the covalent linkage, identifying the correct chemistry, obtaining an appropriate restraint dictionary and ensuring its correct application. The analysis presented herein assesses the extent of problems involving covalent linkages in the Protein Data Bank (PDB). Not only will this facilitate the remediation of existing models, but also, more importantly, it will inform and thus improve the quality of future linkages. By considering linkages of known type in the CCP4 Monomer Library (CCP4-ML), failure to model a covalent linkage is identified to result in inaccurate (systematically longer) interatomic distances. Scanning the PDB for proximal atom pairs that do not have a corresponding type in the CCP4-ML reveals a large number of commonly occurring types of unannotated potential linkages; in general, these may or may not be covalently linked. Manual consideration of the most commonly occurring cases identifies a number of genuine classes of covalent linkages. The recent expansion of the CCP4-ML is discussed, which has involved the addition of over 16 000 and the replacement of over 11 000 component dictionaries using AceDRG. As part of this effort, the CCP4-ML has also been extended using AceDRG link dictionaries for the aforementioned linkage types identified in this analysis. This will facilitate the identification of such linkage types in future modelling efforts, whilst concurrently easing the process involved in their application. The need for a universal standard for maintaining link records corresponding to covalent linkages, and references to the associated dictionaries used during modelling and refinement, following deposition to the PDB is emphasized. The importance of correctly modelling covalent linkages is demonstrated using a case study, which involves the covalent linkage of an inhibitor to the main protease in various viral species, including SARS-CoV-2. This example demonstrates the importance of properly modelling covalent linkages using a comprehensive restraint dictionary, as opposed to just using a single interatomic distance restraint or failing to model the covalent linkage at all.
In this contribution, the current protocols for modelling covalent linkages within the CCP4 suite are considered. The mechanism used for modelling covalent linkages is reviewed: the use of dictionaries for describing changes to stereochemistry as a result of the covalent linkage and the application of link-annotation records to structural models to ensure the correct treatment of individual instances of covalent linkages. Previously, linkage descriptions were lacking in quality compared with those of contemporary component dictionaries. Consequently, AceDRG has been adapted for the generation of link dictionaries of the same quality as for individual components. The approach adopted by AceDRG for the generation of link dictionaries is outlined, which includes associated modifications to the linked components. A number of tools to facilitate the practical modelling of covalent linkages available within the CCP4 suite are described, including a new restraint-dictionary accumulator, the Make Covalent Link tool and AceDRG interface in Coot, the 3D graphical editor JLigand and the mechanisms for dealing with covalent linkages in the CCP4i2 and CCP4 Cloud environments. These integrated solutions streamline and ease the covalent-linkage modelling workflow, seamlessly transferring relevant information between programs. Current recommended practice is elucidated by means of instructive practical examples. By summarizing the different approaches to modelling linkages that are available within the CCP4 suite, limitations and potential pitfalls that may be encountered are highlighted in order to raise awareness, with the intention of improving the quality of future modelled covalent linkages in macromolecular complexes.
The conventional approach to search-model identification in molecular replacement (MR) is to screen a database of known structures using the target sequence. However, this strategy is not always effective, for example when the relationship between sequence and structural similarity fails or when the crystal contents are not those expected. An alternative approach is to identify suitable search models directly from the experimental data. SIMBAD is a sequence-independent MR pipeline that uses either a crystal lattice search or MR functions to directly locate suitable search models from databases. The previous version of SIMBAD used the fast AMoRe rotation-function search. Here, a new version of SIMBAD which makes use of Phaser and its likelihood scoring to improve the sensitivity of the pipeline is presented. It is shown that the additional compute time potentially required by the more sophisticated scoring is counterbalanced by the greater sensitivity, allowing more cases to trigger early-termination criteria, rather than running to completion. Using Phaser solved 17 out of 25 test cases in comparison to the ten solved with AMoRe, and it is shown that use of ensemble search models produces additional performance benefits.
The Collaborative Computational Project Number 4 in Protein Crystallography (CCP4) exists to maintain, develop and provide world-class software that allows researchers to determine macromolecular structures by X-ray crystallography and other biophysical techniques.Over 40 years of existence, CCP4 Software was assembled and distributed as an integrated Suite of programs, traditionally operated via CCP4i (2) GUI in Linux, OSX and Windows platforms.Modern trends in computing suggest a fast-growing interest to mobile platforms and cloud solutions for data management and operations in practically all areas.Answering to these trends, CCP4 releases beta version of CCP4 Cloud, developed for essentially remote and distributed deployment of CCP4 Software.CCP4 Cloud allows a user to keep all necessary data and projects in the cloud and perform all scope of crystallographic computations, from image processing to final refinement, ligand fitting and deposition, remotely via a common web-browser, optionally complemented with CCP4 Cloud Client for interactive model building with Coot.The talk will present architectural solutions and key features of CCP4 Cloud such as ability to seamlessly import data from 3rd party sites (e.g., synchrotrons), high scalability of computational background, convenient (big) data management for multiple users, rich graphical interface with built-in molecular graphics, enhanced data and structure solution pathway provenance.CCP4 Cloud may be used from CCP4 Web portal with any device running a modern web-browser (including tablets and smartphones).All the source code is open and freely available for installation elsewhere to serve local researchers in a lab, or institution, or pharma, or a synchrotron.
This letter announces that PDBx/mmCIF format files will become mandatory for crystallographic depositions to the Protein Data Bank (PDB).
The CCP4 (Collaborative Computational Project, Number 4) software suite for macromolecular structure determination by X-ray crystallography groups brings together many programs and libraries that, by means of well established conventions, interoperate effectively without adhering to strict design guidelines. Because of this inherent flexibility, users are often presented with diverse, even divergent, choices for solving every type of problem. Recently, CCP4 introduced CCP4i2, a modern graphical interface designed to help structural biologists to navigate the process of structure determination, with an emphasis on pipelining and the streamlined presentation of results. In addition, CCP4i2 provides a framework for writing structure-solution scripts that can be built up incrementally to create increasingly automatic procedures.
In protein crystallisation it is not uncommon for contami-
The conventional approach to finding structurally similar search models for use in molecular replacement (MR) is to use the sequence of the target to search against those of a set of known structures. Sequence similarity often correlates with structure similarity. Given sufficient similarity, a known structure correctly positioned in the target cell by the MR process can provide an approximation to the unknown phases of the target. An alternative approach to identifying homologous structures suitable for MR is to exploit the measured data directly, comparing the lattice parameters or the experimentally derived structure-factor amplitudes with those of known structures. Here, SIMBAD, a new sequence-independent MR pipeline which implements these approaches, is presented. SIMBAD can identify cases of contaminant crystallization and other mishaps such as mistaken identity (swapped crystallization trays), as well as solving unsequenced targets and providing a brute-force approach where sequence-dependent search-model identification may be nontrivial, for example because of conformational diversity among identifiable homologues. The program implements a three-step pipeline to efficiently identify a suitable search model in a database of known structures. The first step performs a lattice-parameter search against the entire Protein Data Bank (PDB), rapidly determining whether or not a homologue exists in the same crystal form. The second step is designed to screen the target data for the presence of a crystallized contaminant, a not uncommon occurrence in macromolecular crystallography. Solving structures with MR in such cases can remain problematic for many years, since the search models, which are assumed to be similar to the structure of interest, are not necessarily related to the structures that have actually crystallized. To cater for this eventuality, SIMBAD rapidly screens the data against a database of known contaminant structures. Where the first two steps fail to yield a solution, a final step in SIMBAD can be invoked to perform a brute-force search of a nonredundant PDB database provided by the MoRDa MR software. Through early-access usage of SIMBAD, this approach has solved novel cases that have otherwise proved difficult to solve.
UglyMol (Wojdyr 2016) is a macromolecular viewer specialized in presenting macromolecular models together with the electron density.It uses web technologies (JavaScript and WebGL) and is suitable for embedding in web applications.The project was started as a fork of xtal.js(Echols 2015).
CCP4 and Global Phasing Ltd started in January 2017 a joint open-source project, named Gemmi, to create a new software library that will be used by both organizations.In the first year of the development we focus on handling mmCIF files and monomer library files (ligand CIFs).We started from the lowest level -parsing CIF files and validating them with DDL dictionaries.Next, we will provide interface to manipulate the structure in terms of models, chains, residues and atoms.On top of it we will provide a collection of algorithms, primarily for use in macromolecular refinement programs.Finally, the library will be integrated with BUSTER and Refmac.
CCP4 has been serving the software needs of the protein crystallography community for more than 30 years. In this time the CCP4 Suite of software has been refined through contributions from some of the leading developers in the field of protein crystallographic software and the feedback of both expert and novice users. Today it is a highly comprehensive suite, providing tools and packages covering all aspects from data collection through to structure deposition. Here we will present details of the latest release series of the Suite, version 6.4. This release brings updates to many of the key elements in the Suite. The most obvious of these is the integration of the rolling updates mechanism. This is used to distribute timely fixes, update existing programs and introduce new functionality to users of the suite. Recent updates have seen updates to major programs such as phaser and imosflm/mosflm, and the introduction of a major overhaul of the Experimental Phasing pipeline Crank. An overview is given of the operation behind the updates and releases, including the jhbuild system, repositories and testing, the availability of nightly builds, and work towards the next major release of CCP4. This will see the integration of the CCP4MG package, along with preparations for the introduction of the long awaited CCP4i2.
"In 2013 MX beamlines at the Diamond synchrotron deployed an automated software pipeline, called DIMPLE, for rapid processing of crystals that contain a known protein and possibly a ligand bound. DIMPLE takes the already known ""apo"" structure for the target protein, compares it with the electron density map from X-ray diffraction images, and visualizes areas of the electron density unaccounted for by the structure model. When processing batches of crystals, such feedback allows the user to better decide what to measure next which leads to a more efficient use of the beam time. This year we've enhanced the pipeline to cover more complex cases, including changes in the space group and some changes in conformation. With multiple molecular replacement computations run in parallel, the time from shooting to viewing the difference map is still only a few minutes. While the software is developed primarily for use at synchrotron beamlines, it is included in the CCP4 suite and can be used as well for in-house automation."