The GloSAT project is developing a new observational analysis of global air temperature change over land and ocean since the late 18th century. A new global analysis processing system has been developed that uses a computationally efficient spatial statistical method to estimate air temperature anomaly fields from historical observations. This will be the first presentation of this analysis approach. This method, based on Gaussian Markov Random Fields, jointly estimates temperature anomaly fields over land and ocean based on weather station and ship-based air temperature observations. The increased computational efficiency of the approach compared to conventional kriging-based estimates allows for increased spatial resolution in the analysis. Observational uncertainties are represented within the analysis framework to propagate uncertainty into the output ensemble data set. This accounts for errors arising from uncorrelated effects and structured errors such as residual biases in observations from an individual weather station or ship after correction. Observational error models have been co-developed with project partners providing the input land and marine data products. Initial results from the application of the analysis system to GloSAT air temperature observation data will be demonstrated.
Nucleic acid electron density interpretation after phasing by molecular replacement or other methods remains a difficult problem for computer programs to deal with. Programs tend to rely on time-consuming and computationally exhaustive searches to recognise characteristic features. We present NucleoFind, a deep-learning-based approach to interpreting and segmenting electron density. Using an electron density map from X-ray crystallography obtained after molecular replacement, the positions of the phosphate group, sugar ring and nitrogenous base group can be predicted with high accuracy. On average, 78% of phosphate atoms, 85% of sugar atoms and 83% of base atoms are positioned in predicted density after giving NucleoFind maps produced following successful molecular replacement. NucleoFind can use the wealth of context these predicted maps provide to build more accurate and complete nucleic acid models automatically.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
<p>We present a new data set of air temperature change across land and ocean extending back to the late-18<sup>th</sup> century. This new data set uses marine air temperature observations rather than the sea surface temperature measurements typically used by pre-existing data sets. This allows the new data set to extend further into the past than existing instrumental temperature records, which typically have start dates in the mid-to-late 19<sup>th</sup> century. The new data set brings together advances in understanding of measurement biases affecting all-day marine air temperature observations with a new assessment of the effects of non-standard thermometer enclosures used at land meteorological stations in the early instrumental record. A further innovation is the use of kriging to obtain localised temperature estimates that allow land air temperature records to be converted into anomalies even for stations without observations during the baseline period.&#160;Global and hemispheric series show close agreement with those based on sea-surface temperature for much of the overlapping period of their records, some of the interesting differences will be presented. This data set has been developed under the GloSAT project (https://www.glosat.org/).</p>
The Collaborative Computational Project No. 4 (CCP4) is a UK-led international collective with a mission to develop, test, distribute and promote software for macromolecular crystallography. The CCP4 suite is a multiplatform collection of programs brought together by familiar execution routines, a set of common libraries and graphical interfaces. The CCP4 suite has experienced several considerable changes since its last reference article, involving new infrastructure, original programs and graphical interfaces. This article, which is intended as a general literature citation for the use of the CCP4 software suite in structure determination, will guide the reader through such transformations, offering a general overview of the new features and outlining future developments. As such, it aims to highlight the individual programs that comprise the suite and to provide the latest references to them for perusal by crystallographers around the world.
Tracing the backbone is a critical step in protein model building, as incorrect tracing leads to poor protein models. Here, a neural network trained to identify unfavourable fragments and remove them from the model-building process in order to improve backbone tracing is presented. Moreover, a decision tree was trained to select an optimal threshold to eliminate unfavourable fragments. The neural network was tested on experimental phasing data sets from the Joint Center for Structural Genomics (JCSG), recently deposited experimental phasing data sets (from 2015 to 2021) and molecular-replacement data sets. The experimental results show that using the neural network in the Buccaneer protein-model-building software can produce significantly more complete protein models than those built using Buccaneer alone. In particular, Buccaneer with the neural network built protein models with a completeness that was at least 5% higher for 25% and 50% of the original and truncated resolution JCSG experimental phasing data sets, respectively, for 28% of the recently collected experimental phasing data sets and for 43% of the molecular-replacement data sets.
Recent assessments of climate sensitivity per doubling of atmospheric CO 2 concentration have combined likelihoods derived from multiple lines of evidence. These assessments were very influential in the Intergovernmental Panel on Climate Change Sixth Assessment Report (AR6) assessment of equilibrium climate sensitivity, the likely range lower limit of which was raised to 2.5 °C (from 1.5 °C previously). This study evaluates the methodology of and results from a particularly influential assessment of climate sensitivity that combined multiple lines of evidence, Sherwood et al. (Rev Geophys 58(4):e2019RG000678, 2020 ). That assessment used a subjective Bayesian statistical method, with an investigator-selected prior distribution. This study estimates climate sensitivity using an Objective Bayesian method with computed, mathematical priors, since subjective Bayesian methods may produce uncertainty ranges that poorly match confidence intervals. Identical model equations and, initially, identical input values to those in Sherwood et al. are used. This study corrects Sherwood et al.'s likelihood estimation, producing estimates from three methods that agree closely with each other, but differ from those that they derived. Finally, the selection of input values is revisited, where appropriate adopting values based on more recent evidence or that otherwise appear better justified. The resulting estimates of long-term climate sensitivity are much lower and better constrained (median 2.16 °C, 17–83% range 1.75–2.7 °C, 5–95% range 1.55–3.2 °C) than in Sherwood et al. and in AR6 (central value 3 °C, very likely range 2.0–5.0 °C). This sensitivity to the assumptions employed implies that climate sensitivity remains difficult to ascertain, and that values between 1.5 °C and 2 °C are quite plausible.
Despite the abundance of available software tools, optimal particle selection is still a vital issue in single-particle cryoelectron microscopy (cryo-EM). Regardless of the method used, most pickers struggle when ice thickness varies on a micrograph. IceBreaker allows users to estimate the relative ice gradient and flatten it by equalizing the local contrast. It allows the differentiation of particles from the background and improves overall particle picking performance. Furthermore, we introduce an additional parameter corresponding to local ice thickness for each particle. Particles with a defined ice thickness can be grouped and filtered based on this parameter during processing. These functionalities are especially valuable for on-the-fly processing to automatically pick as many particles as possible from each micrograph and to select optimal regions for data collection. Finally, estimated ice gradient distributions can be stored separately and used to inspect the quality of prepared samples.
Interactive model building can be a difficult and time-consuming step in the structure-solution process. Automated model-building programs such as Buccaneer often make it quicker and easier by completing most of the model in advance. However, they may fail to do so with low-resolution data or a poor initial model or map. The Buccaneer pipeline is a relatively simple program that iterates Buccaneer with REFMAC to refine the model and update the map. A new pipeline called ModelCraft has been developed that expands on this to include shift-field refinement, machine-learned pruning of incorrect residues, classical density modification, addition of water and dummy atoms, building of nucleic acids and final rebuilding of side chains. Testing was performed on 1180 structures solved by experimental phasing, 1338 structures solved by molecular replacement using homologues and 2030 structures solved by molecular replacement using predicted AlphaFold models. Compared with the previous Buccaneer pipeline, ModelCraft increased the mean completeness of the protein models in the experimental phasing cases from 91% to 95%, the molecular-replacement cases from 50% to 78% and the AlphaFold cases from 82% to 91%.
Recently, there has been a dramatic improvement in the quality and quantity of data derived using cryogenic electron microscopy (cryo-EM). This is also associated with a large increase in the number of atomic models built. Although the best resolutions that are achievable are improving, often the local resolution is variable, and a significant majority of data are still resolved at resolutions worse than 3 Å. Model building and refinement is often challenging at these resolutions, and hence atomic model validation becomes even more crucial to identify less reliable regions of the model. Here, a graphical user interface for atomic model validation, implemented in the CCP-EM software suite, is presented. It is aimed to develop this into a platform where users can access multiple complementary validation metrics that work across a range of resolutions and obtain a summary of evaluations. Based on the validation estimates from atomic models associated with cryo-EM structures from SARS-CoV-2, it was observed that models typically favor adopting the most common conformations over fitting the observations when compared with the model agreement with data. At low resolutions, the stereochemical quality may be favored over data fit, but care should be taken to ensure that the model agrees with the data in terms of resolvable features. It is demonstrated that further re-refinement can lead to improvement of the agreement with data without the loss of geometric quality. This also highlights the need for improved resolution-dependent weight optimization in model refinement and an effective test for overfitting that would help to guide the refinement process.
Nowadays, progress in the determination of three-dimensional macromolecular structures from diffraction images is achieved partly at the cost of increasing data volumes. This is due to the deployment of modern high-speed, high-resolution detectors, the increased complexity and variety of crystallographic software, the use of extensive databases and high-performance computing. This limits what can be accomplished with personal, offline, computing equipment in terms of both productivity and maintainability. There is also an issue of long-term data maintenance and availability of structure-solution projects as the links between experimental observations and the final results deposited in the PDB. In this article, CCP4 Cloud, a new front-end of the CCP4 software suite, is presented which mitigates these effects by providing an online, cloud-based environment for crystallographic computation. CCP4 Cloud was developed for the efficient delivery of computing power, database services and seamless integration with web resources. It provides a rich graphical user interface that allows project sharing and long-term storage for structure-solution projects, and can be linked to data-producing facilities. The system is distributed with the CCP4 software suite version 7.1 and higher, and an online publicly available instance of CCP4 Cloud is provided by CCP4.
For half a century the refinement of atomic model parameters to best explain the observed diffraction pattern has been fundamental to the process of crystallographic structure solution. This process has traditionally been carried out by the optimization of individual atomic parameters, with the use of stereochemical restraints to maintain plausible model geometry, particularly when data resolution is poor. However the data are often too poor to reliably indicate how individual atoms should be moved, and as a result the refinement calculation becomes a protracted battle between the noisy data pulling atoms in different directions and the restraints which are trying to maintain model geometry. This limits both the speed and radius of convergence of the calculation. Shift field refinement is a new approach in which shifts to the calculated electron density are determined over extended regions of the unit cell, where the region size may be varied according to the resolution of the data and the type of feature (from whole domain to individual atom) being refined. The enables refinement to capture large domain shifts at low resolution, and to be applied at any resolution with rapid convergence. We have already demonstrated improved molecular replacement results when incorporating this step. We now demonstrate how the method can be used to refine a map against a set of diffraction observations, even in the absence of an atomic model. We also demonstrate how the incorporation of a separate regularization step can be used to improve the refinement results by allowing more cycles of refinement to be run without the risk of model degradation due to accumulated model distortions. This in turn leads to further improvements in the refinement results.
Proteins are macromolecules that perform essential biological functions which depend on their three-dimensional structure. Determining this structure involves complex laboratory and computational work. For the computational work, multiple software pipelines have been developed to build models of the protein structure from crystallographic data. Each of these pipelines performs differently depending on the characteristics of the electron-density map received as input. Identifying the best pipeline to use for a protein structure is difficult, as the pipeline performance differs significantly from one protein structure to another. As such, researchers often select pipelines that do not produce the best possible protein models from the available data. Here, a software tool is introduced which predicts key quality measures of the protein structures that a range of pipelines would generate if supplied with a given crystallographic data set. These measures are crystallographic quality-of-fit indicators based on included and withheld observations, and structure completeness. Extensive experiments carried out using over 2500 data sets show that the tool yields accurate predictions for both experimental phasing data sets (at resolutions between 1.2 and 4.0 Å) and molecular-replacement data sets (at resolutions between 1.0 and 3.5 Å). The tool can therefore provide a recommendation to the user concerning the pipelines that should be run in order to proceed most efficiently to a depositable model.
With many software tools available the optimal particle selection is still a vital issue in the single particle cryoEM. Regardless of the methods used, most pickers struggle when the varying ice thickness is present on the micrograph. We present IceBreaker, which allows us to estimate the relative ice gradient and flatten it based on the K-Means clustering algorithm, thus equalizing the local contrast. It allows differentiation of the particles from the background, and improves the particle pickers performance. Furthermore, a new parameter corresponding to the local ice thickness is introduced for each picked particle. Particles with a defined ice thickness can be grouped, sorted, and filtered based on this parameter during processing. Single particle 3D reconstructions can be made from particles in each ice group to access the effect of ice thickness. These functionalities are especially valuable for on-the-fly processing to automatically pick as many particles as possible from each micrograph and to select optimal ice regions for data collection. The software can be also used to evaluate the quality of the collected data, and also of already refined maps, deposited with the coordinates of the selected particles and to assess how the particles from different ice thickness areas contributed to the final map. Finally, the estimated ice gradient distributions can be stored separately and used to inspect the general quality of prepared samples.
This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.
For the last two decades, researchers have worked independently to automate protein model building, and four widely used software pipelines have been developed for this purpose: ARP/wARP, Buccaneer, Phenix AutoBuild and SHELXE. Here, the usefulness of combining these pipelines to improve the built protein structures by running them in pairwise combinations is examined. The results show that integrating these pipelines can lead to significant improvements in structure completeness and Rfree. In particular, running Phenix AutoBuild after Buccaneer improved structure completeness for 29% and 75% of the data sets that were examined at the original resolution and at a simulated lower resolution, respectively, compared with running Phenix AutoBuild on its own. In contrast, Phenix AutoBuild alone produced better structure completeness than the two pipelines combined for only 7% and 3% of these data sets.
Structural biases, which are intrinsic in the social structures in which we function, play a key role in maintaining boundaries between traditionally privileged and underprivileged groups; however, they are particularly difficult to identify from within those societies. Two instances are highlighted in which the social structures of science appear to have discouraged collaboration, to the disadvantage of software and data users. Possible links are suggested to the strongly hierarchical structure of science and other factors which may in turn also serve to maintain sex and/or gender disparities in participation in the scientific endeavour.
In a recent paper, Kossin et al. showed that during the period from 1979 to 2017, there was a statistically significant increase in the ratio of category 3–5 to category 1–5 tropical storm fixes in the ADT-HURSAT satellite dataset of tropical cyclone observations. The sign of this increase is consistent with previously developed theory and modelling results for how tropical cyclones may change due to climate change. However, without further analysis, it is difficult to understand what the implications of this increase might be for present day tropical cyclone risk. It is also difficult to understand how tropical cyclone risk models might be adjusted to reflect this increase, since this ratio is not typically directly represented in such models. Our goal is therefore to understand the drivers for this increase in terms of changes in the numbers of fixes of different categories of storms in different basins, which are quantities that are more directly related to tropical cyclone risk and risk modelling. We use both heuristic and quantitative methods. We find that the increase in the ratio is mainly driven by a decrease in the denominator (the number of category 1–5 fixes) and to a small extent by a slight increase in the numerator (the number of category 3–5 fixes). The decrease in the denominator is mostly driven by a statistically significant reduction in the number of category 1 fixes outside the North Atlantic. The slight increase in the numerator is mostly driven by a statistically significant increase in the number of category 3–4 fixes in the North Atlantic. Based on these results, we discuss different ways in which the increase in the ratio could be represented in risk models.
The aim of crystallographic structure solution is typically to determine an atomic model which accurately accounts for an observed diffraction pattern. A key step in this process is the refinement of the parameters of an initial model, which is most often determined by molecular replacement using another structure which is broadly similar to the structure of interest. In macromolecular crystallography, the resolution of the data is typically insufficient to determine the positional and uncertainty parameters for each individual atom, and so stereochemical information is used to supplement the observational data. Here, a new approach to refinement is evaluated in which a `shift field' is determined which describes changes to model parameters affecting whole regions of the model rather than individual atoms only, with the size of the affected region being a key parameter of the calculation which can be changed in accordance with the resolution of the data. It is demonstrated that this approach can improve the radius of convergence of the refinement calculation while also dramatically reducing the calculation time.
This is the full dataset of the 2019 Cryo-EM Map-based Model Metrics Challenge sponsored by EMDataResource (www.emdataresource.org, challenges.emdataresource.org, model-compare.emdataresource.org). The goals of this challenge were (1) to assess the quality of models that can be produced using current modeling software, (2) to check the reproducibility of modeling results from different software developers and users, and (3) compare the performance of current metrics used for evaluation of models. The focus was on near-atomic resolution maps with an innovative twist: three of four target maps formed a resolution series (1.8 to 3.1 Å) from the same specimen and imaging experiment. Tools developed in previous challenges were expanded for managing, visualizing and analyzing the 63 submitted coordinate models, and several new metrics were introduced. File Descriptions: 2019-EMDataResource-Challenge-web.pdf: Archive of News, Goals, Timeline, Targets, Modelling Instructions, Process, FAQ, Submission Instructions, Submission Summary Statistics source from the EMDR Challenges website correlation-images.tar.gz: Pairwise correlation tables for selected metric scores from the EMDR Model Compare website maps.tar.gz: The maps used for Fit-to-Map analyses in the Challenge models.tar.gz: The 63 models submitted by the modelling teams results.tar.gz: The output logs for all of the analysis methods Scores.xlsx: Scores for each model and analysis method, compiled into spreadsheet format targets.tar.gz: The reference models used in the analysis Post submission correction to the web archive PDF document: The full list of EMDataResource members on the model committee is as follows: Cathy Lawson, Andriy Kryshtafovych, Greg Pintilie, Mike Schmid, Helen Berman, Wah Chiu.