A Working Group consisting of the co-authors of this paper was established in 2020 to re-evaluate the standard valence geometry used for the validation of nucleic acid structure models in the Protein Data Bank (PDB). This Working Group re-examined the dependence of Cambridge Structural Database (CSD) derived targets on base and sugar type, sugar pucker, and phosphate and glycosidic conformation, before comparing those targets with the geometry of a quality-filtered reference set of nucleic acid crystal structural models held in the PDB. This revealed that the valence bond and angle mean values are close to the CSD targets, but many parameters have highly non-Gaussian or even multimodal distributions. One explanation is the inconsistency of restraints used over time and by different refinement programs. The Working Group recommends a new validation scheme for use by the PDB. For this purpose, we have developed a new three-tier scale for outlier detection-graded as Preferred, Allowed, and Of Concern intervals-based on a combination of quality-curated reference data from the CSD and the PDB. The proposed approach to validation should lead to improved nucleic acid models in (future) PDB-deposited macromolecular structures.
During photosynthetic water oxidation, the Mn4Ca cluster in Photosystem II progresses through five intermediate Si (i = 0-4) states. X-ray crystallography studies have reported the insertion of one new O ligand during the formation of the S3 state, but recent studies question the presence of this additional ligand based on cryo-EM and earlier room-temperature crystallography data. There is also controversy about whether the O-O bond interaction already occurs in the S3 state or in the subsequent S3 to S0 transition. Here we report conventional high-resolution data for the S1, S2, and S3 states to a resolution of ~1.9 Å, and anomalous diffraction data at two energies (9.5 keV and 7 keV), that was used to model the Mn positions, followed by determination of oxygen positions using the high-resolution maps. We show that the new oxygen atom, OX (or O6), in the S3 state is observable as a distinct peak without any restraints, confirming its ligation to Mn1 and Ca. The OX-O5 distance is ~2.1 Å, supporting no strong interaction between them in the S3 state, suggesting that if this is the O-O bond formation site, it is formed during the S3 to S0 transition initiated by the final oxidation of the cluster.
Multipolar scattering models, such as the transferable aspherical atom model, account for atomic chemical interactions and provide a more accurate representation of experimental data. However, the simpler independent atom model (IAM), which assumes non-interacting atoms, is the only model available in the most widely used macromolecular refinement programs. This is primarily because IAM offers a hard-to-beat combination of computational efficiency and modelling power at typical macromolecular resolutions. By contrast, more accurate multipolar modelling has historically been limited due to its computational cost and the absence of an interface between software capable of calculating structure factors and gradients based on multipolar models and software designed for macromolecular refinement. This work introduces pyDiSCaMB, a Python software package designed to integrate between the computational crystallography toolbox (cctbx) and the quantum crystallography library DiSCaMB (Densities in Structural Chemistry and Molecular Biology), thus enabling multipolar scattering models in Phenix's toolkit. The implementation, features and capabilities of pyDiSCaMB are presented, the runtimes for the calculation of structure factor and target gradients with respect to atomic parameters are explored, and Fourier images of electrostatic potential, electron density and deformation maps are computed as illustrative examples. The pyDiSCaMB library will make multipolar modelling widely available to the structural biology community, potentially transforming refinement and model-building for both crystallography and cryogenic electron microscopy (cryoEM).
In macromolecular structure refinement the low observation-to-parameter ratio and the lack of high-resolution data is countered by using a priori information in the form of restraints. Having accurate geometries of the chemical entities in the sample is paramount for generating accurate chemical restraints and, therefore, accurate macromolecular structures. In particular, it is desirable to have accurate restraints for known and novel ligand entities. Quantum Mechanics (QM) can minimise the energy of a ligand by adjusting its geometry, and these geometries can be used to generate restraints macromolecular refinement. We describe here a library of 37,000 small molecules extracted from the Chemical Components Dictionary in the Protein Data Bank and minimized by density functional QM. The library includes restraint files for use in crystallography or cryo-EM refinement, along with files suitable for molecular dynamics simulation. Because the geometries are validated, the restraints library provides users with both functional restraints and minimised geometries. This work also provides procedures for generating new and accurate restraints.
Our work reveals the structure of the active state of Methyl-Coenzyme M Reductase (MCR), the key and rate-limiting enzyme in biological methane formation. We find large differences between the active Ni(I) and inactive Ni(II) proteins and provide insight into how nature makes and breaks the C-H bond of methane. The Ni(II)-F430 center in inactive MCR contains four planar nitrogen ligands, a lower axial glutamine oxo, and an upper axial thiolate. The Ni(I)-enzyme replaces the axial ligands with a single water. The one-electron redox change results in movement of the Ni ion and upward swing of the β-lactam ring in the tetrapyrrole coupled to a domino-like protein quake through second sphere residues, inter-subunit interactions, a substrate tunnel, affecting even the dimensions of the unit cell. These structural changes lead Ni(I)-MCR to release a charge clamp that, in the Ni(II) state, locks down substrate Coenzyme B. Determining the Ni(I)-MCR structure required development of rigorous anaerobic crystallographic techniques. Validation of the MCR redox state was accomplished by in-line and parallel spectroscopic and unit cell analyses. This structure has large implications for developing technologies to limit methane emissions and efficiently produce biofuels. Methodology described here will enhance structural biology for other oxygen-sensitive enzymes. ### Competing Interest Statement The authors have declared no competing interest. Office of Basic Energy SciencesOffice of Basic Energy Sciences, https://ror.org/05mg91w61, DE-FG02-08ER15931, FWP 100593, DE-AC02-05CH11231, DE-SC0014664, DE-AC02-76SF00515, DEAC02-05CH11231 National Institutes of HealthNational Institutes of Health, https://ror.org/01cwqze88, GM149528, GM110501, GM126289, GM117126, GM151988, P30GM133894
Quantum Mechanical methods provide geometries and energies of molecules using just the atomic and electronic positions. Being independent of the experimental data, they provide complementary information. Furthermore, allowing the experimental and QM methods to share information leads to better results.The Quantum Interface (QI) in Phenix (Liebschner et al., 2019) provides close integration with MOPAC (Moussa & Stewart, 2024) allowing the calculation of in situ restraints for drug candidates. Known as QM Restraints (QMR) (Liebschner et al., 2023), this method will provide protein binding pocket-specific restraints during a refinement or in a stand-alone program.Another QI procedure is a novel approach for predicting histidine protonation states using QM methods. Historically, metal coordination has been fraught.The proposed method, Quantum Mechanical Ions (QMI), employs quantum mechanical calculations to predict metal coordination bond lengths and angles. The study demonstrates the effectiveness of QMI through comparisons with existing methods and validation using high-resolution protein structures.Other applications of QI include histidine protonation (Moriarty et al., 2024) and ligand strain energies, including energy comparisons of disordered ligands, all of which are being pursued.
Quantum Mechanical methods provide geometries and energies of molecules using just the atomic and electronic positions. Being independent of the experimental data, they provide complementary information. Furthermore, allowing the experimental and QM methods to share information leads to better results. The Quantum Interface (QI) in Phenix (Liebschner et al., 2019) provides close integration with MOPAC (Moussa & Stewart, 2024) allowing the calculation of in situ restraints for drug candidates. Known as QM Restraints (QMR) (Liebschner et al., 2023), this method will provide protein binding pocket specific restraints during a refinement or in a stand-alone program. Another QI procedure is a novel approach for predicting histidine protonation states using QM methods. Historically, determining histidine protonation has been challenging due to limited resolution in X-ray crystallography and the inherent difficulty in detecting hydrogen atoms. Previous methods relied on empirical or geometric models, which provided some insights but had limitations. The proposed method, Quantum Mechanical Flipping (QMF) (Moriarty et al., In review), employs quantum mechanical calculations to predict protonation states based on the molecular environment. This approach considers all possible configurations of histidine protonation and assesses their feasibility by minimising geometry and energy calculations. QMF accounts for factors such as hydrogen bonding and molecular interactions providing a more accurate prediction of the most likely protonation state particularly in the binding pocket. The study demonstrates the effectiveness of QMF through comparisons with existing methods and validation using high- resolution protein structures. Results show that QMF can accurately predict histidine protonation states, even in cases with limited experimental data. Additionally, QMF's versatility allows it to be applied to various macromolecular environments, including ligand interactions and non-standard amino acids. Other applications of QI include metal coordination and ligand strain energies, both of which are being pursued.
Cryo-EM and X-ray crystallography provide crucial experimental data for obtaining atomic-detail models of biomacromolecules. Refining these models relies on library-based stereochemical data, which, in addition to being limited to known chemical entities, do not include meaningful noncovalent interactions. Quantum mechanical (QM) calculations could alleviate these issues but are too expensive for large molecules. Here we present a novel AI-enabled Quantum Refinement (AQuaRef) based on AIMNet2 machine learned interatomic potential (MLIP) mimicking QM at substantially lower computational costs. By refining 41 cryo-EM and 30 X-ray structures, we show that this approach yields atomic models with superior geometric quality compared to standard techniques, while maintaining an equal or better fit to experimental data. Notably, AQuaRef aids in determining proton positions, as illustrated in the challenging case of short hydrogen bonds in the parkinsonism-associated human protein DJ-1 and its bacterial homolog YajL.
Pyrone-2,4-dicarboxylic acid (PDC) is a valuable polymer precursor that can be derived from the microbial degradation of lignin. The key enzyme in the microbial production of PDC is 4-carboxy-2-hydroxymuconate-6-semialdehyde (CHMS) dehydrogenase, which acts on the substrate CHMS. We present the crystal structure of CHMS dehydrogenase (PmdC from Comamonas testosteroni) bound to the cofactor NADP, shedding light on its three-dimensional architecture, and revealing residues responsible for binding NADP. Using a combination of structural homology, molecular docking, and quantum chemistry calculations, we have predicted the binding site of CHMS. Key histidine residues in a conserved sequence are identified as crucial for binding the hydroxyl group of CHMS and facilitating dehydrogenation with NADP. Mutating these histidine residues results in a loss of enzyme activity, leading to a proposed model for the enzyme's mechanism. These findings are expected to help guide efforts in protein and metabolic engineering to enhance PDC yields in biological routes to polymer feedstock synthesis.
In natural photosynthesis, the light-driven splitting of water into electrons, protons and molecular oxygen forms the first step of the solar-to-chemical energy conversion process. The reaction takes place in photosystem II, where the Mn 4 CaO 5 cluster first stores four oxidizing equivalents, the S 0 to S 4 intermediate states in the Kok cycle, sequentially generated by photochemical charge separations in the reaction center and then catalyzes the O–O bond formation chemistry 1 – 3 . Here, we report room temperature snapshots by serial femtosecond X-ray crystallography to provide structural insights into the final reaction step of Kok’s photosynthetic water oxidation cycle, the S 3 →[S 4 ]→S 0 transition where O 2 is formed and Kok’s water oxidation clock is reset. Our data reveal a complex sequence of events, which occur over micro- to milliseconds, comprising changes at the Mn 4 CaO 5 cluster, its ligands and water pathways as well as controlled proton release through the hydrogen-bonding network of the Cl1 channel. Importantly, the extra O atom O x , which was introduced as a bridging ligand between Ca and Mn1 during the S 2 →S 3 transition 4 – 6 , disappears or relocates in parallel with Y z reduction starting at approximately 700 μs after the third flash. The onset of O 2 evolution, as indicated by the shortening of the Mn1–Mn4 distance, occurs at around 1,200 μs, signifying the presence of a reduced intermediate, possibly a bound peroxide.
Owing to the limited quality of experimental data, such as resolution, atomic model refinement using cryo-EM or crystallographic experimental data is often a challenging task. To make refinement practical, a priori information about model geometry is always used as chemical restraints. These restraints originate from standard libraries that are used across most structure solution software in the field. Geometric restraints derived from these libraries suffer from at least three limitations. First, they lack information about novel molecules, such as ligands. Second, they are agnostic to non-covalent interactions such as hydrogen bonds, salt bridges, pi-interactions, and electrostatics in general. Third, the use of molecule-specific information as a source of geometric restraints, such as polypeptide chain conformations (Ramachandran plot), secondary structure, or rotameric states of amino acid side chains, requires a geometrically sound atomic model in the first place to be defined. Using quantum mechanical (QM) calculations as a source of restraints can eliminate the need for these libraries as well as the use of molecule-specific restraints altogether. However, these calculations are known to be computationally intractable for large molecules unless divide-and-conquer methods are used in conjunction with specialized QM software, some of which (the best one) is not free even for academic users. We present a novel deep learning implementation of a neural network potential that allows performing QM calculations for biomacromolecules in a timescale ranging from seconds to minutes. This, in turn, enables the use of QM-based geometry restraints in standard atomic model refinements. Preliminary results show that these methods produce refined models with much superior geometries compared to refinement using standard restraints while maintaining comparable fit to the experimental data.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
Histidine can be protonated on either or both of the two N atoms of the imidazole moiety. Each of the three possible forms occurs as a result of the stereochemical environment of the histidine side chain. In an atomic model, comparing the possible protonation states in situ, looking at possible hydrogen bonding and metal coordination, it is possible to predict which is most likely to be correct. A more direct method is described that uses quantum-mechanical methods to calculate, also in situ, the minimum geometry and energy for comparison, and therefore to more accurately identify the most likely protonation state.
Neutron diffraction is one of the three crystallographic techniques (X-ray, neutron and electron diffraction) used to determine the atomic structures of molecules. Its particular strengths derive from the fact that H (and D) atoms are strong neutron scatterers, meaning that their positions, and thus protonation states, can be derived from crystallographic maps. However, because of technical limitations and experimental obstacles, the quality of neutron diffraction data is typically much poorer (completeness, resolution and signal to noise) than that of X-ray diffraction data for the same sample. Further, refinement is more complex as it usually requires additional parameters to describe the H (and D) atoms. The increase in the number of parameters may be mitigated by using the `riding hydrogen' refinement strategy, in which the positions of H atoms without a rotational degree of freedom are inferred from their neighboring heavy atoms. However, this does not address the issues related to poor data quality. Therefore, neutron structure determination often relies on the presence of an X-ray data set for joint X-ray and neutron (XN) refinement. In this approach, the X-ray data serve to compensate for the deficiencies of the neutron diffraction data by refining one model simultaneously against the X-ray and neutron data sets. To be applicable, it is assumed that both data sets are highly isomorphous, and preferably collected from the same crystals and at the same temperature. However, the approach has a number of limitations that are discussed in this work by comparing four separately re-refined neutron models. To address the limitations, a new method for joint XN refinement is introduced that optimizes two different models against the different data sets. This approach is tested using neutron models and data deposited in the Protein Data Bank. The efficacy of refining models with H atoms as riding or as individual atoms is also investigated.
Quantum refinement (Q|R) of crystallographic or cryo-EM-derived structures of biomolecules within the Q|R project aims at using ab initio computations instead of library-based chemical restraints. An atomic model refinement requires the calculation of the gradient of the objective function. While it is not a computational bottleneck in classic refinement it is a roadblock if the objective function requires ab initio calculations. A solution to this problem adopted in Q|R is to divide the molecular system into manageable parts and do computations for these parts rather than using the whole macromolecule. This work focuses on the validation and optimization of the automatic divide-and-conquer procedure developed within the Q|R project. Also, we propose an atomic gradient error score that can be easily examined with common molecular visualization programs. While the tool is designed to work within the Q|R setting the error score can be adapted to similar fragmentation methods. The gradient testing tool presented here allows a priori determination of the computationally efficient strategy given available resources for the potentially time-expensive refinement process. The procedure is illustrated using a peptide and small protein models considering different quantum mechanical (QM) methodologies from Hartree-Fock, including basis set and dispersion corrections, to the modern semi-empirical method from the GFN-xTB family. The results obtained provide some general recommendations for the reliable and effective quantum refinement of larger peptides and proteins.
Direct-acting antivirals are needed to combat coronavirus disease 2019 (COVID-19), which is caused by severe acute respiratory syndrome-coronavirus-2 (SARS-CoV-2). The papain-like protease (PLpro) domain of Nsp3 from SARS-CoV-2 is essential for viral replication. In addition, PLpro dysregulates the host immune response by cleaving ubiquitin and interferon-stimulated gene 15 protein (ISG15) from host proteins. As a result, PLpro is a promising target for inhibition by small-molecule therapeutics. Here we have designed a series of covalent inhibitors by introducing a peptidomimetic linker and reactive electrophile onto analogs of the noncovalent PLpro inhibitor GRL0617. The most potent compound inhibited PLpro with k inact /K I = 10,000 M- 1 s- 1, achieved sub-μM EC50 values against three SARS-CoV-2 variants in mammalian cell lines, and did not inhibit a panel of human deubiquitinases at > 30 μM concentrations of inhibitor. An X-ray co-crystal structure of the compound bound to PLpro validated our design strategy and established the molecular basis for covalent inhibition and selectivity against structurally similar human DUBs. These findings present an opportunity for further development of covalent PLpro inhibitors.
Atomic model refinement at low resolution is often a challenging task. This is mostly because the experimental data are not sufficiently detailed to be described by atomic models. To make refinement practical and ensure that a refined atomic model is geometrically meaningful, additional information needs to be used such as restraints on Ramachandran plot distributions or residue side-chain rotameric states. However, using Ramachandran plots or rotameric states as refinement targets diminishes the validating power of these tools. Therefore, finding additional model-validation criteria that are not used or are difficult to use as refinement goals is desirable. Hydrogen bonds are one of the important noncovalent interactions that shape and maintain protein structure. These interactions can be characterized by a specific geometry of hydrogen donor and acceptor atoms. Systematic analysis of these geometries performed for quality-filtered high-resolution models of proteins from the Protein Data Bank shows that they have a distinct and a conserved distribution. Here, it is demonstrated how this information can be used for atomic model validation.
In macromolecular crystallographic structure refinement, ligands present challenges for the generation of geometric restraints due to their large chemical variability, their possible novel nature and their specific interaction with the binding pocket of the protein. Quantum-mechanical approaches are useful for providing accurate ligand geometries, but can be plagued by the number of minima in flexible molecules. In an effort to avoid these issues, the Quantum Mechanical Restraints (QMR) procedure optimizes the ligand geometry in situ, thus accounting for the influence of the macromolecule on the local energy minima of the ligand. The optimized ligand geometry is used to generate target values for geometric restraints during the crystallographic refinement. As demonstrated using a sample of >2330 ligand instances in >1700 protein–ligand models, QMR restraints generally result in lower deviations from the target stereochemistry compared with conventionally generated restraints. In particular, the QMR approach provides accurate torsion restraints for ligands and other entities.