A Working Group consisting of the co-authors of this paper was established in 2020 to re-evaluate the standard valence geometry used for the validation of nucleic acid structure models in the Protein Data Bank (PDB). This Working Group re-examined the dependence of Cambridge Structural Database (CSD) derived targets on base and sugar type, sugar pucker, and phosphate and glycosidic conformation, before comparing those targets with the geometry of a quality-filtered reference set of nucleic acid crystal structural models held in the PDB. This revealed that the valence bond and angle mean values are close to the CSD targets, but many parameters have highly non-Gaussian or even multimodal distributions. One explanation is the inconsistency of restraints used over time and by different refinement programs. The Working Group recommends a new validation scheme for use by the PDB. For this purpose, we have developed a new three-tier scale for outlier detection-graded as Preferred, Allowed, and Of Concern intervals-based on a combination of quality-curated reference data from the CSD and the PDB. The proposed approach to validation should lead to improved nucleic acid models in (future) PDB-deposited macromolecular structures.
L-Asparaginases hydrolyze L-asparagine to L-aspartate with the release of ammonia. Currently, three completely unrelated structural classes of L-asparaginases are known, further subdivided into five types. In each class, the hydrolysis reaction is thought to proceed via a nucleophilic attack of an activated Thr or Ser residue on the carbonyl Csp2 atom of the substrate amide group. With the possible exception of class 2 L-asparaginases, which function as N-terminal nucleophile (Ntn) hydrolases, the identity of the nucleophilic residue is, or at least has been historically, the subject of some controversy, and even in class 2 this issue may not be so entirely obvious. Structural chemistry has, however, excellent tools to figure out reaction mechanisms, based on the application of Bürgi's structure correlation method (SCM). Its principle allows one to predict the reaction trajectory if sufficient structural (crystallographic) examples of the reagents along the reaction path are known. With respect to the nucleophilic attack on a carbonyl group, the stereochemistry is governed by the nucleophile...electrophile distance and the Bürgi–Dunitz angle, later supplemented with the Flippin–Lodge angle. The latter angle is shown to be a poor parameter and is better replaced by the Herschlag dihedral between the planes of the attacking nucleophile and the attacked electrophile. In structural enzymology, applicability of the SCM principle requires the availability of structural examples of the enzyme in question in complex with the substrate or product of the catalytic reaction. In this work, we applied the SCM concept to the three classes of L-asparaginases, identifying in each case the most probable nucleophilic residue as Thr12 in EcAII (class 1), Thr179 in EcAIII (class 2) and Ser48 in ReAV (class 3). In addition, we applied the SCM analysis to the newly identified group of asparaginases without proper classification, called short-chain asparaginases, providing a basis for their proper affiliation in class 1. Finally, the SCM analysis shows that the chirality of the nucleophilic attack in class 2 asparaginases (pro-R) is opposite to that in all other asparaginases.
Isothermal titration calorimetry (ITC) studies of the enzyme kinetics and substrate specificity of Rhizobium etli Class 3 L-asparaginases, ReAIV (constitutive) and ReAV (inducible), showed that despite highly conserved catalytic site, the two isoforms differ significantly in thermostability, zinc affinity, and biochemical properties. As part of a wider investigation of potential non-natural substrates, acrylamide was tested, revealing a pronounced heat effect with ReAIV but none with ReAV. Crystallographic analysis showed the formation of a Michael adduct between acrylamide and a surface-exposed cysteine 183 in ReAIV, while the catalytic activity toward L-asparagine hydrolysis remained unaffected. These findings highlight the unique and multimodal reactivity of ReAIV, suggesting its potential dual application in the food industry: in selective removal of L-asparagine and in covalent sequestration of acrylamide under mild conditions. The acrylamide modification improved crystal morphology of ReAIV, offering practical advantages for structural studies. Additionally, a covalent modification of the catalytic Ser47 residue was observed in the presented crystal structure. Based on B-factor analysis, literature data, and detection of borate contamination in the laboratory water, this modification was interpreted as an orthoborate ester.
We have evaluated the quality of over 1200 crystal structures of hen egg-white lysozyme (HEWL) deposited in the Protein Data Bank (PDB). These structures, collected over nearly 50 years, vary in quality, despite all representing essentially the same small enzyme consisting of 129 amino acid residues. Some of the entries originated from studies of the binding of small-molecule ligands to HEWL, whereas the majority of deposits represent the outcomes of tests of new experimental approaches to crystallization and data collection and/or evaluations of new computational protocols. We found no correlation between Rfree, which is a measure of structure quality, and Rmerge, an indicator of raw data quality, for 136 near-atomic-resolution lysozyme structures. We found out that many of the lysozyme structures deposited as a result of methodology evaluation are not fully or correctly refined. We, therefore, propose that such structures be appropriately flagged in the PDB with a CAVEAT record to prevent their inadvertent inclusion in large-scale data mining analyses or training sets for artificial intelligence methods.
With cryogenic electron microscopy (cryo-EM) on track to surpass X-ray crystallography as the preferred method for determining macromolecular structures, it is important to evaluate and compare the quality of structure models obtained by these methods. This allows us to assess whether the rapidly growing numbers (quantity) correlate with quality and to identify areas in which each method excels or falls short. Selected quality-related parameters were compared for 97 200 crystal structures and 30 139 cryo-EM structures released by the Protein Data Bank (PDB) between 2015 and 2025. Comparison of geometric and stereochemical parameters indicated that, despite significant differences in the resolution of the experimental data, these values were, in the vast majority of cases, close to the expected targets. Nevertheless, we found that crystal structures tend to exhibit more Ramachandran and rotamer outliers than cryo-EM structures, although they unexpectedly have lower clashscore values. Separately, we compared the quality of 612 crystal and 1817 cryo-EM structures in the PDB representing complete ribosomes or their subunits. For this subset of very large, well-defined macromolecules, we found that the quality of many cryo-EM models is higher than that of their crystal counterparts, and that the best cryo-EM structures were also determined at higher resolution. Overall, we conclude that the availability of both techniques has clearly resulted in major advances during the last decade and bodes very well for the future.
It is postulated that the PDB should use the CAVEAT record more prominently to warn scientists using its archives of potential risks and errors.
We have evaluated the quality of all 325 deposits in the PDB (as of December 2024) that correspond to (or contain) the catalytic domain of cAMP-dependent protein kinases (PKA). Detailed analysis was possible for 289 deposits of crystal structures that included not only the atomic coordinates but also structure factors. These structures represent 35 years of studies, and it is not surprising that the more recent structures are generally of better quality than the older ones. We did not encounter deposits with very severe problems, although some minor problems were found. To assess whether a uniform method of structure re-refinement, as implemented in the pipeline and website PDB-REDO, leads to significant improvement of structural models, we compared structure quality indicators for the originally refined structures and their counterparts resulting from PDB-REDO refinement. The re-refinement procedure significantly improved only some older structures, while its success was generally limited. We paid particular attention to the quality of small-molecule ligands, finding that most of them fit the electron density very well. This type of analysis helps identify the highest quality structures among many deposits for certain protein families and, thus, could be extended to other groups of proteins as well.
Common bean ( Phaseolus vulgaris ) encodes three class 2 L-asparaginase enzymes: two potassium-dependent enzymes [PvAIII(K)-1 and PvAIII(K)-2] and a potassium-independent enzyme (PvAIII). Here, we present the crystal structure of PvAIII, which displays a rare P 2 space-group symmetry and a unique pseudosymmetric 4 1 -like double-helical packing. The asymmetric unit contains 32 protein chains (16 αβ units labeled A – P ) organized into two right-handed coiled arrangements, each consisting of four PvAIII (αβ) 2 dimers. Detailed analysis of the crystal structure revealed that this unusual packing originates from three factors: (i) the ability of the PvAIII molecules to form extended intermolecular β-sheets, a feature enabled by the PvAIII sequence and secondary structure, (ii) incomplete degradation of the flexible linker remaining at the C-terminus of α subunits of protein chain C after the autoproteolytic cleavage (maturation) of the PvAIII precursor and (iii) intermolecular entanglement between protein chains from the two helices to create `hydrogen-bond linchpins' that connect adjacent protein chains. The K m value of PvAIII for L-asparagine is approximately five times higher than for β-peptides, suggesting that the physiological role of PvAIII may be more related to the removal of toxic β-peptides than to basic L-asparagine metabolism. A comparison of the active sites of PvAIII and PvAIII(K)-1 shows that the proteins have nearly identical residues in the catalytic center, except for Thr219, which is unique to PvAIII. To test whether the residue type at position 219 affects the enzymatic activity of PvAIII, we designed and produced a T219S mutant. The kinetic parameters determined for L-asparagine hydrolysis indicate that the T/S residue type at position 219 does not affect the L-asparaginase activity of PvAIII.
This work analyzes the rules governing the growth of the numbers of vertices, edges and faces in all possible periodic tessellations of the 2D Euclidean space, and encodes those rules in several types of polynomial growth functions. These encodings map the geometric, combinatorial and topological properties of the tessellations into sets of integer coefficients. Several general statements about these encodings are given with rigorous mathematical proof. The variation of the growth functions is represented graphically and analyzed in orphic diagrams, so named because of their similarity to orphic art. Several examples of 3D space groups are included, to emphasize the complexity of the growth functions in higher dimensions. A freely available Python library is presented to facilitate the discovery of the growth functions and the generation of orphic diagrams.
The ReAV enzyme from Rhizobium etli, a representative of Class 3 L-asparaginases, is sequentially and structurally different from other known L-asparaginases. This distinctiveness makes ReAV a candidate for novel antileukemic therapies. ReAV is a homodimeric protein, with each subunit containing a highly specific zinc-binding site created by two cysteines, a lysine, and a water molecule. Two Ser-Lys tandems (Ser48-Lys51, Ser80-Lys263) are located in the close proximity of the metal binding site, with Ser48 hypothesized to be the catalytic nucleophile. To further investigate the catalytic process of ReAV, site-directed mutagenesis was employed to introduce alanine substitutions at residues from the Ser-Lys tandems and at Arg47, located near the Ser48-Lys51 tandem. These mutational studies, along with enzymatic assays and X-ray structure determinations, demonstrated that substitution of each of these highly conserved residues abolished the catalytic activity, confirming their essential role in enzyme mechanism.
The nutritionally essential sulfur amino acids, methionine and cysteine, are present at suboptimal levels in legumes, such as common bean (Phaseolus vulgaris L.). β-Substituted alanine synthase 4;1 (BSAS4;1) is the major isoform of cytosolic cysteine synthase present in the developing seeds of common bean. There is evidence that in addition to cysteine, this enzyme is also involved in the biosynthesis of the non-proteinogenic amino acid S-methylcysteine, which accumulates in the form of a γ-glutamyl dipeptide. Here, we report the high-resolution structure of recombinant BSAS4;1. Unexpectedly, the crystal structure showed the presence of a molecule of benzoic acid near the active site, which appeared to have been co-purified from Escherichia coli. Kinetic analysis indicated that benzoic acid acts as a competitive inhibitor of BSAS4;1 with respect to O-acetylserine. IC50 values for benzoic acid and the structurally related salicylic acid were both equal to 0.6 mm. Using developing cotyledons grown in vitro, quantification of the incorporation of 13C3- and 15N-labeled serine into cysteine and downstream metabolites indicated that benzoic acid effectively inhibited cysteine biosynthesis in vivo at a concentration of 1.2 mm. The results of experiments tracking the incorporation of 13C-labeled sodium thiomethoxide provided further evidence that BSAS4;1 may be involved in the formation of free S-methylcysteine, through the condensation of O-acetylserine with methanethiol.
The double-layer honeycomb with hexagonal cells, three rhombic faces between the two layers and p3m1 layer space-group symmetry, used universally by honeybees, is often considered to be the most efficient (from the point of view of wax economy) and the only honeycomb manufactured by bees. However, another variant of a symmetric and periodic double-layer hexagonal honeycomb with two hexagons and two rhombi between the two layers and slightly better wax economy was discovered theoretically in 1964 by Fejes Tóth and found in nature some years later. The present work shows that there is yet another possibility, with the interface formed by one hexagon and two quadrangles, in addition to the trivial case with flat hexagonal cell bottoms and very poor wax economy. Moreover, we demonstrate that the geometry of the Fejes Tóth honeycomb can be optimized for even better wax economy. All the theoretical honeycomb types are derived using the principle of Dirichlet-domain construction and shown to have more and less symmetric variants. Wax economy is calculated for each case, confirming that indeed the modified Fejes Tóth honeycomb is the most efficient, while the trivial flat-bottom case is the least.
Rhizobium etli is a nitrogen-fixing bacterium that encodes two l-asparaginases. The structure of the inducible R. etli asparaginase ReAV has been recently determined to reveal a protein with no similarity to known enzymes with l-asparaginase activity, but showing a curious resemblance to glutaminases and β-lactamases. The uniqueness of the ReAV sequence and 3D structure make the enzyme an interesting candidate as potential replacement for the immunogenic bacterial-type asparaginases that are currently in use for the treatment of acute lymphoblastic leukemia. The detailed catalytic mechanism of ReAV is still unknown; therefore, the enzyme was subjected to mutagenetic experiments to investigate its catalytic apparatus. In this work, we generated two ReAV variants of the conserved Lys138 residue (K138A and K138H) that is involved in zinc coordination in the wild-type protein and studied them kinetically and structurally. We established that the activity of wild-type ReAV and the generated variants is significantly reduced in the presence of Cd2+ cations, which slow down the proteins while improving their apparent substrate affinity. Moreover, the inhibitory effect of Cd2+ is enhanced by the substitutions of Lys138, which disrupt the metal coordination sphere. The proteins with impaired activity but increased affinity were cocrystallized with the L-Asn substrate. Here, we present the crystal structures of wild-type ReAV and its K138A and K138H variants, unambiguously revealing bound l-asparagine in the active site. After careful analysis of the stereochemistry of the nucleophilic attack, we assign the role of the primary nucleophile of ReAV to Ser48. Furthermore, we propose that the reaction catalyzed by ReAV proceeds according to a double-displacement mechanism.
A global analysis of protein crystal structures in the Protein Data Bank (PDB) using a newly developed computational approach reveals many pairs with (nearly) identical main-chain coordinates. Such cases are identified and analyzed, showing that duplication is possible since the PDB does not currently have tools or mechanisms that would detect potentially duplicate submissions. Some duplicated entries represent modeling efforts of ligand binding that masquerade as experimentally determined structures. We propose that duplicate entries should either be obsoleted by the PDB or, as a minimum, marked with a clear `CAVEAT' record that would alert potential users to the presence of such problems. We also suggest that using a tool for verifying the uniqueness of the deposited structure, such as that presented in this work, should become part of the routine validation procedure for new depositions.
Ultrahigh-resolution structures provide unprecedented details about protein dynamics, hydrogen bonding and solvent networks. The reported 0.70 Å, room-temperature crystal structure of crambin is the highest-resolution ambient-temperature structure of a protein achieved to date. Sufficient data were collected to enable unrestrained refinement of the protein and associated solvent networks using SHELXL. Dynamic solvent networks resulting from alternative side-chain conformations and shifts in water positions are revealed, demonstrating that polypeptide flexibility and formation of clathrate-type structures at hydrophobic surfaces are the key features endowing crambin crystals with extraordinary diffraction power.