We have evaluated the quality of over 1200 crystal structures of hen egg-white lysozyme (HEWL) deposited in the Protein Data Bank (PDB). These structures, collected over nearly 50 years, vary in quality, despite all representing essentially the same small enzyme consisting of 129 amino acid residues. Some of the entries originated from studies of the binding of small-molecule ligands to HEWL, whereas the majority of deposits represent the outcomes of tests of new experimental approaches to crystallization and data collection and/or evaluations of new computational protocols. We found no correlation between Rfree, which is a measure of structure quality, and Rmerge, an indicator of raw data quality, for 136 near-atomic-resolution lysozyme structures. We found out that many of the lysozyme structures deposited as a result of methodology evaluation are not fully or correctly refined. We, therefore, propose that such structures be appropriately flagged in the PDB with a CAVEAT record to prevent their inadvertent inclusion in large-scale data mining analyses or training sets for artificial intelligence methods.
With cryogenic electron microscopy (cryo-EM) on track to surpass X-ray crystallography as the preferred method for determining macromolecular structures, it is important to evaluate and compare the quality of structure models obtained by these methods. This allows us to assess whether the rapidly growing numbers (quantity) correlate with quality and to identify areas in which each method excels or falls short. Selected quality-related parameters were compared for 97 200 crystal structures and 30 139 cryo-EM structures released by the Protein Data Bank (PDB) between 2015 and 2025. Comparison of geometric and stereochemical parameters indicated that, despite significant differences in the resolution of the experimental data, these values were, in the vast majority of cases, close to the expected targets. Nevertheless, we found that crystal structures tend to exhibit more Ramachandran and rotamer outliers than cryo-EM structures, although they unexpectedly have lower clashscore values. Separately, we compared the quality of 612 crystal and 1817 cryo-EM structures in the PDB representing complete ribosomes or their subunits. For this subset of very large, well-defined macromolecules, we found that the quality of many cryo-EM models is higher than that of their crystal counterparts, and that the best cryo-EM structures were also determined at higher resolution. Overall, we conclude that the availability of both techniques has clearly resulted in major advances during the last decade and bodes very well for the future.
It is postulated that the PDB should use the CAVEAT record more prominently to warn scientists using its archives of potential risks and errors.
We have evaluated the quality of all 325 deposits in the PDB (as of December 2024) that correspond to (or contain) the catalytic domain of cAMP-dependent protein kinases (PKA). Detailed analysis was possible for 289 deposits of crystal structures that included not only the atomic coordinates but also structure factors. These structures represent 35 years of studies, and it is not surprising that the more recent structures are generally of better quality than the older ones. We did not encounter deposits with very severe problems, although some minor problems were found. To assess whether a uniform method of structure re-refinement, as implemented in the pipeline and website PDB-REDO, leads to significant improvement of structural models, we compared structure quality indicators for the originally refined structures and their counterparts resulting from PDB-REDO refinement. The re-refinement procedure significantly improved only some older structures, while its success was generally limited. We paid particular attention to the quality of small-molecule ligands, finding that most of them fit the electron density very well. This type of analysis helps identify the highest quality structures among many deposits for certain protein families and, thus, could be extended to other groups of proteins as well.
This work analyzes the rules governing the growth of the numbers of vertices, edges and faces in all possible periodic tessellations of the 2D Euclidean space, and encodes those rules in several types of polynomial growth functions. These encodings map the geometric, combinatorial and topological properties of the tessellations into sets of integer coefficients. Several general statements about these encodings are given with rigorous mathematical proof. The variation of the growth functions is represented graphically and analyzed in orphic diagrams, so named because of their similarity to orphic art. Several examples of 3D space groups are included, to emphasize the complexity of the growth functions in higher dimensions. A freely available Python library is presented to facilitate the discovery of the growth functions and the generation of orphic diagrams.
The double-layer honeycomb with hexagonal cells, three rhombic faces between the two layers and p3m1 layer space-group symmetry, used universally by honeybees, is often considered to be the most efficient (from the point of view of wax economy) and the only honeycomb manufactured by bees. However, another variant of a symmetric and periodic double-layer hexagonal honeycomb with two hexagons and two rhombi between the two layers and slightly better wax economy was discovered theoretically in 1964 by Fejes Tóth and found in nature some years later. The present work shows that there is yet another possibility, with the interface formed by one hexagon and two quadrangles, in addition to the trivial case with flat hexagonal cell bottoms and very poor wax economy. Moreover, we demonstrate that the geometry of the Fejes Tóth honeycomb can be optimized for even better wax economy. All the theoretical honeycomb types are derived using the principle of Dirichlet-domain construction and shown to have more and less symmetric variants. Wax economy is calculated for each case, confirming that indeed the modified Fejes Tóth honeycomb is the most efficient, while the trivial flat-bottom case is the least.
A global analysis of protein crystal structures in the Protein Data Bank (PDB) using a newly developed computational approach reveals many pairs with (nearly) identical main-chain coordinates. Such cases are identified and analyzed, showing that duplication is possible since the PDB does not currently have tools or mechanisms that would detect potentially duplicate submissions. Some duplicated entries represent modeling efforts of ligand binding that masquerade as experimentally determined structures. We propose that duplicate entries should either be obsoleted by the PDB or, as a minimum, marked with a clear `CAVEAT' record that would alert potential users to the presence of such problems. We also suggest that using a tool for verifying the uniqueness of the deposited structure, such as that presented in this work, should become part of the routine validation procedure for new depositions.
Ultrahigh-resolution structures provide unprecedented details about protein dynamics, hydrogen bonding and solvent networks. The reported 0.70 Å, room-temperature crystal structure of crambin is the highest-resolution ambient-temperature structure of a protein achieved to date. Sufficient data were collected to enable unrestrained refinement of the protein and associated solvent networks using SHELXL. Dynamic solvent networks resulting from alternative side-chain conformations and shifts in water positions are revealed, demonstrating that polypeptide flexibility and formation of clathrate-type structures at hydrophobic surfaces are the key features endowing crambin crystals with extraordinary diffraction power.
The symmetry aspects of close packing of hexagonal layers of equal spheres are usually presented with reference to the two basic patterns: AB… (P63/mmc) and ABC… (Fm3m), of which the latter is the fundament of the Kepler conjecture, published in 1611, stating that no arrangement of equal spheres can achieve higher density. In fact, however, there are infinitely many arrangements of N > 1 layers that, when repeated periodically, will give the same closest packing density of spheres in 3D. One such possibility for N = 4 was discovered by Barlow in 1883, and other authors gave this issue particular attention. The present paper systematically analyzes the symmetry of periodic patterns of hexagonal layers with growing N. It also presents the colorful history of this subject and introduces an ingenious system to describe the symmetry of close stacking of hexagonal layers of spheres.
The absence of solvent molecules in high-resolution protein crystal structure models deposited in the Protein Data Bank (PDB) contradicts the fact that, for proteins crystallized from aqueous media, water molecules are always expected to bind to the protein surface, as well as to some sites in the protein interior. An analysis of the contents of the PDB indicated that the expected ratio of the number of water molecules to the number of amino-acid residues exceeds 1.5 in atomic resolution structures, decreasing to 0.25 at around 2.5 Å resolution. Nevertheless, almost 800 protein crystal structures determined at a resolution of 2.5 Å or higher are found in the current release of the PDB without any water molecules, whereas some other depositions have unusually low or high occupancies of modeled solvent. Detailed analysis of these depositions revealed that the lack of solvent molecules might be an indication of problems with either the diffraction data, the refinement protocol, the deposition process or a combination of these factors. It is postulated that problems with solvent structure should be flagged by the PDB and addressed by the depositors.
The manuscript `Modeling a unit cell: crystallographic refinement procedure using the biomolecular MD simulation platform Amber' presents a novel protein structure refinement method claimed to offer improvements over traditional techniques like Refmac5 and Phenix. Our re-evaluation suggests that while the new method provides improvements, traditional methods achieve comparable results with less computational effort.
The Protein Data Bank (PDB) includes a carefully curated treasury of experimentally derived structural data on biological macromolecules and their various complexes. Such information is fundamental for a multitude of projects that involve large-scale data mining and/or detailed evaluation of individual structures of importance to chemistry, biology and, most of all, to medicine, where it provides the foundation for structure-based drug discovery. However, despite extensive validation mechanisms, it is almost inevitable that among the ∼215 000 entries there will occasionally be suboptimal or incorrect structure models. It is thus vital to apply careful verification procedures to those segments of the PDB that are of direct medicinal interest. Here, such an analysis was carried out for crystallographic models of L-asparaginases, enzymes that include approved drugs for the treatment of certain types of leukemia. The focus was on the adherence of the atomic coordinates to the rules of stereochemistry and their agreement with the experimental electron-density maps. Whereas the current clinical application of L-asparaginases is limited to two bacterial proteins and their chemical modifications, the field of investigations of such enzymes has expanded tremendously in recent years with the discovery of three entirely different structural classes and with numerous reports, not always quite reliable, of the anticancer properties of L-asparaginases of different origins.
The simple Euler polyhedral formula, expressed as an alternating count of the bounding faces, edges and vertices of any polyhedron, V − E + F = 2, is a fundamental concept in several branches of mathematics. Obviously, it is important in geometry, but it is also well known in topology, where a similar telescoping sum is known as the Euler characteristic χ of any finite space. The value of χ can also be computed for the unit polyhedra (such as the unit cell, the asymmetric unit or Dirichlet domain) which build, in a symmetric fashion, the infinite crystal lattices in all space groups. In this application χ has a modified form (χm) and value because the addends have to be weighted according to their symmetry. Although derived in geometry (in fact in crystallography), χm has an elegant topological interpretation through the concept of orbifolds. Alternatively, χm can be illustrated using the theorems of Harriot and Descartes, which predate the discovery made by Euler. Those historical theorems, which focus on angular defects of polyhedra, are beautifully expressed in the formula of de Gua de Malves. In a still more general interpretation, the theorem of Gauss–Bonnet links the Euler characteristic with the general curvature of any closed space. This article presents an overview of these interesting aspects of mathematics with Euler's formula as the leitmotif. Finally, a game is designed, allowing readers to absorb the concept of the Euler characteristic in an entertaining way.
Accurate experimentally determined structure models of biological macromolecules are used by a large and diverse community of researchers. It is an established practice to base the assessment of the structure model quality on both, expectations of correct stereochemistry and, most importantly, on examination of the model's fit to the primary experimental evidence. In the case of X-ray crystallography, the primary evidence is provided by the electron density map. The worldwide Protein Data Bank (wwPDB1) is a global repository of macromolecular models and the accompanying experimental data that allow to examine agreement between the electron density and structural model using programs such as Coot,2 Chimera,3 Pymol,4 or Molstack.5 Throughout its 50-year history, the PDB has accumulated over 180,000 macromolecular structures, and gained the reputation of the gold standard in structural biology and of the most reliable data resource in biomedical research in general.6 Recently, the PDB has seen an influx of many depositions from large-scale crystallographic fragment screening projects using a complex computational procedure called Pan-Dataset Density Analysis (PanDDA7). Based on a sophisticated multi-data-set analysis of reference models and potential ligand-complex crystals,7 such group depositions serve the purpose of identifying very low-occupancy small molecule ligands in macromolecular complexes. In brief, the idea of PanDDA consists of partial background solvent inclusion and subtraction of a virtual “multi-crystal ground state,” to produce so-called “event map,” revealing the supposed ligands in each of a multitude of different ligand data sets. As a result, the PDB accumulates large numbers of “group depositions” of many putative ligand complexes of the same protein target that presently do not conform to the primary objective of the PDB as a repository of high-quality structure models and data. In particular, the maps that can be retrieved for such deposits have dubious agreement with the models (Figure 1). While suited for fragment screening and lead discovery,8 the group deposition models dumped en masse into the PDB do not conform to the quality standards9, 10 expected of PDB entries. In particular, PanDDA deposits confuse most biomedical researchers as their data structure is different than that of other PDB deposits and the quality of the structure models, despite occasional high nominal resolution, is often questionable. The presence of group deposits that do not conform to PDB standards of data retrieval and model quality, but nevertheless are presented on a par with conventional entries, degrades the PDB integrity. For standard entries, the PDB-provided map coefficients or density maps are calculated from the supplied experimental and model data, allowing validation of the model against experiment. This procedure is not possible for PanDDA entries, as the “event maps”7 have a completely different nature and purpose. Consequently, group deposition entries are difficult or even impossible to individually validate and assess. They can mislead PDB users when selected as models to underpin further studies and may also mislead systems like AlphaFold211 during selection of optimal templates for structure prediction. Presence of nonconforming entries is particularly problematic for automated data-mining projects, including applications of Artificial Intelligence, as the presence of such data adds unexpected levels of noise during the learning and testing stages. The gold standard of structural biology, that is, the agreement of a structure model with underlying electron density, fails for group deposition models that severely disagree with the user-accessible electron density maps (Figure 1). If group deposition entries violating accepted quality metrics9, 10 become primary references for their protein families due to their reported very high resolution, the effect could be disastrous (Figure 1b). The presence of such deposits will affect the reliability of the PDB and its reputation as the most reliable repository in life sciences. While even the community of structural biologists may already have problems with understanding the limitations of group deposits, biomedical researchers, and bioinformaticians generally assume that the same quality standards and data structure are consistently applied to all models in the PDB. Nonstructural research communities rarely verify the model by comparison with the electron density, and thus rely entirely on structural biology standard metrics, such as reported resolution, and expect that PDB structure deposits represent uniformly valid and verified models. We postulate that PDB group depositions from large-scale ligand screening projects that do not present fully refined macromolecular models should be, as a minimum, very clearly marked as members of this special category. Ideally, however, such models should be relocated from the PDB into a separate database dedicated to group depositions. The inventors of the PanDDA7 ligand-screening methodology have already developed database capabilities ideally suited to storing and handling fragment screening data.12 They could, therefore, collaborate with the PDB to establish such a database, equipped with tools needed for proper examination and evaluation of fragment-screening entries. Clear annotation and separation of nonstandard entries will minimize contamination of the PDB by suboptimal structures; and importantly, structure-informed biomedical research will remain based on validated and verified experimental evidence. Mariusz Jaskolski, Alexander Wlodawer, Zbigniew Dauter, Wladek Minor, Bernhard Rupp: Conceptualization (equal); formal analysis (equal); writing – original draft (equal); writing – review and editing (equal).
Although the dynamical refinement improves the fit significantly compared to the kinematical refinement, its current implementation still does not describe the electron diffraction data fully.To further improve the fit, it is necessary to properly account for further effects like crystal imperfections, effects of inelastic scattering and also the bonding effects in the electrostatic potential.
Newly discovered, and still uncommon, modulated crystal structure in organic systems require a deeper investigation.No exact and detailed solution of such systems has not been done up-to-date.One possibility is to use an approximation of commensurate modulation which enables constructing a supercell, extending to the case, where translational symmetry (periodicity) is recovered, and simplify the analysis [1].An assumption of commensurateness of the modulation is, however, questionable and rather unverifiable.The goal of our studies was to use a novel, original statistical method of structural modeling which enables a refinement based on the average unit cell with (commensurate or incommensurate) modulation without unclear assumption of commensurateness and supercell approach.The main concept of the statistical method is to express structure in terms of the statistical distribution of atomic positions concerning the periodic reference lattice with lattice constant related to characteristic length-scale present in the structure.The average unit cell, defined as a probability distribution, constructed for periodic crystal is the same as the unit cell.The statistical approach was successfully used for the description of not only periodic crystals or quasicrystals, as well as it can be expanded on modulated structures as well as aperiodic structures with singular continuous components in the diffraction pattern [2].Our model system is a pathogenesis-related protein (Hyp-1) complex with fluorescent probe 8-anilino-1-naphthalene sulfonate (ANS), which is a unique example of a macromolecular system with the modulated crystal structure.Previous studies have shown that Hyp-1/ANS complexes are tetartohedral twinned and crystallized in an asymmetric unit cell containing a repetitive motif of four protein molecules arranged with 7-fold noncrystallographic repetition along the c axis of the C2 space group.Assumption of commensurate structure modulation demanded description of structure in the highly expanded unit cell with 28 unique protein molecules inside [3].The Hyp-1/ANS structure was solved by molecular replacement and refined using maximum-likelihood targets with reliability factors Rwork/Rfree of 22.3/27.8%,respectively.Our approach involved re-integration of raw data, development of the original software in Matlab environment and multidimensional analysis used to build the structure model and perform the refinement for significant improvement of results.The problem of incorporating disorder in the form of phonons into structural analysis was also carried out traditionally by Debye-Waller factor.
The appearance at the end of 2019 of the new SARS-CoV-2 coronavirus led to an unprecedented response by the structural biology community, resulting in the rapid determination of many hundreds of structures of proteins encoded by the virus. As part of an effort to analyze and, if necessary, remediate these structures as deposited in the Protein Data Bank (PDB), this work presents a detailed analysis of 81 crystal structures of the main protease 3CLpro, an important target for the design of drugs against COVID-19. The structures of the unliganded enzyme and its complexes with a number of inhibitors were determined by multiple research groups using different experimental approaches and conditions; the resulting structures span 13 different polymorphs representing seven space groups. The structures of the enzyme itself, all determined by molecular replacement, are highly similar, with the exception of one polymorph with a different inter-domain orientation. However, a number of complexes with bound inhibitors were found to pose significant problems. Some of these could be traced to faulty definitions of geometrical restraints for ligands and to the general problem of a lack of such information in the PDB depositions. Several problems with ligand definition in the PDB itself were also noted. In several cases extensive corrections to the models were necessary to adhere to the evidence of the electron-density maps. Taken together, this analysis of a large number of structures of a single, medically important protein, all determined within less than a year using modern experimental tools, should be useful in future studies of other systems of high interest to the biomedical community.
The fine details of the electron density distribution, typically far beyond the possibilities of standard X-ray diffraction analysis, can be approached when (ultra)high-resolution diffraction data are available, to allow abandoning the standard model of independent, spherically symmetrical atoms. This more complicated, and much more demanding (both experimentally and computationally) method allows, for instance, to analyze the redistribution of electron density into bonds, intermolecular interactions, etc. Moreover, the Atoms-In-Molecules approach, which is based on the analysis of topological features of high-quality electron density distribution, may offer an insight into the hierarchy of interactions, energetic features, etc. Even though such an approach is well-developed for small molecules, its application in macromolecular crystallography is still under development. Ultrahigh resolution in this case means resolution of at least 0.7 – 0.65 Å. Such data are extremely rare for macromolecular crystals. In addition, modelling problems, disorder, high solvent content, etc., severely limit the number of successful studies of experimental electron density distribution in macromolecules, which so far have been reported for proteins only (e.g. crambin [1], aldose reductase [2] and the high-potential iron – sulfur protein [3]). For the present study, ultrahigh-resolution diffraction data (0.55 A) were collected for a Z-DNA hexamer with
This 75th birthday tribute to our Editorial Board member Alexander Wlodawer recounts his decades‐long service to the community of structural biology researchers. His former and current colleagues tell the story of his upbringing and education, followed by an account of his dedication to quality and rigor in crystallography and structural science. The FEBS Journal Editor‐in‐Chief Seamus Martin further highlights Alex's outstanding contributions to the journal's success over many years.