Protein Data Bank in Europe (PDBe) is a founding member of the worldwide Protein Data Bank (wwPDB), delivering open access to experimentally determined macromolecular structures. PDBe also delivers enriched annotations contributed by the PDBe-Knowledge Base (PDBe-KB) consortium. The macromolecular entry pages are the primary interface for millions of users who explore experimental structure data. Here, we describe the redesign of the PDBe entry pages that organize content into logical views, thereby improving usability and facilitating the use of structure models, driving fundamental and applied research, and supporting education. The new design introduces several enhancements, including integrated and central 3D visualization, sequence feature exploration, standardized molecular scenes, AI-driven residue-level annotations from text mining of the literature, and streamlined annotation access via refactored APIs. Importantly, researchers can now upload their own residue-level features and visualize them directly in structural context, displayed alongside PDBe-KB annotations. This functionality enables scientists to interpret unpublished or private data in relation to high-quality structural information, lowering barriers for those without prior expertise in structural biology. Together, these updates create a more accessible, flexible, and scalable framework for interacting with structural data, expanding the resource's value to both domain specialists and the wider life sciences community.
The PDBx/macromolecular crystallographic information file (PDBx/mmCIF) framework is the standard for representing macromolecular structure data in the Protein Data Bank (PDB) and has been mandatory for crystallographic depositions since 2019. The PDBx/mmCIF dictionary, maintained by the PDBx/mmCIF working group established by the Worldwide PDB (wwPDB) consortium, defines the schema, data types, enumerations and relationships that govern the structure and content of mmCIFs. Pre-deposition validation of mmCIFs enables researchers to identify and correct errors, ensuring both dictionary compliance and that data values are meaningful and appropriate for their dataset. This validation step can significantly streamline the deposition process in the wwPDB OneDep system, when depositing to the PDB and Electron Microscopy Data Bank archives or using the PDB-IHM deposition system for integrative/hybrid structures, reducing delays and improving data quality. We present the mmCIF Validator , a comprehensive validation tool available in two complementary implementations: a Visual Studio Code extension for real-time interactive validation during file editing; and a standalone Python script for command-line use, batch processing, and integration into automated workflows and CI/CD (continuous integration and continuous delivery/deployment) pipelines. The validator performs comprehensive checks including item-definition validation, mandatory-item presence (with category-aware checking), enumeration-value validation, data-type validation (including automatic regex-based validation extracted from OneDep deposition-specific categories in the dictionary for types like email, phone, ORCID ID and PDB ID), range constraints with distinction between strictly allowed and advisory boundary conditions, parent/child category relationships, foreign-key integrity, composite-key validation for relationships defined by multiple items, duplicate category and item detection, and complex operation-expression parsing. The tool requires Python 3.7+ and uses only the Python standard library (no pip packages required). It works out of the box with automatic dictionary downloading from the official wwPDB repository, or users can employ a custom distant or local dictionary file. The two implementations share the same validation engine, ensuring consistent results across different usage scenarios. The validator is particularly valuable for structural biologists preparing structures for deposition, biocurators ensuring data quality and software developers building automated quality-control workflows.
We present a novel system that leverages curators in the loop to develop a dataset and model for detecting structure features and functional annotations at residue-level from standard publication text. Our approach involves the integration of data from multiple resources, including PDBe, EuropePMC, PubMedCentral, and PubMed, combined with annotation guidelines from UniProt, and LitSuggest and HuggingFace models as tools in the annotation process. A team of seven annotators manually curated ten articles for named entities, which we utilized to train a starting PubmedBert model from HuggingFace. Using a human-in-the-loop annotation system, we iteratively developed the best model with commendable performance metrics of 0.90 for precision, 0.92 for recall, and 0.91 for F1-measure. Our proposed system showcases a successful synergy of machine learning techniques and human expertise in curating a dataset for residue-level functional annotations and protein structure features. The results demonstrate the potential for broader applications in protein research, bridging the gap between advanced machine learning models and the indispensable insights of domain experts.
Mitochondria are membrane-bound organelles of endosymbiotic origin with limited protein-coding capacity. The import of nuclear-encoded proteins and nucleic acids is required and essential for maintaining organelle mass, number, and activity. As plant mitochondria do not encode all the necessary tRNA types required, the import of cytosolic tRNA is vital for organelle maintenance. Recently, two mitochondrial outer membrane proteins, named Tric1 and Tric2, for tRNA import component, were shown to be involved in the import of cytosolic tRNA. Tric1/2 binds tRNAala via conserved residues in the C-terminal Sterile Alpha Motif (SAM) domain. Here we report the X-ray crystal structure of the Tric1 SAM domain. We identified the ability of the SAM domain to form a helical superstructure with six monomers per helical turn and key amino acid residues responsible for its formation. We determined that the oligomerization of the Tric1 SAM domain may play a role in protein function whereby mutation of Gly241 introducing a larger side chain at this position disrupted the oligomer and resulted in the loss of RNA binding capability. Furthermore, complementation of Arabidopsis thaliana Tric1/ 2 knockout lines with a mutated Tric1 failed to restore the defective plant phenotype. AlphaFold2 structure prediction of both the SAM domain and Tric1 support a cyclic pentameric or hexameric structure. In the case of a hexameric structure, a pore of sufficient dimensions to transfer tRNA across the mitochondrial membrane is observed. Our results highlight the importance of oligomerization of Tric1 for protein function.
Mitochondria are membrane bound organelles of endosymbiotic origin with limited protein coding capacity. As a consequence, the continual import of nuclear-encoded protein and nucleic acids such as DNA and small non-coding RNA is required and essential for maintaining organelle mass, number and activity. As plant mitochondria do not encode all the necessary tRNA types required, the import of cytosolic tRNA is vital for organelle maintenance. Recently, two mitochondrial outer membrane proteins, named Tric1 and Tric2, for tRNA import component, were shown to be involved in the import of cytosolic tRNA. Tric1/2 binds tRNAala via conserved residues in the C-terminal Sterile Alpha Motif (SAM) domain. Here we report the X-ray crystal structure of the Tric1 SAM domain. We identified the ability of the SAM domain to form a helical superstructure with 6 SAM domains per helical turn and key amino acid residues responsible for its formation. We determined that the oligomerization of Tric1 SAM domain was essential for protein function whereby mutation of Gly241 resulted in the disruption of the oligomer and the loss of RNA binding capability in Tric1. Furthermore, complementation of Arabidopsis thaliana Tric1/2 knockout lines with a mutated Tric1 failed to restore the defective plant phenotype suggesting the oligomerization is essential for function in planta. AlphaFold2 structure prediction of the SAM domain and Tric1 support a cyclic hexamer generating a pore of sufficient dimensions to transfer tRNA across the mitochondrial membrane. Our results highlight the importance of oligomerization of Tric1 for protein function.
The archiving and dissemination of protein and nucleic acid structures as well as their structural, functional and biophysical annotations is an essential task that enables the broader scientific community to conduct impactful research in multiple fields of the life sciences. The Protein Data Bank in Europe (PDBe; pdbe.org) team develops and maintains several databases and web services to address this fundamental need. From data archiving as a member of the Worldwide PDB consortium (wwPDB; wwpdb.org), to the PDBe Knowledge Base (PDBe-KB; pdbekb.org), we provide data, data-access mechanisms, and visualizations that facilitate basic and applied research and education across the life sciences. Here, we provide an overview of the structural data and annotations that we integrate and make freely available. We describe the web services and data visualization tools we offer, and provide information on how to effectively use or even further develop them. Finally, we discuss the direction of our data services, and how we aim to tackle new challenges that arise from the recent, unprecedented advances in the field of structure determination and protein structure modeling.
Many pathogenic gram-negative bacteria have developed mechanisms to increase resistance to cationic antimicrobial peptides by modifying the lipid A moiety. One modification is the addition of phospho-ethano-lamine to lipid A by the enzyme phospho-ethano-lamine transferase (EptA). Previously we reported the structure of EptA from Neisseria, revealing a two-domain architecture consisting of a periplasmic facing soluble domain and a transmembrane domain, linked together by a bridging helix. Here, the conformational flexibility of EptA in different detergent environments is probed by solution scattering and intrinsic fluorescence-quenching studies. The solution scattering studies reveal the enzyme in a more compact state with the two domains positioned close together in an n-do-decyl-β-d-maltoside micelle environment and an open extended structure in an n-do-decyl-phospho-choline micelle environment. Intrinsic fluorescence quenching studies localize the domain movements to the bridging helix. These results provide important insights into substrate binding and the molecular mechanism of endotoxin modification by EptA.
Phospholipids are key components of cellular membranes and are emerging as important functional regulators of different membrane proteins, including pentameric ligand-gated ion channels (pLGICs). Here, we take advantage of the prokaryote channel ELIC as a model to understand the determinants of phospholipid interactions in this family of receptors. A high-resolution crystal structure of ELIC in a lipid-bound state reveals a phospholipid site at the lower half of pore-forming transmembrane helices M1 and M4 and at a nearby site for neurosteroids, cholesterol or general anesthetics. This site is shaped by an M4 helix kink and a Trp-Arg-Pro triad that is highly conserved in eukaryote GABAA/C and glycine receptors. A combined approach reveals that M4 is intrinsically flexible and that M4 deletions or disruptions of the lipid binding site accelerate desensitization in ELIC, suggesting that lipid interactions shape the agonist response. Our data offer a structural context for understanding lipid modulation in the family of pLGICs.
Neonicotinoids, sometimes called ‘neonics’, are a class of insecticides targeting neuronal nicotinic acetylcholine receptors, which belong to the family of pentameric ligand-gated ion channels (pLGICs) or Cys-loop receptors. The widespread application of these neurotoxic insecticides in agriculture is tied to the worldwide decline of bee populations, due to their sublethal effects on bee memory, behavior and reproduction. In 2018, the member states of the European Union, agreed on a ban of most neonicotinoids, with exemption for greenhouses. However, a new generation of sulfoximine-based insecticides, such as sulfoxaflor, has recently been introduced to the agricultural market claiming they are chemically distinct from neonicotinoids. Using the acetylcholine binding protein (AChBP) as a well-established tool for structural studies of nicotinic receptors, we engineered AChBP variants that mimic the neurotransmitter binding site of honeybee nicotinic receptors. In parallel, we pursued functional characterization of these receptors expressed in Xenopus oocytes and using electrophysiological techniques. Three-dimensional structures of sulfoxaflor-bound AChBPs reveal that the molecule is recognized through receptor interactions that are virtually identical to the neonicotinoid thiacloprid. Surprisingly, binding assays reveal a ∼100-fold difference in affinity between sulfoxaflor and thiacloprid, leaving open important questions with respect to the precise molecular mechanism of sulfoxaflor recognition. Our study will attempt to begin addressing these central questions.
Pentameric ligand-gated ion channels (pLGICs) or Cys-loop receptors are involved in fast synaptic signaling in the nervous system. Allosteric modulators bind to sites that are remote from the neurotransmitter binding site, but modify coupling of ligand binding to channel opening. In this study, we developed nanobodies (single domain antibodies), which are functionally active as allosteric modulators, and solved co-crystal structures of the prokaryote (Erwinia) channel ELIC bound either to a positive or a negative allosteric modulator. The allosteric nanobody binding sites partially overlap with those of small molecule modulators, including a vestibule binding site that is not accessible in some pLGICs. Using mutagenesis, we extrapolate the functional importance of the vestibule binding site to the human 5-HT3 receptor, suggesting a common mechanism of modulation in this protein and ELIC. Thus we identify key elements of allosteric binding sites, and extend drug design possibilities in pLGICs with an accessible vestibule site.
Pentameric ligand-gated ion channels (pLGICs) belong to a class of ion channels involved in fast synaptic signaling in the central and peripheral nervous systems. Molecules acting as allosteric modulators target binding sites that are remote from the neurotransmitter binding site, but functionally affect coupling of ligand binding to channel opening. Here, we investigated an allosteric binding site in the ion channel vestibule, which has converged from a series of studies on prokaryote and eukaryote channel homologs. We discovered single domain antibodies, called nanobodies, which are functionally active as allosteric modulators, and solved co-crystal structures of the prokaryote channel ELIC bound either to a positive (PAM) or a negative (NAM) allosteric modulator. We extrapolate the functional importance of the vestibule binding site to eukaryote ion channels, suggesting a conserved mechanism of allosteric modulation. This work identifies key elements of allosteric binding sites and extends drug design possibilities in pLGICs using nanobodies.
Phospholipids are key components of cellular membranes and are emerging as important functional regulators of different membrane proteins, including pentameric ligand-gated ion channels (pLGICs). Here, we take advantage of the prokaryote channel ELIC (Erwinia ligand-gated ion channel) as a model to understand the determinants of phospholipid interactions in this family of receptors. A high-resolution structure of ELIC in a lipid-bound state reveals a phospholipid site at the lower half of pore-forming transmembrane helices M1 and M4 and at a nearby site for neurosteroids, cholesterol or general anesthetics. This site is shaped by an M4-helix kink and a Trp-Arg-Pro triad that is highly conserved in eukaryote GABA(A/C) and glycine receptors. A combined approach reveals that M4 is intrinsically flexible and that M4 deletions or disruptions of the lipid-binding site accelerate desensitization in ELIC, suggesting that lipid interactions shape the agonist response. Our data offer a structural context for understanding lipid modulation in pLGICs.
Neonicotinoids, sometimes called ‘neonics’, are a class of insecticides targeting neuronal nicotinic acetylcholine receptors, which belong to the family of pentameric ligand-gated ion channels (pLGICs) or Cys-loop receptors. The widespread application of these neurotoxic insecticides in agriculture is tied to the worldwide decline of bee populations, due to their sublethal effects on bee memory, behavior and reproduction. In 2018, the member states of the European Union, agreed on a ban of most neonicotinoids, with exemption for greenhouses. However, a new generation of sulfoximine-based insecticides, such as sulfoxaflor, has recently been introduced to the agricultural market claiming they are chemically distinct from neonicotinoids. Using the acetylcholine binding protein (AChBP) as a well-established tool for structural studies of nicotinic receptors, we engineered AChBP variants that reliably mimic the neurotransmitter binding site of honeybee nicotinic receptors. Three-dimensional structures of sulfoxaflor-bound receptors reveal, despite the distinct chemistry claim, that the molecule is recognized through receptor interactions that are virtually identical to neonicotinoids. This observation suggests that sulfoximine- and neonicotinoid-type insecticides share a common mode of action, and that the legislative ban on certain insecticides should, as a precaution, be expanded based on pharmacological activity at their target receptors rather than on chemical composition. Crystal structure of an insect nicotinic receptor mimic (iAChBP) in complex with the insecticide sulfoxaflor. This abstract is from the Experimental Biology 2019 Meeting. There is no full text article associated with this abstract published in The FASEB Journal.
Phosphoribosyltransferases (PRTs) bind 5'-phospho-α-d-ribosyl-1'-pyrophosphate (PRPP) and transfer its phosphoribosyl group (PRib) to specific nucleophiles. Anthranilate PRT (AnPRT) is a promiscuous PRT that can phosphoribosylate both anthranilate and alternative substrates, and is the only example of a type III PRT. Comparison of the PRPP binding mode in type I, II and III PRTs indicates that AnPRT does not bind PRPP, or nearby metals, in the same conformation as other PRTs. A structure with a stereoisomer of PRPP bound to AnPRT from Mycobacterium tuberculosis (Mtb) suggests a catalytic or post-catalytic state that links PRib movement to metal movement. Crystal structures of Mtb-AnPRT in complex with PRPP and with varying occupancies of the two metal binding sites, complemented by activity assay data, indicate that this type III PRT binds a single metal-coordinated species of PRPP, while an adjacent second metal site can be occupied due to a separate binding event. A series of compounds were synthesized that included a phosphonate group to probe PRPP binding site. Compounds containing a "bianthranilate"-like moiety are inhibitors with IC50 values of 10-60μM, and Ki values of 1.3-15μM. Structures of Mtb-AnPRT in complex with these compounds indicate that their phosphonate moieties are unable to mimic the binding modes of the PRib or pyrophosphate moieties of PRPP. The AnPRT structures presented herein indicated that PRPP binds a surface cleft and becomes enclosed due to re-positioning of two mobile loops.
There are twenty-five published structures of Mycobacterium tuberculosis anthranilate phosphoribosyltransferase (Mtb-AnPRT) that use the same crystallization protocol. The structures include protein complexed with natural and alternative substrates, protein:inhibitor complexes, and variants with mutations of substrate-binding residues. Amongst these are varying space groups (i.e. P21, C2, P21212, P212121). This article outlines experimental details for 3 additional Mtb-AnPRT:inhibitor structures. For one protein:inhibitor complex, two datasets are presented - one generated by crystallization of protein in the presence of the inhibitor and another where a protein crystal was soaked with the inhibitor. Automatic and manual processing of these datasets indicated the same space group for both datasets and thus indicate that the space group differences between structures of Mtb-AnPRT:ligand complexes are not related to the method used to introduce the ligand.
Significance At this time, multidrug-resistant gram-negative bacteria are estimated to cause approximately 700,000 deaths per year globally, with a prediction that this figure could reach 10 million a year by 2050. Antivirulence therapy, in which virulence mechanisms of a pathogen are chemically inactivated, represents a promising approach to the development of treatment options. The family of lipid A phosphoethanolamine transferases in gram-negative bacteria confers bacterial resistance to innate immune defensins and colistin antibiotics. The development of inhibitors to block lipid A phosphoethanolamine transferase could improve innate immune clearance and extend the usefulness of colistin antibiotics. The solved crystal structure and biophysical studies suggest that the enzyme undergoes large conformational changes to enable binding and catalysis of two very differently sized substrates.
Mycobacterium tuberculosis(Mtb) is the causative agent of tuberculosis. Access to iron in host macrophages depends on iron-chelating siderophores called mycobactins and is strongly correlated withMtbvirulence. Here, the crystal structure of anMtbenzyme involved in mycobactin biosynthesis, MbtN, in complex with its FAD cofactor is presented at 2.30 Å resolution. The polypeptide fold of MbtN conforms to that of the acyl-CoA dehydrogenase (ACAD) family, consistent with its predicted role of introducing a double bond into the acyl chain of mycobactin. Structural comparisons and the presence of an acyl carrier protein, MbtL, in the same gene locus suggest that MbtN acts on an acyl-(acyl carrier protein) rather than an acyl-CoA. A notable feature of the crystal structure is the tubular density projecting from N(5) of FAD. This was interpreted as a covalently bound polyethylene glycol (PEG) fragment and resides in a hydrophobic pocket where the substrate acyl group is likely to bind. The pocket could accommodate an acyl chain of 14–21 C atoms, consistent with the expected length of the mycobactin acyl chain. Supporting this, steady-state kinetics show that MbtN has ACAD activity, preferring acyl chains of at least 16 C atoms. The acyl-binding pocket adopts a different orientation (relative to the FAD) to other structurally characterized ACADs. This difference may be correlated with the apparent ability of MbtN to catalyse the formation of an unusualcisdouble bond in the mycobactin acyl chain.
Anthranilate phosphoribosyltransferase (AnPRT) is essential for the biosynthesis of tryptophan in Mycobacterium tuberculosis (Mtb). This enzyme catalyzes the second committed step in tryptophan biosynthesis, the Mg²⁺-dependent reaction between 5'-phosphoribosyl-1'-pyrophosphate (PRPP) and anthranilate. The roles of residues predicted to be involved in anthranilate binding have been tested by the analysis of six Mtb-AnPRT variant proteins. Kinetic analysis showed that five of six variants were active and identified the conserved residue R193 as being crucial for both anthranilate binding and catalytic function. Crystal structures of these Mtb-AnPRT variants reveal the ability of anthranilate to bind in three sites along an extended anthranilate tunnel and expose the role of the mobile β2-α6 loop in facilitating the enzyme's sequential reaction mechanism. The β2-α6 loop moves sequentially between a "folded" conformation, partially occluding the anthranilate tunnel, via an "open" position to a "closed" conformation, which supports PRPP binding and allows anthranilate access via the tunnel to the active site. The return of the β2-α6 loop to the "folded" conformation completes the catalytic cycle, concordantly allowing the active site to eject the product PRA and rebind anthranilate at the opening of the anthranilate tunnel for subsequent reactions. Multiple anthranilate molecules blocking the anthranilate tunnel prevent the β2-α6 loop from undergoing the conformational changes required for catalysis, thus accounting for the unusual substrate inhibition of this enzyme.
The tryptophan-biosynthesis pathway is essential for Mycobacterium tuberculosis (Mtb) to cause disease, but not all of the enzymes that catalyse this pathway in this organism have been identified. The structure and function of the enzyme complex that catalyses the first committed step in the pathway, the anthranilate synthase (AS) complex, have been analysed. It is shown that the open reading frames Rv1609 (trpE) and Rv0013 (trpG) encode the chorismate-utilizing (AS-I) and glutamine amidotransferase (AS-II) subunits of the AS complex, respectively. Biochemical assays show that when these subunits are co-expressed a bifunctional AS complex is obtained. Crystallization trials on Mtb-AS unexpectedly gave crystals containing only AS-I, presumably owing to its selective crystallization from solutions containing a mixture of the AS complex and free AS-I. The three-dimensional structure reveals that Mtb-AS-I dimerizes via an interface that has not previously been seen in AS complexes. As is the case in other bacteria, it is demonstrated that Mtb-AS shows cooperative allosteric inhibition by tryptophan, which can be rationalized based on interactions at this interface. Comparative inhibition studies on Mtb-AS-I and related enzymes highlight the potential for single inhibitory compounds to target multiple chorismate-utilizing enzymes for TB drug discovery.
AnPRT (anthranilate phosphoribosyltransferase), required for the biosynthesis of tryptophan, is essential for the virulence of Mycobacterium tuberculosis (Mtb). AnPRT catalyses the Mg2+-dependent transfer of a phosphoribosyl group from PRPP (5'-phosphoribosyl-1'-pyrophosphate) to anthranilate to form PRA (5'-phosphoribosyl anthranilate). Mtb-AnPRT was shown to catalyse a sequential reaction and significant substrate inhibition by anthranilate was observed. Antimycobacterial fluoroanthranilates and methyl-substituted analogues were shown to act as alternative substrates for Mtb-AnPRT, producing the corresponding substituted PRA products. Structures of the enzyme complexed with anthranilate analogues reveal two distinct binding sites for anthranilate. One site is located over 8 Å (1 Å=0.1 nm) from PRPP at the entrance to a tunnel leading to the active site, whereas in the second, inner, site anthranilate is adjacent to PRPP, in a catalytically relevant position. Soaking the analogues for variable periods of time provides evidence for anthranilate located at transient positions during transfer from the outer site to the inner catalytic site. PRPP and Mg2+ binding have been shown to be associated with the rearrangement of two flexible loops, which is required to complete the inner anthranilate-binding site. It is proposed that anthranilate first binds to the outer site, providing an unusual mechanism for substrate capture and efficient transfer to the catalytic site following the binding of PRPP.