BACKGROUND:Studying a new species using high-throughput sequencing requires a high-quality reference genome. However, assembling chromosome length sequences remains challenging. Recent advances in chromatin conformation capture (Hi-C) have provided a new approach to scaffolding genome assemblies, and the last ten years have seen a proliferation of such methods. However, to our knowledge no comprehensive benchmarking of Hi-C scaffolders has been conducted to date. RESULTS:Through a literature review we identify the most popular Hi-C scaffolders - Lachesis, HiRise, 3d-dna, SALSA, and AllHiC. We test their ability to scaffold four well studied genomes - S. cerevisiae, L. tarentolae, A. thaliana, and H. sapiens. Scaffolders are tasked with both scaffolding fragmented versions of the reference genome as well as de novo assemblies derived from long read datasets. We find that all scaffolders can exceed 80% accuracy under ideal circumstances but that their performance quickly deteriorates under more challenging conditions. Surprisingly, many scaffolders also show poor performance on the best assemblies, where contigs are near chromosome length. Overall, we find that HiRise and Lachesis offer the best performance on average across all conditions. CONCLUSIONS:We compare the performance of five Hi-C scaffolders using multiple reference species under both ideal and real-life conditions, thereby illuminating their strengths and weaknesses.
Identification of ligands targeting essential enzymes in Mycobacterium species remains an important strategy for anti-tuberculosis drug discovery. Here, a native mass spectrometry approach was employed using pooled 100-compound mixtures, enabling the direct detection of intact HPPK-ligand complexes in solution. Dual-mode MS acquisitions (low collision energy for complex detection and high collision energy for ligand confirmation), combined with an automated data analysis workflow, ensured robust identification of binding events from these complex samples. This strategy led to the identification of several HPPK-binding small molecules, all belonging to the dammarane triterpene glycoside (ginsenoside) class. Subsequent analysis of the hits revealed clear structure-affinity relationships, highlighting how specific aglycone modifications and glycosylation patterns influence binding to HPPK. Our findings expand the known chemical space of HPPK ligands and demonstrate the utility of native MS-based screening coupled with automated data analysis to uncover new ligand scaffolds for challenging enzyme targets.
Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.
Trypanosoma cruzi, the etiological agent of Chagas disease, depends on glycolysis for ATP production, rendering its glycolytic enzymes attractive targets for therapeutic development. Here, we report the high-resolution crystal structures of two essential glycolytic enzymes, glucose-6-phosphate isomerase (Tc PGI, 1.8 Å) and enolase (Tc enolase, 2.4 Å) and provide structural and computational analyses to support structure-based drug design. Tc PGI adopts a dimeric αβα sandwich fold and features a parasite-specific 53-residue N-terminal extension and a unique C-terminal hook region which both distinguish it from its human ortholog. Tc enolase exhibits the conserved (α/β) 8 TIM barrel fold but harbors minor distinct structural deviations, including an extended α17 helix and a structured α1 region, which differentiate it from human isoforms. Both enzymes exhibited high thermal stability, consistent with adaptation to the parasite's complex life cycle. Structure-based virtual screening using a scaffold with known multi-target potential identified distinct high-affinity inhibitors for each enzyme. Molecular dynamics simulations further confirmed stable enzyme-inhibitor interactions and favorable binding energetics. Collectively, these findings reveal structural signatures unique to T. cruzi glycolytic enzymes and lay the groundwork for the development of antiparasitic therapeutics.
Carbon-nitrogen hydrolases (CNHs) are members of the diverse nitrilase superfamily of enzymes that facilitate cellular adaptation to environmental stress by metabolizing nitrogen, detoxifying xenobiotics and catabolizing environmentally derived metabolites. Helicobacter pylori CNH (HpCNH) may contribute to metabolic flexibility under acid stress, detoxification of reactive nitrogen species or nutrient scavenging in the nutrient-limited gastric environment. Here, we report the 2.1 Å resolution crystal structure of a CNH from H. pylori strain G27 (PDB entry 6mg6). HpCNH adopts the characteristic nitrilase-superfamily αββα-sandwich core and contains the conserved catalytic cysteine typical of enzymatically active CNHs. The overall structure and active site of HpCNH are most similar to those of carbamoylputrescine amidohydrolase from the plant Medicago truncatula. Despite structural variations in loop regions, including near the active site, HpCNH retains the key residues required to bind putrescine and the prototypical N-carbamoylputrescine amidase active site.
Trypanosoma brucei, the causative agent of Human African Trypanosomiasis (HAT), relies exclusively on purine salvage for nucleotide biosynthesis, making its nucleotide-processing enzymes attractive drug targets. Here, we present a comprehensive structural and functional characterization of T. brucei's nucleoside diphosphate kinase B (TbNDPK), a key enzyme in nucleotide homeostasis. Circular dichroism and fluorescence spectroscopy revealed that TbNDPK is highly stable under thermal and chemical stress and undergoes nucleotide-induced conformational changes. This study also presents high-resolution crystal structures of the apo enzyme and complexes with UDP, CDP, and GDP, showing a conserved hexameric fold, with induced-fit binding via a flexible loop involving Phe59 and key active-site residues. Enzymatic assays revealed substrate preferences for UDP and GDP, while deoxyribonucleotide diphosphates were processed with significantly reduced efficiency. Molecular dynamics simulations revealed ligand-dependent flexibility and subunit-specific nucleotide dynamics, indicating potential asymmetry and cooperative communication within the hexamer. Collectively, these findings position TbNDPK as a thermostable, catalytically efficient, and structurally distinct enzyme optimized for ribonucleotide metabolism and support its potential as a selective target for future antitrypanosomal drug discovery.
Recent advances in Large Language Models (LLMs) present new opportunities for automating critical bottlenecks in scientific workflows such as literature reviews or protocol design. One such bottleneck is the purification of recombinant proteins, a vital aspect of biomedical research that frequently fails. To improve success rates, researchers must manually define optimal large-scale purification conditions and establish robust rescue protocols for proteins with low stability or solubility - a time-intensive process. To address this gap, we introduce a multi-agent LLM system that automates the creation and optimization of protein purification protocols to facilitate the production of high-concentration, high-purity protein samples. Our application streamlines the labor-intensive manual process of sequence similarity searches, literature reviews, and protocol comparison. Operating in a tool-like constrained workflow, the system identifies analogous proteins, leverages specialized LLM agents to extract successful purification methodologies from primary source literature, and cross-references them against failed protocols to generate optimization recommendations. Evaluation on a select number of targets demonstrated high accuracy in protocol extraction and the generation of scientifically sound, expert-validated optimization recommendations. While this system reduces complex analysis time from hours to minutes, we identify the lack of programmatic open access to literature, specifically primary citations in the Protein Data Bank, as a fundamental limitation to LLM agent-based scientific workflows. Ultimately, this system demonstrates the feasibility of using LLM agents to streamline wet-lab workflows while preserving methodological transparency and reproducibility.
Plasmodium vivax is a major cause of malaria globally and has recently been transmitted locally in the USA. P. vivax produces homologs of host proteins, including cytokines such as macrophage migration inhibitory factor (MIF). MIF regulates both adaptive and innate immune responses and contributes to the pathogenesis of parasitic infections, including malaria. Plasma concentrations of P. vivax MIF (PvMIF) correlate with the severity of P. vivax malaria. Plasmodium spp. MIFs have been recognized as candidate malaria vaccines. PvMIF, like other protozoan MIFs, binds to host CD74 and can suppress host MIF-CD74 signaling. The production, crystallization and 1.8 Å resolution structure of PvMIF (PDB entry 9b0m, pdb_00009b0m) are reported. PvMIF crystallized in space group P63 with a single molecule in the asymmetric unit. The biological unit of PvMIF is the prototypical MIF trimer.
Trypanosoma cruzi, the causative agent of Chagas disease, relies heavily on glycolysis for ATP production, making glycolytic enzymes attractive targets for therapeutic intervention. In this study, we report high-resolution crystal structures of two essential glycolytic enzymes, glucose-6-phosphate isomerase (Tc PGI; 1.8 A) and enolase (Tc enolase; 2.4 A), and integrate structural, biophysical, and computational analyses to evaluate their drug-target potential. Tc PGI adopts a dimeric alpha-beta-alpha sandwich fold and contains a parasite-specific 53-residue N-terminal extension and a distinctive C-terminal hook region that are absent in the human ortholog. Tc enolase displays the conserved (alpha/beta)8 TIM-barrel architecture but exhibits localized structural differences, including an extended alpha17 helix and a structured alpha1 region, relative to human isoforms. Both enzymes exhibit high thermal stability, consistent with adaptation to the parasite's complex life cycle. Structure-based virtual screening using a scaffold previously associated with multi-target inhibition identified candidate ligands with favorable docking scores for each enzyme. Subsequent molecular dynamics simulations and binding free-energy analyses supported stable enzyme-ligand interactions and favorable energetic profiles. Together, these results define parasite-specific structural features of two key glycolytic enzymes and provide a structural framework for future experimental validation and structure-guided development of selective antiparasitic inhibitors. ### Competing Interest Statement The authors have declared no competing interest. NIH Common Fund, 75N93022C00036
Trichomonas vaginalis causes trichomoniasis, the most common non-viral sexually transmitted disease in humans. T. vaginalis pyrophosphate-dependent phosphofructokinase ( Tv PPi-PFK) is a putative target for rational, structure-based drug discovery, given its absence in mammals and its importance for parasite survival. Tv PPi-PFK is a cytosolic enzyme that catalyzes the phosphorylation of fructose-6-phosphate using pyrophosphate (PPi) as the phosphoryl donor. This reversible reaction, catalyzed by Tv PPi-PFK, is the first committed step in glycolysis. Its reverse reaction is vital for gluconeogenesis in T. vaginalis . The purification, crystallization, structure determination, and crystal structures of Tv PPi-PFK are reported. Tv PPi-PFK is the first reported eukaryotic PPi-PFK structure. Tv PPi-PFK retains the overall PPi-PFK topology observed in bacterial PPi-PFK including conserved motifs essential for pyrophosphate binding and PPi-PFK catalytic activity. In addition to the catalytic PPi-PFK binding sites, Tv PPi-PFK has two additional ligand binding sites. The first binds AMP usurped during protein production and helps stabilize the Tv PPi-PFK tetramer. A second ligand binding site was observed in proximity to the AMP-binding site and accommodates sugar phosphates soaked into preformed crystals. This sugar phosphates binding site is distinct from the Tv PPi-PFK active site that binds fructose-6-phosphate. Future mutagenesis and activity studies are planned to determine the relevance of both sites. Synopsis:The production, crystallization, and crystal structures of a pyrophosphate-dependent phosphofructokinase from Trichomonas vaginalis ( Tv PPi-PFK) are reported. Tv PPi-PFK has a prototypical PPi-PFK active site as well as unexpected AMP and sugar-phosphate binding sites at the dimer interface.
We report high-quality long-read genome assemblies and annotations for two widely studied reference strains of Leishmania (Viannia) braziliensis, a primary agent of cutaneous and mucocutaneous leishmaniasis. These genomes should facilitate studies of animal infectivity and pathogenesis of cutaneous and severe mucocutaneous leishmaniasis.
The aminoacyl-tRNA synthetases (AaRSs) are an ancient family of structurally diverse enzymes that are divided into two major classes. The functionalities of most AaRSs are inextricably linked to their oligomeric states. While GluRSs were previously classified as monomers, the current investigation reveals that the form expressed in Pseudomonas aeruginosa is a rotationally pseudosymmetrical homodimer featuring intersubunit tRNA binding sites. Both subunits display a highly bent, "pipe strap" conformation, with the anticodon binding domain directed toward the active site. The tRNA binding sites are similar in shape to those of the monomeric GluRSs, but are formed through an approximately 180-degree rotation of the anticodon binding domains and dimerization via the anticodon and D-arm binding domains. As a result, each anticodon binding domain is poised to recognize the anticodon loop of a tRNA bound to the adjacent protomer. Additionally, the anticodon binding domain has an α-helical C-terminal extension containing a conserved lysine-rich consensus motif positioned near the predicted location of the acceptor arm, suggesting dual functions in tRNA recognition. The unique architecture of PaGluRS broadens the structural diversity of the GluRS family, and member synthetases of all bacterial AaRS subclasses have now been identified that exhibit oligomerization.
Helicobacter pylori is the primary causative agent of peptic ulcer disease, among other gastrointestinal ailments, and currently affects over half of the global population. Although some treatments exist, growing resistance to these drugs has prompted efforts to develop novel approaches to fighting this pathogen. To generate many of the nucleotides essential to biochemical processes, H. pylori relies exclusively on the de novo biosynthesis of these molecules. Recent drug-discovery efforts have targeted the first committed step of this pathway, catalysed by a class 2 dihydroorotate dehydrogenase (DHODH). However, these initiatives have been limited by the lack of a crystal structure. Here, we detail the crystal structure of H. pylori DHODH (HpDHODH) at 2.25 Å resolution (PDB entry 6b8s). We performed a large-scale bioinformatics search to find evolutionary homologs. Our results indicate that HpDHODH shows high conservation of both sequence and structure in its active site. We identified key polar interactions between the HpDHODH protein and its requisite flavin mononucleotide (FMN) cofactor, identifying amino-acid residues that are critical to its function. Most notably, we found that HpDHODH maintains several structural features that allow it to associate with the inner membrane and utilize ubiquinone to achieve catalytic turnover. We discovered a hydrophobic channel that runs from the putative membrane interface on the N-terminal microdomain to the core of the protein. We predict that this channel establishes a connection between the ubiquinone pool in the membrane and the FMN in the active site. These findings provide a structural explanation for the competitive inhibition of ubiquinone by pyrazole-based compounds that was determined biochemically in other studies. Understanding this mechanism may facilitate the development of new drugs targeting this enzyme and push the effort to find a resistance-free treatment for H. pylori.
Glycyl tRNA synthetases (GlyRSs) are prospective drug targets for combating Mycobacterium tuberculosis (Mtb) and cancer in humans. These synthetases are of the α2-subtype, with the ortholog in humans being dual targeted to the cytosol and mitochondria. Whereas the human enzyme has been structurally characterized previously in several liganded states, no structures of MtbGlyRS have thus far been reported. Here, we describe our recent work with MtbGlyRS and the closely-related Mycobacterium thermoresitibile GlyRS (MtrGlyRS), which progressed through all phases of the structural genomics pipeline, for the purpose of facilitating structure-based drug discovery. MtbGlyRS was expressed in Mycobacterium smegmatis and MtrGlyRS in Escherichia coli. Crystal structures are described for complexes of the two enzymes with adenosine monophosphate (AMP) and glycyl-sulfamoyl-adenylate (glycyl-AMS) at resolutions of 1.65/2.90 and 2.25/1.95 Å, respectively, and for MtrGlyRS in its apo state at 2.85 Å. Despite crystallizing in the dimeric state characteristic of many class II synthetases, the two enzymes elute predominantly as monomers during size exclusion chromatography. Strikingly, significant portions of the dimer interface and active site are unstructured in the MtrGlyRS apoenzyme crystal. AMP orders two tRNA recognition loops and a section of the insertion domain, and glycyl-AMS further stabilizes the structure, including the closure of a lid motif. Both the active and anticodon binding sites display structural differences with the human GlyRS and thus the collection of crystal structures should be useful for guiding drug development efforts targeting the various characterized structural states.
We report a high-quality hybrid genome assembly and annotation for a Leishmania (Viannia) guyanensis line (M4147) expressing a firefly luciferase reporter gene and bearing the totivirus LRV1 strongly implicated in parasite hypervirulence. This reference genome should facilitate the study of infectivity and pathogenesis of cutaneous and severe leishmaniasis.
The methylerythritol phosphate (MEP) pathway is a metabolic pathway that produces the isoprenoids isopentyl pyrophosphate and dimethylallyl pyrophosphate. Notably, the MEP pathway is present in bacteria and not in mammals, which makes the enzymes of the MEP pathway attractive targets for the discovery of new anti-infective agents due to the reduced chances of off-target interactions leading to side effects. There are seven enzymes in the MEP pathway, the fifth of which is IspF. Crystal structures of Burkholderia pseudomallei IspF were determined with five different sulfonamide ligands bound. The sulfonamide-containing ligands were ethoxzolamide, acetazolamide, sulfapyridine and sulfamonomethoxine. The fifth bound ligand was a synthetic analog of acetazolamide. All ligands coordinated to the active-site Zn +2 ion through the sulfonamide group, although sulfapyridine and sulfamonomethoxine, both of which are known antibacterial agents, possess similar binding interactions that are distinct from the other three sulfonamides. These structural data will aid in the discovery of new IspF inhibitors.
Helicobacter pylori, a type 1 carcinogen that causes human gastric ulcers and cancer, is a priority target of the Seattle Structural Genomics Center for Infectious Disease (SSGCID). These efforts include determining the structures of potential H. pylori therapeutic targets. Here, the purification, crystallization and X-ray structure of one such target, H. pylori biotin protein ligase (HpBPL), are reported. HpBPL catalyzes the activation of various biotin-dependent metabolic pathways, including fatty-acid synthesis, gluconeogenesis and amino-acid catabolism, and may facilitate the survival of H. pylori in the high-pH gastric mucosa. HpBPL is a prototypical bacterial biotin protein ligase, despite having less than 35% sequence identity to any reported structure in the Protein Data Bank. A biotinyl-5-ATP molecule sits in a well conserved cavity. HpBPL shares extensive tertiary-structural similarity with Mycobacterium tuberculosis biotin protein ligase (MtBPL), despite having less than 22% sequence identity. The active site of HpBPL is very similar to that of MtBPL and has the necessary residues to bind inhibitors developed for MtBPL.
Entamoeba histolytica causes amebiasis, a neglected disease that kills ∼100 000 people globally each year. Due to emerging drug resistance, E. histolytica is one of the target organisms for structure-based drug discovery by the Seattle Structural Genomics Center for Infectious Disease (SSGCID). Purification, crystallization and three structures of the putative drug target endoribonuclease L-PSP from E. histolytica (EhL-PSP) are presented. EhL-PSP has a two-layer α/β-sandwich with structural homology to endoribonuclease L-PSP. All three structures reveal the prototypical YjgF/YER057c/UK114 family trimer topology with accessible allosteric active sites. Citrate molecules from the crystallization solution are bound to the allosteric site in two of the three reported structures. The large allosteric site of EhL-PSP is well conserved with bacterial YjgF/YER057c/UK114 family members and could be targeted for inhibition, drug discovery or repurposing.
Mycobacterium tuberculosis is a Gram-positive bacillus that causes tuberculosis and is a leading cause of mortality worldwide. This disease is a growing health threat due to the occurrence of multidrug resistance. Mycolic acids are essential for generating cell walls and their modification is important to the virulence and persistence of M. tuberculosis. A family of S -adenosylmethionine-dependent mycolic acid synthases modify mycolic acids and represent promising drug targets. UmaA is currently the least-understood member of this family. This paper describes the crystal structure of UmaA. UmaA is a monomer composed of two domains: a structurally conserved SAM-binding domain and a variable substrate-binding auxiliary domain. Fortuitously, our structure contains a nitrate in the active site, a structural mimic of carbonate, which is a known general base in cyclopropane-adding synthases. Further investigation indicated that the structure of the N-terminus is highly flexible. Finally, we have identified S -adenosyl- N -decyl-aminoethyl as a promising potential inhibitor.
The focused issue on Empowering education through structural genomics is introduced. The virtual issue is available at https://journals.iucr.org/special_issues/2024/educationsg.