Abstract Tuberculosis (TB), caused by the Mycobacterium tuberculosis complex (MTBC), remains a pressing global health challenge, with a high burden in West Africa, including The Gambia. Understanding the genetic diversity of circulating MTBC strains is essential for improving diagnosis, surveillance and treatment strategies. In this study, we characterise the population structure and drug resistance landscape of MTBC strains circulating in The Gambia over nearly two decades (2002–2021). We analysed whole-genome sequencing (WGS) data from 1,803 TB isolates. Lineage 4 (L4) was predominant (67.2%), followed by the West Africa-restricted lineage 6 (L6, 26.6%), with L4 exhibiting greater genetic diversification over time. Drug susceptibility profiling of these isolates revealed that 78% (1421/1803) were drug-susceptible, while 6.5% (119/1803) harboured resistance to first-line drugs, primarily to isoniazid, rifampicin, or both. Notably, 15.5% (282/1803) isolates carried mutations classified as having uncertain significance according to the WHO resistance catalogue. Comparative analyses revealed a lineage 6-specific ethambutol-associated mutation of uncertain significance (embC Ala307Thr) occurring at a higher frequency in Gambian isolates than in the broader West Africa region or globally. Structural modelling demonstrated that many first-line drug resistance mutations are located in highly conserved, solvent-inaccessible regions of target proteins, often impacting protein stability, suggesting a trade-off between drug resistance, bacterial fitness, and evolutionary adaptation. Together, these findings highlight the coexistence of globally widespread and regionally restricted MTBC lineages in The Gambia and reveal a substantial burden of resistance-associated mutations of uncertain significance in the WHO catalogue. Sustained genomic surveillance and region-specific interpretation of resistance mutations are essential to support End TB strategies in high-burden settings.
Trypanosoma brucei is the causal agent of African trypanosomiasis in humans and animals, the latter resulting in significant negative economic impacts in afflicted areas of the world. Resistance has arisen to the trypanocidal drugs pentamidine and melarsoprol through mutations in the aquaglyceroporin TbAQP2 that prevent their uptake. Here, we use cryogenic electron microscopy to determine the structure of TbAQP2 from T. brucei, bound to either the substrate glycerol or to the sleeping sickness drugs, pentamidine or melarsoprol. The drugs bind within the AQP2 channel at a site completely overlapping that of glycerol. Mutations leading to a drug-resistant phenotype were found in the channel lining. Molecular dynamics (MD) simulations showed the channel can be traversed by pentamidine, with a low energy binding site at the centre of the channel, flanked by regions of high energy association at the extracellular and intracellular ends. Drug-resistant TbAQP2 mutants are still predicted to bind pentamidine, but the much weaker binding in the centre of the channel observed in the MD simulations would be insufficient to compensate for the high energy processes of ingress and egress, hence impairing transport at pharmacologically relevant concentrations. The structures of drug-bound TbAQP2 represent a novel paradigm for drug–transporter interactions and are a new mechanism for targeting drugs in pathogens and human cells.
Tuberculosis is an infectious disease caused by Mycobacterium tuberculosis (Mtb) and is one of the leading causes of death worldwide. This disease is typically treated by combining several antimicrobials for extended periods, which can lead to treatment interruptions by patients and promote the emergence of multidrug-resistant strains, necessitating the use of alternative or second-line drugs. In this perspective, dihydrodipicolinate synthase (DapA) from M. tuberculosis (MtDapA), which catalyzes the aldol condensation between pyruvate and aspartate-semialdehyde (ASA) to produce dihydrodipicolinate, is an essential enzyme in Mtb for the production of l-lysine and meso-diaminopimelate (mDAP). Through crystallographic assays, we have determined the structure of MtDapA in complex with its substrate, pyruvate, covalently bonded through a Schiff base to the catalytic l-lysine at a resolution of 1.5 Å. Through structural analysis, we describe the arrangement of interactions between the active site amino acid residues and pyruvate, providing insight into the binding mode of this molecule. In addition, we performed further biophysical assays, including differential scanning fluorimetry (DSF) and isothermal titration calorimetry (ITC), to obtain insights into the pyruvate affinity and the potential role of l-lysine and mDAP as allosteric regulators of MtDapA. However, in contrast to those observed in other orthologous enzymes, particularly those from Gram-negative bacteria, MtDapA does not have an affinity for l-lysine or mDAP. Consequently, this enzyme is not allosterically regulated by the products of this pathway. The results shown here provide evidence regarding the functioning of the enzyme regulatory mechanism and valuable structural features to aid in the future development of MtDapA inhibitors, which may be further explored in drug discovery campaigns against tuberculosis.
Mycobacterium abscessus is one of the leading causes of pulmonary infections caused by non-tuberculous mycobacteria. The ability of M. abscessus to establish a chronic infection in the lung relies on a series of adaptive mutations impacting, in part, global regulators and cell envelope biosynthetic enzymes. One of the genes under strong evolutionary pressure during host adaptation is ubiA, which participates in the elaboration of the arabinan domains of two major cell envelope polysaccharides: arabinogalactan (AG) and lipoarabinomannan (LAM). We here show that patient-derived UbiA mutations not only cause alterations in the AG, LAM, and mycolic acid contents of M. abscessus but also tend to render the bacterium more prone to forming biofilms while evading uptake by innate immune cells and enhancing their pro-inflammatory properties. The fact that the effects of UbiA mutations on the physiology and pathogenicity of M. abscessus were impacted by the rough or smooth morphotype of the strain suggests that the timing of their selection relative to morphotype switching may be key to their ability to promote chronic persistence in the host.IMPORTANCEMultidrug-resistant pulmonary infections caused by Mycobacterium abscessus and subspecies are increasing in the U.S.A. and globally. Little is known of the mechanisms of pathogenicity of these microorganisms. We have identified single-nucleotide polymorphisms (SNPs) in a gene involved in the biosynthesis of two major cell envelope polysaccharides, arabinogalactan and lipoarabinomannan, in lung-adapted isolates from 13 patients. Introduction of these individual SNPs in a reference M. abscessus strain allowed us to study their impact on the physiology of the bacterium and its interactions with immune cells. The significance of our work is in identifying some of the mechanisms used by M. abscessus to colonize and persist in the human lung, which will facilitate the early detection of potentially more virulent clinical isolates and lead to new therapeutic strategies. Our findings may further have broader biomedical impacts, as the ubiA gene is conserved in other tuberculous and non-tuberculous mycobacterial pathogens.
Similarity between candidate drug targets and human proteins is commonly assessed to minimize the occurrence of side effects. Although numerous drugs have been found to disrupt the health of the human microbiome, no comprehensive comparison between established drug targets and the human microbiome metaproteome has yet been conducted. Therefore, herein, sequence and structure alignments between human and pathogen drug targets and representative human gut, oral, and vaginal microbiome metaproteomes were performed. Both human and pathogen drug targets were found to be similar in sequence, function, structure, and drug binding capacity to proteins in diverse pathogenic and non-pathogenic bacteria from all three microbiomes. The gut metaproteome was identified as particularly susceptible overall to off-target effects. Certain symptoms, such as infections and immune disorders, may be more common among drugs that non-selectively target host microbiota. These findings suggest that similarities between human microbiome metaproteomes and drug target candidates should be routinely checked.
Enzymes of the GNAT (GCN5-relate N-acetyltransferases) superfamily are important regulators of cell growth and development. They are functionally diverse and share low amino acid sequence identity, making functional annotation difficult. In this study, we report the function and structure of a new ribosomal enzyme, Nα-acetyl transferase from Bacillus cereus (RimLBC), a protein that was previously wrongly annotated as an aminoglycosyltransferase. Firstly, extensive comparative amino acid sequence analyses suggested RimLBC belongs to a cluster of proteins mediating acetylation of the ribosomal protein L7/L12. To assess if this was the case, several well established substrates of aminoglycosyltransferases were screened. The results of these studies did not support an aminoglycoside acetylating function for RimLBC. To gain further insight into RimLBC biological role, a series of studies that included MALDI-TOF, isothermal titration calorimetry, NMR, X-ray protein crystallography, and site-directed mutagenesis confirmed RimLBC affinity for Acetyl-CoA and that the ribosomal protein L7/L12 is a substrate of RimLBC. Last, we advance a mechanistic model of RimLBC mode of recognition of its protein substrates. Taken together, our studies confirmed RimLBC as a new ribosomal Nα-acetyltransferase and provide structural and functional insights into substrate recognition by Nα-acetyltransferases and protein acetylation in bacteria.
The number of antibiotic resistant pathogens is increasing rapidly, and with this comes a substantial socioeconomic cost that threatens much of the world. To alleviate this problem, we must use antibiotics in a more responsible and informed way, further our understanding of the molecular basis of drug resistance, and design new antibiotics. Here, we focus on a key drug-resistant pathogen, Mycobacterium tuberculosis, and computationally analyze trends in drug-resistant mutations in genes of the proteins embA, embB, embC, and katG, which play essential roles in the action of the first-line drugs ethambutol and isoniazid. We use docking to predict binding modes of isoniazid to katG that agree with suggested binding sites found in our laboratory using cryo-EM. Using mutant stability predictions, we recapitulate the idea that resistance occurs when katG's heme cofactor is destabilized rather than due to a decrease in affinity to isoniazid. Conversely, we have identified resistance mutations that affect the affinity of ethambutol more drastically than the affinity of the natural substrate of embB. With this, we illustrate that we can distinguish between the two types of drug resistance-cofactor destabilization and drug affinity reduction-suggesting potential uses in the prediction of novel drug-resistant mutations.
Ligand binding hotspots are regions of protein surfaces that form particularly favourable interactions with small molecule pharmacophores. Targeting interactions with these hotspots maximises the efficiency of ligand binding. Existing methods are capable of identifying hotspots but often lack assays to quantify ligand binding and direct elaboration at these sites. Herein, we describe a fragment-based competitive 19 F Ligand Based NMR (LB-NMR) screening platform that enables routine, quantitative ligand profiling focused at ligand-binding hotspots. As a proof of concept, the method was applied to 4′-phosphopantetheine adenylyltransferase (PPAT) from Mycobacterium abscessus ( Mabs ). X-ray crystallographic characterisation of the hits from a 960-member fragment screen identified three ligand-binding hotspots across the PPAT active site. From the fragment hits a collection of 19 F reporter candidates were designed and synthesised. By rigorous prioritisation and use of optimisation workflows, a single 19 F reporter molecule was generated for each hotspot. Profiling the binding of a set of structurally characterised ligands by competitive 19 F LB-NMR with this suite of 19 F reporters recapitulated the binding affinity and site ID assignments made by ITC and X-ray crystallography. This quantitative mapping of ligand binding events at hotspot level resolution establishes the utility of the fragment-based competitive 19 F LB-NMR screening platform for hotspot-directed ligand profiling
Machine learning has provided a means to accelerate early-stage drug discovery by combining molecule generation and filtering steps in a single architecture that leverages the experience and design preferences of medicinal chemists. However, designing machine learning models that can achieve this on the fly to the satisfaction of medicinal chemists remains a challenge owing to the enormous search space. Researchers have addressed de novo design of molecules by decomposing the problem into a series of tasks determined by design criteria. Here we provide a comprehensive overview of the current state of the art in molecular design using machine learning models as well as important design decisions, such as the choice of molecular representations, generative methods and optimization strategies. Subsequently, we present a collection of practical applications in which the reviewed methodologies have been experimentally validated, encompassing both academic and industrial efforts. Finally, we draw attention to the theoretical, computational and empirical challenges in deploying generative machine learning and highlight future opportunities to better align such approaches to achieve realistic drug discovery end points. Data-driven generative methods have the potential to greatly facilitate molecular design tasks for drug design.
The major human bacterial pathogen Pseudomonas aeruginosa causes multidrug-resistant infections in people with underlying immunodeficiencies or structural lung diseases such as cystic fibrosis (CF). We show that a few environmental isolates, driven by horizontal gene acquisition, have become dominant epidemic clones that have sequentially emerged and spread through global transmission networks over the past 200 years. These clones demonstrate varying intrinsic propensities for infecting CF or non-CF individuals (linked to specific transcriptional changes enabling survival within macrophages); have undergone multiple rounds of convergent, host-specific adaptation; and have eventually lost their ability to transmit between different patient groups. Our findings thus explain the pathogenic evolution of P. aeruginosa and highlight the importance of global surveillance and cross-infection prevention in averting the emergence of future epidemic clones.
We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models. We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop.
Mycobacterium abscessus is increasingly recognized as the causative agent of chronic pulmonary infections in humans. One of the genes found to be under strong evolutionary pressure during adaptation of M. abscessus to the human lung is embC which encodes an arabinosyltransferase required for the biosynthesis of the cell envelope lipoglycan, lipoarabinomannan (LAM). To assess the impact of patient-derived embC mutations on the physiology and virulence of M. abscessus, mutations were introduced in the isogenic background of M. abscessus ATCC 19977 and the resulting strains probed for phenotypic changes in a variety of in vitro and host cell-based assays relevant to infection. We show that patient-derived mutational variations in EmbC result in an unexpectedly large number of changes in the physiology of M. abscessus, and its interactions with innate immune cells. Not only did the mutants produce previously unknown forms of LAM with a truncated arabinan domain and 3-linked oligomannoside chains, they also displayed significantly altered cording, sliding motility, and biofilm-forming capacities. The mutants further differed from wild-type M. abscessus in their ability to replicate and induce inflammatory responses in human monocyte-derived macrophages and epithelial cells. The fact that different embC mutations were associated with distinct physiologic and pathogenic outcomes indicates that structural alterations in LAM caused by nonsynonymous nucleotide polymorphisms in embC may be a rapid, one-step, way for M. abscessus to generate broad-spectrum diversity beneficial to survival within the heterogeneous and constantly evolving environment of the infected human airway.
Summary Analysing protein structure similarities is an important step in protein engineering and drug discovery. Methodologies that are more advanced than simple RMSD are available but often require extensive mathematical or computational knowledge for implementation. Grouping and optimizing such tools in an efficient open-source library increases accessibility and encourages the adoption of more advanced metrics. Melodia is a Python library with a complete set of components devised for describing, comparing and analysing the shape of protein structures using differential geometry of 3D curves and knot theory. It can generate robust geometric descriptors for thousands of shapes in just a few minutes. Those descriptors are more sensitive to structural feature variation than RMSD deviation. Melodia also incorporates sequence structural annotation and 3D visualizations. Availability and implementation Melodia is an open-source Python library freely available on https://github.com/rwmontalvao/Melodia_py, along with interactive Jupyter Notebook tutorials.
Aurora A kinase, a cell division regulator, is frequently overexpressed in various cancers, provoking genome instability and resistance to antimitotic chemotherapy. Localization and enzymatic activity of Aurora A are regulated by its interaction with the spindle assembly factor TPX2. We have used fragment-based, structure-guided lead discovery to develop small molecule inhibitors of the Aurora A-TPX2 protein-protein interaction (PPI). Our lead compound, CAM2602, inhibits Aurora A:TPX2 interaction, binding Aurora A with 19 nM affinity. CAM2602 exhibits oral bioavailability, causes pharmacodynamic biomarker modulation, and arrests the growth of tumor xenografts. CAM2602 acts by a novel mechanism compared to ATP-competitive inhibitors and is highly specific to Aurora A over Aurora B. Consistent with our finding that Aurora A overexpression drives taxane resistance, these inhibitors synergize with paclitaxel to suppress the outgrowth of pancreatic cancer cells. Our results provide a blueprint for targeting the Aurora A-TPX2 PPI for cancer therapy and suggest a promising clinical utility for this mode of action.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
Ostrinia furnacalis is a species of moth in the Crambidae family that is harmful to maize and other corn crops in Southeast Asia and the Western Pacific regions. Ostrinia furnacalis causes devastating losses to economically important corn fields. The β-N-acetyl-D-hexosaminidase is an essential enzyme in O. furnacalis and its substrate binding +1 active site is different from that of the plants and humans β-N-acetyl-D-hexosaminidases. To develop environment-friendly insecticides against OfHex1, we conducted structure-guided computational insecticide discovery to identify potential inhibitors that can bind the active site and inhibit the substrate binding and activity of the enzyme. We adopted a three-pronged strategy to conduct virtual screening using Glide and virtual screening workflow (VSW) in Schrödinger Suite-2022-3, against crystal structures of OfHex1 (PDB Id:3NSN), its homologue in humans (PDB Id: 1NP0) and Alphafold model of β-N-acetyl-D-hexosaminidase from Trichogramma pretiosum, an egg parasitoid that protects the crops from O. furnacalis. A library of 20,313 commercially available and "insecticide-like" compounds was extracted from published literature. LigPrep enabled 44,943 ready-to-dock conformers generation. Glide docking revealed 18 OfHex1-specific hits that were absent in human and T. pretiosum screens. Reference docking was conducted using inhibitors/natural ligands in the crystal structures and hits with better docking scores than the reference were selected for MD simulations using Desmond to understand the stability of hit-target interactions. We noted five compounds that bound to OfHex1 TMX active-site based on their docking scores, consistent binding as noted by MD simulations and their insecticide/pesticide likeliness as noted by the Comprehensive Pesticide Likeness Analysis.Communicated by Ramaswamy H. Sarma.
The coenzyme A (CoA) biosynthesis pathway has attracted attention as a potential target for much-needed novel antimicrobial drugs, including for the treatment of tuberculosis (TB), the lethal disease caused by Mycobacterium tuberculosis (Mtb). Seeking to identify inhibitors of Mtb phosphopantetheine adenylyltransferase (MtbPPAT), the enzyme that catalyses the penultimate step in CoA biosynthesis, we performed a fragment screen. In doing so, we discovered three series of fragments that occupy distinct regions of the MtbPPAT active site, presenting a unique opportunity for fragment linking. Here we show how, guided by X-ray crystal structures, we could link weakly-binding fragments to produce an active site binder with a KD <20 μM and on-target anti-Mtb activity, as demonstrated using CRISPR interference. This study represents a big step toward validating MtbPPAT as a potential drug target and designing a MtbPPAT-targeting anti-TB drug.
AbstractThe protein kinase Aurora A, and its close relative, Aurora B, regulate human cell division. Aurora A is frequently overexpressed in cancers of the breast, ovary, pancreas and blood, provoking genome instability and resistance to anti-mitotic chemotherapy. Intracellular localization and enzymatic activity of Aurora A are regulated by its interaction with the spindle assembly factor TPX2. Here, we have used fragment-based, structure-guided lead discovery to develop small-molecule inhibitors of the Aurora A-TPX2 protein-protein interaction (PPI). These compounds act by novel mechanism compared to existing Aurora A inhibitors and they are highly specific to Aurora A over Aurora B. Our biophysically, structurally and phenotypically validated lead compound,CAM2602, exhibits oral bioavailability, favourable pharmacokinetics, pharmacodynamic biomarker modulation, and arrest of growth in tumour xenografts. Consistent with our original finding that Aurora A overexpression drives taxane-resistance in cancer cells, inhibition of Aurora A-TPX2 PPI synergizes with paclitaxel to suppress the outgrowth of pancreatic cancer cells. Our results provide a blueprint for targeting the Aurora A-TPX2 PPI for cancer therapy and suggest a promising clinical utility for this mode of action.
Deep generative models for structure-based drug design (SBDD), where molecule generation is conditioned on a 3D protein pocket, have received considerable interest in recent years. These methods offer the promise of higher-quality molecule generation by explicitly modelling the 3D interaction between a potential drug and a protein receptor. However, previous work has primarily focused on the quality of the generated molecules themselves, with limited evaluation of the 3D molecule \emph{poses} that these methods produce, with most work simply discarding the generated pose and only reporting a "corrected" pose after redocking with traditional methods. Little is known about whether generated molecules satisfy known physical constraints for binding and the extent to which redocking alters the generated interactions. We introduce PoseCheck, an extensive analysis of multiple state-of-the-art methods and find that generated molecules have significantly more physical violations and fewer key interactions compared to baselines, calling into question the implicit assumption that providing rich 3D structure information improves molecule complementarity. We make recommendations for future research tackling identified failure modes and hope our benchmark can serve as a springboard for future SBDD generative modelling work to have a real-world impact.
Abstract The classical Non-Homologous End Joining (c-NHEJ) pathway is the predominant process in mammals for repairing endogenous, accidental or programmed DNA Double-Strand Breaks. c-NHEJ is regulated by several accessory factors, post-translational modifications, endogenous chemical agents and metabolites. The metabolite inositol-hexaphosphate (IP6) stimulates c-NHEJ by interacting with the Ku70–Ku80 heterodimer (Ku). We report cryo-EM structures of apo- and DNA-bound Ku in complex with IP6, at 3.5 Å and 2.74 Å resolutions respectively, and an X-ray crystallography structure of a Ku in complex with DNA and IP6 at 3.7 Å. The Ku-IP6 interaction is mediated predominantly via salt bridges at the interface of the Ku70 and Ku80 subunits. This interaction is distant from the DNA, DNA-PKcs, APLF and PAXX binding sites and in close proximity to XLF binding site. Biophysical experiments show that IP6 binding increases the thermal stability of Ku by 2°C in a DNA-dependent manner, stabilizes Ku on DNA and enhances XLF affinity for Ku. In cells, selected mutagenesis of the IP6 binding pocket reduces both Ku accrual at damaged sites and XLF enrolment in the NHEJ complex, which translate into a lower end-joining efficiency. Thus, this study defines the molecular bases of the IP6 metabolite stimulatory effect on the c-NHEJ repair activity.