
Patients with inflammatory bowel disease (IBD) have an increased risk of colorectal cancer (CRC), but how chronic intestinal inflammation drives malignant transformation remains unclear. We retrospectively reanalyzed published single-cell transcriptomic datasets from intestinal biopsies of healthy individuals and patients with IBD; differential expression was assessed using independent t tests with Benjamini–Hochberg false discovery rate correction. We then integrated those single-cell findings with the Cancer Genome Atlas bulk transcriptomes and pharmacogenomic cohorts to trace stromal programs across the IBD-to-cancer continuum. ARHGEF15 emerged as a stromal gene enriched in CD74 hi HLA-DRB1 hi arterial pericytes within inflamed tissue. Its expression rose steadily from IBD to CRC and tracked with epithelial–mesenchymal transition (EMT) activity. In CRC, higher ARHGEF15 expression was associated with shorter overall and progression-free survival. These retrospective, in silico findings identify ARHGEF15 as an exploratory stromal biomarker associated with inflammatory EMT and stromal–immune remodeling during IBD-to-CRC progression. Prospective experimental and clinical validation is required to establish its prognostic or therapeutic relevance.
Aromatase (CYP19A1) is the rate-limiting enzyme in estrogen biosynthesis and the principal therapeutic target for estrogen receptor-positive (ER+) breast cancer. Despite the clinical success of third-generation aromatase inhibitors (AIs) such as letrozole and anastrozole, acquired resistance and systemic adverse effects, including musculoskeletal pain, bone density loss, and cardiovascular complications, which necessitate the identification of novel scaffolds. This study employs an integrated computational pipeline combining ligand-based quantitative structure-activity relationship (QSAR) modelling and structure-based molecular docking to screen the AfroDB and EANPDB natural product libraries for putative aromatase inhibitors. A curated dataset of 3,068 unique compounds derived from ChEMBL (v33) was used to train Random Forest models, with the binary classification model achieving an area under the receiver operating characteristic curve (AUC) of 0.95 on the independent holdout set, an accuracy of 0.90, and a Matthews correlation coefficient (MCC) of 0.76. The regression model for pIC 50 prediction attained an R 2 of 0.65 and an RMSE of 0.81 pIC 50 units on the test set. The trained classifiers were applied to screen 1,871 AfroDB compounds; 361 were predicted active, of which 71 resided within the defined Applicability Domain. Molecular docking into the human aromatase crystal structure (PDB ID: 3S79) revealed that the top candidate, a prenylated dihydroflavone (5,7-dihydroxy-4′-methoxy-3′-(3-hydroxy-3-methylbut-1-enyl)-5′-(3-methylbut-2-enyl)flavanone), achieved a binding affinity of −9.598 kcal/mol, outperforming both letrozole (−7.11 kcal/mol) and anastrozole (−7.617 kcal/mol) under identical docking conditions. The docking protocol was validated by redocking of the co-crystallised ligand (RMSD = 1.35 Å). Comprehensive interaction analysis identified contacts with key active-site residues Arg115, Phe221, Thr310, Val370, Met374, and Phe430. ADMET profiling indicated that 65% of prioritised hits satisfy Lipinski’s Rule of Five, while 95% meet Veber bioavailability criteria. This study provides a reproducible computational framework for next-generation AI discovery, adhering to the TRIPOD+AI reporting guidelines for machine learning in biomedical research.
Plasmodium falciparum Calcium-Dependent Protein Kinase 4 ( Pf CDPK4) is a validated target for malaria transmission-blocking interventions, as its inhibition disrupts male gametocyte exflagellation. Hexahydroquinolines (HHQs) have emerged as promising gametocytocidal agents. This study aimed to design novel HHQs as potential Pf CDPK4 inhibitors with good binding affinity, favourable pharmacokinetics, and structural stability. A library of 20,000 novel HHQ analogs was generated using genetic algorithm-driven de novo molecular design in AlvaBuilder and systematically filtered through a pipeline comprising machine-learning-based bioactivity prediction, PAINS removal, pharmacophore modeling, ADMET screening, and structure-based virtual screening. Top-ranking compounds underwent bioisosteric optimization and were evaluated using 300 ns molecular dynamics simulations and density functional theory (DFT) calculations. Comparative in silico validation against known inhibitor Bumped Kinase Inhibitor-1 (BKI-1) and the co-crystallized ligand DXR as controls demonstrated that four HHQ analogs exhibited comparable binding affinities and stable interactions with key residues within the ATP-binding pocket. Favourable ADMET profiles, dynamic stability, and supportive electronic properties further reinforced their inhibition potential. These findings provide biologically meaningful computational evidence supporting HHQ scaffolds as potential Pf CDPK4 inhibitors and demonstrate the utility of integrated bioinformatics approaches for malaria transmission-blocking drug discovery. Further experimental validation is required to confirm the inhibitory activities of the identified hexahydroquinoline compounds.
Bacillus thuringiensis (Bt) is a well-known entomopathogenic bacterium widely used in biopesticide formulations due to its diverse arsenal of insecticidal toxins and environmental adaptability. However, the genetic diversity and virulence potential of native Bt strains from South Asia remain underexplored. In this study, we performed whole-genome sequencing and comparative genomic analysis of B. thuringiensis strain JSd1, isolated from Bangladeshi soil. The high-quality draft genome, assembled at 75x coverage, comprises 5.39 Mb in 64 contigs with a GC content of 35.28% and 99.26% completeness. Genome annotation revealed 5,833 genes, including 5,756 protein-coding sequences and numerous non-coding RNAs. It also harbors 4 different plasmids. Importantly, we identified 25 genomic islands harboring mobile elements and hypothetical proteins, highlighting the strain’s dynamic genome evolution. The genome encodes a diverse array of virulence factors linked to insecticidal activity. Notable genes include cry22A and vip3A homologs, multiple bmp1 -like metalloproteases, enhancin , cytK , and chiA , as well as various chitinases and serine proteases. The co-occurrence of chromosomal and plasmid-encoded virulence factors suggests modular acquisition mechanisms. Secondary metabolite biosynthetic clusters, such as those for petrobactin, bacillibactin, thuricin CD, and other novel RiPPs and NRPs, were detected, supporting the strain’s potential in biological control. Phylogenetic analysis positioned JSd1 within the B. cereus sensu lato group, forming a highly supported clade with B. thuringiensis serovar konkukian and B. anthracis . Comparative genomic and pan-genome analyses revealed substantial genomic diversity among B. thuringiensis strains. The strain’s genome contains 2,237 core genes and a large accessory genome, reflecting its ecological adaptability. Our findings suggest that B. thuringiensis JSd1 is a promising candidate for development as a biopesticide targeting insect pests in Bangladeshi agriculture. The comprehensive genomic insights lay the groundwork for further functional validation and field applications, contributing to sustainable pest management strategies in the region.
Bacillus thuringiensis (Bt) is a well-known entomopathogenic bacterium widely used in biopesticide formulations due to its diverse arsenal of insecticidal toxins and environmental adaptability. However, the genetic diversity and virulence potential of native Bt strains from South Asia remain underexplored. In this study, we performed whole-genome sequencing and comparative genomic analysis of B. thuringiensis strain JSd1, isolated from Bangladeshi soil. The high-quality draft genome, assembled at 75x coverage, comprises 5.39 Mb in 64 contigs with a GC content of 35.28% and 99.26% completeness. Genome annotation revealed 5,833 genes, including 5,756 protein-coding sequences and numerous non-coding RNAs. It also harbors 4 different plasmids. Importantly, we identified 25 genomic islands harboring mobile elements and hypothetical proteins, highlighting the strain’s dynamic genome evolution. The genome encodes a diverse array of virulence factors linked to insecticidal activity. Notable genes include cry22A and vip3A homologs, multiple bmp1 -like metalloproteases, enhancin , cytK , and chiA , as well as various chitinases and serine proteases. The co-occurrence of chromosomal and plasmid-encoded virulence factors suggests modular acquisition mechanisms. Secondary metabolite biosynthetic clusters, such as those for petrobactin, bacillibactin, thuricin CD, and other novel RiPPs and NRPs, were detected, supporting the strain’s potential in biological control. Phylogenetic analysis positioned JSd1 within the B. cereus sensu lato group, forming a highly supported clade with B. thuringiensis serovar konkukian and B. anthracis . Comparative genomic and pan-genome analyses revealed substantial genomic diversity among B. thuringiensis strains. The strain’s genome contains 2,237 core genes and a large accessory genome, reflecting its ecological adaptability. Our findings suggest that B. thuringiensis JSd1 is a promising candidate for development as a biopesticide targeting insect pests in Bangladeshi agriculture. The comprehensive genomic insights lay the groundwork for further functional validation and field applications, contributing to sustainable pest management strategies in the region.
Functional archetype analysis of single-cell RNA-sequencing (scRNA-seq) data is important because clustering of cell types is often nuanced and inexact. Intermediate phenotypes exist that are difficult to account for in these analyses, particularly in the setting of infant development. Further, different cell types work together to achieve biological functions, thereby broadly supporting tissue function at homeostasis and through environmental challenges via phenotypic plasticity. Currently, the process of assigning archetypes to cell clusters is labor-intensive because annotation requires manual upload of expression data to multiple programs. ArchetypeShift is an R-based pipeline that integrates existing archetypal analysis methods with annotation by Ingenuity Pathway Analysis (IPA) and Kyoto Encyclopedia of Genes and Genomes (KEGG), graphics visualization, and trajectory analysis for scRNA-seq data. Generating IPA- and/or KEGG-informed dot plots, UMAPs (uniform manifold approximation and projections) of archetype weights, archetype maps, heatmaps defining the top genes for each archetype program, and trajectory analysis graphics, ArchetypeShift streamlines the biological interpretation of archetype programs within a unified and efficient analytical framework. The source code for ArchetypeShift is available on GitHub ( https://github.com/Neo-NEC-Lab/ArchetypeShift ).
The rapid emergence of multidrug-resistant Klebsiella pneumoniae has significantly reduced the effectiveness of conventional antibiotics, highlighting the need for alternative therapeutic strategies. This study employed a comprehensive in silico pipeline to identify antimicrobial peptides (AMPs) targeting the essential DNA replication initiator protein DnaA. A total of 28,361 peptide sequences were collected from publicly available AMP databases and sequentially filtered based on peptide length, net charge, GRAVY score, instability index, antimicrobial activity, toxicity, hemolytic potential, aggregation propensity, sequence similarity and favorable amphipathic properties. Four peptides satisfied all selection criteria and were subjected to structural prediction, membrane-binding analysis, protein–DNA docking, protein–peptide docking, and Normal Mode Analysis. Protein–DNA docking identified the functional DNA-binding residues of DnaA, while peptide docking demonstrated that all four peptides interacted within this region. Peptide 3 exhibited the strongest predicted interaction, with a binding energy of −61.8±5.1, a buried surface area of 1151.7±30.8 Å 2 , and seven hydrogen bonds with key DnaA residues. Normal Mode Analysis further supported the structural stability of the peptide–protein complexes. These findings identify four promising AMP candidates targeting DnaA and provide a computational framework for peptide prioritization against multidrug-resistant K. pneumoniae . However, the proposed interactions remain computational predictions and require experimental validation.
Intestinal transcriptomic data encompass substantial value in studying inflammatory bowel disease. By analyzing intestinal gene expression profiles, this study aimed to discover transcripts exhibiting diagnostic efficacy for IBD. Intestinal transcriptomic data from 1,458 patients with Crohn’s disease (CD), 1,211 patients with ulcerative colitis (UC), and 674 healthy controls were acquired by integrating seven datasets with accession numbers GSE66407, GSE193677, GSE126124, GSE83687, GSE75214, GSE36807, and GSE16879. Then, the DEGs in IBD were analyzed by ML methods to discern transcripts that hold diagnostic power. ROC analysis identified potential biomarkers, which were subsequently tested in 16 external datasets. C2 , NLRC5 , S100P , PGAP3 , and GPR15 emerged as prominent genes based on both RF and LASSO methods. Meanwhile, only NLRC5 and C2 had an AUC of the ROC curve greater than the 0.7 threshold in the integrated data. ROC analysis across 23 cohorts demonstrated that the AUC of C2 and NLRC5 were above 0.7 in the majority of the cohorts, particularly that for C2 . However, upregulated C2 also exhibited diagnostic efficacy for autoimmune gastritis (AIG), eosinophilic esophagitis (EoE), and colorectal cancer (CRC), reflecting the substantial activation of the complement cascade in these disorders manifesting with gastrointestinal inflammation.
The increasing global burden of Type 2 diabetes mellitus (T2DM) and the side effects of conventional drugs have driven the search for natural therapeutic alternatives. This study investigates the antidiabetic potential of South African essential oils — specifically Artemisia afra, Agathosma betulina, and Cymbopogon citratus — which are traditionally used in diabetes treatment but lack comprehensive scientific validation. Using an integrated research framework: mining the phytochemical metabolites of the plant source from Dr. Duke’s Phytochemical and Ethnobotanical Database and relevant literature; screening of metabolites using molecular docking and molecular dynamics (MD) to identify potential leads; and in vitro enzymatic assays of the oils for α-amylase (AA) and α-glucosidase (AG) inhibition. Computational results identified quercetin-3,7-diglucoside (Q37DG), quercetin-7-O-glucoside (Q7G), and rutin as promising lead compounds. Notably, Q37DG demonstrated favorable inhibitory potential, with ΔG bind values of -52.73 kcal/mol for AA and -36.60 kcal/mol for AG, exceeding the known inhibitor acarbose (-43.59 and -34.41 kcal/mol, respectively). MD simulations further supported Q37DG’s role in enhancing enzyme stability upon binding, indicating a competitive inhibition mechanism through strong interactions with catalytic residues such as ASP469 (AG) and ASP197 (AA). Genetic Algorithm-based Quantitative Structure-Activity Relationship (GA-QSAR) models predict impressive IC 50 values for Q37DG (3.47 μM for AG and 0.68 μM for AA) compared to acarbose (6.92 and 2.51 μM). Experimental in vitro assays confirmed A. betulina essential oil as the oil with the most activity (IC 50 = 6.02 μg/ml for AA and 6.70 μg/ml for AG). Surprisingly, the lead compounds identified via computational screening are reported constituents of A. betulina , thereby validating the integrated data mining and in silico approach. This work promotes green chemistry by valuing indigenous plant resources and suggests that these leads warrant further investigation for nano-formulation and therapeutic development.
Carbonic anhydrases are metalloenzymes found both in vertebrates and invertebrates. The CAs catalyze the reversible hydration of CO 2 to bicarbonate and H+ ions and play a significant role in respiration, transport of CO 2 , pH homeostasis, electrolyte secretion, and biosynthetic reactions. The enzymatic activity of CAs is due to the coordination of Zn 2+ in the active site by three histidine residues; however, in humans, there are three CAs known as CA-related proteins (CARPs) that are catalytically inactive due to the absence of one or more of the three histidine residues required for the coordination of the Zn 2+ in the active site. Studies have shown that CARPs are expressed in all parts of the brain and are overexpressed in some cancers suggesting that the CARPs play a crucial role in neurological disorders and the development of cancer. However, the precise physiological roles of CARPs are still an enigma. In this study, we present a comprehensive biological workflow that employs various machine learning methodologies and statistical procedures to assess the similarity across CARPs by evaluating shared biological parameters. This approach enabled us to identify potential biomarkers, including transcription factors, co-expressed genes, and phenotypes, that may influence the expression of CARPs within disease pathways. Furthermore, we proposed a computational human health model by analyzing drugs and chemical candidates to prioritize compounds that may modulate regulatory networks associated with CARPs. These computational analyses identified candidate compounds for future experimental investigation in the context of neurological disorders and cancer.
Hepatocellular carcinoma (HCC) is a common malignant tumor with a poor prognosis for advanced metastasis while there is currently a lack of prognostic markers and effective therapeutic targets in this area. The PSD gene family, comprising four genes, has been implicated in tumor metastasis in breast cancer, but its role in HCC remains unclear. To elucidate the role-played by the PSD gene family in the prognosis of HCC and explore the relationship between the PSD gene family and tumor metastasis, we conducted a bioinformatics analysis using TCGA and gene expression omnibus (GEO) databases to investigate PSD gene family expression in HCC and its correlation with patient survival. Furthermore, we performed quantitative real-time polymerase chain reaction (qPCR) validation of the expression levels of the PSD gene family in clinical samples. Gene correlation and enrichment analyses were also performed to explore its relationship with tumor metastasis. In conclusion, the expression levels of the PSD gene family are significantly downregulated in HCC ( P < 0.01), and the lower expressions are associated with poor survival. In addition, there is a link between the PSD gene family and RAS pathway activation, suggesting its potential role in metastasis.
Terpenes represent a diverse class of plant secondary metabolites that play important roles in development, ecological interactions, and aroma, and are synthesized by the terpene synthase (TPS) gene family. To investigate the evolutionary history and regulatory diversification of this family in coffee, we performed a genome-wide analysis of TPS genes in Coffea arabica and its diploid progenitors, C. canephora and C. eugenioides . Using domain-based searches targeting the PF01397 module, we identified 69 TPS genes in C. arabica , 40 in C. canephora , and 55 in C. eugenioides . Phylogenetic classification, duplication analysis, and synteny mapping revealed that the tandem duplication is the major driver of TPS expansion, with C. arabica retaining progenitor-derived copies alongside limited lineage-specific duplications. Selection analyses indicated that most TPS duplicates from C. arabica are under strong purifying constraint, while only a single orthogroup (OG0000013) showed limited evidence of episodic positive selection with a minor fraction of inferred sites and a lack of strong BEB support. Indicating that adaptive divergence was highly localized against a predominantly conserved evolutionary background. To link evolutionary divergence with transcriptional regulation, RNA-seq data from multiple tissues and developmental stages were analyzed. A substantial proportion of TPS genes were transcriptionally active and exhibited tissue- and stage-specific regulation. While many duplicates showed conserved or asymmetric expression, the OG0000013 duplicates combined localized adaptive sequence divergence with significant upregulation during late seed development. In conclusion, the retention and asymmetric transcription of progenitor-derived homeologs indicate that regulatory partitioning, rather than extensive gene-family innovation, has the dominant evolutionary outcome following allopolyploidization in C. arabica .
Antimicrobial resistance poses an increasing global challenge, driving the urgent need for alternative strategies to identify novel therapeutic agents. Microbial natural products encoded by biosynthetic gene clusters (BGCs) remain among the most promising sources of bioactive compounds. Although Corynebacterium glutamicum is best known as an industrial producer of amino acids, its potential as a producer of secondary metabolites has not been comprehensively assessed, despite the availability of numerous high-quality genome sequences. In this study, we carried out a comparative pangenome analysis of 36 complete C. glutamicum genomes and systematically mined for BGCs to explore the species’ biosynthetic repertoire. Our analysis revealed variation in BGC content among strains, with several isolates harboring more hybrid clusters than others, suggesting metabolic diversity across the species. In addition to conserved terpene biosynthetic pathways, we detected polyketide-associated clusters not previously reported in C. glutamicum , expanding its recognized metabolic potential. RiPP-like clusters, including Lactococcin-related variants, were also identified, highlighting an underexplored reservoir of antimicrobials. To prioritize candidates for future validation, Support Vector Machine, Random Forest, and k-Nearest Neighbor models were trained on Composition, Transition, and Distribution (CTD) physicochemical sequence features and applied to genome-mined small open reading frames. The models demonstrated strong predictive performance, with the Support Vector Machine achieving the highest accuracy (84.1%), F1 score (83.8%), and area under the ROC curve (AUC = 0.920). After removing duplicate sequence IDs and applying a high-confidence AMP probability threshold (≥0.95), 18 unique AMP-like candidates were identified as promising. Overall, this study presents C. glutamicum as a promising source of bioactive metabolite candidates and shows how pangenome-scale mining combined with machine learning can support antimicrobial peptide discovery while still requiring experimental validation of the predicted leads.
Epidemic louse-borne typhus caused by Rickettsia prowazekii and transmitted by the human body louse, remains a major public health threat in many developing regions. Historical records indicate that outbreaks have resulted in up to million cases annually, with mortality ranging from thousands to millions. Early and accurate diagnosis is critical for improving the efficacy of antibiotic therapies. However, existing serological diagnostic methods often suffer from limited sensitivity, specificity, and reliability. In this study, we applied integrated in silico immunoinformatics and structural bioinformatics approaches to identify antigenic targets and to design and characterize engineered proteins for diagnostic applications. Genomic and proteomic analyses identified Sca4 and OmpB as highly antigenic proteins of Rickettsia prowazekii. Computational epitope prediction mapped linear B-cell epitopes within multi-epitope regions of Sca4 and immunodominant regions of OmpB. Two engineered proteins (MA1 and MA2) were computationally designed: MA1 using selected linear B-cell epitopes linked via flexible peptide linkers with a fusion tag, and MA2 using truncated immunodominant domains with a fusion tag. The engineered proteins were computationally evaluated for physicochemical properties, antigenicity, solubility, stability, and sequence homology, and they demonstrated favorable characteristics with minimal similarity to human proteins. Structural modeling, followed by molecular docking and molecular dynamics simulations with the TLR4 receptor, suggested preserved structural integrity and stable protein-receptor interactions. Overall, these in silico findings underscore the potential of MA1 and MA2 constructs as potential diagnostic candidates and provide a basis for further experimental validation, including protein expression and purification, antibody production in animal models, and immunodiagnostic assay development.
Antimicrobial resistance poses an increasing global challenge, driving the urgent need for alternative strategies to identify novel therapeutic agents. Microbial natural products encoded by biosynthetic gene clusters (BGCs) remain among the most promising sources of bioactive compounds. Although Corynebacterium glutamicum is best known as an industrial producer of amino acids, its potential as a producer of secondary metabolites has not been comprehensively assessed, despite the availability of numerous high-quality genome sequences. In this study, we carried out a comparative pangenome analysis of 36 complete C. glutamicum genomes and systematically mined for BGCs to explore the species’ biosynthetic repertoire. Our analysis revealed variation in BGC content among strains, with several isolates harboring more hybrid clusters than others, suggesting metabolic diversity across the species. In addition to conserved terpene biosynthetic pathways, we detected polyketide-associated clusters not previously reported in C. glutamicum , expanding its recognized metabolic potential. RiPP-like clusters, including Lactococcin-related variants, were also identified, highlighting an underexplored reservoir of antimicrobials. To prioritize candidates for future validation, Support Vector Machine, Random Forest, and k-Nearest Neighbor models were trained on Composition, Transition, and Distribution (CTD) physicochemical sequence features and applied to genome-mined small open reading frames. The models demonstrated strong predictive performance, with the Support Vector Machine achieving the highest accuracy (84.1%), F1 score (83.8%), and area under the ROC curve (AUC = 0.920). After removing duplicate sequence IDs and applying a high-confidence AMP probability threshold (≥0.95), 18 unique AMP-like candidates were identified as promising. Overall, this study presents C. glutamicum as a promising source of bioactive metabolite candidates and shows how pangenome-scale mining combined with machine learning can support antimicrobial peptide discovery while still requiring experimental validation of the predicted leads.
Epidemic louse-borne typhus caused by Rickettsia prowazekii and transmitted by the human body louse, remains a major public health threat in many developing regions. Historical records indicate that outbreaks have resulted in up to million cases annually, with mortality ranging from thousands to millions. Early and accurate diagnosis is critical for improving the efficacy of antibiotic therapies. However, existing serological diagnostic methods often suffer from limited sensitivity, specificity, and reliability. In this study, we applied integrated in silico i mmunoinformatics and structural bioinformatics approaches to identify antigenic targets and to design and characterize engineered proteins for diagnostic applications. Genomic and proteomic analyses identified Sca4 and OmpB as highly antigenic proteins of Rickettsia prowazekii . Computational epitope prediction mapped linear B-cell epitopes within multi-epitope regions of Sca4 and immunodominant regions of OmpB. Two engineered proteins (MA1 and MA2) were computationally designed: MA1 using selected linear B-cell epitopes linked via flexible peptide linkers with a fusion tag, and MA2 using truncated immunodominant domains with a fusion tag. The engineered proteins were computationally evaluated for physicochemical properties, antigenicity, solubility, stability, and sequence homology, and they demonstrated favorable characteristics with minimal similarity to human proteins. Structural modeling, followed by molecular docking and molecular dynamics simulations with the TLR4 receptor, suggested preserved structural integrity and stable protein–receptor interactions. Overall, these in silico findings underscore the potential of MA1 and MA2 constructs as potential diagnostic candidates and provide a basis for further experimental validation, including protein expression and purification, antibody production in animal models, and immunodiagnostic assay development.
The incidence of antibiotic resistance is critically jeopardizing the public health of the new millennium, affecting both clinical and therapeutic outcomes. Typhoid perforation in low- and middle-income countries (LMICs) is mostly caused by Salmonella typhi. The development of antibacterial therapies is drawn to the enzyme 3-dehydroquinate dehydratase type 1 (DHQ1), which is an important enzyme of the shikimate pathway in bacteria. Its role in the synthesis of chorismite, a natural precursor of the route leading to the synthesis of aromatic amino acids, validates it as an important therapeutic target. Therefore, the current work is based on the investigation of bioactive compounds that inhibit the crucial DHQ1 enzyme. Structure-based virtual screening (SBVS) and docking studies were done using AutoDock Vina tools to retrieve potential hit molecules with high binding affinity against DHQ1 from several ligand repositories. From a total of 551 compounds, 206 were filtered out for physicochemical properties, the Lipinski rule and subjected to molecular docking. A total of 10 ligands were selected with the highest binding affinities ranging from -9.8 to -9.1 (kcal/mol), ie, Cabotegravir, Imatinib, Prulifloxacin, Limonin, Silibinin, Atovaquone, Betamethasone Valerate, GSK1324726A, Isavuconazole, and Raltegravir. In addition, ADMET analysis, bioactivity prediction, and pharmacokinetic property prediction were carried out as prediction of activity spectra for substance (PASS) analysis to further prioritize the best inhibitory compound. Molecular dynamic simulations were used to verify the stability of the protein "DHQ1" docked with 3 hit ligands Cabotegravir, Limonin, and Silibinin by calculating root mean square fluctuation (RMSF), root mean square deviation (RMSD), radius of gyration (Rg), H-bond interaction, and molecular mechanics/generalized Born surface area (MM/GBSA) scores. Based on molecular dynamic (MD) simulations, it is reported that Cabotegravir and Limonin showed the optimal binding features with bacterial DHQ1 enzyme and can be deemed as a repurposed drug candidate as compared to Silibinin, which did not show favorable interactions with DHQ1 and needs to be explored further against typhoidal Salmonella.
Prompt and accurate identification of Listeria monocytogenes at the whole-genome level is essential for food safety surveillance and outbreak prevention, yet existing culture-based, PCR, and next-generation sequencing (NGS) workflows are either slow, labor-intensive, or rely on opaque machine learning models with limited interpretability. This study proposes an explainable genomic classification framework that couples transformer-based DNA embeddings with gradient boosting to distinguish Listeria from non-Listeria genomes. A curated dataset of 700 complete bacterial genomes (350 L. monocytogenes and 350 biologically related non-Listeria genomes) was assembled from the NCBI Assembly database and rigorously filtered to retain high-quality assemblies. Each genome was tokenized into overlapping 6-mers and encoded using the pretrained DNABERT model to obtain contextual genome-level embeddings, which were then classified with a LightGBM classifier. Model decisions were interpreted using SHapley Additive exPlanations (SHAP) to quantify the contribution of individual embedding dimensions and associated k-mer patterns to each prediction. After final verification of the confusion matrix, the proposed DNABERT + LightGBM + SHAP pipeline correctly classified 335/350 Listeria genomes and 330/350 non-Listeria genomes, corresponding to 665/700 correct classifications and a corrected accuracy of 95.00%. The same evaluation yielded precision of 94.37%, recall of 95.71%, F1-score of 95.03%, and an AUC of 0.9976. Comparative experiments show that the framework consistently outperforms conventional k-mer-based Random Forest, TF-IDF + SVM, CNN, XGBoost, and DNABERT + Logistic Regression baselines across all metrics. Beyond its strong predictive performance, SHAP-based analysis reveals discriminative sequence patterns that may correspond to putative genomic signatures of Listeria. Overall, the proposed approach provides a high-performance, interpretable tool for genome-scale Listeria identification and offers a transferable template for explainable pathogen genomics in food safety and public health applications.
Overuse of antimicrobial drugs is known to cause an increase in bacterial resistance among surviving pathogens, reducing the effectiveness of future treatments. Publicly available sequencing collections, such as the National Center for Biotechnology Information Sequence Read Archive (SRA), allow for global investigation of antimicrobial resistance across pathogens. In this study, we developed a pipeline to process 10 803 globally sourced unique bacterial isolates from publicly available SRA datasets (9 pathogens, 4 antibiotics), representing sequencing data generated by multiple laboratories using diverse sequencing platforms and protocols. The pipeline extracted SRA metadata to determine read layout and length, applied quality control, trimming, and decontamination. Preprocessed isolates were mapped reads to an antimicrobial resistance gene class and to a strain-level genome reference library constructed from complete genomes on the SRA submitted between 2009 and 2020. Three classifiers—L1-penalized logistic regression, random forest, and extreme gradient boosting—were trained on the resulting feature matrices, and their outputs were combined in a majority-vote ensemble. Internal training resulted in 83.8% balanced accuracy on average, and external testing yielded 80.2%. Variable importance analyses identified known resistance gene classes and strain markers such as Acinetobacter baumannii LAC-4 and Campylobacter jejuni 81-176 and NCTC13255, confirming biological relevance. This work demonstrates a scalable approach for antimicrobial resistance prediction using heterogeneous sequencing data.