
Background:CXCL12 is a critical chemokine involved in immune cell trafficking and tumor metastasis through its interaction with CXCR4 and CXCR7. The C55Y mutation in CXCL12 is hypothesized to disrupt its structure and function. Unlike previously reported CXCL12 mutations that affect transcriptional regulation or single-receptor binding, C55Y targets a conserved cysteine residue essential for a structural disulfide bridge, potentially inducing unique global destabilization. This study investigates the molecular consequences of the C55Y mutation using computational tools. Objectives:This study aimed to evaluate the structural and functional impact of the C55Y mutation in CXCL12, particularly its interactions with CXCR4 and CXCR7, and its relevance to prostate cancer pathogenesis. Methods:Molecular dynamics simulations (MDS), evolutionary conservation analysis, and molecular docking with CXCR4 and CXCR7 receptors were performed to assess the impact of the C55Y mutation. Tools such as I-TASSER, Meta-SNP, and PolyPhen-2 were used to analyze structural stability, while HDOCK and Schrödinger simulations assessed receptor binding. RMSD and RMSF calculations were used to evaluate protein dynamics. Results:The C55Y mutation induces significant structural instability in CXCL12, with increased RMSD and altered secondary structure. Molecular dynamics simulation further demonstrated the stability of the complexes through RMSD, RMSF, hydrogen bond, and Radius of Gyration (RoG) analyses, where the normal complexes exhibited comparatively stable structural compactness during the simulation period. Docking studies showed reduced binding affinity to CXCR4 and CXCR7, indicating disrupted receptor interactions. Notably, this dual-receptor binding loss distinguishes C55Y from other mutations that typically impair only CXCR4 signaling. These findings suggest that the mutation impairs CXCL12's function in prostate cancer, metastasis of prostate cancer, and immune response. Conclusion:The C55Y mutation in CXCL12 disrupts its structural integrity and receptor binding, uniquely compromising both the CXCR4 and CXCR7 pathways. This highlights its potential role in prostate cancer progression through mechanisms distinct from previously characterized CXCL12 variants. This study provides insights into the molecular mechanisms of CXCL12-mediated signaling and its implications for future therapeutic strategies targeting chemokine signaling to inhibit prostate cancer.
Background: Prime mass amino acids residues, assigned based upon the nominal mass of their repeating structure in peptides and proteins, have been found to occur more often than by chance. These 9 such residues are predominantly hydrophobic in character and thus play an important role in protein stability and folding. This study investigates their prevalence in human proteins and the function of such proteins. Objectives: To identify human proteins rich in the prime mass residues and assess their functions. Results: Prime-rich proteins are found to proliferate among those important to transcriptional events, cell adhesion, and the formation of cell surface membranes. Eighteen prime-rich proteins are identified that are abundant in 3 or 4 prime residues, the latter notably contain alanine, cysteine, histidine or threonine, and proline. The percentage of prime residues in these prime-rich proteins exceeds one-half (>50%) of all residues in many cases, and these residues are found to be concentrated in particular regions of the protein. Conclusion: Proteins rich in prime residues are found to play an important role in biological events important to cellular life. Such prime residues represent those associated with the highest hydrophobicity, and those essential to protein folding and function. Many prime residues were recruited early into the genetic code according to the composition of the last universal common ancestor (LUCA) proteins. Consideration is given to whether the formation of evolution of these proteins is predicated by the primality of the residues or their physicochemical and structural properties. This remains an open question and awaits further inquiry and debate.
Background: Severe dengue (SD) represents a life-threatening progression of dengue virus infection. Early identification of patients at risk of transitioning from dengue fever (DF) to SD remains a major clinical challenge. Unraveling the transcriptomic changes underlying this progression may aid in developing timely therapeutic interventions. Methods: RNA-seq datasets comprising 103 samples (62 SD and 41 DF) were retrieved from the GEO repository. Following normalization using DESeq2, differentially expressed genes (DEGs) were identified between the 2 disease stages. Functional enrichment analysis was performed to uncover dysregulated biological processes. A hybrid computational framework combining classical machine learning (Logistic Regression, Support Vector Machine, and Random Forest) and deep learning models (Artificial Neural Network, Convolutional Neural Network, and Transformer-based architectures) were applied to classify SD and DF samples. Model performance was evaluated using ROC-AUC and balanced accuracy metrics. Results: Differential expression analysis identified 55 significantly dysregulated genes distinguishing severe dengue from dengue fever. These genes were enriched in pathways related to metal ion homeostasis, platelet signaling, ferroptosis, and oxidative stress. Among multiple machine learning and deep learning models, the Transformer-CNN achieved the best performance (test AUC = 0.85; balanced accuracy = 0.89). SHAP-based interpretation highlighted ILDR2, TCP1, HNRNPUL1, SEC14L5, ATP2C2, LOXL3, ACVRL1, STEAP3, and ST8SSIA5 as key discriminative features. Integrated network analyses further implicated coordinated regulation of iron metabolism, calcium signaling, and platelet dysfunction in severe dengue. Conclusion: This study integrates RNA-seq and hybrid Artificial Intelligence modeling to identify transcriptomic signatures associated with dengue severity. The study highlights candidate genes and pathways that provide a hypothesis-generating foundation; further increasing the sample size and experimental validation will support early risk stratification in severe dengue.
Objective: Mycobacterium leprae causes leprosy, an infectious disease that has persisted over centuries and is still an issue in public health in many nations. Diverse bioinformatics techniques have been effectively employed to annotate the functions of hypothetical proteins (HPs) originating from different pathogenic bacteria. The objective of the current research was to elucidate the functions of an HP obtained from M. leprae.Methods: A variety of in silico tools were utilized to make predictions regarding the structure and function of this protein. To identify homologous proteins, the BLASTp program was used to search for sequence similarity across the available biological databases. Additionally, using the proper bioinformatics methods, a number of properties were determined, including physicochemical characteristics, subcellular localization, phylogenetic analysis, functional annotation, pathway analysis, protein-protein interaction, secondary and tertiary structure determination, active site detection, quality assessment analysis, molecular docking, pharmacokinetic and toxicity profiling, and further molecular dynamics simulations.Results: The HP exhibited putative biological activity associated with a conserved functional domain, the CT_C_D superfamily domain. The allophanate hydrolase activity of the chosen HP was predicted by the functional annotation. Pathway analysis demonstrated the protein's involvement in cellular and metabolic processes. Numerous functional partners that play a crucial role in bacterial survival were identified through the chosen HP's protein-protein interactions. Furthermore, active site prediction and molecular docking analysis of the HP with ligands indicated that it could be a therapeutic target for M. leprae. ADMET analysis indicated that the selected compound has favorable bioavailability, drug-likeness, and safety. The stability of these complexes was verified by molecular dynamics simulations, which suggests they have therapeutic potential.Conclusion: This study emphasizes the effectiveness of in silico methods in predicting the biological functions of HP and generating hypotheses for potential therapeutic targets.
Objective:Species of glossiphoniid leech provide model systems for exploring fundamental questions in evolutionary and developmental biology. In this study, we took advantage of large embryonic cells and stereotypical cleavage patterns in the leech, Helobdella austinensis, to identify genes associated with the births of teloblasts and their lineage specifications. Methods:Using staged embryos and pools of dissected precursor cells and teloblasts, a systematic computational comparison of staged and cell type-specific transcriptomic data was employed to identify unique gene sets associated with mesodermal (M), neuroectodermal (N) and teloblast (M + N) cell formation. Results and Conclusions:Our predicted candidate genes comprised sets of nucleic acid binding factors as anticipated but displayed little similarity with differentiation factors in other metazoans (eg, mammals, cnidarians), suggesting that stem cell/lineage induction processes have diverged across the Animalia and may reflect fundamentally different modes of embryonic development.
Background:Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has triggered a global health crisis, emphasizing the urgent need for accurate and rapid diagnostic tools. Modern molecular biology technologies, including CRISPR-Cas systems, provide highly efficient strategies for viral detection. Bioinformatic pipelines are essential for identifying conserved genomic regions and enabling rational single-guide RNA (sgRNA) design. Methods:This study aimed to design specific sgRNAs targeting the spike gene of SARS-CoV-2 isolates from Iranian patients using the SHERLOCK diagnostic platform. Complete genomes of the RefSeq virus and 470 SARS-CoV-2 isolates, representing all variants of concern (VOCs) detected in Iran, were retrieved from the NCBI and GISAID databases. Multiple sequence alignment with ClustalW identified conserved sequences within the receptor-binding domain (RBD) that differ from the RBD of SARS-CoV and MERS-CoV RefSeq genomes. Based on these regions, sgRNAs and isothermal amplification primers were designed using ADAPT, OLIGO7, and the UCSC Genome Browser to maximize diagnostic sensitivity and specificity. Secondary and tertiary structures of sgRNA-target complexes were analyzed via RNAfold and RNAup to select the most efficient sgRNA-amplicon combination. Results:Twenty-two-nucleotide sgRNA candidates were initially selected based on sequence alignment, showing high similarity to the SARS-CoV-2 RefSeq and low homology to SARS-CoV and MERS-CoV genomes. Analyses of secondary structures, RNA-RNA interactions, and free energy identified 6 sgRNAs with favorable 2-dimensional conformations and strong interaction profiles. Among these, the sgRNA1-Amplicon2 sequence exhibited the most stable 3-dimensional structure and a molecular docking score of -309.67, indicating high sensitivity and specificity for viral detection. Conclusion:This study successfully designed an sgRNA with high sensitivity and specificity for rapid SARS-CoV-2 detection using the CRISPR-Cas13a system, informed by genomic analysis of Iranian isolates. The proposed approach provides an efficient framework for the rapid design and deployment of CRISPR-based diagnostic tools applicable to diverse viral pathogens.
The increasing integration of bioinformatics with cloud infrastructure, artificial intelligence, and real-time data analytics has introduced unprecedented cybersecurity challenges. Sensitive genomic and clinical data, when exposed to evolving cyber threats such as data exfiltration, injection attacks, and model poisoning, require more adaptive and resilient defense mechanisms. This research presents an Evolution-Inspired Cyber Defense Architecture (EICDA) that mimics biological immune and evolutionary processes to detect, respond to, and adapt against cyber intrusions targeting bioinformatics-informed systems. The proposed architecture employs a genetic algorithm-based detection core combined with reinforcement-driven adaptation, enabling real-time learning and decision reconfiguration. EICDA is validated on multiple datasets including NSL-KDD, CICIDS2017, and a simulated genomic API log environment, achieving a detection accuracy of 96.2%, a false positive rate of 2.8%, and sub-150 ms response latency. Comparative analyses with SVM, Random Forest, CNN, and traditional AIS highlight EICDA's superior adaptability and robustness. The framework also demonstrates resilience under threat drift conditions, positioning it as a viable defense model for next-generation bioinformatics platforms. This research provides a novel contribution by fusing evolutionary intelligence with cybersecurity to protect critical biomedical infrastructures.
Background: During leptospiral infection, the host innate immune response is initiated through recognition of pathogen-associated molecular patterns (PAMPs) by Toll-like receptor 2 (TLR2). Among these PAMPs, LipL32, Loa22, Lsa21, and lipopolysaccharide biosynthesis genes, are of particular interest. Objective: This study aimed to investigate these molecules’ genetic variability and evolutionary conservation in recently isolated clinical Leptospira strains from Sri Lanka. Results: We analyzed the whole-genome sequences of 25 clinical Leptospira isolates obtained from patients across Sri Lanka, sequenced using long-read technology and annotated using a standardized pipeline. Genes encoding LipL32, Loa22, Lsa21, and enzymes within the lipopolysaccharide biosynthesis locus were extracted and analyzed for phylogenetic relationships and sequence variation. LipL32 and Loa22 were highly conserved across all isolates, with only a single amino acid substitution observed in each. In contrast, genes associated with lipopolysaccharide biosynthesis, specifically those encoding glycosyl transferase and a sodium-dependent anion transporter, exhibited notable genetic variation, including multiple single nucleotide polymorphisms leading to amino acid changes. The Lsa21 gene was present only in Leptospira interrogans strains and showed no protein-level variation. Leptospira borgpetersenii isolates demonstrated strong conservation across all gene targets at both nucleotide and protein levels. Conclusion: Our findings highlight the high conservation of LipL32 and Loa22, reinforcing their potential as stable targets for molecular diagnostics and serological assays. In contrast, the variability observed in lipopolysaccharide biosynthesis genes suggests a possible role in immune evasion or adaptation, warranting further functional investigation. The restricted presence of Lsa21 in specific species also raises questions about its contribution to pathogenicity.
Background:Cronobacter sakazakii, a foodborne pathogen with a fatality rate of 33%, is a rod-shaped, Gram-negative, non-spore-forming bacterium responsible for causing meningitis, bacteremia, and necrotizing enterocolitis. Despite many unknown functions of hypothetical proteins in bacterial genomes, bioinformatic techniques have successfully annotated their roles in various pathogens. Objectives:The aim of this investigation is to identify and annotate the structural and functional properties of a hypothetical protein (HP) from Cronobacter sakazakii 7G strain (accession no. WP_004386962.1, 277 residues) using computational tools. Methods:Multiple bioinformatic tools were used to identify the homologous protein and to construct and validate its 3D structure. A 3D model was generated using SWISS-MODEL and validated using tools, developing a reliable 3D structure. The STRING and CASTp servers provided information on protein-protein interactions and active sites, identifying functional partners. Results:The putative protein was soluble, stable, and localized in the cytoplasmic membranes, indicating its biological activity. Functional annotation identified TagJ (HsiE1) within the protein, a member of the ImpE superfamily involved in the transport of toxins and a part of the bacterial type VI secretion system (T6SS). The 3-dimensional structure of this protein was validated through molecular docking involving 6 different compounds. Among these, ceforanide demonstrated the strongest binding scores, -7.5 kcal/mol for the hypothetical protein and -7.2 kcal/mol for its main template protein (PDB ID: 4UQX.1). Conclusion:Comparative genomics study suggests that the protein found in C. sakazakii may be a viable therapeutic target because it seems distinctive and different from human proteins. The results of multiple sequence alignment (MSA) and molecular docking supported HP's potential involvement as a T6SS. These in silico results represent that the examined HP could be valuable for studying C. sakazakii infections and creating medicines to treat C. sakazakii-mediated disorders.
Background:The WRKY gene family is identified as one of the most prominent transcription factor families in plants and is involved in various biological processes such as metabolism, growth and development, and response to biotic and abiotic stresses. In many plant species, the WRKY gene family was widely studied and analyzed but little to no information for Fortunella hindsii. However, the completion of the whole genome sequencing of Fortunella hindsii allowed us to investigate the genome-wide analysis of WRKY proteins. Objective:The main objective of this study was to analyze and identify the WRKY gene family in Fortunella hindsii genome. Methodology:Various bioinformatics approaches have been used to conduct this study. Results:We constituted 46 members of the Fortunella hindsii WRKY gene family, which were unevenly distributed on all nine chromosomes. The phylogenetic relationship of predicted WRKY proteins of Fortunella hindsii with the WRKY proteins of Arabidopsis showed that 46 FhWRKY genes were divided into three main groups (G1, G2, G3) with five subgroups (2A, 2B, 2C, 2D, and 2E) of G2 group. Domain, conserved motif identification, and gene structure were conducted and the results found that these FhWRKY proteins have conserved identical characteristics within groups and maintain differences between groups. In silico subcellular localization, results showed that FhWRKY genes are located in the nucleus. The cis-regulatory element analysis identified several key CREs that are significantly associated with light, hormone responses, and stress. The gene ontology analysis of these predicted FhWRKY genes showed that these genes are significantly enriched in sequence-specific DNA binding, transcriptional activity, cellular biosynthesis, and metabolic processes. Conclusion:Therefore, overall, our results provided an excellent foundation for further functional characterization of WRKY genes with an aim of Fortunella hindsii citrus crop improvement.
Background: The WRKY gene family is identified as one of the most prominent transcription factor families in plants and is involved in various biological processes such as metabolism, growth and development, and response to biotic and abiotic stresses. In many plant species, the WRKY gene family was widely studied and analyzed but little to no information for Fortunella hindsii . However, the completion of the whole genome sequencing of Fortunella hindsii allowed us to investigate the genome-wide analysis of WRKY proteins. Objective: The main objective of this study was to analyze and identify the WRKY gene family in Fortunella hindsii genome. Methodology: Various bioinformatics approaches have been used to conduct this study. Results: We constituted 46 members of the Fortunella hindsii WRKY gene family, which were unevenly distributed on all nine chromosomes. The phylogenetic relationship of predicted WRKY proteins of Fortunella hindsii with the WRKY proteins of Arabidopsis showed that 46 FhWRKY genes were divided into three main groups (G1, G2, G3) with five subgroups (2A, 2B, 2C, 2D, and 2E) of G2 group. Domain, conserved motif identification, and gene structure were conducted and the results found that these FhWRKY proteins have conserved identical characteristics within groups and maintain differences between groups. In silico subcellular localization, results showed that FhWRKY genes are located in the nucleus. The cis -regulatory element analysis identified several key CREs that are significantly associated with light, hormone responses, and stress. The gene ontology analysis of these predicted FhWRKY genes showed that these genes are significantly enriched in sequence-specific DNA binding, transcriptional activity, cellular biosynthesis, and metabolic processes. Conclusion: Therefore, overall, our results provided an excellent foundation for further functional characterization of WRKY genes with an aim of Fortunella hindsii citrus crop improvement.
Pantoea sp. strain MHSD4 is a bacterial endophyte isolated from the leaves of the medicinal plant Pellaea calomelanos. Here, we report on strain MHSD4 draft whole genome sequence and annotation. The draft genome size of Pantoea sp. strain MHSD4 is 4 647 677 bp with a G+C content of 54.2% and 41 contigs. The National Center for Biotechnology Information Prokaryotic Genome Annotation Pipeline tool predicted a total of 4395 genes inclusive of 4235 protein-coding genes, 87 total RNA genes, 14 non-coding (nc) RNAs and 70 tRNAs, and 73 pseudogenes. Biosynthesis pathways for naphthalene and anthracene degradation were identified. Putative genes involved in bioremediation such as copA, copD, cueO, cueR, glnGm , and trxC were identified. Putative genes involved in copper homeostasis and tolerance were identified which may suggest that Pantoea sp. strain MHSD4 has biotechnological potential for bioremediation of heavy metals. Keywords , , whole genome sequencing , bioremediation , bacterial endophyte
Mycobacterium tuberculosis (Mtb) is the causative agent of tuberculosis (TB), an infectious disease that is a major killer worldwide. Due to selection pressure caused by the use of antibacterial drugs, Mtb is characterised by mutational events that have given rise to multi drug resistant (MDR) and extensively drug resistant (XDR) phenotypes. The rate at which mutations occur is an important factor in the study of molecular evolution, and it helps understand gene evolution. Within the same species, different protein-coding genes evolve at different rates. To estimate the rates of molecular evolution of protein-coding genes, a commonly used parameter is the ratio d N/ d S, where d N is the rate of non-synonymous substitutions and d S is the rate of synonymous substitutions. Here, we determined the estimated rates of molecular evolution of select biological processes and molecular functions across 264 strains of Mtb. We also investigated the molecular evolutionary rates of core genes of Mtb by computing the d N/ d S values, and estimated the pan genome of the 264 strains of Mtb. Our results show that the cellular amino acid metabolic process and the kinase activity function evolve at a significantly higher rate, while the carbohydrate metabolic process evolves at a significantly lower rate for M. tuberculosi s. These high rates of evolution correlate well with Mtb physiology and pathogenicity. We further propose that the core genome of M. tuberculosis likely experiences varying rates of molecular evolution which may drive an interplay between core genome and accessory genome during M. tuberculosis evolution. Keywords Comparative genomics , N/ , S , molecular evolution , ,
Mycobacterium orygis , a subspecies of the Mycobacterium tuberculosis complex (MTBC), has emerged as a significant concern in the context of One Health, with implications for zoonosis or zooanthroponosis or both. MTBC strains are characterized by the unique insertion element IS 6110 , which is widely used as a diagnostic marker. IS 6110 transposition drives genetic modifications in MTBC, imparting genome plasticity and profound biological consequences. While IS 6110 insertions are customarily found in the MTBC genomes, the evolutionary trajectory of strains seems to correlate with the number of IS 6110 copies, indicating enhanced adaptability with increasing copy numbers. Here, we present a comprehensive analysis of IS 6110 insertions in the M. orygis genome, utilizing ISMapper, and elucidate their genetic consequences in promoting successful host adaptation. Our study encompasses a panel of 67 paired-end reads, comprising 11 isolates from our laboratory and 56 sequences downloaded from public databases. Among these sequences, 91% exhibited high-copy, 4.5% low-copy, and 4.5% lacked IS 6110 insertions. We identified 255 insertion loci, including 141 intragenic and 114 intergenic insertions. Most of these loci were either unique or shared among a limited number of isolates, potentially influencing strain behavior. Furthermore, we conducted gene ontology and pathway analysis, using eggNOG-mapper 5.0, on the protein sequences disrupted by IS 6110 insertions, revealing 63 genes involved in diverse functions of Gene Ontology and 45 genes participating in various KEGG pathways. Our findings offer novel insights into IS 6110 insertions, their preferential insertion regions, and their impact on metabolic processes and pathways, providing valuable knowledge on the genetic changes underpinning IS 6110 transposition in M. orygis . Keywords , , IS , bovine tuberculosis , zoonosis , complex , One Health
Background: Molecular epidemiology has shown the presence of four genotypes circulating across Africa, a paucity of data exists regarding phylogeography of the African Yellow fever (YF) genotypes. The need to fill this gap with spatiotemporal data from continuous YF outbreaks in Africa conceptualized this study; which aims to investigate the most recent transmission events and directional spread of yellow fever virus (YFV) using updated genomic sequence data. Methods: Yellow fever sequence data was utilized along with epidemiologic data from outbreaks in Africa, to analyze the case/fatality distribution and genetic diversity. Phylodynamic and phylogeographic were utilized to investigate ancestral history, virus population dynamics, and geographic dispersal of yellow fever across Africa. Results: There was a sharp increase in laboratory confirmed cases after year 2015, with Nigeria and the Democratic Republic of Congo having the highest numbers of cases. Phylogeny of the YF genotypes followed a previously reported pattern with distinct geographic clustering. Historical dispersal of YFV was discovered to have occurred from West into Central/East Africa, with recent introductions occurring in West Africa. Conclusions: We have shown the continuous circulation of YF in Africa, with distinct genotype distributions within the west and central African sub-regions. We have also shown the potential contribution of African genotypes, in the historical dispersal of yellow fever. We advocate for expanded and integrated molecular surveillance of YFV and other Arboviruses in Africa.
Objective: Neisseria meningitidis is an encapsulated, diplococcus, kidney bean-shaped bacteria that causes bacterial meningitis. Our study hopes to advance our understanding of disease progression, the spread frequency of the bacteria in people, and the interactions between the bacteria and human body by identifying a functional protein, potentially serving as a target for meningococcal medicine in the future. Methods: A hypothetical protein HP (PBJ89160.1) from N. meningitidis was employed in this study for extensive structural and functional characterization. In the predictive functional role of HP, several constitutive bioinformatics approaches are applied, such as prediction of physiological properties, domain and motif family function, secondary and tertiary structure prediction, energy minimization, quality validation, docking, and ADMET analysis. To create the protein’s three-dimensional (3D) structure, a template protein (PDB_ID: 3GXA) is used with 99% sequence identity by homology modeling technique with the HHpred server. To mitigate the pathogenicity associated with the HP function, it was docked with the natural ligand methionine and five other drug compounds like Verapamil, Loperamide, Thioridazine, Chlorpromazine, and Auranofine. Results: The protein is predicted to be acidic, soluble and hydrophilic by physicochemical properties analysis. Subcellular localization analysis demonstrated the protein to be periplasmic. The HP has an ATP-binding cassette transporter (also known as ABC transporter) involved in uptake of methionine (MetQ) that creates nutritional virulence in host. Energy minimization, multiple quality assessments, and validation value determination led to the conclusion that the HP model had a workable and acceptable quality. Following ADMET analysis and binding affinity assessments from the docking studies, Loperamide emerged as the most promising therapeutic compound, effectively inhibiting the ATP transporter activity of the HP. Conclusion: Comparative genomic analysis revealed that this protein is specific to N. meningitidis and has no homologs in human proteins, thereby identifying it as a potential target for therapeutic intervention.
Background: Clostridium botulinum and Clostridium perfringens, 2 major foodborne pathogenic fusobacteria, have a variety of virulent protein types with nervous and enterotoxic pathogenic potential, respectively. Objective: The relationship between the molecular evolution of the 2 Clostridium genomes and virulence proteins was studied via a bioinformatics prediction method. The genetic stability, main features of gene coding and structural characteristics of virulence proteins were compared and analyzed to reveal the phylogenetic characteristics, diversity, and distribution of virulence factors of foodborne Clostridium strains. Methods: The phylogenetic analysis was performed via composition vector and average nucleotide identity based methods. Evolutionary distances of virulence genes relative to those of housekeeping genes were calculated via multilocus sequence analysis. Bioinformatics software and tools were used to predict and compare the main functional features of genes encoding virulence proteins, and the structures of virulence proteins were predicted and analyzed through homology modeling and a deep learning algorithm. Results: According to the diversity of toxins, genome evolution tended to cluster based on the protein-coding virulence genes. The evolutionary transfer distances of virulence genes relative to those of housekeeping genes in C. botulinum strains were greater than those in C. perfringens strains, and BoNTs and alpha toxin proteins were located extracellularly. The BoNTs have highly similar structures, but BoNT/A/B and BoNT/E/F have significantly different conformations. The beta2 toxin monomer structure is similar to but simpler than the alpha toxin monomer structure, which has 2 mobile loops in the N-terminal domain. The C-terminal domain of the CPE trimer forms a “claudin-binding pocket” shape, which suggests biological relevance, such as in pore formation. Conclusions: According to the genotype of protein-coding virulence genes, the evolution of Clostridium showed a clustering trend. The genetic stability, functional and structural characteristics of foodborne Clostridium virulence proteins reveal the taxonomy and diverse distribution of virulence factors.
Introduction: Predicting Self-interacting proteins (SIPs) is a crucial area of research in predicting protein functions, as well as in understanding gene-disease and disease-drug associations. These interactions are integral to numerous cellular processes and play pivotal roles within cells. However, traditional methods for identifying SIPs through biological experiments are often expensive, time-consuming, and have long cycles. Therefore, the development of effective computational methods for accurately predicting SIPs is not only necessary but also presents a significant challenge. Results: In this research, we introduce a novel computational prediction technique, VGGNGLCM, which leverages protein sequence data. This method integrates the VGGNet deep convolutional neural network (VGGN) with the Gray-Level Co-occurrence Matrix (GLCM) to detect Self-interacting proteins associations. Specifically, we initially utilized Position Specific Scoring Matrix (PSSM) to capture protein evolutionary information and integrated key features from PSSM using GLCM. We then employed VGGNet as a predictive classifier, leveraging its capabilities for powerful learning and classification prediction. Subsequently, the extracted features were input into the VGGNet deep convolutional neural network to identify Self-interacting proteins. To evaluate the performance of the VGGNGLCM model, we conducted experiments using yeast and human datasets, achieving average accuracies of 95.68% and 97.72% respectively. Additionally, we compared the prediction performance of the VGGNet classifier with that of the Convolutional Neural Network (CNN) and the state-of-the-art Support Vector Machine (SVM) using the same feature extraction method. We also compared the prediction ability of VGGNGLCM with other existing approaches. The comparison results further demonstrate the superior performance of VGGNGLCM over other prediction models in this domain. Conclusion: The experimental verification further strengthens the evidence that VGGNGLCM is effective and robust compared to existing methods. It also highlights the high accuracy and robustness of the VGGNGLCM model in predicting Self-interacting proteins (SIPs). Consequently, we believe that the VGGNGLCM method serves as a valuable computational tool and can catalyze extensive bioinformatics research related to SIPs prediction.