
As saliva-based genomic analyses expand, the effect of common pre-sampling behaviors on salivary genomic DNA (gDNA) recovery remains poorly characterized. In this single-donor pilot study, we tested eight pre-sampling conditions combining tooth brushing, gargling, and water intake. Saliva was collected in triplicate under each condition, cells were isolated, and gDNA was extracted using the Qiagen Kit. Three-way ANOVA showed that brushing was the dominant factor, accounting for 62.9
Abstract Purpose Precision oncology depends on identifying cancer driver genes and linking them to targeted therapies. Current methods using curated gene sets or generic classifiers often miss biologically relevant patterns in complex gene interaction networks. Methods We developed the Precision Medicine Gene Network Analyser, integrating network topology analysis with machine learning for cancer gene identification. The dataset included 699 cancer driver genes (COSMIC Cancer Gene Census) and 15,050 background genes, mapped to high-confidence protein–protein interaction networks from STRING (456,300 edges, 15,749 nodes). Network features such as degree, betweenness, PageRank, k-core, and clustering coefficients were extracted. Imbalance Aware Network Integrator (IANI) was proposed to address class imbalance, where balanced resampling and ensemble models (logistic regression, random forest, gradient boosting) were combined with deep neural networks using focal loss, optimising thresholds for maximum F1-score. Hub genes were defined using a statistical cutoff of mean outdegree + 2 × SD (standard deviation). Results On a test set of 3150 samples (140 cancer, 3010 non-cancer genes), the optimised ensemble improved ROC-AUC from 0.84 to 0.96, precision from 0.78 to 0.90, and recall from 0.42 to 0.81 (F1 = 0.85) at a threshold of 0.466. Hub analysis identified 689 hubs with fourfold enrichment of cancer genes (16.1% vs. 4.4%, p < 10 − 20), showing higher betweenness centrality (p < 0.001). Key features such as degree (0.32), betweenness (0.24), and PageRank (0.19) contributed 75% of the model’s performance. Top hubs (TP53: 758, EGFR: 512, AKT1: 415 connections) showed 60–67% cancer gene enrichment, with pathway clustering in p53 signalling (75%) and cell cycle regulation (67.7%). Conclusion Integrating protein interaction topology with imbalance-aware machine learning achieved 96% discrimination accuracy. This work forms a base for the upcoming phases of drug-gene mapping and patient-specific therapy prediction within the Precision Medicine Gene Network Analyser.
Abstract Background and objectives Squamous cell carcinoma antigen recognized by T-cells 3 (SART3) has emerged as a promising target for cancer immunotherapy, given its overexpression in various malignancies and low or absent expression in non-tumorous tissues. This study aimed to design rationally and in silico evaluate a multi-epitope T cell vaccine targeting SART3, incorporating a TLR4 agonist adjuvant. The vaccine’s predicted immunogenicity, physicochemical properties, structural stability, and interaction with TLR4 were comprehensively assessed. Additional assessments of cytokine-inducing potential, B-cell epitopes, and disulfide engineering opportunities were also executed. Methods Potential T-cell epitopes from SART3 were identified using IEDB and screened for antigenicity (VaxiJen), toxicity (ToxinPred), and MHC-I/II binding affinity. Cytokine-inducing epitopes were evaluated using IL4pred, IL-10Pred, and IFNepitope servers. B-cell epitopes were predicted using ElliPro. The vaccine underwent comprehensive physicochemical, structural (I-TASSER/GalaxyRefine), molecular docking (HDOCK), molecular dynamics simulations, and disulfide engineering (Disulfide by Design 2.0) analyses. Results The optimized 344-residue vaccine demonstrated non-allergenicity, high stability (instability index 17.16), antigenicity (Vaxijen 0.67), and solubility (SOLpro 0.96). HDOCK predicted favorable vaccine–TLR4 binding (ΔG = − 265.61 kcal/mol, confidence 91%). MD simulations confirmed complex stability. Cytokine analysis revealed the potential to induce IL-4 and IL-10. The Val80–Ala123 pair exhibited the lowest bond energy (1.16 kcal/mol), indicating the optimal geometry for disulfide bond formation. The in silico immune simulations demonstrated a robust immune response following vaccine administration. Conclusion This rationally designed SART3-targeted multi-epitope vaccine exhibits promising in silico characteristics across immunogenicity, physicochemical, cytokine-inducing, B-cell epitope, structural, and disulfide engineering profiles, warranting experimental validation for cancer immunotherapy development.
Clinical text de-identification enables the use of electronic health records while protecting patient privacy, but public training data remain scarce and often have mismatched documentation styles. Recent works have proposed using large language models (LLMs) to generate synthetic clinical notes, but it remains unclear if they reflect distributions of real clinical notes. We examine how lexical and semantic drift across training and evaluation corpora affects protected health information (PHI) tagger performance. We generated synthetic notes from scratch for four categories using five generator LLMs and one judge LLM. Next, we fine-tuned small de-identification models on real, synthetic, and mixed corpora, and evaluated them on three external benchmarks under a harmonized label schema. Models trained on broad, clinically oriented sources transfer better than those on legal or narrowly synthetic data. These results suggest that although synthetic data lacks some real-world distributional properties, it remains useful in low-resource settings. We found that compact distributional and embedding-based drift measures moderately correlate with out-of-distribution F1 score, a practically important result because drift estimation can improve synthetic-data quality control and alignment.
Abstract The rapid growth of biological resources and associated research outputs has increased the complexity of resource discovery, access, and reuse across heterogeneous repositories. However, fragmented metadata schemas, limited interoperability, and siloed access mechanisms continue to hinder the integrated exploration of biological resources and their associated knowledge. Herein, we present BioOne (Biological resources One-Stop service platform), a unified informatics framework designed to address the fragmentation of national biological resources by integrating heterogeneous metadata from 14 distinct clusters and establishing a seamless resource-to-knowledge pipeline to enhance the discoverability and practical utilization of biological assets. BioOne is a national-scale web-based discovery platform that harmonizes biological resource metadata across 14 domain-specific biological resource clusters in Korea and systematically links these resources with external knowledge objects, including research papers, patents, biological datasets, and disease–drug–target information, through a unified discovery interface. To achieve this, we adopted a standard-aligned metadata integration framework, interoperable identifier mapping, and a modular system architecture to support scalable indexing, cross-domain search, and association-based navigation. By extending conventional catalog-based biological resource databases with an integrated discovery and access layer connecting distributed biorepositories to evidence-oriented knowledge resources, BioOne provides an informatics infrastructure for data-driven discovery, translational research, and coordinated utilization of biological resources at the national scale. The BioOne also offers a transferable implementation model for the large-scale integration of distributed biological resource systems.
Abstract Large-scale genomic rearrangements are prevalent in cancer genomes and can profoundly rewire three-dimensional (3D) genome architecture, leading to aberrant oncogene activation through enhancer hijacking. The rewired 3D organization generates unique chromatin contact signatures, which can be detected using deep learning-based approaches. However, extending such analyses to single-cell resolution, which is critical to delineate clonal heterogeneity in cancer, remains a major challenge, due to the limited number of training sets as single-cell Hi-C techniques are not standardized and only limited datasets are available across different methods. Here, we introduce scCAPReSE, a few-shot learning-based framework that adopts representations from a pre-trained image foundation model, CLIP, to enable robust classification of structural variation (SV) patterns in single-cell Hi-C data. By extracting and fine-tuning base weights from the foundation model, scCAPReSE enables effective training of deep learning classifiers using only a few hundred large-scale SV examples derived from a single cancer cell line while adapting classification tasks to heterogeneous single-cell Hi-C libraries. scCAPReSE achieved over 90% classification accuracy when evaluated on sci-Hi-C datasets. When further applied to scNanoHi-C data from the K562 chronic myeloid leukemia cell line, scCAPReSE correctly identified the Philadelphia chromosome translocation but also revealed substantial cell-to-cell variability in the contribution of SV-mediated chromatin interactions, highlighting previously inaccessible heterogeneity in cancer 3D genome organization. In summary, scCAPReSE provides a broadly applicable and data-efficient framework for detecting SV-driven 3D genome reorganization at single-cell resolution, enabling quantitative dissection of cancer-specific chromatin architecture and clonal heterogeneity. The developed method is freely available at https://github.com/kaistcbfg/CAPReSE.
As large language models (LLMs) become increasingly popular for information extraction (IE), concerns persist regarding the stability and reliability of their outputs. While accuracy has traditionally been the main evaluation metric, consistency-defined as the stability of model outputs across repeated runs-has recently been proposed as a complementary signal of reliability. In this work, we examine the relationship between accuracy and consistency in hard-prompted generative LLMs applied to entity and relation extraction. We conduct a systematic evaluation using four LLMs (GPT, DeepSeek, Qwen, Kimi) on the EPOP corpus, a plant-health dataset with rich entity types, long-range relations, overlapping relations, and strong argument constraints. To refine the interpretation of consistency, we distinguish between recoverable output variations-those that preserve the meaning of the extracted information-and critical ones that result in semantic errors. Our results show that while some positive correlation between accuracy and consistency exists, it is model-dependent and varies with task complexity. In structured prediction tasks, we show that consistency should be measured at the semantic level, ignoring superficial variations in format or wording. These insights have important implications for using self-consistency as a confidence filter and for designing reliable generative IE pipelines in specialized domains.
Antimicrobial peptides (AMPs) are universally found in both intracellular and extracellular settings and have significant antibiotic-resistant bacteria are becoming a bigger problem. In medical laboratories, it has shown notable anti-bacterial effectiveness in treating diabetic foot infections and related issues. New medication development frequently targets (AMPs), which are certainly ensuing components of adaptive immune system. The findings of this research employs deep learning to identify antibiotic activity. Numerous computational methods have been established to detect antimicrobial peptides via deep learning algorithms. We introduced a novel deep learning approach called antimicrobial peptides using Capsule Neural Network (AMP-CapsNet) to precisely forecast them and evaluated its efficacy against deep learning and baseline models. AMPs prediction using capsule neural networks, a type of next generation neural network, to build prediction models. Additionally, we utilized Amino Acid Composition (AAC) for effective features encoded method and as well as dipeptide composition (DPC). Every model underwent independent cross-validation and external testing. The findings indicate that the enhanced AMP-CapsNet deep learning model surpassed its counterparts, achieving an accuracy of 97.29% and an AUC score of 98.91% on the test set using with dipeptide Composition (DPC). The proposed AMP-CapsNet demonstrates superior performance of the testing set achieved accuracy 97.29% score with DPC and accuracy 84.42% score with AAC approach. Consequently, the technique we advocate is anticipated to enhance the accuracy of antimicrobial peptide predictions in the future. By producing powerful peptides for medication development and application, this study advances deep learning-based AMP drug discovery approaches. This finding has important ramifications for how biological data is processed and how pharmacology is calculated.
Intercellular mitochondrial transfer (MT) is emerging as a transformative communication axis in cancer biology. Intact mitochondria or mitochondrial components can be exchanged between tumor cells, stromal elements, and immune cells via tunneling nanotubes, extracellular vesicles, cell fusion, or phagocytic uptake. This organelle exchange enables metabolic adaptation by restoring OXPHOS (oxidative phosphorylation), increasing ATP production, and enhancing survival in hostile environments. Conversely, tumor cells also hijack mitochondria from cytotoxic lymphocytes thereby undermining immune function and contributing to immune escape and tumor progression. These converging metabolic exchanges fuel immune evasion, metastatic potential, and resistance to chemotherapy, radiation, and immunotherapy. Cutting-edge tracing tools, including mitochondrial reporter proteins and single-cell mitochondrial genome lineage mapping, have uncovered MT events both in vitro and in vivo. Therapeutic strategies designed to block mitochondrial trafficking, inhibit nanotube formation or vesicle uptake, or enhance immune cell mitochondrial resilience hold promise for tumor sensitization and restoration of antitumor immunity. A deeper understanding of MT provides novel insight into cancer metabolism and intercellular communication, offering a foundation for future therapeutic innovation and potential clinical application as both a biomarker and a therapeutic target.
Thyroid cancer (THCA) is a common malignant tumor of the endocrine system, and significant clinical challenges remain in its diagnosis and prognostic evaluation. This study aims to elucidate the role of AGPAT4 in thyroid cancer by investigating its expression, involvement in metabolic pathways, and potential as a prognostic biomarker. We analyzed data from 512 thyroid cancer patients and 279 controls, performed differential expression analysis of AGPAT4 in thyroid cancer, analyzed the gene expression correlation of AGPAT4 in thyroid cancer, and the protein–protein interaction (PPI) network and functional enrichment analysis of AGPAT4 and its differentially expressed genes (DEGs) were constructed. The Kruskal–Wallis test and receiver operating characteristic (ROC) curve analysis were used to investigate the correlation between AGPAT4 expression and clinicopathological characteristics as well as its diagnostic efficacy. Cox regression analysis and Kaplan–Meier analysis were employed to evaluate its prognostic value. Additionally, single-sample gene set enrichment analysis (ssGSEA) was utilized to explore the association between AGPAT4 expression and the level of immune infiltration in the tumor microenvironment. Our findings revealed that AGPAT4 was significantly downregulated in thyroid cancer (THCA) tissues (P < 0.001), suggesting a potential tumor-suppressive role of AGPAT4 in thyroid cancer. AGPAT4 exhibited robust efficacy in distinguishing tumor tissues from normal tissues, with an area under the receiver operating characteristic curve (AUC) of 0.973. Furthermore, AGPAT4 expression levels were significantly correlated with pathological stage and survival rate (P < 0.05). Kaplan–Meier survival analysis showed that patients with high AGPAT4 expression had better progression-free interval (PFI) (HR = 0.45, P = 0.007). Protein–protein interaction (PPI) network and functional enrichment analyses revealed that AGPAT4 is involved in key pathways associated with thyroid cancer progression. Immune infiltration analysis suggested an association between AGPAT4 expression and immune responses in the tumor microenvironment. AGPAT4 holds promise as a potential biomarker for the differential diagnosis and prognostic assessment of thyroid cancer, thereby providing a possible reference for the further exploration of therapeutic strategies against this disease.
Fusion genes are key oncogenic drivers in various cancers; however, their role in hepatocellular carcinoma (HCC) remains underexplored. Here, we analyzed RNA-seq data from 68 HCC patients and identified several fusion products where SLC39A14-PIWIL2 stood out a putative driver. Functional assays revealed that the promoter of SLC39A14 potentially drives the overexpression of a truncated PIWIL2 protein (tPIWIL2), which retains its oncogenic MID and PIWI domains, in liver tissues. Both the wild-type and tPIWIL2 were found to interact with oncogenic partners HDAC3 and NME2 through these domains, as demonstrated by structural modeling and molecular dynamics simulations. To disrupt these interactions, we designed novel decoy peptides that potentially competes with both HDAC3 and NME2, effectively inhibiting PIWIL2-driven tumor activity in Huh7, HepG2, SNU449, and SNU398 HCC cell lines. Among the tested candidates, NEP1 markedly suppressed PIWIL2-driven oncogenic activity, and its co-administration with 5-fluorouracil (5-FU) significantly reduced PIWIL2-induced chemoresistance, thereby enhancing therapeutic efficacy. Collectively, these findings establish SLC39A14-PIWIL2 as a novel oncogenic fusion in HCC and highlight fusion protein–targeted peptide therapeutics as a promising avenue for precision treatment in HCC. Graphical Abstract
BACKGROUND:Environmental pollutants have a profound impact on microbial dynamics. This study highlights the influence of anthropogenic activity on the shift in bacterial diversity in the catchment area compared to upstream and downstream at Kathajodi, using a metagenomic approach for the first time in River Kathajodi. METHODS:Water samples were collected from upstream, catchment, and downstream locations and transported at 4°C to the laboratory for DNA extraction, library preparation, sequencing, and physicochemical analysis employing inductively coupled plasma. The extracted DNA was sequenced via the Illumina HiSeq platform and analyzed through MG-RAST for taxonomic and functional classification using KEGG and COG annotations. Statistical diversity analysis, including rarefaction curves, alpha- and beta-diversity indices, and Venn diagrams, provided insights into microbial composition and community variations across sites. RESULTS:A significant abundance of pollution indicator members of phylum Bacteroidetes (29.82%) in the catchment (CM), highly contaminated with metals, fecal, and other organic pollutants, could be attributed to their high metabolic capabilities to degrade them. The pristine upstream (US) exhibited an abundance of Shewanella (25.04%), Pseudomonas (17.35%), and Synechococcus (5.62%). The CM, influenced by high anthropogenic activity, showed higher abundances of Flavobacterium (5.20%), Arcobacter (4.05%), and Bacteroides (3.88%). In contrast, downstream (DS), with fewer anthropogenic activities, displayed higher abundances of Aeromonas (4.40%), Acidovorax (0.52%), and Acidimicrobium (0.32%). The highest bacterial diversity of CM could be due to the influence of the physicochemical properties of city waste effluent. From the Venn diagram, 73 common OTUs at the genera level were observed in all three sites, which indicates that the native microflora of the river water niche remains unaffected irrespective of the temporary changes in the vicinity. The functional profiling through KEGG and COG revealed that CM was enriched in carbohydrate metabolism (12.11%), while DS exhibited higher contributions to amino acid metabolism, along with the highest relative abundance of general function prediction (R) (12.89%), all indicative of stress adaptation and metabolic flexibility under polluted conditions. The clean upstream is home to oxygen-loving helpful bacteria, the catchment supports nutrient-hungry and sewage-linked microbes, while the downstream is dominated by metal-tolerant and possibly harmful bacteria, showing the clear impact of human activities along the river. CONCLUSIONS:The marked shift in bacterial diversity between US, CM, and DS regions highlights the ecological consequences of anthropogenic impact. These findings emphasize the need for effective environmental management to safeguard water quality and prevent undesirable health issues.
Infective endocarditis (IE) is a serious infection of the heart valves, and standard culture methods often miss the bacteria responsible, especially in culture-negative cases. To address this, we used 16S rRNA gene-based next-generation sequencing (NGS) on heart valve tissue. This approach allowed us to map out the bacterial communities present and evaluate their potential role in IE. We identified six key bacterial genera—Enterococcus, Streptococcus, Coxiella, Staphylococcus, Haemophilus, and Cutibacterium—plus three specific species: Streptococcus troglodytae, Haemophilus parainfluenzae, and Coxiella burnetii. Our co-occurrence analysis showed that these bacteria tend to exist independently within infected valve tissue, with no significant correlations between them. We detected bacterial taxa, including Cutibacterium and Streptococcus troglodytae. Although S. troglodytae is rarely associated with IE, and Cutibacterium comprises low-abundance bacteria not typically linked to this condition. These findings demonstrate the value of NGS in identifying pathogens that standard culture methods may overlook. As these results are based on computational analyses, further laboratory validation is required. Incorporating NGS into diagnostic protocols may enhance pathogen detection in culture-negative IE and support more targeted treatment and prevention strategies.
Artificial intelligence (AI)-assisted scientific writing is now a common practice in academic publishing, yet concerns persist regarding the authenticity and reproducibility of AI-generated content. While AI tools offer significant advantages, particularly for non-native English speakers who face substantial linguistic barriers in scientific communication, the risk of AI hallucinations and fabricated citations threatens the integrity of scholarly discourse. Journals often require disclosure of the entire AI prompt rather than meaningful intellectual contributions, but this is becoming increasingly impractical as AI prompts are getting longer and more complex. In this paper, I argue that transparency in AI-assisted writing should focus on capturing the author’s core research perspective and section-specific key points—the foundational elements that drive meaningful scientific communication. To address this challenge, I developed a web-based tool that implements a human-in-the-loop approach requiring authors to define their research perspective and create detailed outlines with key points before any AI text generation occurs. The tool mitigates AI hallucination by only allowing the use of user-provided citations and generating transparency reports documenting the key elements used for text generation. I validated this approach by writing this paper using the tool itself, demonstrating how the transparency reporting method works in practice. This methodology ensures that AI serves as a linguistic tool rather than a content generator, preserving scientific integrity while democratizing access to high-quality academic writing across linguistic and cultural boundaries.
Imaging-based spatial transcriptomics (ST) enables the quantification of gene expression at single-cell resolution while preserving spatial context, but its utility is limited by small gene panels and challenges in accurate cell segmentation. To address these limitations, we present a graph autoencoder framework that integrates subcellular transcript distribution patterns with cell-level gene expression profiles for enhanced cell clustering in imaging-based ST (SPICEiST). The clustering performance of SPICEiST was systematically evaluated across several cancer datasets and gene panel sizes. The results demonstrate that SPICEiST consistently outperforms the conventional cell-level gene expression-based methods in distinguishing subtle differences in cell states, as measured by the number of cell clusters and clustering indices, such as the CHI and DBI. Moreover, the findings indicate that SPICEiST can further enhance the performance, even with advancements in cell segmentation, particularly for datasets with small gene panels. Overall, these improvements in cell clustering indices, CHI and DBI, were more pronounced in datasets with small gene panels of around 300 genes, in contrast to those with large panels containing over a thousand genes. Notably, SPICEiST also reveals more spatially intermixed and less compartmentalized cell clusters, a characteristic that better reflects the complex and heterogeneous nature of tumor microenvironments. This effect was especially evident in the datasets with large panels. These findings highlight the value of leveraging subcellular transcript patterns to overcome the inherent limitations of imaging-based ST, particularly for small gene panels, and may provide new insights into tumor heterogeneity.
Ontology frameworks are essential for organizing complex biological knowledge, such as genes, phenotypes, and pathways, and for ensuring consistent data annotation and retrieval. In biological research, ontologies like the Gene Ontology (GO) and crop-specific trait ontologies (TO) for Oryza sativa (rice) standardize terminology across studies, supporting cross-study comparison and hypothesis generation. However, ontology annotations usually rely on expert manual review of the literature, a process that is accurate but time-consuming, labor-intensive, and difficult to scale as biological data grows. Manual approaches are also prone to inconsistencies and errors. The emergence of large language models (LLMs) such as ChatGPT, DeepSeek, and KIMI, along with curated databases like Rice-Alterome and PubAnnotation, offers new opportunities for semi-automated ontology curation. This study explores how these technologies can be integrated to develop an efficient literature-based curation system for rice trait ontology. We developed a curation system that integrates Rice-Alterome—a comprehensive database of rice genomic variations, mutations, and sentence-level literature evidence linked to GO and TO terms with PubAnnotation, an open-source platform for collaborative text annotation. LLMs (DeepSeek and KIMI) were integrated via APIs to automate the extraction, annotation, and validation of trait-related information via prompt engineering. The system was evaluated through use cases designed to demonstrate its performance and functionality compared to manual curation. The proposed system substantially enhanced the retrieval and organization of literature evidence compared to manual methods. The integrated platform, available through a dedicated website, connects Rice-Alterome, PubAnnotation, and LLMs to streamline ontology curation and evidence discovery. This framework reduces the time domain experts need to locate and validate relevant information and provides interactive tools for users to add, merge, or refine trait annotations. The LLM-driven prompt-based querying also improved the identification of implicit or missing information that may be overlooked during manual curation. Integrating LLMs with Rice-Alterome and PubAnnotation offers a promising solution for automating rice trait ontology curation. This approach accelerates evidence collection and enhances data consistency and accessibility. Future extensions of this framework will target additional crops such as wheat and maize and focus on refining LLM-based retrieval and annotation mechanisms for broader agricultural genomics applications.
Spatial transcriptomics technologies have significantly enhanced the analysis of gene expression profiles by retaining the spatial information of intact tissue sections and enabling the possibility of a more profound comprehension of tissue structures and cellular relationships. Despite this, most platforms have limited resolution, and at numerous capture spots, multiple signals from various cells are present, requiring deconvolution, a set of computational steps to deduce the underlying cellular composition. Over the last few years, a range of algorithms has been proposed to address this problem, each employing distinct computational principles and processing paradigms. The present review seeks to present a comprehensive analysis of twenty such algorithms, focusing on their methodological foundations. We contrast the underlying computational algorithms, modeling methods, and data processing pipelines that underlie them, and how they deal with external references, noise and sparsity in the data. By drawing out the conceptual as well as technical foundations of each algorithm, we aim to provide researchers a complete and hands-on grasp of the computational landscape of spatial transcriptomics deconvolution. This review is a methodological handbook to enable deep understanding of current deconvolution methods to develop novel strategies and help in selecting or applying these existing tools for different biological contexts.
Schistosomiasis remains a significant public health burden, necessitating the development of effective vaccines against it. In this study, a multi-epitope subunit vaccine was designed against adenylate kinase 2 protein and evaluated for its potential to elicit protective immunity against three Schistosoma species. CTL, HTL, and B-cell epitopes were identified using immunoinformatics tools and linked using AAY and KK linkers, respectively. The 50S ribosomal protein L7/L12, a known TLR4 agonist, was incorporated as an adjuvant to enhance immune activation in the vaccine. Molecular docking and molecular dynamics (MD) simulations demonstrated a strong binding affinity between the vaccine and human TLR4, supported by low RMSD and Rg values, indicating structural stability. The negative binding energy further validated the vaccine’s potential for engaging TLR4. The immunogenic profile predicted robust activation of CD4+ and CD8+ T cells, as well as neutralizing antibody responses. These findings highlight the potential of the vaccine to stimulate both cell-mediated and humoral immunity, making it a promising candidate for further experimental validation against schistosomiasis.