Ultra-multiplex PCR has become essential for targeted amplicon sequencing in molecular biology and clinical genetics. We implemented PCRpanel, a command-line Java application with a companion web interface that jointly optimises primer thermodynamics, linguistic sequence complexity, primer-dimer interactions, and multiplex-aware, reference-guided specificity and repeat masking for automated ultra-multiplex panel design. We benchmarked PCRpanel in silico against NGS-PrimerPlex and Olivar and validated it experimentally by designing tiled primer pools targeting all exonic regions of COL4A3, COL4A4, COL4A5 and COL4A6 and amplifying DNA from four clinical Alport syndrome samples. As this was a pilot run aimed at initial validation of the design, sequencing is reported for one representative patient sample, in which each primer pool was amplified separately and sequenced at two template dilutions, giving 14 libraries on the Illumina MiSeq platform; pool 1.1 was run as a 50-primer sub-pool during protocol optimisation. Reads were aligned to GRCh38 with the DRAGEN Bio-IT Platform v4.4.4 and coverage was quantified over an exon-restricted target interval of 45,372 bp derived from GENCODE v49. Duplicate marking was disabled, because amplicon molecules generated by a common primer pair are positionally indistinguishable from PCR duplicates. Here we present PCRpanel, a command-line Java application for the automated design of thermodynamically optimised primer panels for short- and long-read amplicon sequencing. PCRpanel supports workflows from simple two-primer assays to ultra-high-plex designs containing hundreds of amplicons and accommodates targets ranging from complete viral genomes and eukaryotic gene families to structural-variant breakpoints and environmental metagenomes. The software jointly optimises sequence complexity, thermodynamic stability, primer-dimer interactions, and off-target amplification, enabling robust primer selection even in repetitive or homologous regions. Efficient algorithms generate ultra-high-plex panels within seconds to minutes on standard hardware, up to 15 min when the full human reference genome is screened, and both gene-specific and universal designs are supported across homologous gene families. We validated the approach by designing 237 primer pairs (474 primers) targeting all exonic regions of the three established Alport syndrome genes COL4A3, COL4A4 and COL4A5, together with the adjacent COL4A6. Experimental validation on four clinical Alport syndrome DNA samples confirmed successful amplification of all 237 primer pairs. As this was a pilot run, sequencing of one representative patient sample, in which each primer pool was amplified separately and sequenced at two template dilutions as 14 libraries (pool 1.1 as a 50-primer sub-pool during protocol optimisation), was used to evaluate analytical performance: read alignment rates were 94.2–99.3
Background/Objectives: Cardiac arrhythmias are among the leading causes of sudden cardiac death (SCD). Pathogenic variants in potassium channel genes play a key role in inherited arrhythmia syndromes, yet their contribution in Central Asian populations remains poorly characterized. Methods: We performed targeted next-generation sequencing (NGS) using a 96-gene custom Haloplex panel in 79 Kazakhstani patients with clinically diagnosed arrhythmias, including atrioventricular block, sick sinus syndrome, and atrial fibrillation. Detected variants in potassium channel genes were classified according to ACMG guidelines and correlated with clinical phenotypes. Results: A total of 52 variants were identified across 11 potassium channel genes. Two likely pathogenic variants (KCNH2 p.Cys66Gly and p.Arg176Trp) and six variants of uncertain significance (VUS) in KCNQ1, KCNE2, KCNE3, and KCNJ8 were detected. Two novel previously unreported variants were found in KCNE5 and KCND3. Patients harboring pathogenic variants commonly presented with early-onset arrhythmias or a positive family history of cardiovascular disease. Carriers of KCNH2 variants exhibited mild QT prolongation and recurrent syncope. Conclusions: This is the first genetic study of potassium channel gene mutations in Kazakhstani patients with cardiac arrhythmias. The detection of pathogenic and novel variants highlights the clinical utility of integrating genetic testing into diagnostic and management pathways for arrhythmia syndromes. Population-specific genomic data are essential for improving risk stratification, guiding medication safety, and enabling cascade family screening in Central Asia.
To ensure that the benefits of biomedical research are shared globally, we must overcome the barriers preventing the establishment of national genomic projects in underrepresented regions by navigating the ethical, legal and financial complexities of launching these initiatives. We identify seven major challenges — ranging from policy awareness to data privacy regulations — and conclude that successful implementation requires strategic local capacity building and a commitment to open data standards that ensure genomic discoveries are accessible to all.
We thank the authors of the Comment [...].
Background: NR1H3 encodes Liver X Receptor- α (LXRα), which is a nuclear receptor regulating cholesterol efflux, lipid metabolism and inflammation. Impaired function of NR1H3 in the liver leads to hypertriglyceridemia and hepatic steatosis. Moreover, inactivation of the gene in macrophages increases inflammatory signaling that causes atherogenesis. This dual action of LXRα becomes a reason of fatty liver and cardiovascular disease. Human studies confirm that rare damaging NR1H3 mutations cause hepatic cholesterol accumulation, inflammation and fibrosis, whereas common variants have been associated with carotid atherosclerosis. Materials and Methods: The aim of this study is to investigate the role of NR1H3 variants in the development of hepatic steatosis and atherosclerosis in a Kazakhstani cohort. 401 subjects were included with documented CVD. Whole exome sequencing was performed to analyze LXRα variants prevalence. Furthermore, clinical, biochemical, and imaging data (lipid profile, FibroScan, echocardiography, angiography) were analyzed to characterize phenotype– genotype correlations. Results: Whole exome sequencing identified two patients aged 57 and 58, with likely pathogenic NR1H3 variants (rs74842890, rs1591127791). Patients have type 2 diabetes and serious hepatic abnormalities. FibroScan revealed severe steatosis in both patients, with moderate fibrosis in one and mild in the other. The lipid profiles showed elevated LDL-C and high triglycerides level. Moreover, one patient exhibited an extremely high Lp(a) level at 178 mg/dL. Both patients demonstrated normal ejection fraction, though one case showed reduced strain. Clinically, two of them had documented atherosclerosis disease with very high risk. Conclusion: These findings show correlation of NR1H3 variants with phenotype of fatty liver, dyslipidemia, diabetes and atherosclerosis. Even with a small number of subjects, the results suggest impaired LXRα function may cause combined hepatic and cardiovascular pathology. Acknowledgements: Supported by the Committee of Science of the Ministry of Science and Higher Education of Republic of Kazakhstan (AP23490249), (AR19677442), (BR24993023), (BR24992841), (BR21881970) and Nazarbayev University CRP (211123CRP1608). Key words: Atherosclerosis, Liver X Receptor- α, hepatotoxicity, inflammation
This study aimed to investigate the association between VDR gene polymorphisms and serum vitamin D levels, as well as injury predisposition, in elite athletes from Kazakhstan. We recruited 137 elite athletes from Kazakhstan across nine different sports. Serum vitamin D levels were measured by Access 25(OH)D Vitamin D Total assay. The VDR gene polymorphisms were identified using quantitative PCR (qPCR). The association between VDR gene polymorphisms and serum vitamin D levels, as well as injury predisposition, was assessed using statistical methods. Over 60
Background: Channelopathies are among the major causes of sudden cardiac death (SCD). Mutations in ion channel genes have been associated with life-threatening arrhythmias. This study aimed to characterize genetic spectrum of Kazakhstani patients with channelopathies using targeted next-generation sequencing (NGS), which may facilitate personalized therapeutic strategies. Material and methods: Genomic DNA was extracted from whole blood samples of patients with clinically diagnosed channelopathies. Targeted NGS was performed on the NovaSeq 6000 platform (Illumina) using the TruSight Cardio panel, which covers 174 genes associated with inherited cardiovascular disorders. Bioinformatic analysis was performed using standard pipelines, and variants were classified according to the American College of Medical Genetics and Genomics (ACMG) guidelines. International databases (ClinVar, MutationTaster, gnomAD, Polyphen-2) were additionally used for variant interpretation. Results: Genetic screening revealed 23 rare variants. We found five disease-causing pathogenic (P) and likely pathogenic (LP) variants in our study. Namely, disease-causing mutations are detected in SCN5A, KCNQ2, CACNB2, KCNQ1. In addition, 11 variants of unknown significance (VUS) with suggestive evidence of pathogenicity were identified in patients. Moreover, several variants were located in genes previously implicated in arrhythmia susceptibility, warranting further investigation. The results of genetic screening were reported to clinicians to support further clinical decision-making and patient management. Conclusion: Our study demonstrated genetic profiles of patients with channelopathies by using the TruSight Cardio panel. Genetic screening revealed clinically significant variants within ion channel genes. Acknowledgement: The study was funded by grants from the Committee of the Ministry of Science and Higher Education, Republic of Kazakhstan (AP23490249, BR24993023) and Nazarbayev University CRP - 211123CRP1608. Keywords: channelopathies, mutations, targeted sequencing
Cancer cells can sustain survival independently of exogenous growth factors. To investigate their adaptation to serum deprivation, we analyzed transcriptomic responses in two cancer cell lines. Transcriptome analysis revealed upregulation of mRNAs encoding cholesterol biosynthesis enzymes. This was a critical adaptive response, as a pharmacological inhibition of the pathway with statin triggered a robust apoptotic cell death accompanied by generation of a mitochondrial reactive oxygen species. The mechanistic target of rapamycin complex 1 (mTORC1), a master regulator of cell growth, is known to be engaged in controlling lipid biosynthesis. We detected the high polysomal and preribosomal peaks not only in serum-containing medium but also under serum deprivation, indicating a high rate of protein synthesis and ribosomal biogenesis independent of serum. In addition, the inhibition of mTOR kinase activity substantially reduced polysome abundance, with a more pronounced effect in serum-deprived cancer cells. Notably, the mTOR kinase inhibition also prevented the upregulation of the cholesterol synthesis enzyme that established a direct link between mTOR activity, protein synthesis, and cholesterol biosynthesis. Together, our results show that cancer cells adapt to serum withdrawal by activating the cholesterol synthesis pathway through mTOR-dependent regulation of gene expression and protein synthesis, underscoring a critical mechanism of survival under serum withdrawal.
Background: Epilepsy is a common neurological disorder with diverse etiologies, among which genetics plays a significant role. Advances in next-generation sequencing (NGS) have enabled the identification of numerous genes and variants associated with epilepsy, improving our understanding of its molecular mechanisms. Determining the genetic basis of epilepsy is crucial for optimizing treatment strategies, particularly in drug-resistant cases where certain variants can influence responsiveness to anti-epileptic drugs (AEDs). This study aimed to explore the genetic landscape of pediatric epilepsy patients in Kazakhstan using whole-genome sequencing (WGS). Materials and methods: Ten pediatric epilepsy patients of Asian ancestry with epilepsy of unknown etiology were recruited. gDNA was sequenced using the NovaSeq6000 platform. Bioinformatic processing and variant annotation were performed using ANNOVAR and public databases such as ClinVar and Franklin Genoox. Variants were classified according to the American College of Medical Genetics and Genomics (ACMG) guidelines. Results: We identified 10 pathogenic, 3 likely pathogenic, and 8 VUS variants, several in genes previously linked to epilepsy and neurodevelopmental disorders. Notably, a pathogenic variant in KCNT1 (rs397515404) and a likely pathogenic variant in ARX (rs1064794843) were associated with infantile epileptic encephalopathy. Additionally, a polymorphism in the ABCB1 gene (rs2032582), which encodes P-glycoprotein, was detected in seven patients (five homozygous C/C, two heterozygous A/C). ABCB1 is known to influence AED response and may play a role in drug resistance. However, this variant is currently classified as VUS, and its direct involvement in epilepsy remains uncertain. Conclusion: This study provides the first comprehensive WGS-based genetic profiling of pediatric epilepsy patients in Kazakhstan, identifying both disease-associated variants and potential pharmacogenomic markers. The discovery of ABCB1 rs2032582 highlights the importance of integrating genetic testing into epilepsy management and underscores the need for larger studies to validate these findings. Acknowledgement: This research has been funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Program targeted funding No. BR27199879). Keywords: Epilepsy, NGS, genetic associations
Background: Esophageal squamous cell carcinoma (ESCC) is one of the most lethal cancers in Central Asia, yet its molecular landscape in Kazakhstan remains poorly explored. However, the fusion-gene landscape in Kazakhstan remains undercharacterized. Fusion transcripts can serve as diagnostic biomarkers or therapeutic targets, motivating a focused survey in a local cohort. Materials and methods: We analyzed paired tumor and adjacent normal tissues from ESCC patients collected in Kazakhstan (34 normal, 33 tumor samples). Tumor specimens were obtained from patients who had undergone Ivor-Lewis esophagectomy without receiving prior chemotherapy or radiotherapy. Total RNA was sequenced (paired-end), reads aligned to GRCh38 (Gencode v44 / CTAT genome library). Fusion calling was performed with two orthogonal tools (STAR-Fusion v1.13.0 and Arriba v2.4.0). Callsets were intersected to prioritize concordant events; candidate junctions were manually inspected in IGV to remove artifacts. Primers for prioritized fusions were designed and PCR followed by Sanger sequencing is underway for orthogonal validation. Results: Intersection of STAR-Fusion and Arriba reduced noise and produced a concise set of high-confidence fusion candidates detected in multiple tumors and absent from matched normals. Manual IGV review confirmed strong split- and spanning-read support at exon boundaries for several recurrent events. Notable recurrent partners include MAP4K5-L2HGDH, ADAMTS2-ENSG00000253652, RAB8A-CIB3, BAZ2B-WDSUB1, HUWE1-SUPT3H, and PLEKHA5-KRAS. Conclusion: Using complementary fusion callers and rigorous manual curation, we generated an assay-ready shortlist of recurrent fusion transcripts in Kazakh ESCC. Pending PCR/Sanger confirmation, these candidates warrant follow-up as potential biomarkers or mechanistic leads for ESCC in Central Asia. Acknowledgement: We thank all the participants in this study. This research has been funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grants No. AP23490594, BR18574184, BR24993023, BR24992841, BR27199879), Nazarbayev University funding CRP grants 021220CRP2222, 211123CRP1608. Key words: esophageal squamous cell carcinoma, gene fusions, RNA-seq, Kazakhstan
Genomic studies on Central Asian populations are still limited, and this study presents the first comprehensive, high-quality genotyping dataset for 224 healthy individuals of Kazakh origin. The data were generated using the Illumina Infinium SNP Genotyping Array GSA MG v2, covering 665,608 SNPs, with 523,630 SNPs retained after quality control. This dataset serves as a valuable reference for comparative studies focused on human genomics, supporting research in population clustering, ethnicity validation, and diverse biomedical investigations on Kazakh individuals. This resource marks a significant contribution to genomic research on individuals from underrepresented region of Central Asia.
Background: Vitamin D plays a vital role in musculoskeletal health, immune function, and athletic performance. Its deficiency is common among elite athletes due to factors such as training environment, season, diet and high physiological demands. Genetic variation may influence individual susceptibility to vitamin D deficiency by affecting its metabolism, transport, and signaling pathways. This study aimed to investigate the association between VDR gene polymorphisms and vitamin D status in elite Kazakhstani athletes. Materials and methods: A total of 92 elite male power athletes, involved in several Olympic disciplines, were recruited during the summer training period. Serum 25(OH)D concentrations were measured using a chemiluminescent immunoassay. Four common VDR single-nucleotide polymorphisms - FokI (rs2228570), TaqI (rs731236), BsmI (rs1544410), and ApaI (rs7975232) were genotyped using TaqMan real-time PCR. Associations between genotypes and vitamin D were analyzed with linear regression, adjusted for age, BMI, and sports experience. Results: Vitamin D insufficiency (<30 ng/mL) was observed in 63% of participants, including 38% with deficiency (<20 ng/mL). Among the four studied polymorphisms, only the FokI A/A genotype showed a strong association with vitamin D insufficiency (OR = 9.25, 95% CI: 2.01– 42.51, p < 0.01), while no significant associations were found for TaqI, BsmI, or ApaI. Conclusion: Our findings reveal a high prevalence of vitamin D inadequacy among elite male power athletes in Kazakhstan. The VDR FokI A/A genotype is a potential genetic biomarker for vitamin D insufficiency, highlighting the importance of personalized monitoring and targeted supplementation strategies to optimize athlete health and performance. Future studies should investigate additional genetic factors that influence vitamin D metabolism in athletic populations. Acknowledgement: This research has been funded by the Science Committee of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. AP19680003) Keywords: Vitamin D deficiency, genetic polymorphisms, athletic performance References: Ammerman, B.M., Ling, D., Callahan, L.R., Hannafin, J.A., Goolsby, M.A. Prevalence of Vitamin D Insufficiency and Deficiency in Young, Female Patients with Lower Extremity Musculoskeletal Complaints. Sports Health: A Multidisciplinary Approach 13, 173–180 (2021). Cui, A., Zhang, T., Xiao, P., Fan, Z., Wang, H., Zhuang, Y. Global and Regional Prevalence of Vitamin D Deficiency in Population Based Studies from 2000 to 2022: A Pooled Analysis of 7.9 Million Participants. Front. Nutr. 10, 1070808 (2023).
Background: Population-specific reference genomes are crucial for accurate genomic analysis and precision medicine applications. Currently, there is a lack of high-quality chromosome- level reference genomes specifically representing the Kazakh population, which limits the effectiveness of biomedical investigations and genomic studies in this demographic. This project aims to address this gap by creating a comprehensive chromosomal-level assembly of Kazakh individuals' whole genomes using advanced genomic technologies, including next- generation sequencing (NGS, Illumina), third-generation sequencing (TGS, Oxford Nanopore), optical genome mapping, and Hi-C chromosomal conformation capture. Materials and methods: We generated chromosome-scale assemblies of Kazakh individuals by integrating multiple modern genomic technologies. Whole-genome sequencing was performed using both long-read and short-read platforms. The assemblies were polished with short-read data, scaffolded with Bionano optical genome maps, and organized at the chromosomal level using Hi-C chromosomal conformation data. Quality assessment was conducted at each stage, and comparative analyses will be performed against other global reference genomes. Results: The project delivered high-quality, chromosome-level assemblies for individuals of the Kazakh population. These assemblies provide insights into unique genetic features, improve the accuracy of population-specific genomic analyses, and will be deposited in open- access repositories for use in biomedical and bioinformatics research worldwide. Acknowledgement: This study is supported by a grant AP23490594 from the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan. Keywords: Chromosome-scale assembly; Kazakh genome; long-read sequencing; Bionano optical mapping; Hi-C; de novo assembly; population genomics.
Background: Accurate variant annotation plays a crucial role in whole-genome sequencing (WGS) studies, enabling researchers to identify pathogenic variants, prioritize candidate genes, and better understand population-specific genomic diversity. Several widely used annotation tools, including ANNOVAR1, OpenCRAVAT2, and Nirvana3 differ in supported databases, classification frameworks, and computational efficiency. However, systematic comparisons of these tools on large-scale population WGS data remain limited, particularly for underrepresented populations such as the Kazakh cohort. Materials and methods: We compared the performance of ANNOVAR, OpenCRAVAT, and Nirvana using a high- coverage WGS dataset (~4.9M SNVs and InDels) from a Kazakh population sample. Each VCF file was processed independently with standardized database configurations relevant for clinical interpretation. All analyses were performed on an Illumina DRAGEN v4.3.13 server (Oracle Linux 8.9) equipped with 2× Intel Xeon Gold 6226R CPUs @ 2.90 GHz (64 threads), 512 GB RAM, and FPGA-based hardware acceleration. This unified computational environment ensured consistent performance measurements across tools. The comparison focused on the number of annotated variants, supported databases, detection of rare pathogenic variants, and computational runtime. Results: Our analysis revealed notable differences between the tools. ANNOVAR annotated 4.93M variants in 2h 21m (4,391,599 SNPs, 539,696 INDELs) and reported 9 pathogenic and 4 likely pathogenic variants. OpenCRAVAT processed 4,913,713 variants in 7h 39m (4,376,625 SNPs, 537,088 INDELs), identifying 18 pathogenic and 13 likely pathogenic variants. Nirvana completed the analysis in 9m 21s, capturing 4,840,343 variants (3,907,526 SNPs, 932,817 INDELs) and reported substantially more clinically relevant findings, including 726 pathogenic and 277 likely pathogenic variants. Notably, ~22% of pathogenic variants were uniquely identified by a single tool, highlighting the complementarity of different annotation strategies (Figure 1). Conclusion: No single annotation tool provides complete variant coverage. ANNOVAR offers speed and efficiency, OpenCRAVAT provides deeper predictive insights, and Nirvana enhances clinical interpretation, particularly for structural variants. Combining results from multiple workflows significantly improves annotation depth and clinical relevance, especially in large-scale population WGS studies. Key words: whole-genome sequencing, variant annotation, ANNOVAR, OpenCRAVAT, Nirvana, population genomics References: Wang, K., Li, M. & Hakonarson, H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing Nucleic acids research 38, e164-e164, (2010). Pagel, A. et al. Integrated Informatics Analysis of Cancer-Related Variants. JCO Clin Cancer Inform 4, 310-317, (2020). Stromberg, et al. in Proceedings of the 8th ACM International Conference on Bioinformatics, Computational Biology,and Health Informatics 596, (2017).
Kazakhstan is on the list of countries with a high burden of MDR TB according to the WHO for 2021-2025. The purpose of the study is to determine the relationship of MDR with the Beijing M.tuberculosis family, as well as further genomic analysis of the data of whole-genome sequencing of clinical MDR M.tuberculosis isolates common in Kazakhstan [1]. The objective of this study was to investigate the association between multidrug resistance (MDR) and the Beijing lineage of Mycobacterium tuberculosis, as well as to conduct a comprehensive genomic analysis of MDR clinical isolates circulating in Kazakhstan using whole genome sequencing (WGS) [2]. A total of 76 clinical MDR M. tuberculosis isolates were collected from various regions of Kazakhstan. Genotyping was performed using spoligotyping, and drug susceptibility testing (DST) to first-line anti-TB drugs was conducted using the BACTEC MGIT 960 system. DNA library preparation for WGS was carried out in accordance with the Illumina DNA Prep Reference Guide (Illumina, USA), and sequencing was performed using the Illumina NovaSeq 6000 platform. Genotyping results revealed a high prevalence of the Beijing lineage, identified in 85.5% (65/76) of MDR isolates. The remaining isolates belonged to the following lineages: LAM9 – 3 isolates (3.9%), LAM-RUS – 3 (3.9%), T1 – 2 (2.6%), H3 – 1 (1.3%), MANU2 – 1 (1.3%), and S– 1 (1.3%). The most frequent resistance pattern included resistance to isoniazid, rifampicin, and ethambutol, observed in 23 isolates (30.2%). Resistance to isoniazid, rifampicin, pyrazinamide, and ethambutol was detected in 21 isolates (27.6%), while resistance to isoniazid, rifampicin, and pyrazinamide was found in 17 cases (22.4%). Resistance to only isoniazid and rifampicin was identified in 15 isolates (19.7%). Beijing family strains of M.tuberculosis were identified as a dominant family and associated with a high risk of MDR-TB in Kazakhstan. Genomic analysis using databases CASTB, MTBseq, Mykrobe, and TB Profiler is in process and will allow determining the genetic profile of whole drug resistance. Grant references: This study was funded by a grant from Nazarbayev University under Collaborative Research Program №11022021CRP1511, U.K. Keywords: tuberculosis, drug resistance, MIRU-VNTR analysis, genotyping, M. tuberculosis. Global tuberculosis World Health Organization; 2025. WHO Publication; 2025/ A.Daniyarov, A.Molkenov, et all. Genomic Analysis of Multidrug-Resistant Mycobacterium tuberculosis Strains From Patients in Kazakhstan. Front Genet. 2021.
Aeromonas spp. are opportunistic pathogens that are widely distributed in water sources, with several species being associated with fish and human diseases. We have previously identified an Aeromonas AB005 isolate from diseased Acipencer baerii. This isolate was identified as A. hydrophila based on the 16S rRNA and gyrB gene sequences. However, this novel strain does not produce indole and tested negative for ornithine decarboxylase and d-xylose fermentation—differences that set it apart from typical A. hydrophila strains. In the present study, this strain was subjected to whole-genome sequencing and compared with the genomes of the type strain (Aeromonas hydrophila ATCC 7966T) and other Aeromonas spp. Comprehensive genome analysis suggests that AB005 represents a distinct species within the genus. The draft genome of the AB005 strain comprises 4,780,815 base pairs with a GC content of 61.2% and contains 6104 predicted protein-coding sequences along with numerous genes implicated in antibiotic resistance. The core/pan-genome analysis reveals extensive genetic diversity, indicative of a dynamic genomic structure. These findings collectively underscore the taxonomic distinction of the AB005 strain as a novel species and highlight its potential pathogenic implications in aquaculture and public health settings.