Central Asians are underrepresented in genomic research, limiting insights into their genetic history and disease risk. We established the Central Asian Genomic Diversity Project and sequenced whole genomes from 166 individuals across 20 Central Asian and Afghan Hazara groups. We identify marked differentiation driven by varying West/East Eurasian ancestry; Tajiks align with West Eurasians, Dungans with East Asians, and we report four geographically structured Turkic-related clusters, two Indo-European clines, and long-range migration events, including Siberian links in Hazaras and Sino-Tibetan ties in Dungans. Admixture dates cluster ∼650-1000 years ago, coinciding with the Song-Yuan era and Mongol expansion. We characterize distinct distributions of medically relevant variants and population-specific adaptation signatures across metabolic, immune, and neurological pathways, and illuminate shifts in subsistence practices correlated with trait-associated variation. We also detect Neanderthal-like and Denisovan-like segments that show group-specific associations with immunity, psychiatric risk, drug metabolism, and diabetes, underscoring the scientific imperative for a broader characterization of Central Asian evolutionary history and informing precision medicine.
A search for the modern descendants of the Neolithic population has been conducted using two datasets of Y-chromosome polymorphisms: literature data on the ancient population of Northeast Eurasia and our own data on 256 whole genomes of 11 indigenous peoples of the Russian Far East (Aleuts, Chukchi, Evens, Evenks, Itelmens, Koryaks, Nanais, Negidals, Nivkhs, Orochi, and Ulchi). Y-SNP analysis revealed that both Kyordyughen I (4200 YBP) and Kyordyughen II (4600 YВР) samples of the Yakutian Late Neolithic belong to the haplogroup N-L708 and lie on its two branches that diverged at 6200 YBP. A quarter (67 of 256) of the analysed samples, including Chukchi, Evens, Evenks, Itelmens, Koryaks, Nanais, Nivkhs, Orochi, and Ulchi, belong to the haplogroup N-L708 and are, to various extents, genetically related to the Neolithic Yakutian individuals. The most direct descendants of the famous Kyordyughen warrior (Kyordyughen I) are the indigenous peoples of Kamchatka and Chukotka (Chukchi, Koryaks, Evens). The divergence time of their Y-lineages (4300 ± 1000 YВР) is consistent with the radiocarbon dates for Kyordyughen I. Literary data on Y-STR polymorphism suggest that descendants of the Kyordyughen warrior dispersed far across North Asia. Ancient N-L708 carriers started to expand from Transbaikalia across Northeast Eurasia at ~7000 YBP. About 4000-3000 YВР, the lineages close to the Kyordyughen samples arrived in today’s Krasnoyarsk territory, Yakutia, Mongolia and Chukotka. Our findings are in good agreement with archaeological data and the autosomal genome-based modelling of the Kyordyughen warrior’s origin.
This article presents a population dataset of 27 Y- chromosomal short tandem repeat (Y-STR) loci obtained from 501 unrelated Uzbek males residing in southern Kazakhstan. Samples were collected from four geographic locations: Saryagash District (N = 79), Sayram District (N = 213), Taraz (N = 70), and Turkistan (N = 139). Genotyping was performed using the Yfiler™ Plus PCR Amplification Kit, and the resulting profiles were used to estimate haplotype frequencies, forensic parameters, and interpopulation genetic relationships. A total of 424 distinct haplotypes were identified, and summary diversity metrics were calculated for each regional group and for the pooled dataset. The overall haplotype diversity reached 0.9991, while discrimination capacity was 0.85. Allele frequency distributions for 23 single-copy loci and allele combination frequencies for multi-copy loci are provided, together with information on microvariants, null alleles, and copy number variations. Predicted haplogroup assignments derived from Y-STR haplotypes, AMOVA outputs, pairwise Rst matrices, median-joining networks, and multidimensional scaling plots based on overlapping marker sets for comparative populations, are included as part of the dataset. Haplogroup prediction based on Y-STR haplotypes showed that the paternal gene pool is mainly represented by haplogroups R1a (21.6%) and C2 (21.4%), followed by J2, R1b, N1, O2, and Q. These data expand the representation of Central Asian populations in forensic and population-genetic reference datasets and provide a resource for future studies of paternal lineage diversity, forensic reference data, interpopulation comparison, and comparative analyses of Y-chromosomal variation in Central Asia.
Although the social structure of Central Eurasian pastoral nomads has been described as fluid and ad hoc by historians and anthropologists, its potential genetic basis remains poorly understood. To evaluate whether kinship based social organization has biological foundations, we surveyed Kazak populations in Jetisuu, Kazakstan, for genomic diversity with special reference to their kinship structure. We generated genome-wide SNP data (∼750K) using GenoChip microarrays for 80 individuals from four Kazak clans in Jetisuu and 10 Kazaks from other regions. Our results reveal substantial concordance between genetic and genealogical data, with ∼64% of Jetisuu Kazaks sharing the same Y-chromosome haplotype, indicating they have a common paternal ancestry aligning with clan genealogies. By contrast, maternal lineages show remarkable heterogeneity, thereby reflecting female exogamy. Despite this sex-biased admixture pattern, autosomal SNPs reveal no pronounced population structure in Kazaks, suggesting genetic homogenization through ancient admixture followed by continuous gene flow between Kazak populations. Notably, Jetisuu Kazaks exhibit significantly reduced levels of runs of homozygosity (ROH) compared to their sedentary neighbors such as Turkmens, Tajiks, Uyghurs, and Uzbeks. This genomic signature likely results from clan exogamy, which functions as a biological mechanism to prevent inbreeding and maintain genetic diversity. Our findings demonstrate that nomadic social organization represents a sophisticated example of gene-culture co-evolution, where cultural practices have systematically shaped genetic patterns over centuries. These data provide new insights into historical pastoral nomadic societies and offer a nuanced perspective on anthropological debates about kinship authenticity in Central Eurasian nomads.
The underrepresentation of Central Asian genomic data has constrained our understanding of their demographic history and hindered advancements in precision medicine and health equity. Despite the region’s rich historical tapestry, characterized by numerous trans-Eurasian migrations following the advent of agriculture and pastoralism, the genetic contributions of ancient Eurasians to modern Central Asians remain poorly understood. To address this gap, we performed an anthropologically informed Central Asian Genomic Diversity Project and reported the results of pilot whole-genome sequencing work on 166 Central Asians and Afghanistan Hazaras (CAAH) from 20 populations to investigate their demographic history, local adaptation, medical relevance, and archaic introgression. Significant genetic differentiation among CAAH populations was revealed. Tajik, Karluks, Turkmen, and Uzbek individuals exhibited higher proportions of West Eurasian ancestry, whereas the Kyrgyz, Karakalpak, Uyghur, and Hazara populations presented increased ancestry related to ancient Northeast Asians. In contrast, Dungans demonstrated a predominance of East Asian-derived ancestry. Four Turkic-related genetic clusters corresponding to geographic distribution were identified, supporting the “Northeast Asia origin” hypothesis for Turkic groups. Additionally, two Indo-European genetic clines were detected, with Hazaras being notably isolated. Strong genetic affinities were observed between Hazaras and Altaic groups in Siberia and between Dungans and Sino-Tibetan-speaking East Asians, underscoring the impact of ancient long-distance migrations on Eurasian genetic diversity. The recent east-west admixture in CAAH was estimated to have occurred 23-31 generations ago, aligning with the Song and Yuan dynasties and the Mongol Empire period. The mutation spectra of candidate disease-causing variants and pharmacogenomic genes were characterized, indicating that differentiated demographic histories significantly influence the genetic architecture of diseases among different Central Asians. Differential post-admixture adaptation signatures identified in the four genetically distinct groups have substantial effects on immune, metabolic, neural, and physical traits. Shifts in subsistence strategies significantly shaped the genetic architecture of complex traits in Central Asians. Neanderthal-like sequences exhibited varying phenotypic effects across genetically distinct CAAH strains, including susceptibility to immune and psychiatric conditions in West Eurasian-biased CAAH individuals and drug metabolism in East Eurasian-biased CAAH individuals. Denisovan-like segments were primarily linked to type 2 diabetes, etc. This research on Central Asian genomic diversity enhances the understanding of their evolutionary history and admixture events, promoting health equity and advancing precision medicine initiatives. He et al. conducted a pilot study on the Central Asian Genomic Diversity Project, utilizing whole-genome sequencing of 166 individuals from 20 Central Asian populations. They identified fine-scale population substructures shaped by complex ancient trans-Eurasian migration and admixture processes. Their comprehensive analysis revealed post-admixture adaptations and archaic introgressions, shedding light on demographic events that influenced medically relevant mutation spectra and adaptations affecting immune, metabolic, neural, and physical traits. Neanderthal introgression segments significantly influence phenotypic traits, including susceptibility to immune and psychiatric disorders, whereas Denisovan-derived sequences have effects on disease susceptibility. This work advances our understanding of Central Asian genomic diversity and evolutionary history as well as their implications for health.
For decades, there has been scientific interest in the variation and geographic distribution of paternal lineages associated with the human Y chromosome. However, the relevant data have been dispersed across numerous publications, making it difficult to consolidate. Additionally, understanding the relationships between different variants, and the tools used to analyze them, have evolved over time, further complicating efforts to harmonize this information. The Universal Y-SNP Database (UYSD) marks a substantial advancement by providing a comprehensive and accessible platform for Y-SNP and haplogroup data from populations around the world. UYSD harmonizes diverse datasets into a unified repository, facilitating the exploration of global Y-chromosomal variation. The platform handles data generated with both high- and low-throughput technology and is compatible with the automated analysis software tool, Yleaf v3. Key functionalities include the ability to: i) visualize haplogroup distributions on an interactive world map, ii) estimate haplogroup frequencies in geographic regions with sparse data through interpolation, and iii) display detailed phylogenetic trees of Y-chromosomal haplogroups. Currently, UYSD encompasses data from over 6,600 males across 27 populations. This dataset largely aligns with known global Y-haplogroup patterns, but also reveals unexplored finer-scale geographic variations. While the present dataset is largely European-centered, UYSD is designed for ongoing expansion by the scientific community, aiming to include more global data and higher-resolution population sequencing data. The platform thus offers valuable insights into human genetic diversity and migration patterns, serving several fields of research such as: human population genetics, genetic anthropology, ancient DNA analysis and forensic genetics.
The study of genetic history through Y-chromosome polymorphism in modern and ancient populations helps reconstruct ethnogenesis. Research on the Nogai gene pool, as descendants of ancient nomads, allows tracing migration routs in the Eurasian steppe corridor. The Nogai ethnic group formed in the Middle Ages, as a result of interactions of the populations of the Central Asian, West Siberian, Ural-Volga, North Caucasian, and East European historical and cultural regions. A unified expanded panel of Y-chromosome haplogroups was used to study four main populations of the Nogai people: Astrakhan, Kuban, Stavropol, and Karanonogais. The total sample (N=368) revealed a diverse spectrum of 32 Y-haplogroups, with no single haplogroup dominating in any of these groups. Astrakhan Nogais show high frequencies of C-M217 and J-M172, while Kuban Nogais have an exceptionally high frequency of R-M198 and its “European” subclade R-M458. Karanogais display a distinct combination of R-M198(xM458), N-M178, R-Y13887, and J-M172(xM67, M12). Stavropol Nogais combine R-M198(xM458) and C-M217. The Nogai gene pool is characterized by the accumulation of “steppe” Y-chromosome variants, which is obviously related to the tribes of the steppe belt of Eurasia, probably of pre-Turkic origin. Increased J-M172 frequency may result from ancient Neolithic migrations along the Caspian coast and later nomadic movements. “Steppe” variants of the Nogai Y-chromosome population were also identified in samples from burials at the Mamai-Gora burial ground in the Northern Black Sea region (15th century AD). The greatest genetic similarity is observed between the Nogais and other Turkic peoples of the Caucasus, who are located on the periphery of the “Nogai” cluster within the framework of the Caucasian genetic landscape.
Central Asians are underrepresented in genomic research, limiting insights into their history and medical risk. We established the Central Asian Genomic Diversity Project and sequenced whole genomes of 166 individuals from 20 Central Asian and Afghanistan Hazara groups. Analyses reveal marked differentiation driven by varying West/East Eurasian ancestry; Tajiks align with West Eurasians, Dungans with East Asians. We identify four geographically structured Turkic-related clusters, two Indo-European clines, and long-range migration events, including Siberian links in Hazaras and Sino-Tibetan ties in Dungans. Admixture dates cluster ~650-1,000 years ago, coinciding with the Song-Yuan era and Mongol expansion. Variant discovery reveals distinct distributions of medically relevant variants and group-specific selection signals in metabolic, immune, and neurological pathways. Subsistence practice shifts correlate with trait-associated variation. Neanderthal-like and Denisovan-like segments show group-specific associations with immunity, psychiatric risk, drug metabolism, and type 2 diabetes. These data clarify Central Asian evolutionary history and inform precision medicine. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Protocols ### Funding Statement National Natural Science Foundation of China grant 82402203 (GLH) Science Committee of the Ministry of Education and Science of the Republic of Kazakhstan grant BR18574101 (MZ) National Natural Science Foundation of China grant 82202078 (MGW) Faculty Development Competitive Research Grants Programs of Nazarbayev University grant SST2019012 (MZ) Major Project of the National Social Science Foundation of China grant 23&ZD203 (LHW, GLH) Open Project of the Key Laboratory of Forensic Genetics of the Ministry of Public Security grant 2022FGKFKT05 (GLH) Open Project of the Key Laboratory of Forensic Genetics of the Ministry of Public Security grant 2024FGKFKT02 (MGW) Center for Archaeological Science of Sichuan University grant 23SASA01 (GLH) Center for Archaeological Science of Sichuan University grant 24SASB03 (GLH) Sichuan Science and Technology Program grant 2024NSFSC1518 (GLH) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This work was developed through collaborations among anthropological, genetic, and linguistic scientists within the CAGDP Consortium, aiming to improve the representation of the genomic diversity of ethnolinguistically diverse populations and to provide comprehensive insights into their population history, biological adaptation, archaic introgression, and medical implications. The study was approved by the Medical Ethics Committee of the National Centre for Biotechnology of Kazakhstan (protocol No. 2 of 10 June 2020) and the Nazarbayev University Institutional Research Ethics Committee (protocol No. 17 of 16 April 2019), and all procedures were conducted in accordance with the Declaration of Helsinki (90). The participants were selected based on the criterion that all four grandparents were indigenous to the respective ethnic groups. Informed consent was obtained from all participants. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The raw data of 166 individuals were submitted to the Genome Variation Map database under accession number GVM000900 (https://bigd.big.ac.cn/gvm/getProjectDetail?Project=GVM000900). All the data used here are included in the supplementary materials. Reference populations can be found in publicly available datasets, such as the Human Genetic Diversity Project dataset, Oceania genomic resource, and the Allen Ancient DNA Resource (HO and 1240K datasets) from the David Reich Lab (https://reich.hms.harvard.edu/datasets).
Because of limited availability of ancient genomes, the genetic history of prehistoric Inner Asian hunter-gatherers remains incomplete, especially for eastern Kazakhstan where the Eurasian Steppe meets mountain forests of Inner Asia. Here we report genome-wide data of two Early Neolithic (EN) hunter-gatherers and 19 Middle-Late Bronze Age (MLBA) pastoralists, from the site of Koken in the Upper Irtysh River region in eastern Kazakhstan. We find that the two EN individuals differed in their genetic profiles and yet were second-degree relatives. They were genetically most similar to subsequent Neolithic individuals in the Irtysh region, while contemporaneous hunter-gatherers from the Tobol-Ishim and Upper Ob River regions had distinct genetic profiles, likely influenced by riverine geography. The Koken MLBA individuals were genetically similar to other MLBA steppe pastoralists, while genetic outliers provide evidence of two distinct trajectories of admixture with local hunter-gatherer populations. These findings illuminate the dynamic population structure of Inner Asian hunter-gatherers and their genetic legacy in subsequent pastoralist populations.
The North Eurasian forest and forest-steppe zones have sustained millennia of sociocultural connections among northern peoples, but much of their history is poorly understood. In particular, the genomic formation of populations that speak Uralic and Yeniseian languages today is unknown. Here, by generating genome-wide data for 180 ancient individuals spanning this region, we show that the Early-to-Mid-Holocene hunter-gatherers harboured a continuous gradient of ancestry from fully European-related in the Baltic, to fully East Asian-related in the Transbaikal. Contemporaneous groups in Northeast Siberia were off-gradient and descended from a population that was the primary source for Native Americans, which then mixed with populations of Inland East Asia and the Amur River Basin to produce two populations whose expansion coincided with the collapse of pre-Bronze Age population structure. Ancestry from the first population, Cis-Baikal Late Neolithic-Bronze Age (Cisbaikal_LNBA), is associated with Yeniseian-speaking groups and those that admixed with them, and ancestry from the second, Yakutia Late Neolithic-Bronze Age (Yakutia_LNBA), is associated with migrations of prehistoric Uralic speakers. We show that Yakutia_LNBA first dispersed westwards from the Lena River Basin around 4,000 years ago into the Altai-Sayan region and into West Siberian communities associated with Seima-Turbino metallurgy-a suite of advanced bronze casting techniques that expanded explosively from the Altai1. The 16 Seima-Turbino period individuals were diverse in their ancestry, also harbouring DNA from Indo-Iranian-associated pastoralists and from a range of hunter-gatherer groups. Thus, both cultural transmission and migration were key to the Seima-Turbino phenomenon, which was involved in the initial spread of early Uralic-speaking communities.
IntroductionThe Y chromosome, transmitted exclusively through the paternal line, is a well-established tool for verifying genealogical data. The Kazakh tribe Zhetiru in Kazakhstan, comprising seven clans, has conflicting historical and genealogical narratives regarding its origin—either as a union of seven independent clans or as descendants of a single common ancestor. A detailed genetic investigation has not yet addressed this question.Methods350 male volunteers from the Zhetiru tribe were analyzed using 23 Y-STR loci and 17 Y-SNPs. We calculated genetic distances using Arlequin and STRAF, and explored genetic structure with median-joining networks using a comparative dataset of over 3,000 Kazakh individuals.ResultsAt the tribal level, haplotype diversity (0.997) and haplogroup diversity (0.91) are high. However, at the clan level, haplotypic diversity decreases, revealing clear founder effects in the main haplogroups of Kerderi (R1a1a), Kereit (N1a2), Tama (C2a1a3), and Teleu (J2a2). The genetic structures of Zhagalbaily, Ramadan, and Tabyn indicate additional sub-clan founders. The ages of key clusters suggest stable genetic lineages for over 1,000 years. Zhetiru clans do not form a distinct genetic cluster among Kazakh tribes but demonstrate genetic affinities with others.ConclusionThis study demonstrates the effective application of genetic genealogy approaches in verifying historical and genealogical records concerning the Zhetiru tribe and determining its origin from distinct, genetically independent clans.
Congenital spinal deformities (CSDs) are rare but severe conditions caused by abnormalities in vertebral development during embryogenesis. These deformities, including scoliosis, kyphosis, and lordosis, significantly impair patients’ quality of life and present challenges in diagnosis and treatment. This review integrates genetic, molecular, and developmental insights to provide a comprehensive framework for classifying and understanding CSDs. Traditional classification systems based on morphological criteria, such as failures in vertebral formation, segmentation, or mixed defects, are evaluated alongside newer molecular-genetic approaches. Advances in genetic technologies, including whole-exome sequencing, have identified critical genes and pathways involved in somitogenesis and sclerotome differentiation, such as TBX6, DLL3, and PAX1, as well as key signaling pathways like Wnt, Notch, Hedgehog, BMP, and TGF-β. These pathways regulate vertebral development, and their disruption leads to skeletal abnormalities. The review highlights the potential of molecular classifications based on genetic mutations and developmental stage-specific defects to enhance diagnostic precision and therapeutic strategies. Early diagnosis using non-invasive prenatal testing (NIPT) and emerging tools like CRISPR-Cas9 gene editing offer promising but ethically complex avenues for intervention. Limitations in current classifications and the need for further research into epigenetic and environmental factors are discussed. This study underscores the importance of integrating molecular genetics into clinical practice to improve outcomes for patients with CSDs.
Y-chromosomal short tandem repeats (Y-STRs) serve as essential markers in forensic genetics, population studies, and paternal lineage reconstruction due to their strict uniparental inheritance and high discriminatory power. Despite their global relevance, Central Asian populations, particularly the Kyrgyz, remain underrepresented in major Y-STR reference databases. These population data represent 23 Y-STR loci from 514 unrelated Kyrgyz males sampled from three geographically distinct regions: Northern East (N = 134), Northern West (N = 183), and Southern Kyrgyzstan (N = 197). Genotyping was conducted using the PowerPlex Y23 System, and the resulting dataset has been submitted to the Y Chromosome Haplotype Reference Database (YHRD) to strengthen forensic and anthropological research in the region. A total of 346 unique haplotypes were identified, demonstrating high haplotype diversity (HD = 0.981–0.990) and discrimination capacity (64–70 %). AMOVA analysis indicates that the division of Kyrgyz populations into northern and southern groups does not accurately represent their genetic structure, as over 99 % of genetic variation is distributed within subpopulations, indicating weak differentiation and substantial shared paternal ancestry among the regional Kyrgyz groups. The analysis also identifies four dominant haplogroup clusters (R1a, C2a, N1, and R1b), providing valuable insights into the historical and demographic dynamics of the Kyrgyz people. This dataset enhances our understanding of Kyrgyz genetic diversity, contributes to forensic applications, and fills a critical gap in population genetic research on Central Asian lineages.
This study presents a comprehensive analysis of 23 Y-STR data for the Merkit clan, a subgroup within the Kerey tribe of the Kazakh people. A total of 64 complete haplotypes were generated using the PowerPlex Y23 System. The data obtained using 23 Y-STR markers has been submitted to the Y Chromosome Haplotype Reference Database (YHRD) at yhrd.org, which will significantly enhance the forensic database for the Kazakh population in Kazakhstan. The research focuses on the distribution of haplotypes within the clan and their genealogical lines, which were visualized using a Median-joining network and Multidimensional scaling plot. The study identifies four distinct haplogroup clusters, revealing important insights into the genetic makeup and historical lineage of the Merkits. This dataset not only enriches our understanding of Kazakh genetic structure but also holds significant value for anthropological and population genetic research, as well as for forensic genetics. This work bridges a notable gap in genetic research on the Merkit clan, contributing to a deeper understanding of Central Asian nomadic tribes.
This study investigates the Y-chromosome genetic diversity of the Turkmen population in Turkmenistan, analyzing 23 Y-STR loci for the first time in a sample of 100 individuals. Combined with comparative data from Turkmen populations in Afghanistan, Iran, Iraq, Russia, and Uzbekistan, this analysis offers insights into the genetic structure and relationships among Turkmen populations across regions across Central Asia and the Near East. High haplotype diversity in the Turkmen of Turkmenistan is shaped by founder effects (lineage expansions) from distinct haplogroups, with haplogroups Q and R1a predominating. Subhaplogroups Q1a and Q1b identified in Turkmenistan trace back to ancient Y-chromosome lineages from the Bronze Age. Comparative analyses, including genetic distance (RST), median-joining network, and multidimensional scaling (MDS), highlight the genetic proximity of the Turkmen in Turkmenistan to those in Afghanistan and Iran, while Iraqi Turkmen display unique characteristics, aligning with Near Eastern populations. This study underscores the Central Asian genetic affinity across most Turkmen populations. It demonstrates the value of deep-sequencing Y-chromosome data in tracing the patrilineal history of Central Asia for future studies. These findings contribute to a more comprehensive understanding of Turkmen genetic ancestry and add new data to the ongoing study of Central Asian population genetics.
Objectives The collection of genotype data was conducted as an essential part of a pivotal research project with the goal of examining the genetic variability of skin, hair, and iris color among the Kazakh population. The data has practical application in the field of forensic DNA phenotyping (FDA). Due to the limited size of forensic databases from Central Asia (Kazakhstan), it is practically impossible to obtain an individual identification result based on forensic profiling of short tandem repeats (STRs). However, the pervasive use of the FDA necessitates validation of the currently employed set of genetic markers in a variety of global populations. No such data existed for the Kazakhs. The Phenotype Expert kit (DNA Research Center, LLC, Russia) was used for the first time in this study to collect data. Data description The present study provides genotype data for a total of 60 SNP genetic markers, which were analyzed in a sample of 515 ethnic Kazakhs. The dataset comprises a total of 41 single nucleotide polymorphisms (SNPs) obtained from the HIrisPlex-S panel. Additionally, there are 4 SNPs specifically related to the AB0 gene, 1 marker associated with the AMELX/Y genes, and 14 SNPs corresponding to the primary haplogroups of the Y chromosome. The aforementioned data could prove valuable to researchers with an interest in investigating genetic variability and making predictions about phenotype based on eye color, hair color, skin color, AB0 blood group, gender, and biogeographic origin within the male lineage.