Ubykh people inhabited the area of what now is Greater Sochi, one time being neighbors of the Abkhaz and Adyghe. In 1864, after the Caucasian War, a small group of Ubykh survivors migrated to Turkey. The Ubykh language is a branch of the Northwest Caucasian (Abkhazo-Adyghe) languages, which is either placed in the Abkhaz or Adyghe subgroup or distinguished in a separate subgroup, according to various classifications. All this makes Ubykhs an important but genetically unexplored part of the gene pool of the Caucasus population. We collected an exclusive sample (N = 36) of Ubykh individuals, mainly from their diaspora in Turkey. Application of a high resolution panel of Y-chromosome markers (59 SNP and 17 STR) revealed the major West Caucasus haplogroup G2 (75% of the Ubykh population) and pan-Eurasian haplogroup R1a (19%). Within the G2 haplogroup, we identified 14 branches and genotyped 14 branch-defining markers. The frequencies of the same 14 markers in 19 other populations of the Caucasus were included for comparison. Both the genetic distances and the gene geographical map of the most frequent subhaplogroup (defined by the YY1215 marker at the GRCh37 position 8903699, comprising 50% of the Ubykh sample) revealed similarity of the Ubykh and Adyghe populations. Other Abkhaz-Adyghe-speaking populations appear to be at a greater genetic distance from the Ubykhs, similar to the distance to the Turkic-speaking groups of the North Caucasus. The Ossetian gene pool has little in common with Ubykhs, thus contradicting the hypothesis of the Alanian substrate in Ubykhs. Other populations of the Caucasus and Transcaucasia (of Nakh-Daghestanian and Kartvelian linguistic groups) do not show apparent genetic similarities to Ubykhs.
We studied the Y-chromosome pool of the ethnic Russian population of Novgorod oblast (Russia) by 49 SNP and 17 STR markers. The total sample (N = 191) consists of four populations of the Novgorod region, including its southwestern (Shelon Pyatina) and eastern (Bezhetsk Pyatina) parts. Altogether, these four populations represent both the area of the Sopki archaeological culture (supposedly linked with the Novgorod Slovens tribe known from the chronicles) and the area of the Long Barrows culture (supposedly linked with the Krivichi Slavic tribe or with Balts). The pronounced genetic differences between southern and northern Russian populations are well known from previous studies; however, the Novgorod gene pool turned out to be neither northern nor southern, but a representative of the intermediate buffer zone. This zone was identified in this study and included a set of regional Russian populations from Pskov in the west to Kostroma in the east. All four studied populations of Novgorod region are genetically similar. The minor differences among them might represent the medieval Slavic migrations along the rivers, which survived despite the massive demographic shifts during the following history. Haplogroup N3 comprises one-fifth of the Novgorod pool of paternal lineages, with conditionally “Finnic” N3a4 and conditionally “East Baltic Sea Coast” N3a3 clades being almost equally frequent. The N3a3 phylogenetic network revealed the specific “Balto-Slavic” cluster of STR haplotypes, which is frequent in Baltic-speaking Lithuanians but infrequent in Finno-Ugric speaking Estonians. The Novgorod haplotypes lie outside this cluster, indicating that the Novgorod population received both N3a3 and N3a4 from Finno-Ugric speaking populations of the region, which, in turn, acquired the Mesolithic gene pool of the Northeastern Europe.
The Upper Volga region was an area of contacts of Finno-Ugric, Slavic, and Scandinavian speaking populations in the 8th–10th centuries AD. However, their role in the formation of the contemporary gene pool of the Russian population of the region is largely unknown. To answer this question, we studied four populations of Yaroslavl oblast (N = 132) by a wide panel of STR and SNP markers of the Y-chromosome. Two of the studied populations appear to be genetically similar: the indigenous Russian population of Yaroslavl oblast and population of Katskari are characterized by the same major haplogroup, R-M198 (xM458). Haplogroup R-M458 composes more than half of Sitskari’s gene pool. The major haplogroup in the gene pool of the population of the ancient town of Mologa is N-M178. Subtyping N-M178 by newest “genomeera” Y-SNP markers showed different pathways of entering this haplogroup into the gene pools of Yaroslavl Volga region populations. The majority of Russian populations have subvariant N3a3-CTS10760; the regular sample of Yaroslavl oblast is equally represented by subvariants N3a3-CTS10760 and N3a4-Z1936, while subvariant N3a4-Z1936 predominates in the gene pool of population of Mologa. This N3a4-Z1936 haplogroup is common among the population of the north of Eastern Europe and the Volga-Ural region. The obtained results indicate preservation of the Finno-Ugric component in the gene pool of population of Mologa and a contribution of Slavic colonization in the formation of the gene pool of the Yaroslavl Volga region populations and make it possible to hypothesize the genetic contribution of the “downstream” (Rostov- Suzdal) rather than “upstream” (Novgorod) Slavic migration wave.
Population biobanks are collections of thoroughly annotated biological material stored for many years. Population biobanks are a valuable resource for both basic science and applied research and are essential for extensive analysis of gene pools. Population biobanks make it possible to carry out fundamental studies of the genetic structure of populations, explore their genetic processes, and reconstruct their genetic history. The importance of biobanks for applied research is no less significant: they are essential for development of personalized medicine and genetic ecological monitoring of populations and are in high demand in forensic science. Establishment of an efficient and representative biobank requires strict observance of the principles of sample selection in populations, protocols of DNA extraction, quality control, and storage and documentation of biological materials. We reviewed regional biobanks and presented the organizational model of population biobank establishment based on the Biobank of Indigenous Population of Northern Eurasia created under supervision of E.V. Balanovska and O.P. Balanovsky. The results obtained using the biobanks in transdisciplinary research and prospective applications for the purposes of genogeography, genomic medicine, and forensic science are presented.
Nonrecombinant portions of the genome, Y chromosome and mitochondrial DNA, are widely used for research on human population gene pools and reconstruction of their history. These systems allow the genetic dating of clusters of emerging haplotypes. The main method for age estimations is ρ statistics, which is an average number of mutations from founder haplotype to all modern-day haplotypes. A researcher can estimate the age of the cluster by multiplying this number by the mutation rate. The second method of estimation, ASD, is used for STR haplotypes of the Y chromosome and is based on the squared difference in the number of repeats. In addition to the methods of calculation, methods of Bayesian modeling assume a new significance. They have greater computational cost and complexity, but they allow obtaining an a posteriori distribution of the value of interest that is the most consistent with experimental data. The mutation rate must be known for both calculation methods and modeling methods. It can be determined either during the analysis of lineages or by providing calibration points based on populations with known formation time. These two approaches resulted in rate estimations for Y-chromosomal STR haplotypes with threefold difference. This contradiction was only recently refuted through the use of sequence data for the complete Y chromosome; “whole-genomic” rates of single nucleotide mutations obtained by both methods are mutually consistent and mark the area of application for different rates of STR markers. An issue even more crucial than that of the rates is correlation of the reconstructed history of the haplogroup (a cluster of haplotypes) and the history of the population. Although the need for distinguishing “lineage history” and “population history” arose in the earliest days of phylogeographic research, reconstructing the population history using genetic dating requires a number of methods and conditions. It is known that population history events leave distinct traces in the history of haplogroups only under certain demographic conditions. Direct identification of national history with the history of its occurring haplogroups is inappropriate and is avoided in population genetic studies, although because of its simplicity and attractiveness it is a constant temptation for researchers. An example of DNA genealogy, an amateur field that went beyond the borders of even citizen science and is consistently using the principle of equating haplogroup with lineage and population, which leads to absurd results (e.g., Eurasia as an origin of humankind), can serve as a warning against a simplified approach for interpretation of genetic dating results.
Siberian Tatars form the largest Turkic-speaking ethnic group in Western Siberia. The group has a complex hierarchical system of ethnographically diverse populations. Five subethnic groups of Tobol-Irtysh Siberian Tatars (N = 388 samples) have been analyzed for 50 informative Y-chromosomal SNPs. The subethnic groups have been found to be extremely genetically diverse (FST = 21%), so the Siberian Tatars form one of the strongly differentiated ethnic gene pools in Siberia and Central Asia. Every method employed in our studies indicates that different subethnic groups formed in different ways. The gene pool of Isker-Tobol Tatars descended from the local Siberian indigenous population and an intense, albeit relatively recent gene influx from Northeastern Europe. The gene pool of Yalutorovsky Tatars is determined by the Western Asian genetic component. The subethnic group of Siberian Bukhar Tatars is the closest to the gene pool of the Western Caucasus population. Ishtyak-Tokuz Tatars have preserved the genetic legacy of Paleo-Siberians, which connects them with populations from Southern, Western, and Central Siberia. The gene pool of the most isolated Zabolotny (Yaskolbinsky) Tatars is closest to Ugric peoples of Western Siberia and Samoyeds of the Northern Urals. Only two out of five Siberian Tatar groups studied show partial genetic similarity to other populations calling themselves Tatars: Isker-Tobol Siberian Tatars are slightly similar to Kazan Tatars, and Yalutorovsky Siberian Tatars, to Crimean Tatars. The approach based on the full sequencing of the Y chromosome reveals only a weak (2%) Central Asian genetic trace in the Siberian Tatar gene pool, dated to 900 years ago. Hence, the Mongolian hypothesis of the origin of Siberian Tatars is not supported in genetic perspective.
Yu. P. Altukhov suggested that heterozygosity is an indicator of the state of the gene pool. The idea and a linked concept of genetic ecological monitoring were applied to a new dataset on mtDNA variation in East European ethnic groups. Haplotype diversity (an analog of the average heterozygosity) was shown to gradually decrease northwards. Since a similar trend is known for population density, interlinked changes were assumed for a set of parameters, which were ordered to form a causative chain: latitude increases, land productivity decreases, population density decreases, effective population size decreases, isolation of subpopulations increases, genetic drift increases, and mtDNA haplotype diversity decreases. An increase in genetic drift increases the random inbreeding rate and, consequently, the genetic load. This was confirmed by a significant correlation observed between the incidence of autosomal recessive hereditary diseases and mtDNA haplotype diversity. Based on the findings, mtDNA was assumed to provide an informative genetic system for genetic ecological monitoring; e.g., analyzing the ecology-driven changes in the gene pool.