The indigenous populations of Udmurtia (five populations of Udmurts and Besermyans) have been studied for the first time by the genome-wide array. The populations were studied in the context of all Finno-Ugric and other surrounding ethnic groups. The total analyzed dataset included 728 individual genomes from 76 populations. The ADMIXTURE analysis of ancestral components identified the specific “Udmurt” component comprising nearby 100 % of genomes of all studied Udmurt individuals and the lion’s share of genomes of the studied Besermyans. The second component in Besermyan was the “White Sea” one predominating in Komi Zyryans and northernmost populations of ethnic Russians. The three independent methods - ADMIXTURE, F, and PCA - demonstrated that Udmurts and Besermyans are genetically closest to Komi-Permyaks, followed by Mari, Chuvash, Bashkirs, and Volga Tatars. We also characterized the pharmacogenetic status of Udmurtia population by estimating frequencies of 45 pharmacogenetic markers. The cartographic analysis revealed that genetic markers used in pharmacogenetic recommendations were observed in Besermyan and Udmurt populations at frequencies close to the frequencies in indigenous groups from Volga-Ural region but not from the more remote regions. Thus, within the revealed area of genetic similarity the same pharmacogenetic protocols can be applied. Here we publish the maps of distribution of the “Udmurt” ancestral components in the area of European Russia and Ural. The maps can be used by scholars in humanitarian fields in their own studies to trace genetically the interactions of Udmurts and Besermyans with other ethnic groups.
Introduction. Historical sources, as well as ethnographic, anthropological and linguistic data, speak of a significant influence of the Mongol-speaking tribes on the ethnogenesis of the Tuvans. Instead, the degree of Mongolian influence on the gene pool can only be assessed in molecular genetic studies. In this work, according to the data of complete sequencing of the C2-M217 haplogroup, a population screening of the Y-gene pool of the most numerous Tuvan tribal group Mongush was carried out. Materials and methods. DNA isolated from venous blood samples of 98 representatives of the Mongush tribal group collected in four regions of the Republic of Tyva was analyzed. Based on the full sequencing of the C2-M217 haplogroup and bioinformatic analysis, a phylogenetic analysis was carried out, a phylogenetic tree was built, and the age of the subhaplogroups was calculated. Results and discussion. It was found that the Central Asian haplogroup C2-M217 is represented in all representatives of the Mongush tribal group by only one line – the subhaplogroup C2a1a2a2a2-SK1066. Its presence in the gene pool may be associated with the mass migration of Mongol-speaking tribes in the 12th–14th centuries, when the territory of Tuva came under the rule of Genghis Khan. At the same time, this subhaplogroup was found only in the samples of the Chaa-Khol and Barun-Khemchik kozhuuns with frequencies of 12% and 2%, respectively; in the gene pools of the Tandyn and Erzin kozhuuns, the Central Asian haplogroup C2-M217 was not found. Phylogenetic analysis based on the full sequencing of the C2-M217 haplogroup made it possible to calculate the age of the C2a1a2a2a2-SK1066 sub-haplogroup – it was about 900 years. The predominance of South Siberian Neolithic haplogroups (Q1b-YP1691, N1a2-L666, N3a5a-F4205) in the gene pool of Tuvans is consistent with the data of anthropologists that it was the South Siberian layer that played the main role in their ethnogenesis, the Central Asian contribution belongs to a later time. Conclusions. According to the full sequencing of the Central Asian haplogroup C2-M217 among representatives of the Tuvan Mongush tribal group, the expansion of the Mongol-speaking tribes into Central Asia, which had a great cultural, economic and linguistic impact on the population of Tuva, was not so significant reflected in the gene pool of Tuvans. The study confirmed the data of anthropologists about the later Central Asian contribution to the ethnogenesis of the Tuvans in comparison with the earlier and much more significant South Siberian one.
Currently available genetic tools effectively distinguish between different continental origins. However, North Eurasia, which constitutes one-third of the world’s largest continent, remains severely underrepresented. The dataset used in this study represents 266 populations from 12 North Eurasian countries, including most of the ethnic diversity across Russia’s vast territory. A total of 1,883 samples were genotyped using the Illumina Infinium Omni5Exome-4 v1.3 BeadChip. Three principal components were computed for the entire dataset using three iterations for outlier removal. It allowed the merging of 266 populations into larger groups while maintaining intragroup homogeneity, so 29 ethnic geographic groups were formed that were genetically distinguishable enough to trace individual ancestry. Several feature selection methods, including the random forest algorithm, were tested to estimate the number of genetic markers needed to differentiate between the groups; 5,229 ancestry-informative SNPs were selected. We tested various classifiers supporting multiple classes and output values for each class that could be interpreted as probabilities. The logistic regression was chosen as the best mathematical model for predicting ancestral populations. The machine learning algorithm for inferring an ancestral ethnic geographic group was implemented in the original software “Homeland” fitted with the interface module, the prediction module, and the cartographic module. Examples of geographic maps showing the likelihood of geographic ancestry for individuals from different regions of North Eurasia are provided. Validating methods show that the highest number of ethnic geographic group predictions with almost absolute accuracy and sensitivity was observed for South and Central Siberia, Far East, and Kamchatka. The total accuracy of prediction of one of 29 ethnic geographic groups reached 71%. The proposed method can be employed to predict ancestries from the populations of Russia and its neighbor states. It can be used for the needs of forensic science and genetic genealogy.
The study of population frequencies of rare clinically significant alleles is a prerequisite of the development of personalized medicine. We performed genotyping of 1785 DNA samples from representatives of Russian populations according to 10 benign polymorphic markers of two genes involved in oncogenesis: 3 variants of the CDKN2A gene ( rs3731249 , rs116150891 , and rs6413464 ) and 7 markers of the RB1 gene ( rs149800437 , rs147754935 , rs183898408 , rs146897002 , rs4151539 , rs187912365 , and rs144668210 ). Genotyping was performed using the Illumina biochip test system. The sample covered 28 populations of the Russian Federation and neighboring countries, which were later combined into 3 groups (Asian, European, and Caucasian). The information from the ALFA (NCBI) project was used as reference for the frequencies of these polymorphisms in the Asian and European populations. It was shown that rare alleles in 8 of 10 studied polymorphic markers are presented in Russian populations of European and Caucasian origin with frequencies that are tens and hundreds of times higher than the available data for Western European populations, and in Russian Asian populations, alternative alleles of 5 markers absent in the Asian population of the ALFA project were found. In the subpopulation of Astrakhan Tatars, exceptionally high frequencies of rare alleles were identified; this requires further study.
Results and discussion. Y-chromosomes of the Kamchatkan Chukchi show a low diversity, being dominated by two variants N3a5b-B202 (57%) and N3*-M178* (19%), both virtually absent in the neighboring groups of Kamchatka. Phylogenetic analysis of the major lineages reveals some similarity with the populations of Chukotka (including Yupik Eskimo) at the level of clusters and closely matching haplotypes. More likely, this similarity can be explained by the gene flow from the Chukchi homeland to the south (i.e. to the northern regions of Kamchatka) over the past two centuries. The phylogenetic analysis of the N3a5b-B202 haplogroup revealed a correspondence between its two NGS-identified sub-branches (N3a5b1-B203, N3a5b2-B204) and two clusters found using STR markers. This concordance made it possible to compare the age estimates of these haplogroups obtained by different methods; the ages fell within the interval of 0.5-1.5 thousand years ago. N3*-M178* lineage is represented by a single STR-cluster (~800 years old) for which no “full-genome” analogue has yet been identified. The rest of the Chukchi paternal gene pool includes haplogroups C2-M48x (SK1066) (7%) and Q-M242 (7%), which are also present in the neighboring populations, primarily Koryaks, as well as Evens and Evenks. Conclusions. As expected, due to historical considerations, Chukchi of Kamchatka are most similar in Y-chromosome to the Chukotkan populations. The dating of the major branches of haplogroup N3 found in the Chukchi indicates a likely population growth within the past 0.5-1.5 thousand years. There is also a minor component in Chukchi, shared with the populations of Kamchatka, but its strict dating and origin is still unclear: it can be attributed both to the shared ancient NE Asian ancestry and to a recent admixture with local Kamchatkan groups.
Due to the low specificity and sensitivity of non-invasive clinical tests trehalose malabsorption remained out of sight of gastroenterologists. Therefore, the specialists regard this disorder as rare. Trehalose became widely used in the food industry as a harmless sucrose substitute, sweetener and stabilizer. After the discovery of the trehalase gene (rs2276064 TREH), it was found that the A*TREH allele is the determinant of the disaccharide absorption disorders, and the allele's carriership may be high in some groups. There is not enough information on the A*TREH frequency in the population of Russia. The aim of the study was to analyze the allele and genotype frequencies of the trehalase gene (rs2276064 TREH) in the main population groups of the Russian Federation and neighboring countries. Methods. DNA samples from 1146 unrelated subjects belonging to 21 population groups of Russia, Azerbaijan, Tajikistan and Mongolia were genotyped by the two following methods: 1) using the Infinium iSelect HD Custom Genotyping BeadChip (Illumina, USA) on the iScan platform; 2) by the real time polymer-chain reaction (PCR) method on the Bio-Rad CFX96 Touch amplifier. Results. It has been found that on the territory of the Russian Federation the frequency of the A*TREH allele increases from the west to the east. The frequencies are lowest in the groups of Russians and Finns of the Northwest (0.01-0.03), up to 0.07 in the populations of Central Russia and the Volga region, and even higher toward the Southern Urals (Bashkirs 0.15), in the Transurals and Southern Siberia (0.19 in the Altai people, 0.30 in the Tuvinians and Mongols). Up to 1% of the population of the European part of the Russian Federation have the AA*TREH genotype (i.e. trehalose intolerance in phenotype), and up to 15% (GA*TREH genotype) have a reduced ability to absorb the disaccharide. In the Asian part of the country (Siberia, Altai, Baikal) the genotypes carriers constitute up to 12 and 46% respectively. Conclusion. Trehalose malabsorbtion is an underappreciated problem of particular practical importance for regions with high concentrations of indigenous population (Yakutia-Sakha, Buryatia, Tyva, etc.). It would be feasible to consider food labelling of trehalose.
We studied the diversity of Y-chromosomal haplogroup N-L666 in Southern Siberia, where this lineage itself is an approximate equivalent to N2-P43. The whole sample included 590 representatives of western, central, southern, southeastern and northeastern (Tojin) Tuvans, who identified themselves as one of 21 tribal groups, as well as Tofalars. The N-L666 subsample consisted of 138 individuals and was studied using 15 STR markers in two scales: local (considering the areal groups and tribal clans of Tuvans) and regional (in comparison with the populations of Southern and Western Siberia). Results. Two clusters of N-L666 STR haplotypes were identified: cluster A, specific for Tuvans and Tofalars (covering 19% and 16% of their gene pools respectively) and cluster B, widely scattered throughout Siberia from the West to the Transbaikalia (reaching ~30% in Tofalars). The ubiquity and a greater age of the cluster B favor the idea of its origin in the ancestral population – the ultimate source of the haplogroup N-L666 in Siberia – commonly alleged to be Samoyedic by language. On the contrary, the narrow geographic range and a relatively recent age of the cluster A indicate its formation in the area inhabited by Tuvans and Tofalars during the last thousand years. The emergence of subclusters A1, A2, B1 may be the result of demographic growth in the populations of Tuvans, southern Altaians and Khakas about 300-450 years ago. The spread of the same haplotypes, clusters and subclusters among different regional groups and clans of Tuvans indicates a common source of haplogroup N-L666 for them, which existed in the gene pool long before the separation of the studied populations. Conclusions. A specific cluster of haplogroup N-L666 in Tuvans was presumably founded by a representative of one of the Samoyedic tribes, whose numerous descendants participated in the formation of the Tuvans, Tofalars and southern Altaians over the last thousand years.
Western Kazakhstan is populated by three clans totaling 2 million people. Since the clans are patrilineal, the Y-chromosome is the most informative genetic system for tracing their origin. We genotyped 40 Y-SNP and 17 Y-STR markers in 330 Western Kazakhs. High phylogenetic resolution within haplogroup C2a1a2-M48 was achieved by using additional SNPs. Three lines of evidence indicate that the Alimuly and Baiuly clans (but not the Zhetiru clan) have a common founder placed 700 ± 200 years back by the STR data and 500 ± 200 years back by the sequencing data. This supports traditional genealogy claims about the descent of these clans from Emir Alau, who lived 650 years ago and whose lineage might be carried by two-thirds of Western Kazakhs. There is accumulation of specific haplogroups in the subclans representing other lineages, confirming that the clan structure corresponds with the paternal genetic structure of the steppe population.
Hepatitis B virus (HBV) has been infecting humans for millennia and remains a global health problem, but its past diversity and dispersal routes are largely unknown. We generated HBV genomic data from 137 Eurasians and Native Americans dated between ~10,500 and ~400 years ago. We date the most recent common ancestor of all HBV lineages to between ~20,000 and 12,000 years ago, with the virus present in European and South American hunter-gatherers during the early Holocene. After the European Neolithic transition, Mesolithic HBV strains were replaced by a lineage likely disseminated by early farmers that prevailed throughout western Eurasia for ~4000 years, declining around the end of the 2nd millennium BCE. The only remnant of this prehistoric HBV diversity is the rare genotype G, which appears to have reemerged during the HIV pandemic.
BACKGROUND:Information about the distribution of clinically significant genetic markers in different populations may be helpful in elaborating personalized approaches to the clinical management of COVID-19 in the absence of consensus guidelines.AIM:Analyze frequencies and distribution patterns of two markers associated with severe COVID-19 (rs11385942 and rs657152) and look for potential correlations between these markers and deaths from COVID-19 among populations in Russia and across the world.METHODS:We genotyped 1883 samples from 91 ethnic groups pooled into 28 populations representing Russia and its neighbor states. We also compiled a dataset on 32 populations from other regions using genotypes extracted or imputed from the available databases. Geographic maps showing the frequency distribution of the analyzed markers were constructed using the obtained data.RESULTS:The cartographic analysis revealed that rs11385942 distribution follows the West Eurasian pattern: the marker is frequent among the populations of Europe, West Asia and South Asia but rare or absent in all other parts of the globe. Notably, the transition from high to low rs11385942 frequencies across Eurasia is not abrupt but follows the clinal variation pattern instead. The distribution of rs657152 is more homogeneous. The analysis of correlations between the frequencies of the studied markers and the epidemiological characteristics of COVID-19 in a population revealed that higher frequencies of both risk alleles correlated positively with mortality from this disease. For rs657152, the correlation was especially strong (r = 0.59, p = 0.02). These reasonable correlations were observed for the "Russian" dataset only: no such correlations were established for the "world" dataset. This could be attributed to the differences in methodology used to collect COVID-19 statistics in different countries.CONCLUSION:Our findings suggest that genetic differences between populations make a small yet tangible contribution to the heterogeneity of the pandemic worldwide.
Additional file 10: Table S5. Y-chromosome SNP and STR data of Konyrat tribe: update up to C2b1a3a-M407
Nekhvatka informatsii o rasprostranennosti v RF farmakogeneticheskikh markerov privodit k nevozmozhnosti vnedreniya algoritmov personalizatsii, razrabotannykh dlya Zapadnoy Yevropy. Tsel'yu raboty bylo sistematicheskoye izucheniye rasprostranennosti ryada znachimykh farmakogeneticheskikh markerov po vsey territorii Rossii. Iz neskol'kikh massivov populyatsionno-geneticheskikh dannykh otobrany 45 markerov (ADME-genov; genov, kodiruyushchikh farmakodinamicheskiye misheni lekarstvennykh sredstv; genov, kodiruyushchikh komponenty sistemy gemostaza), genotipirovannykh summarno dlya 2197 individov. Opredeleny chastoty etikh markerov v 50 populyatsiyakh, vklyuchayushchikh informatsiyu o 137 etnicheskikh i subetnicheskikh gruppakh. V rezul'tate sozdan farmakogeneticheskiy atlas — sistematicheskoye sobraniye genogeograficheskikh kart rasprostranennosti farmakogeneticheskikh DNK markerov po vsey territorii Rossii i sopredel'nykh stran. Atlas vyyavil tri patterna prostranstvennoy izmenchivosti. Pattern klinal'noy izmenchivosti (gradiyentnogo izmeneniya chastot po osi «vostok–zapad») ob"yedinyayet markery, sleduyushchiye osnovnoy zakonomernosti vsego genofonda naseleniya Severnoy Yevrazii (13% kart atlasa). Pattern ravnomernogo raspredeleniya vydelyayet markery, srednyaya chastota kotorykh kharakterna dlya bol'shinstva regionov Rossii (27% kart atlasa). Pattern «ochagovoy» izmenchivosti ob"yedinyayet farmakogeneticheskiye markery, kharakternyye tol'ko dlya opredelennoy gruppy etnosov i otsutstvuyushchiye v drugikh regionakh (60% kart atlasa). Atlas pokazyvayet, chto srednyaya chastota markera i informatsiya o yego vstrechayemosti v otdel'nykh populyatsiyakh ne mogut sluzhit' ukazaniyem na tip yego raspredeleniya v prostranstve RF — dlya vyyavleniya patterna izmenchivosti neobkhodima genogeograficheskaya karta.
The lack of information about the frequency of pharmacogenetic markers in Russia impedes the adoption of personalized treatment algorithms originally developed for West European populations. The aim of this paper was to study the distribution of some clinically significant pharmacogenetic markers across Russia. A total of 45 pharmacogenetic markers were selected from a few population genetic datasets, including ADME, drug target and hemostasis-controlling genes. The total number of donors genotyped for these markers was 2,197. The frequencies of these markers were determined for 50 different populations, comprised of 137 ethnic and subethnic groups. A comprehensive pharmacogenetic atlas was created, i.e. a systematic collection of gene geographic maps of frequency variation for 45 pharmacogenetic DNA markers in Russia and its neighbor states. The maps revealed 3 patterns of geographic variation. Clinal variation (a gradient change in frequency along the East-West axis) is observed in the pharmacogenetic markers that follow the main pattern of variation for North Eurasia (13% of the maps). Uniform distribution singles out a group of markers that occur at average frequency in most Russian regions (27% of the maps). Focal variation is observed in the markers that are specific to a certain group of populations and are absent in other regions (60% of the maps). The atlas reveals that the average frequency of the marker and its frequency in individual populations do not indicate the type of its distribution in Russia: a gene geographic map is needed to uncover the pattern of its variation.
According to the data on the Y-chromosome polymorphism, an assessment of the influence of two factors (tribal (clan) and administrative (kozhuuns) subdivision) on the structure of the gene pool of Tuvans is provided. The ten most common Tuvan clans (Ak, Baraan, Irgit, Kol, Kyrgyz, Mongush, Oorzhak, Oyun, Khertek, Choodu) are included in this analysis. They cover two-thirds of the total sample (N = 545) of Tuvans from six kozhuuns (Barun-Khemchiksky, Tandinsky, Tere-Kholsky, Todzhinsky, Chaa-Kholsky, Erzinsky). Genetic portraits of clans created on the basis of 52 SNP markers of the Y chromosome reveal the founder effect for all clans, except the two largest—Kyrgyz and Mongush. Making up one-third of the sample, these two clans are conglomerates. Analysis of molecular variance (AMOVA) indicates an equal degree of genetic separation of clans (5.2%) and kozhuuns (6.8%). But the Mantel test detects a high correlation (r = 0.50) between the genetic and clan structure against the low correlation (r = 0.25) between the genetic and geographical distances. The reason for the difference between the AMOVA and Mantel test lies in the fact that the four most numerous kozhuuns are “monoclan,” in which one clan prevails, and only two kozhuuns (Todzhinsky and Tere-Kholsky) include representatives of three different clans. The selection of the western and eastern clusters on the graph of multidimensional scaling fits in well with the anthropological data (Sayansky and Katangsky anthropological types), but does not confirm any of the ethnographic versions of the Tuvan ethnogenesis (“Samoyedic,” “Mongolian,” “Turkic”). The sum of results indicates that the tribal structure most fully reflects the architectonics of the Tuvan gene pool, and for this reason, it must be taken into account in population genetic research.
The genomes of present-day humans outside Africa originated almost entirely from a single migration out ~50,000-70,000 years ago[1, 2], followed by mixture with Neanderthals contributing ~2% to all non-Africans [3, 4]. However, the details of this initial migration remain poorly-understood because no ancient DNA analyses are available from this key time period, and interpretation of present-day autosomal data is complicated due to subsequent population movements/reshaping [5]. One locus, however, does retain male-specific information from this early period: the Y-chromosome, where a detailed calibrated phylogeny has been constructed [6]. Three present-day Y lineages were carried by the initial migration: the rare haplogroup D, the moderately rare C, and the very common FT lineage which now dominates most non-African populations [6, 7]. We show that phylogenetic analyses of haplogroup C, D and FT sequences, including very rare deep-rooting lineages, together with phylogeographic analyses of ancient and present-day non-African Y-chromosomes, all point to East/South-east Asia as the origin 50,000-55,000 years ago of all known non-African male lineages (apart from recent migrants). This observation contrasts with the expectation of a West Eurasian origin predicted by a simple model of expansion from a source near Africa [8, 9], and can be interpreted as resulting from extensive genetic drift in the initial population or replacement of early western lineages from the east, thus informing and constraining models of the initial expansion.
ВведениеÃåíîôîíä ëþáîãî íàðîäà ÿâëÿåò ñîáîé ñëîaeíûé êîâåð
ВведениеÍà ñåâåðå Êàì÷àòñêîãî êðàÿ, ñîåäèíÿþùèì ïîëóîñòðîâ ñ ìàòåðèêîì, ïåðåñåêëèñü òðè ïîòîêà êîðåííîãî íàñåëåíèÿ ñåâåðî-âîñòî÷íîé îêîíå÷íîñòè Åâðàçèè: äðåâíåå íàñåëåíèå ×óêîòêè (÷óê÷è), äðåâíåå íàñåëåíèå Êàì÷àòêè (êîðÿêè) è î÷åíü ïîçäíèé (XIX âåêà) ïîòîê òóíãóñîâ (ýâåíû).Äåìîãðà-ôè÷åñêèå ïîðòðåòû íàðîäîâ, íûíå aeèâóùèõ íà ïåðåñå÷åíèè ýòèõ òðåõ ïîòîêîâ, ìîãóò ñëóaeèòü ïðèáëèçèòåëüíîé àíàëîãîâîé ìîäåëüþ âçàèìîäåéñòâèÿ ìåaeäó êîðåííûìè íàðîäàìè ñåâåðà Åâðàçèè â óñëîâèÿõ ðàñòóùåãî
One of the tasks of population-based biobanks is to determine the frequencies of clinically relevant genetic polymorphisms in the population. The population of Russia is very heterogeneous both ethnically and genetically. Therefore, the frequencies of genetic markers are in demand not in one sample, but in a series of samples reflecting the heterogeneity of the gene pool of different peoples and regions. Aim. To divide the population of Russia and neighboring countries into population groups meeting certain conditions, as well as having a representative sample in existing data and biobanks. Material and methods . We developed a method for combining populations into larger groups with maintaining intragroup homogeneity based on the principal components analysis with K-means clustering, followed by refinement of clustering for higher homogeneity and a more equal distribution of group sizes using FST distances. The technology has been adjusted using the example of the Biobank of Northern Eurasia. Therefore, the material was the genome-wide data on 4.5 million genetic markers for 1,883 samples representing 247 populations of Russia and neighboring countries from this biobank. The developed approach, the resulting set of populations and related map can be applied for other collections of biomaterials from Russian populations. Results . Application of this approach made it possible to divide the entire population of Russia and neighboring countries into 29 ethnogeographic groups, characterized by relative genetic homogeneity. This set of populations is recommended as a baseline for population screenings to identify the frequency of any genetic markers among the population of Russia. A map has been constructed showing the division of population into 29 ethnogeographic areas. Conclusion . On the basis of a reliable genome-wide data, the zoning of gene pool of the Russian population was carried out. We identified ethnogeographic groups with intergroup contrasting allele frequencies, but at the same time with relatively homogeneous intragroup parameters. The resulting map and register of groups can be used in population genetic, medical genetic and pharmacogenetic studies.
Background The knowledge of clinically relevant markers distribution might become a useful tool in COVID-19 therapy using personalized approach in the lack of unified recommendations for COVID-19 patients management during pandemic. We aimed to identify the frequencies and distribution patterns of rs11385942 and rs657152 polymorphic markers, associated with severe COVID-19, among populations of the world, as well at the national level within Russia. The study was also dedicated to reveal whether population frequencies of both polymorphic markers are associated with COVID-19 cases, recovery and death rates. Methods We genotyped 1883 samples from 91 ethnic populations from Russia and neighboring countries by rs11385942 and rs657152 markers. Local populations which were geographically close and genetically similar were pooled into 28 larger groups. In the similar way we compiled a dataset on the other regions of the globe using genotypes extracted or imputed from the available datasets (32 populations worldwide). The differences in alleles frequencies between groups were estimated and the frequency distribution geographic maps have been constructed. We run the correlation analysis of both markers frequencies in various populations with the COVID-19 epidemiological data on the same populations. Findings The cartographic analysis revealed that distribution of rs11385942 follows the West Eurasian pattern: it is frequent in Europeans, West Asians, and particularly in South Asians but rare or absent in all other parts of the globe. Notably, there is no abrupt changes in frequency across Eurasia but the clinal variation instead. The distribution of rs657152 is more homogeneous. Higher population frequencies of both risk alleles correlated positively with the death rate. For the rs11385942 we can state the tendency only (r=0,13, p=0.65), while for rs657152 the correlation was significantly high (r=0,59, p=0,02). These reasonable correlations were obtained on the Russian dataset, but not on the world dataset. Interpretation Using epidemiological statistics on Russia and neighboring countries we revealed the evident correlation of the risk alleles frequencies with the death rate from COVID-19. The lack of such correlations at the world level should be attributed to the differences in the ways epidemiological data have been counted in different countries. So that, we believe that genetic differences between populations make small but real contribution into the heterogeneity of the pandemic worldwide. New studies on the correlations between COVID-19 recovery/mortality rates and population’s gene pool are urgently needed. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by the Ministry of Science and Education of the Russian Federation (State assignments for the Research Centre for Medical Genetics and Vavilov Institute of General Genetics) and by the Ministry of Health (State assignment for the Russian Medical Academy of Continuous Professional Education). The authors have no other relevant affiliations or financial involvement with any organization or entity with a financial interest in or financial conflict with the subject matter or materials discussed in the manuscript apart from those disclosed. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study has been performed in accordance with the Declaration of Helsinki and was approved by the Ethical Committee of the Research Centre for Medical Genetics, Moscow. Written informed consent was obtained from all sample donor. All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes Authors declare that all data provided in this manuscript is available and could be submitted on request.
This study explored the gene pools of Russian and Karelian populations of Tver region. Forty-one samples representing Tver Karels (n = 11) and Russians residing in the Western, Central and Eastern districts of Tver region (n = 30) were genotyped using a genome-wide panel of 4,559,465 SNPs. In order to investigate the phenomenon of genetic admixture between Slavic and Finnish-speaking populations, the obtained results were compared to the data on the Russian populations inhabiting the neighboring territories, Karels from Karelia and other North Eastern Europeans. Studying the gene pools of Russian populations with a genome-wide SNP panel is essential for cataloging their genetic diversity and identifying the distinct features of regional gene pools; in addition, it provides valuable data for practical pharmacogenomics and forensics. Using the principal component analysis, the ADMIXTURE method and D- and f3-statistics, we demonstrated that the gene pool of Tver Karels is closest to the gene pool of Karelian Karels, despite a long (300 to 500 years) history of living among the larger Russian population and the twentyfold population decline during the 20th century. At the same time, the gene pool of Tver Karels exhibits more pronounced similarity to the gene pool of the studied Russian populations than does any other Karelian population. The genetic admixture between Tver Russians and Tver Karels occurred due to a more intense gene flow from Russians to Karels whereas the gene flow from Karels to Russians was much weaker: Tver Russians turned out to be as genetically different from Karels as Pskov Russians. The genetic similarity of Tver Karels to Karelian Karels assessed with the autosomal SNP panel exhibits a slight shift towards the Russian gene pool and is consistent with the previously published analysis of Y-chromosome lineages in these populations that detected no admixture between Tver Karels and Russians.