The North Eurasian forest and forest-steppe zones have sustained millennia of sociocultural connections among northern peoples, but much of their history is poorly understood. In particular, the genomic formation of populations that speak Uralic and Yeniseian languages today is unknown. Here, by generating genome-wide data for 180 ancient individuals spanning this region, we show that the Early-to-Mid-Holocene hunter-gatherers harboured a continuous gradient of ancestry from fully European-related in the Baltic, to fully East Asian-related in the Transbaikal. Contemporaneous groups in Northeast Siberia were off-gradient and descended from a population that was the primary source for Native Americans, which then mixed with populations of Inland East Asia and the Amur River Basin to produce two populations whose expansion coincided with the collapse of pre-Bronze Age population structure. Ancestry from the first population, Cis-Baikal Late Neolithic-Bronze Age (Cisbaikal_LNBA), is associated with Yeniseian-speaking groups and those that admixed with them, and ancestry from the second, Yakutia Late Neolithic-Bronze Age (Yakutia_LNBA), is associated with migrations of prehistoric Uralic speakers. We show that Yakutia_LNBA first dispersed westwards from the Lena River Basin around 4,000 years ago into the Altai-Sayan region and into West Siberian communities associated with Seima-Turbino metallurgy-a suite of advanced bronze casting techniques that expanded explosively from the Altai1. The 16 Seima-Turbino period individuals were diverse in their ancestry, also harbouring DNA from Indo-Iranian-associated pastoralists and from a range of hunter-gatherer groups. Thus, both cultural transmission and migration were key to the Seima-Turbino phenomenon, which was involved in the initial spread of early Uralic-speaking communities.
Microhaplotypes are genetic markers that are short DNA sequences, typically consisting of two or more single nucleotide polymorphisms (SNPs) sufficiently close molecularly (<300 basepairs) that their alleles are inherited together from parent to child. Microhaplotypes can be considered a special class of haplotypes, “micro” because they extend for only a few dozens of basepairs, not thousands of basepairs. The advantage of microhaplotypes is that, with massively parallel sequencing of a small segment, the phase is known. Microhaps have been shown to be highly informative in demonstrating uniqueness of DNA profiles and in determining the biogeographic ancestry of an individual. Tests of biologic relationships, such as paternity tests, can be done with microhaplotypes. Deconvolution of DNA mixtures is an area in which microhaplotypes appear to be especially promising. This brief introductory review presents the scientific background for the development of microhaplotypes as a new type of genetic marker and examples of some of the recent applications.
The COMT gene encodes for catechol-O-methyl-transferase, an enzyme playing a major role in regulation of synaptic catecholamine neurotransmitters. Investigating 4 markers of the COMT gene (rs2020917, rs4818, rs4680, rs9332377) in 6 Tunisian populations and a pool of Libyans. Our objective was to determine the distribution of allelic, genotypic and haplotypic frequencies by comparison to other populations of the 1000 genomes project and 59 populations from the Kidd Lab dataset. The allelic frequencies established for these SNPs in the North African populations are similar to those of Europeans and South Asians. Linkage disequilibrium between these SNPs and haplotypes frequencies are different between populations whose clustering in principal components analysis (PCA) according to their geographic origin was more significant using haplotypic frequencies. COMT activity prediction by haplotypes genotyping could be limited to rs4818-rs4680 micro-haplotypes. The Low activity haplotype (CG) displays the highest frequency in African populations (55%), in the 59 Kidd Lab populations we found also that Sub-Saharan Africans, Native Americans, and some East Asian and Pacific Island populations all have frequencies in the 50-81% range for (CG) where as its lowest frequency was found in Europeans (10%), this results have been also confirmed for Southwest Asians. North Africans and South Asians with intermediate frequencies have approximately similar values (20% and 25%). Europeans show the highest frequencies of haplotypes with predicted High and Medium activity in contrast to Africans. North Africans and South Asians present similar results for all the category of the COMT activity prediction by haplotypes genotyping. The high level of genetic diversity of COMT haplotypes, not only allows distinction between populations according to their history settlement, origin and ethnicity, it constitutes a basis for studies of association of the COMT gene polymorphism with pathologies, drugs response and for forensic investigation in North African populations.
The demographic history of human populations in North Africa has been characterized by complex migration processes that have determined the current genetic structure of these populations. We examined the autosomal markers of eight sampled populations in northern Africa (Tunisia and Libya) to explore their genetic structure and to place them in a global context. We genotyped a set of 30 autosomal single-nucleotide polymorphisms (SNPs) extending 9.5 Mb and encompassing the 17q21 inversion region. Our data include 403 individuals from Tunisia and Libya. To put our populations in the global context, we analyzed our data in comparison with other populations, including those of the 1000 Genomes Project. To evaluate the data, we conducted genetic diversity, principal component, STRUCTURE, and haplotype analyses. The analysis of genetic composition revealed the genetic heterogeneity of North African populations. The principal component and STRUCTURE analyses converged and revealed the intermediate position of North Africans between Europeans and Asians. Haplotypic analysis demonstrated that the normal (H1) and inverted (H2) polymorphisms in the chromosome 17q21 region occur in North Africa at frequencies similar to those found in European and Southwest Asian populations. The results highlight the complex demographic history of North Africa, reflecting the influence of genetic flow from Europe and the Near East that dates to the prehistoric period. These gene flows added to demographic factors (inbreeding, endogamy), natural factors (topography, Sahara), and cultural factors that play a role in the emergence of the diverse and heterogeneous genetic structures of North African populations. This study contributes to a better understanding of the complex structure of North African populations.
Supplementary Figure 1, Tables 1-4 from A SNP in a let-7 microRNA Complementary Site in the KRAS 3′ Untranslated Region Increases Non–Small Cell Lung Cancer Risk
The North Eurasian forest and forest-steppe zones have sustained millennia of sociocultural connections among northern peoples. We present genome-wide ancient DNA data for 181 individuals from this region spanning the Mesolithic, Neolithic and Bronze Age. We find that Early to Mid-Holocene hunter-gatherer populations from across the southern forest and forest-steppes of Northern Eurasia can be characterized by a continuous gradient of ancestry that remained stable for millennia, ranging from fully West Eurasian in the Baltic region to fully East Asian in the Transbaikal region. In contrast, cotemporaneous groups in far Northeast Siberia were genetically distinct, retaining high levels of continuity from a population that was the primary source of ancestry for Native Americans. By the mid-Holocene, admixture between this early Northeastern Siberian population and groups from Inland East Asia and the Amur River Basin produced two distinctive populations in eastern Siberia that played an important role in the genetic formation of later people. Ancestry from the first population, Cis-Baikal Late Neolithic–Bronze Age (Cisbaikal_LNBA), is found substantially only among Yeniseian-speaking groups and those known to have admixed with them. Ancestry from the second, Yakutian Late Neolithic–Bronze Age (Yakutia_LNBA), is strongly associated with present-day Uralic speakers. We show how Yakutia_LNBA ancestry spread from an east Siberian origin ∼4.5kya, along with subclades of Y-chromosome haplogroup N occurring at high frequencies among present-day Uralic speakers, into Western and Central Siberia in communities associated with Seima-Turbino metallurgy: a suite of advanced bronze casting techniques that spread explosively across an enormous region of Northern Eurasia ∼4.0kya. However, the ancestry of the 16 Seima-Turbino-period individuals—the first reported from sites with this metallurgy—was otherwise extraordinarily diverse, with partial descent from Indo-Iranian-speaking pastoralists and multiple hunter-gatherer populations from widely separated regions of Eurasia. Our results provide support for theories suggesting that early Uralic speakers at the beginning of their westward dispersal where involved in the expansion of Seima-Turbino metallurgical traditions, and suggests that both cultural transmission and migration were important in the spread of Seima-Turbino material culture.
Precision Medicine is an emerging approach for disease treatment and prevention that takes into account individual variability in genes, environment, and lifestyle. Autoimmune diseases are those in which the body’s natural defense system loses discriminating power between its own cells and foreign cells, causing the body to mistakenly attack healthy tissues. These conditions are very heterogeneous in their presentation and therefore difficult to diagnose and treat. Achieving precision medicine in autoimmune diseases has been challenging due to the complex etiologies of these conditions, involving an interplay between genetic, epigenetic, and environmental factors. However, recent technological and computational advances in molecular profiling have helped identify patient subtypes and molecular pathways which can be used to improve diagnostics and therapeutics. This review discusses the current understanding of the disease mechanisms, heterogeneity, and pathogenic autoantigens in autoimmune diseases gained from genomic and transcriptomic studies and highlights how these findings can be applied to better understand disease heterogeneity in the context of disease diagnostics and therapeutics.
Background: The single nucleotide polymorphisms (SNPs) of the dopamine D3 receptor (DRD3), the CUB and sushi multiple domains 1 (CSMD1) and the neuregulin 1 (NRG1) genes were used to study the genetic diversity and affinity among North African populations and to examine their genetic relationships in worldwide populations. Methods: The rs3773678, rs3732783 and rs6280 SNPs of the DRD3 gene located on chromosome 3, the rs10108270 SNP of the CSMD1 gene and the rs383632, rs385396 and rs1462906 SNPs of the NRG1 gene located on chromosome 8 were analysed in 366 individuals from seven North African populations (Libya, Kairouan, Mehdia, Sousse, Kesra, Smar and Kerkennah). Results: The low values of F-ST indicated that only 0.27%-1.65% of the genetic variability was due to the differences between the populations. The Kairouan population has the lowest average heterozygosity among the North African populations. Haplotypes composed of the ancestral alleles ACC and ACAT were more frequent in the Kairouan population than in other North African populations. The PCA and the haplotypic analysis showed that the genetic structure of populations in North Africa was closer to that of Europeans, Admixed Americans, South Asians and East Asians. However, analysis of the rs3732783 and rs6280 SNPs revealed that the CT microhaplotype was specific to the North African population. Conclusions: The Kairouan population exhibited a relatively low rate of genetic variability. The North African population has undergone significant gene flow but also evolutionary forces that have made it genetically distinct from other populations.
A small panel of highly informative loci that can be genotyped on the same equipment as the standard CODIS short tandem repeat (STR) markers has strong potential for application in forensic casework. Single nucleotide polymorphisms (SNPs) can be typed by a couple of methods on capillary electrophoresis (CE) machines and on sequencers, but the amount of information relative to the laboratory effort has hindered use of SNPs in actual casework. Insertion-deletion markers (InDels) suffer from similar problems. Microhaplotypes (MHs) are much more informative per locus but have similar technical difficulties unless they are typed by massively parallel sequencing (MPS). As forensic labs are acquiring sequencing machines, MHs become more likely to be used in casework, especially if multiplexed with STRs. Here we present the details of a multipurpose panel of 24 MHs with the highest effective number of alleles (Ae) from previous work. An augmented STR panel of 24 loci (20 CODIS markers plus four commonly typed STRs) is also considered. The Ae and ancestry informativeness (In) distributions of these two datasets are compared. The MH panel is shown to have better individualization and population distinction than the augmented CODIS STRs. We note that the 24 MHs should be better for mixture analyses than the STRs. Finally, we suggest that a commercial kit including both the standard CODIS markers and this set of 24 MH would greatly improve the discrimination power over that of current commercial assays.
In recent years, the number of publications on microhaplotypes has averaged more than a dozen papers annually. Many have contributed to a significant increase in the number of highly polymorphic microhaplotype loci. This increase allows microhaplotypes to be very informative in four main areas of forensic uses of DNA: individualization, ancestry inference, kinship analysis, and mixture deconvolution. The random match Probability (RMP) can be as small as 10−100 for a large panel of microhaplotypes. It is possible to measure the heterozygosity of an MH as the effective number of alleles (Ae). Ae > 7.5 exists for African populations and >4.5 exists for Native American populations for a smaller panel of two dozen selected microhaplotypes. Using STRUCTURE, at least 10 different ancestral clusters can be defined by microhaplotypes. The Ae for a locus is also identical to the Paternity Index (PI), the measure of how informative a locus will be in parentage testing. High Ae loci can also be useful in missing persons cases. Finally, high Ae microhaplotypes allow the near certainty of seeing multiple additional alleles in a mixture of two or more individuals in a DNA sample. In summary, a panel of higher Ae microhaplotypes can outperform the standard CODIS markers.
Population genetic studies of North Asian ethnic groups have focused on genetic variation of sex chromosomes and mitochondria. Studies of the extensive variation available from autosomal variation have appeared infrequently. We focus on relationships among population samples using new North Asia microhaplotype data. We combined genotypes from our laboratory on 58 microhaplotypes, distributed across 18 autosomes, on 3945 individuals from 75 populations with corresponding data extracted for 26 populations from the Thousand Genomes consortium and for 22 populations from the GenomeAsia 100 K project. A total of 7107 individuals in 122 total populations are analyzed using STRUCTURE, Principal Component Analysis, and phylogenetic tree analyses. North Asia populations sampled in Mongolia include: Buryats, Mongolians, Altai Kazakhs, and Tsaatans. Available Siberians include samples of Yakut, Khanty, and Komi Zyriane. Analyses of all 122 populations confirm many known relationships and show that most populations from North Asia form a cluster distinct from all other groups. Refinement of analyses on smaller subsets of populations reinforces the distinctiveness of North Asia and shows that the North Asia cluster identifies a region that is ancestral to Native Americans.
Micronesia began to be peopled earlier than other parts of Remote Oceania, but the origins of its inhabitants remain unclear. We generated genome-wide data from 164 ancient and 112 modern individuals. Analysis reveals five migratory streams into Micronesia. Three are East Asian related, one is Polynesian, and a fifth is a Papuan source related to mainland New Guineans that is different from the New Britain-related Papuan source for southwest Pacific populations but is similarly derived from male migrants ~2500 to 2000 years ago. People of the Mariana Archipelago may derive all of their precolonial ancestry from East Asian sources, making them the only Remote Oceanians without Papuan ancestry. Female-inherited mitochondrial DNA was highly differentiated across early Remote Oceanian communities but homogeneous within, implying matrilocal practices whereby women almost never raised their children in communities different from the ones in which they grew up.
DNA polymorphic markers and self-defined ethnicity groupings are used to group individuals with shared ancient geographic ancestry. Here we studied whether ancestral relationships between individuals could be identified from metabolic screening data reported by the California newborn screening (NBS) program. NBS data includes 41 blood metabolites measured by tandem mass spectrometry from singleton babies in 17 parent-reported ethnicity groupings. Ethnicity-associated differences identified for 71% of NBS metabolites (29 of 41, Cohen's d > 0.5) showed larger differences in blood levels of acylcarnitines than of amino acids (P < 1e-4). A metabolic distance measure, developed to compare ethnic groupings based on metabolic differences, showed low positive correlation with genetic and ancient geographic distances between the groups' ancestral world populations. Several outlier group pairs were identified with larger genetic and smaller metabolic distances (Black versus White) or with smaller genetic and larger metabolic distances (Chinese versus Japanese) indicating the influence of genetic and of environmental factors on metabolism. Using machine learning, comparison of metabolic profiles between all pairs of ethnic groupings distinguished individuals with larger genetic distance (Black versus Chinese, AUC = 0.96), while genetically more similar individuals could not be separated metabolically (Hispanic versus Native American, AUC = 0.51). Additionally, we identified metabolites informative for inferring metabolic ancestry in individuals from genetically similar populations, which included biomarkers for inborn metabolic disorders (C10:1, C12:1, C3, C5OH, Leucine-Isoleucine). This work sheds new light on metabolic differences in healthy newborns in diverse populations, which could have implications for improving genetic disease screening.
Being able to determine the biogeographic ancestry of an individual has been important in many forensic efforts to determine the identity of a partial or decayed body. Now a DNA sample can be the forensic equivalent of an unidentified individual using just some blood left at the scene of a crime or some semen left from a rape. The standard DNA markers, short tandem repeats, used globally to tie an individual to the crime scene DNA, are not good for inferring biogeographic ancestry because they do not vary much among populations. But they do vary greatly among individuals - every population has many alleles at each locus, but it is the same “many” in most populations. Single nucleotide polymorphisms (SNPs), on the other hand, can vary greatly in their frequencies among populations. Recent efforts by many researchers have begun to provide those markers. Panels of Ancestry Informative SNPs (AISNPs) have been developed by many researchers. These have been shown to vary in frequency among populations. Databases have been developed for estimating how likely it is for a DNA profile for those SNPs to occur in the different populations. Data requirements, analytic methods, and criteria for interpreting the results are all discussed.
The dopamine - related genes, like dopamine D2 receptor (DRD2) gene and ankyrin repeat and kinase domain containing 1 (ANKK1) gene are implicated in neurological functions. Some polymorphisms of the DRD2/ANKK1 locus (TaqIA, TaqIB, TaqID) have been used to study genetic diversity and the evolution of human populations. The present investigation aims to assess the genetic diversity in seven North African populations in order to explore their genetic structure and to compare them to others worldwide populations studied for the same locus. Nine single nucleotide polymorphisms (SNPs) from the DRD2/ANKK1 locus (rs1800497 TaqIA, rs2242592, rs1124492, rs6277, rs6275, rs1079727, rs2002453, rs2234690 and rs1079597 TaqIB) were typed in 366 individuals from seven North African populations: six from Tunisia (Sousse, Smar, Kesra, Kairouan, Mehdia and Kerkennah) and one from Libya. The allelic frequencies of rs2002453 and rs2234690 were higher in the Smar population than in the other North African populations. More, the Smar population showed the lowest average heterozygosity (0.313). The principal component analysis (PCA) showed that the Smar population was clearly separated from others. Furthermore, linkage disequilibrium analysis shown a high linkage disequilibrium in the North African population and essentially in Smar population. Comparison with other world populations has shown that the heterozygosity of North African population was very close to that of the African and European populations. The PCA and the haplotypic analysis suggested the presence of an important Eurasian genetic component for the North African population. These results suggested that the Smar population was isolated from the others North Africans ones by its peculiar genetic structure because of isolation, endogamy and genetic drift. On the other hand, the North African population is characterized by a multi ancestral gene pool from Eurasia and sub-Saharan Africa due to human migration since prehistoric times.
AbstractBackgroundOnly a few studies have investigated the association of single nucleotide polymorphisms in STAT3 gene with the susceptibility to cancer and response to chemotherapy. Our aim was to determine the allele frequencies of rs3869550, rs957971, and rs7211777 at the STAT3 gene in North African populations and compare them to 1000 genomes populations, and to investigate their relation with cancer.MethodsThe targeted SNPs have been analyzed in six Tunisian populations and a sample of Libyans using TaqMan® Assay. The results were compared to 1000 Genomes Project population samples. Targeting of the regions encompassing the three SNPs by micro‐ARN was assessed using miR databases.ResultsThe analysis of the 3 SNPs showed that North African populations were close to South Asians. As expected, African populations presented a significant frequency of the ancestral CCG haplotype in contrast to other populations where the fully derived TGA haplotype was more frequent. The presence and diversity of rare haplotypes at STAT3 in North African populations could have been generated by recombination between the two major haplotypes. A screening of the micro‐RNA databases showed that the STAT3 region with the mutated allele of rs7211777 (G>A) could be targeted by miR hsa‐miR‐3606‐5p, which also targets genes involved in breast cancer.
The TAS2R38 gene is involved in bitter taste perception. This study documents the distinctive diversity patterns in northern Africa of functional single-nucleotide polymorphisms (SNPs) rs713598 and rs1726866 at the TAS2R38 locus and places those patterns in the context of global TAS2R38 diversity. Data previously genotyped with TaqMan assay were analyzed for rs713598 and rs1726866 for 375 unrelated subjects (305 Tunisians from seven locations: Mahdia, Sousse, Kesra, Nebeur, Kairouan, Smar, and Kerkennah; plus 70 Libyans). Data were analyzed to present haplotypes and genotypes before comparison with data from worldwide populations. This study provides information about TAS2R38 diversity in a part of the world that is relatively understudied. Considering the two SNPs rs713598 and rs1726866, the CA nucleotide haplotype leading to the PV amino acid haplotype is extremely rare almost everywhere, but it is relatively frequent (between 6% and 15%) in northern Africa, where it coexists with the globally common amino acid haplotypes PA, AA, and AV. Given its higher frequency in North Africa, the authors propose the CA nucleotide haplotype as a biogeographic marker for forensic purposes.
Single-nucleotide polymorphisms (SNPs) and small genomic regions with multiple SNPs (microhaplotypes, MHs) are rapidly emerging as novel forensic investigative tools to assist in individual identification, kinship analyses, ancestry inference, and deconvolution of DNA mixtures. Here, we analyzed information for 90 microhaplotype loci in 4009 individuals from 79 world populations in 6 major biogeographic regions. The study included multiplex microhaplotype sequencing (mMHseq) data analyzed for 524 individuals from 16 populations and genotype data for 3485 individuals from 63 populations curated from public repositories. Analyses of the 79 populations revealed excellent characteristics for this 90-plex MH panel for various forensic applications achieving an overall average effective number of allele values (A(e)) of 4.55 (range 1.04-19.27) for individualization and mixture deconvolution. Population-specific random match probabilities ranged from a low of 10(-115) to a maximum of 10(-66). Mean informativeness (I-n) for ancestry inference was 0.355 (range 0.117-0.883). 65 novel SNPs were detected in 39 of the MHs using mMHseq. Of the 3018 different microhaplotype alleles identified, 1337 occurred at frequencies > 5% in at least one of the populations studied. The 90-plex MH panel enables effective differentiation of population groupings for major biogeographic regions as well as delineation of distinct subgroupings within regions. Open-source, web-based software is available to support validation of this technology for forensic case work analysis and to tailor MH analysis for specific geographical regions.
EDITORIAL article Front. Genet., 23 June 2021 | https://doi.org/10.3389/fgene.2021.708222
Microhaplotypes are emerging biomarkers for forensic applications. In this study, a sequence-based multiplex assay of 74 microhaplotypes (230 SNPs) was developed on the Ion Torrent S5™ (Thermo Fisher Scientific) system and the potential for its application to mixture deconvolution was explored. The 74 loci are distributed across the autosomal human genome and have Ae (i.e., effective number of alleles) values ranging from 1.307 to 6.010 (median = 2.706) and In (i.e., informativeness) values ranging from 0.096 to 0.660 (median = 0.251); the amplicon sizes range between 157 and 325 bp. The typing performance of the panel was evaluated on a series of in-silico two to five-person DNA mixtures and results were compared to fragment and sequence-based STRs. The 74plex-locus assay was found sensitive down to 0.05 ng of input DNA and effective for the analysis of mixtures at different contributor ratios and input DNA amounts. As expected, none or very partial minor CE-STR profile(s) were reported for highly imbalanced two-person and high-order DNA mixtures while sequencing of STRs enabled the detection of more individual minor alleles. For microhaplotypes, a full minor profile was detected down to a 20:1 ratio at 10 ng and minimal allele dropout at 1 ng of input DNA. A higher rate of allele dropout from the minor donor(s) was reported at 1 ng than 10 ng for three-person mixtures while for four- and five-person mixtures, the same number of dropouts was observed for almost all minor donors. Overall this microhaplotype panel is a powerful tool that can complement and enhance size- and sequence-based STR analysis of forensic DNA mixtures.