High-altitude environments pose significant challenges to human survival and reproduction, drawing considerable attention to the demographic and adaptive histories of populations in these regions. Here, we present whole-genome sequences from diverse Himalayan populations, offering new insights into the genomic history of this region. We find that population structure in the Himalayas began as early as 10,000 years ago, predating archaeological evidence of permanent habitation above 2,500 meters by ∼6,000 years. The widespread presence of the introgressed adaptive EPAS1 haplotype across all high-altitude populations highlights a shared genetic origin and its importance for survival in this region. We identify additional selection signals in genes linked to hypoxia, physical activity, immunity, and metabolism, all of which could have facilitated adaptation to the harsh environment. Over time, increasing genetic structure led to the emergence of the strongly differentiated ethnic groups observed today, many of which maintained small effective population sizes throughout their history or experienced severe bottlenecks. Between 6,000 and 3,000 years ago, with the advent of agriculture, a few uniparental lineages became predominant; however, significant population growth was not observed in the Himalayas except in the Tibetans. In recent history, we detect bidirectional gene flow between high-altitude and lowland groups on both sides of the Himalayan range, coinciding with the rise and expansion of historical regional powers, particularly during the Tibetan and northern Indian Gupta Empires. In the past few centuries, migrations to the Himalayas seem to have occurred alongside conflicts and population displacements in nearby regions and show some sex bias.
High-altitude environments pose substantial challenges for human survival and reproduction, attracting considerable attention to the demographic and adaptive histories of high-altitude populations. Previous work focused mainly on Tibetans, establishing their genetic relatedness to East Asians and their genetic adaptation to high altitude, especially at EPAS1. Here, we present 87 new whole-genome sequences from 16 Himalayan populations and the insight they provide into the genomic history of the region. We show that population structure in the Himalayas began to emerge as early as 10,000 years ago, predating archaeological evidence of permanent habitation above 2,500 meters by approximately 6,000 years. The high prevalence of the introgressed adaptive EPAS1 haplotype in all high-altitude populations today supports a shared genetic origin and its importance for survival in this region. We also identify additional selection signals in genes associated with hypoxia, physical activity, immunity and metabolism which could have facilitated adaptation to the harsh environment. Over time, increasing genetic structure led to the diverse and strongly differentiated ethnic groups observed today, most of which maintained small population sizes throughout their history or experienced severe bottlenecks. Between 6,000 and 3,000 years ago, a few uniparental lineages became predominant, likely coinciding with the advent of agriculture, although significant population growth was not observed in the Himalayas except in the Tibetans. In more recent times, we detect bidirectional gene flow between high-altitude and lowland groups, occurring on both sides of the Himalayan range. The timing of this admixture aligns with the rise and expansion of historical regional powers, particularly during the Tibetan Empire and the northern Indian Gupta Empire. In the past few centuries, migrations to the Himalayas seem to have occurred alongside conflicts and population displacements in nearby regions and show some sex bias. ### Competing Interest Statement The authors have declared no competing interest.
Archaic admixture has had a substantial impact on human evolution with multiple events across different clades, including from extinct hominins such as Neanderthals and Denisovans into modern humans. In great apes, archaic admixture has been identified in chimpanzees and bonobos but the possibility of such events has not been explored in other species. Here, we address this question using high-coverage whole-genome sequences from all four extant gorilla subspecies, including six newly sequenced eastern gorillas from previously unsampled geographic regions. Using approximate Bayesian computation with neural networks to model the demographic history of gorillas, we find a signature of admixture from an archaic ‘ghost’ lineage into the common ancestor of eastern gorillas but not western gorillas. We infer that up to 3% of the genome of these individuals is introgressed from an archaic lineage that diverged more than 3 million years ago from the common ancestor of all extant gorillas. This introgression event took place before the split of mountain and eastern lowland gorillas, probably more than 40 thousand years ago and may have influenced perception of bitter taste in eastern gorillas. When comparing the introgression landscapes of gorillas, humans and bonobos, we find a consistent depletion of introgressed fragments on the X chromosome across these species. However, depletion in protein-coding content is not detectable in eastern gorillas, possibly as a consequence of stronger genetic drift in this species.
Southeast Asia comprises 11 countries that span mainland Asia across to numerous islands that stretch from the Andaman Sea to the South China Sea and Indian Ocean. This region harbors an impressive diversity of history, culture, religion and biology. Indigenous people of Malaysia display substantial phenotypic, linguistic, and anthropological diversity. Despite this remarkable diversity which has been documented for centuries, the genetic history and structure of indigenous Malaysians remain under-studied. To have a better understanding about the genetic history of these people, especially Malaysian Negritos, we sequenced whole genomes of 15 individuals belonging to five indigenous groups from Peninsular Malaysia and one from North Borneo to high coverage (30X). Our results demonstrate that indigenous populations of Malaysia are genetically close to East Asian populations. We show that present-day Malaysian Negritos can be modeled as an admixture of ancient Hoabinhian hunter-gatherers and Neolithic farmers. We observe gene flow from South Asian populations into the Malaysian indigenous groups, but not into Dusun of North Borneo. Our study proposes that Malaysian indigenous people originated from at least three distinct ancestral populations related to the Hoabinhian hunter-gatherers, Neolithic farmers and Austronesian speakers.
The Middle East region is important to understand human evolution and migrations but is underrepresented in genomic studies. Here, we generated 137 high-coverage physically phased genome sequences from eight Middle Eastern populations using linked-read sequencing. We found no genetic traces of early expansions out-of-Africa in present-day populations but found Arabians have elevated Basal Eurasian ancestry that dilutes their Neanderthal ancestry. Population sizes within the region started diverging 15-20 kya, when Levantines expanded while Arabians maintained smaller populations that derived ancestry from local hunter-gatherers. Arabians suffered a population bottleneck around the aridification of Arabia 6 kya, while Levantines had a distinct bottleneck overlapping the 4.2 kya aridification event. We found an association between movement and admixture of populations in the region and the spread of Semitic languages. Finally, we identify variants that show evidence of selection, including polygenic selection. Our results provide detailed insights into the genomic and selective histories of the Middle East.
A nonsense allele at rs1343879 in human MAGEE2 on chromosome X has previously been reported as a strong candidate for positive selection in East Asia. This premature stop codon causing ∼80% protein truncation is characterized by a striking geographical pattern of high population differentiation: common in Asia and the Americas (up to 84% in the 1000 Genomes Project East Asians) but rare elsewhere. Here, we generated a Magee2 mouse knockout mimicking the human loss-of-function mutation to study its functional consequences. The Magee2 null mice did not exhibit gross abnormalities apart from enlarged brain structures (13% increased total brain area, P = 0.0022) in hemizygous males. The area of the granular retrosplenial cortex responsible for memory, navigation, and spatial information processing was the most severely affected, exhibiting an enlargement of 34% (P = 3.4×10-6). The brain size in homozygous females showed the opposite trend of reduced brain size, although this did not reach statistical significance. With these insights, we performed human association analyses between brain size measurements and rs1343879 genotypes in 141 Chinese volunteers with brain MRI scans, replicating the sexual dimorphism seen in the knockout mouse model. The derived stop gain allele was significantly associated with a larger volume of gray matter in males (P = 0.00094), and smaller volumes of gray (P = 0.00021) and white (P = 0.0015) matter in females. It is unclear whether or not the observed neuroanatomical phenotypes affect behavior or cognition, but it might have been the driving force underlying the positive selection in humans.
The number and distribution of recessive alleles in the population for various diseases are not known at genome-wide-scale. Based on 6447 exome-sequences of healthy, genetically-unrelated Europeans of two distinct ancestries, we estimate that every individual is a carrier of at least 2 pathogenic variants in currently known autosomal recessive (AR) genes, and that 0.8-1% of European couples are at-risk of having a child affected with a severe AR genetic disorder. This risk is 16.5-fold higher for first cousins, but is significantly more increased for skeletal disorders and intellectual disabilities due to their distinct genetic architecture.
Male infertility is a prevalent condition, affecting 5–10% of men. So far, few genetic factors have been described as contributors to spermatogenic failure. Here, we report the first re-sequencing study of the Y-chromosomal Azoospermia Factor c ( AZFc ) region, combined with gene dosage analysis of the multicopy DAZ, BPY2 , and CDY genes and Y-haplogroup determination. In analysing 2324 Estonian men, we uncovered a novel structural variant as a high-penetrance risk factor for male infertility. The Y lineage R1a1-M458, reported at >20% frequency in several European populations, carries a fixed ~1.6 Mb r2/r3 inversion, destabilizing the AZFc region and predisposing to large recurrent microdeletions. Such complex rearrangements were significantly enriched among severe oligozoospermia cases. The carrier vs non-carrier risk for spermatogenic failure was increased 8.6-fold (p=6.0×10 −4 ). This finding contributes to improved molecular diagnostics and clinical management of infertility. Carrier identification at young age will facilitate timely counselling and reproductive decision-making.
Structural variants contribute substantially to genetic diversity and are important evolutionarily and medically, but they are still understudied. Here we present a comprehensive analysis of structural variation in the Human Genome Diversity panel, a high-coverage dataset of 911 samples from 54 diverse worldwide populations. We identify, in total, 126,018 variants, 78% of which were not identified in previous global sequencing projects. Some reach high frequency and are private to continental groups or even individual populations, including regionally restricted runaway duplications and putatively introgressed variants from archaic hominins. By de novo assembly of 25 genomes using linked-read sequencing, we discover 1,643 breakpoint-resolved unique insertions, in aggregate accounting for 1.9 Mb of sequence absent from the GRCh38 reference. Our results illustrate the limitation of a single human reference and the need for high-quality genomes from diverse populations to fully discover and understand human genetic variation.
The genomes of present-day humans outside Africa originated almost entirely from a single out-migration ~ 50,000–70,000 years ago, followed by mixture with Neanderthals contributing ~ 2% to all non-Africans. However, the details of this initial migration remain poorly understood because no ancient DNA analyses are available from this key time period, and interpretation of present-day autosomal data is complicated due to subsequent population movements/reshaping. One locus, however, does retain male-specific information from this early period: the Y chromosome, where a detailed calibrated phylogeny has been constructed. Three present-day Y lineages were carried by the initial migration: the rare haplogroup D, the moderately rare C, and the very common FT lineage which now dominates most non-African populations. Here, we show that phylogenetic analyses of haplogroup C, D and FT sequences, including very rare deep-rooting lineages, together with phylogeographic analyses of ancient and present-day non-African Y chromosomes, all point to East/Southeast Asia as the origin 50,000–55,000 years ago of all known surviving non-African male lineages (apart from recent migrants). This observation contrasts with the expectation of a West Eurasian origin predicted by a simple model of expansion from a source near Africa, and can be interpreted as resulting from extensive genetic drift in the initial population or replacement of early western Y lineages from the east, thus informing and constraining models of the initial expansion.
Genome sequences from diverse human groups are needed to understand the structure of genetic variation in our species and the history of, and relationships between, different populations. We present 929 high-coverage genome sequences from 54 diverse human populations, 26 of which are physically phased using linked-read sequencing. Analyses of these genomes reveal an excess of previously undocumented common genetic variation private to southern Africa, central Africa, Oceania, and the Americas, but an absence of such variants fixed between major geographical regions. We also find deep and gradual population separations within Africa, contrasting population size histories between hunter-gatherer and agriculturalist groups in the past 10,000 years, and a contrast between single Neanderthal but multiple Denisovan source populations contributing to present-day human populations.
The Iron and Classical Ages in the Near East were marked by population expansions carrying cultural transformations that shaped human history, but the genetic impact of these events on the people who lived through them is little-known. Here, we sequenced the whole genomes of 19 individuals who each lived during one of four time periods between 800 BCE and 200 CE in Beirut on the Eastern Mediterranean coast at the center of the ancient world's great civilizations. We combined these data with published data to traverse eight archaeological periods and observed any genetic changes as they arose. During the Iron Age (∼1000 BCE), people with Anatolian and South-East European ancestry admixed with people in the Near East. The region was then conquered by the Persians (539 BCE), who facilitated movement exemplified in Beirut by an ancient family with Egyptian-Lebanese admixed members. But the genetic impact at a population level does not appear until the time of Alexander the Great (beginning 330 BCE), when a fusion of Asian and Near Easterner ancestry can be seen, paralleling the cultural fusion that appears in the archaeological records from this period. The Romans then conquered the region (31 BCE) but had little genetic impact over their 600 years of rule. Finally, during the Ottoman rule (beginning 1516 CE), Caucasus-related ancestry penetrated the Near East. Thus, in the past 4,000 years, three limited admixture events detectably impacted the population, complementing the historical records of this culturally complex region dominated by the elite with genetic insights from the general population.
The genomes of present-day humans outside Africa originated almost entirely from a single migration out ~50,000-70,000 years ago[1, 2], followed by mixture with Neanderthals contributing ~2% to all non-Africans [3, 4]. However, the details of this initial migration remain poorly-understood because no ancient DNA analyses are available from this key time period, and interpretation of present-day autosomal data is complicated due to subsequent population movements/reshaping [5]. One locus, however, does retain male-specific information from this early period: the Y-chromosome, where a detailed calibrated phylogeny has been constructed [6]. Three present-day Y lineages were carried by the initial migration: the rare haplogroup D, the moderately rare C, and the very common FT lineage which now dominates most non-African populations [6, 7]. We show that phylogenetic analyses of haplogroup C, D and FT sequences, including very rare deep-rooting lineages, together with phylogeographic analyses of ancient and present-day non-African Y-chromosomes, all point to East/South-east Asia as the origin 50,000-55,000 years ago of all known non-African male lineages (apart from recent migrants). This observation contrasts with the expectation of a West Eurasian origin predicted by a simple model of expansion from a source near Africa [8, 9], and can be interpreted as resulting from extensive genetic drift in the initial population or replacement of early western lineages from the east, thus informing and constraining models of the initial expansion.
Background In the process of adaptation of humans to their environment, positive or adaptive selection has played a main role. Positive selection has, however, been under-studied in African populations, despite their diversity and importance for understanding human history. Results Here, we have used 119 available whole-genome sequences from five Ethiopian populations (Amhara, Oromo, Somali, Wolayta and Gumuz) to investigate the modes and targets of positive selection in this part of the world. The site frequency spectrum-based test SFselect was applied to idfentify a wide range of events of selection (old and recent), and the haplotype-based statistic integrated haplotype score to detect more recent events, in each case with evaluation of the significance of candidate signals by extensive simulations. Additional insights were provided by considering admixture proportions and functional categories of genes. We identified both individual loci that are likely targets of classic sweeps and groups of genes that may have experienced polygenic adaptation. We found population-specific as well as shared signals of selection, with folate metabolism and the related ultraviolet response and skin pigmentation standing out as a shared pathway, perhaps as a response to the high levels of ultraviolet irradiation, and in addition strong signals in genes such as IFNA, MRC1, immunoglobulins and T-cell receptors which contribute to defend against pathogens. Conclusions Signals of positive selection were detected in Ethiopian populations revealing novel adaptations in East Africa, and abundant targets for functional follow-up.
Natural selection acts on genetic variants by increasing the frequency of alleles responsible for a cellular function that is favorable in a certain environment. In a previous genome-wide scan for positive selection in contemporary humans, we identified a signal of positive selection in European and Asians at the genetic variant rs10180970. The variant is located in the second intron of the ABCA12 gene, which is implicated in the lipid barrier formation and down-regulated by UVB radiation. We studied the signal of selection in the genomic region surrounding rs10180970 in a larger dataset that includes DNA sequences from ancient samples. We also investigated the functional consequences of gene expression of the alleles of rs10180970 and another genetic variant in its proximity in healthy volunteers exposed to similar UV radiation. We confirmed the selection signal and refine its location that extends over 35 kb and includes the first intron, the first two exons and the transcription starting site of ABCA12 . We found no obvious effect of rs10180970 alleles on ABCA12 gene expression. We reconstructed the trajectory of the T allele over the last 80,000 years to discover that it was specific to H. sapiens and present in non-Africans 45,000 years ago.
Classic selective sweeps occur when positive selection increases a variant's frequency from low to high in a population, and underlie some long-studied human characteristics such as variation in skin, hair or eye colour. In such well-studied 'gold standard' examples, a known variant has been associated with a plausible phenotype and underlying selective force. Signatures of classic sweeps have more recently been detected in population-genetic data independently of any prior information about the corresponding phenotype or selective force, and usually without suggesting any insights into these. Motivated by the need to understand such candidates, we first review the gold standards and show that our understanding of them is often incomplete or unconvincing; only two of the examples we consider are compellingly explained. We assess approaches for large-scale association of classic sweep candidate variants to phenotypes and selective forces, test these on the gold standards, and discuss the standards of evidence needed to adequately understand a selective sweep.
We report high-coverage whole-genome sequencing data from 46 Yemeni individuals as well as genome-wide genotyping data from 169 Yemenis from diverse locations. We use this dataset to define the genetic diversity in Yemen and how it relates to people elsewhere in the Near East. Yemen is a vast region with substantial cultural and geographic diversity, but we found little genetic structure correlating with geography among the Yemenis – probably reflecting continuous movement of people between the regions. African ancestry from admixture in the past 800 years is widespread in Yemen and is the main contributor to the country’s limited genetic structure, with some individuals in Hudayda and Hadramout having up to 20% of their genetic ancestry from Africa. In contrast, individuals from Maarib appear to have been genetically isolated from the African gene flow and thus have genomes likely to reflect Yemen’s ancestry before the admixture. This ancestry was comparable to the ancestry present during the Bronze Age in the distant Northern regions of the Near East. After the Bronze Age, the South and North of the Near East therefore followed different genetic trajectories: in the North the Levantines admixed with a Eurasian population carrying steppe ancestry whose impact never reached as far south as the Yemen, where people instead admixed with Africans leading to the genetic structure observed in the Near East today.