Insights into the admixture history between modern and archaic humans require accurately inferred introgressed fragments within modern genomes. Here, we introduce two enhancements to hidden Markov models (HMMs) implemented in hmmix. First, we develop a method for sampling hidden state sequences conditional on observed genomic data, enabling robust estimation of admixture summary statistics-such as admixture proportion and fragment length distributions. This represents an improvement compared to relying solely on point estimates as provided by classical decoding methods. Additionally, we integrate the Finite Markov Chain Imbedding (FMCI) framework, allowing exact analytical calculation of these admixture statistics, tailored to large scale human genomes. Second, we implement a novel hybrid decoding method which combines the strengths of Viterbi and Posterior decoding methods, substantially improving the reliability of archaic fragments identified. We validate these improvements on data from the 1000 Genomes Project and demonstrate that our sampling method yields more accurate admixture estimates from single individuals compared to existing approaches requiring extensive population-level datasets. Moreover, we show how hybrid decoding can be instrumental in resolving the inference of local archaic haplotype structure in modern human genomes. These methodological advancements will enhance HMM-based analyses in any field of science and will provide deeper insight into the complex history of genetic interactions between archaic and modern human populations.
Summary Homo erectus is a species that occupies a central role in the study of human evolution. To investigate its taxonomic identity, Fu et al. 2026 extracted and sequenced enamel proteins from 6 fossils assigned to Homo erectus , originating from 3 localities in China and dated to around 400 thousand years ago. Unexpectedly, all samples possess an amino acid variant on the enamel protein ameloblastin (AMBN) which is uniquely shared with more recent Denisovan fossils and a subset of present day humans who are known to carry introgressed Denisovan ancestry. The samples also showcase an additional, novel variant on the same enamel protein. To explain this result, Fu et al. 2026 propose a model where the sampled Homo erectus population (or its recent ancestors) interbred with later arriving Denisovans, introducing one of the two AMBN variants into the Denisovan population. While this model fits our overall understanding of the interactions between these archaic populations, it rests on the assumption that the Denisovan AMBN variant has an archaic, Homo erectus- like, source. Here we show that the AMBN gene of late Denisovans has no signal of introgression from a ‘super-archaic’ source, making the suggested model unlikely for this gene. We also show that the AMBN gene of Denisova 25, an earlier Denisova, does show a signal of introgression, but one that better matches Neanderthals and modern humans, rather than a super-archaic source. We propose a number of alternative models that could help explain the observed affinity between the sampled H. erectus and Denisovans without requiring the introgression of AMBN from an archaic source into Denisovans. We discuss their strengths and weaknesses and how new data could help resolve them.
Two major tasks in applications of hidden Markov models are to (i) compute distributions of summary statistics of the hidden state sequence, and (ii) decode the hidden state sequence. We describe finite Markov chain imbedding (FMCI) and hybrid decoding to solve each of these two tasks. In the first part of our paper we use FMCI to compute posterior distributions of summary statistics such as the number of visits to a hidden state, the total time spent in a hidden state, the dwell time in a hidden state, and the longest run length. We use simulations from the hidden state sequence, conditional on the observed sequence, to establish the FMCI framework. In the second part of our paper we apply hybrid segmentation for improved decoding of a HMM. We demonstrate that hybrid decoding shows increased performance compared to Viterbi or Posterior decoding (often also referred to as global or local decoding), and we introduce a novel procedure for choosing the tuning parameter in the hybrid procedure. Furthermore, we provide an alternative derivation of the hybrid loss function based on weighted geometric means. We demonstrate and apply FMCI and hybrid decoding on various classical data sets, and supply accompanying code for reproducibility.
Archaic introgression from Neanderthals and Denisovans has contributed to the genomes of present-day humans, impacting traits such as our immune system. Most methods for detecting archaic ancestry rely on sequenced archaic reference genomes, which limits the detection of ancestry from unsampled or highly diverged archaic populations. Here, we discuss hmmix, a hidden Markov model–based method that infers introgressed genomic segments without requiring archaic reference data, but assumes the availability of an introgression-free outgroup population. We give a brief overview of hmmix; how it works, how to run it and which precomputed resources are available. We also discuss its imitations and recommendations on when and when not to use it and how to run it on species other than humans.
India has been underrepresented in genomic surveys. We generated whole-genome sequences from 2,762 individuals in India, capturing the genetic diversity across most geographic regions, linguistic groups, and historically underrepresented communities. We find most Indians harbor ancestry primarily from three ancestral groups: South Asian hunter-gatherers, Eurasian Steppe pastoralists, and Neolithic farmers related to Iranian and Central Asian cultures. The extensive homozygosity and identity-by-descent sharing among individuals reflects strong founder events due to a recent shift toward endogamy. We uncover that most of the genetic variation in Indians stems from a single major migration out of Africa that occurred around 50,000 years ago, followed by 1%-2% gene flow from Neanderthals and Denisovans. Notably, Indians exhibit the largest variation and possess the highest amount of population-specific Neanderthal ancestry segments among worldwide groups. Finally, we discuss how this complex evolutionary history has shaped the functional and disease variation on the subcontinent.
High-altitude environments pose significant challenges to human survival and reproduction, drawing considerable attention to the demographic and adaptive histories of populations in these regions. Here, we present whole-genome sequences from diverse Himalayan populations, offering new insights into the genomic history of this region. We find that population structure in the Himalayas began as early as 10,000 years ago, predating archaeological evidence of permanent habitation above 2,500 meters by ∼6,000 years. The widespread presence of the introgressed adaptive EPAS1 haplotype across all high-altitude populations highlights a shared genetic origin and its importance for survival in this region. We identify additional selection signals in genes linked to hypoxia, physical activity, immunity, and metabolism, all of which could have facilitated adaptation to the harsh environment. Over time, increasing genetic structure led to the emergence of the strongly differentiated ethnic groups observed today, many of which maintained small effective population sizes throughout their history or experienced severe bottlenecks. Between 6,000 and 3,000 years ago, with the advent of agriculture, a few uniparental lineages became predominant; however, significant population growth was not observed in the Himalayas except in the Tibetans. In recent history, we detect bidirectional gene flow between high-altitude and lowland groups on both sides of the Himalayan range, coinciding with the rise and expansion of historical regional powers, particularly during the Tibetan and northern Indian Gupta Empires. In the past few centuries, migrations to the Himalayas seem to have occurred alongside conflicts and population displacements in nearby regions and show some sex bias.
Spinocerebellar ataxia type 10 (SCA10) is a rare autosomal dominant ataxia caused by a large expansion of the (ATTCT)n repeat in ATXN10. SCA10 was described in Native American and Asian individuals which prompted a search for an expanded haplotype to confirm a common ancestral origin for the expansion event. All patients with SCA10 expansions in our cohort share a single haplotype defined at the 5'-end by the minor allele of rs41524547, located ~35 kb upstream of the SCA10 expansion. Intriguingly, rs41524547 is located within the miRNA gene, MIR4762, within its DROSHA cleavage site and just outside the seed sequence for mir4792-5p. The world-wide frequency of rs41524547-G is less than 5% and found almost exclusively in the Americas and East Asia-a geographic distribution that mirrors reported SCA10 cases. We identified rs41524547-G(+) DNA from the 1000 Genomes/International Genome Sample Resource and our own general population samples and identified SCA10 repeat expansions in up to 25% of these samples. The reduced penetrance of these SCA10 expansions may be explained by a young (pre-onset) age at sample collection, a small repeat size, purity of repeat units, or the disruption of miR4762-5p function. We conclude that rs41524547-G is the most robust at-risk SNP allele for SCA10, is useful for screening of SCA10 expansions in population genetics studies and provides the most compelling evidence to date for a single, prehistoric origin of SCA10 expansions sometime prior to or during the migration of individuals across the Bering Land Bridge into the Americas.
India has been underrepresented in whole genome sequencing studies. We generated 2,762 high coverage genomes from India-including individuals from most geographic regions, speakers of all major languages, and tribal and caste groups-providing a comprehensive survey of genetic variation in India. With these data, we reconstruct the evolutionary history of India through space and time at fine scales. We show that most Indians derive ancestry from three ancestral groups related to ancient Iranian farmers, Eurasian Steppe pastoralists and South Asian hunter-gatherers. We uncover a common source of Iranian-related ancestry from early Neolithic cultures of Central Asia into the ancestors of Ancestral South Indians (ASI), Ancestral North Indians (ANI), Austro-asiatic-related and East Asian-related groups in India. Following these admixtures, India experienced a major demographic shift towards endogamy, resulting in extensive homozygosity and identity-by-descent sharing among individuals. At deep time scales, Indians derive around 1-2% of their ancestry from gene flow from archaic hominins, Neanderthals and Denisovans. By assembling the surviving fragments of archaic ancestry in modern Indians, we recover ~1.5 Gb (or 50%) of the introgressing Neanderthal and ~0.6 Gb (or 20%) of the introgressing Denisovan genomes, more than any other previous archaic ancestry study. Moreover, Indians have the largest variation in Neanderthal ancestry, as well as the highest amount of population-specific Neanderthal segments among worldwide groups. Finally, we demonstrate that most of the genetic variation in Indians stems from a single major migration out of Africa that occurred around 50,000 years ago, with minimal contribution from earlier migration waves. Together, these analyses provide a detailed view of the population history of India and underscore the value of expanding genomic surveys to diverse groups outside Europe.
Gene flow from Neanderthals has shaped genetic and phenotypic variation in modern humans. We generated a catalog of Neanderthal ancestry segments in more than 300 genomes spanning the past 50,000 years. We examined how Neanderthal ancestry is shared among individuals over time. Our analysis revealed that the vast majority of Neanderthal gene flow is attributable to a single, shared extended period of gene flow that occurred between 50,500 to 43,500 years ago, as evidenced by ancestry correlation, colocalization of Neanderthal segments across individuals, and divergence from the sequenced Neanderthals. Most natural selection—positive and negative—on Neanderthal variants occurred rapidly after the gene flow. Our findings provide new insights into how contact with Neanderthals shaped modern human origins and adaptation.
The X chromosome in non-African humans shows less diversity and less Neanderthal introgression than ex-pected from neutral evolution. Analyzing 162 human male X chromosomes worldwide, we identified fourteen chromosomal regions where nearly identical haplotypes spanning several hundred kilobases are found at high frequencies in non-Africans. Genetic drift alone cannot explain the existence of these haplotypes, which must have been associated with strong positive selection in partial selective sweeps. Moreover, the swept haplotypes are entirely devoid of archaic ancestry as opposed to the non-swept haplotypes in the same genomic regions. The ancient Ust'-Ishim male dated at 45,000 before the present (BP) also carries the swept haplotypes, implying that selection on the haplotypes must have occurred between 45,000 and 55,000 years ago. Finally, we find that the chromosomal positions of sweeps overlap previously reported hotspots of selective sweeps in great ape evolution, suggesting a mechanism of selection unique to X chromosomes.
A major part of the human Y chromosome consists of palindromes with multiple copies of genes primarily expressed in testis, many of which have been claimed to affect male fertility. Here we examine copy number variation in these palindromes based on whole genome sequence data from 11,527 Icelandic men. Using a subset of 7947 men grouped into 1449 patrilineal genealogies, we infer 57 large scale de novo copy number mutations affecting palindrome 1. This corresponds to a mutation rate of 2.34 × 10 −3 mutations per meiosis, which is 4.1 times larger than our phylogenetic estimate of the mutation rate (5.72 × 10 −4 ), suggesting that de novo mutations on the Y are lost faster than expected under neutral evolution. Although simulations indicate a selection coefficient of 1.8% against non-reference copy number carriers, we do not observe differences in fertility among sequenced men associated with their copy number genotype, but we lack statistical power to detect differences resulting from weak negative selection. We also perform association testing of a diverse set of 341 traits to palindromic copy number without any significant associations. We conclude that large-scale palindrome copy number variation on the Y chromosome has little impact on human phenotype diversity.
Genomic analyses of Neanderthals have previously provided insights into their population history and relationship to modern humans 1 – 8 , but the social organization of Neanderthal communities remains poorly understood. Here we present genetic data for 13 Neanderthals from two Middle Palaeolithic sites in the Altai Mountains of southern Siberia: 11 from Chagyrskaya Cave 9 , 10 and 2 from Okladnikov Cave 11 —making this one of the largest genetic studies of a Neanderthal population to date. We used hybridization capture to obtain genome-wide nuclear data, as well as mitochondrial and Y-chromosome sequences. Some Chagyrskaya individuals were closely related, including a father–daughter pair and a pair of second-degree relatives, indicating that at least some of the individuals lived at the same time. Up to one-third of these individuals’ genomes had long segments of homozygosity, suggesting that the Chagyrskaya Neanderthals were part of a small community. In addition, the Y-chromosome diversity is an order of magnitude lower than the mitochondrial diversity, a pattern that we found is best explained by female migration between communities. Thus, the genetic data presented here provide a detailed documentation of the social organization of an isolated Neanderthal community at the easternmost extent of their known range.
Much remains unknown about the population history of early modern humans in southeast Asia, where the archaeological record is sparse and the tropical climate is inimical to the preservation of ancient human DNA 1 . So far, only two low-coverage pre-Neolithic human genomes have been sequenced from this region. Both are from mainland Hòabìnhian hunter-gatherer sites: Pha Faen in Laos, dated to 7939–7751 calibrated years before present (yr cal bp; present taken as ad 1950), and Gua Cha in Malaysia (4.4–4.2 kyr cal bp ) 1 . Here we report, to our knowledge, the first ancient human genome from Wallacea, the oceanic island zone between the Sunda Shelf (comprising mainland southeast Asia and the continental islands of western Indonesia) and Pleistocene Sahul (Australia–New Guinea). We extracted DNA from the petrous bone of a young female hunter-gatherer buried 7.3–7.2 kyr cal bp at the limestone cave of Leang Panninge 2 in South Sulawesi, Indonesia. Genetic analyses show that this pre-Neolithic forager, who is associated with the ‘Toalean’ technocomplex 3 , 4 , shares most genetic drift and morphological similarities with present-day Papuan and Indigenous Australian groups, yet represents a previously unknown divergent human lineage that branched off around the time of the split between these populations approximately 37,000 years ago 5 . We also describe Denisovan and deep Asian-related ancestries in the Leang Panninge genome, and infer their large-scale displacement from the region today.
After the main out-of-Africa event, humans interbred with Neanderthals leaving 1-2% of Neanderthal DNA scattered in small fragments in all non-African genomes today1,2. Here we investigate the size distribution of these fragments in non-African genomes3. We find consistent differences in fragment length distributions across Eurasia with 11% longer fragments in East Asians than in West Eurasians. By comparing extant populations and ancient samples, we show that these differences are due to a different rate of decay in length by recombination since the Neanderthal admixture. In line with this, we observe a strong correlation between the average fragment length and the accumulation of derived mutations, similar to what is expected by changing the ages at reproduction as estimated from trio studies4. Altogether, our results suggest consistent differences in the generation interval across Eurasia, by up to 20% (e.g. 25 versus 30 years), over the past 40,000 years. We use sex-specific accumulations of derived alleles to infer how these changes in generation intervals between geographical regions could have been mainly driven by shifts in either male or female age of reproduction, or both. We also find that previously reported variation in the mutational spectrum5 may be largely explained by changes to the generation interval and not by changes to the underlying mutational mechanism. We conclude that Neanderthal fragment lengths provide unique insight into differences of a key demographic parameter among human populations over the recent history.
Modern humans appeared in Europe by at least 45,000 years ago 1 – 5 , but the extent of their interactions with Neanderthals, who disappeared by about 40,000 years ago 6 , and their relationship to the broader expansion of modern humans outside Africa are poorly understood. Here we present genome-wide data from three individuals dated to between 45,930 and 42,580 years ago from Bacho Kiro Cave, Bulgaria 1 , 2 . They are the earliest Late Pleistocene modern humans known to have been recovered in Europe so far, and were found in association with an Initial Upper Palaeolithic artefact assemblage. Unlike two previously studied individuals of similar ages from Romania 7 and Siberia 8 who did not contribute detectably to later populations, these individuals are more closely related to present-day and ancient populations in East Asia and the Americas than to later west Eurasian populations. This indicates that they belonged to a modern human migration into Europe that was not previously known from the genetic record, and provides evidence that there was at least some continuity between the earliest modern humans in Europe and later people in Eurasia. Moreover, we find that all three individuals had Neanderthal ancestors a few generations back in their family history, confirming that the first European modern humans mixed with Neanderthals and suggesting that such mixing could have been common.
AbstractWe sequenced the genome of a Neandertal from Chagyrskaya Cave in the Altai Mountains, Russia, to 27-fold genomic coverage. We estimate that this individual lived ~80,000 years ago and was more closely related to Neandertals in western Eurasia (1,2) than to Neandertals who lived earlier in Denisova Cave (3), which is located about 100 km away. About 12.9% of theChagyrskayagenome is spanned by homozygous regions that are between 2.5 and 10 centiMorgans (cM) long. This is consistent with that Siberian Neandertals lived in relatively isolated populations of less than 60 individuals. In contrast, a Neandertal from Europe, a Denisovan from the Altai Mountains and ancient modern humans seem to have lived in populations of larger sizes. The availability of three Neandertal genomes of high quality allows a first view of genetic features that were unique to Neandertals and that are likely to have been at high frequency among them. We find that genes highly expressed in the striatum in the basal ganglia of the brain carry more amino acid-changing substitutions than genes expressed elsewhere in the brain, suggesting that the striatum may have evolved unique functions in Neandertals.
We present analyses of the genome of a ~34,000-year-old hominin skull cap discovered in the Salkhit Valley in northeastern Mongolia. We show that this individual was a female member of a modern human population that, following the split between East and West Eurasians, experienced substantial gene flow from West Eurasians. Both she and a 40,000-year-old individual from Tianyuan outside Beijing carried genomic segments of Denisovan ancestry. These segments derive from the same Denisovan admixture event(s) that contributed to present-day mainland Asians but are distinct from the Denisovan DNA segments in present-day Papuans and Aboriginal Australians.
Human evolutionary history is rich with the interbreeding of divergent populations. Most humans outside of Africa trace about 2% of their genomes to admixture from Neanderthals, which occurred 50–60 thousand years ago 1 . Here we examine the effect of this event using 14.4 million putative archaic chromosome fragments that were detected in fully phased whole-genome sequences from 27,566 Icelanders, corresponding to a range of 56,388–112,709 unique archaic fragments that cover 38.0–48.2% of the callable genome. On the basis of the similarity with known archaic genomes, we assign 84.5% of fragments to an Altai or Vindija Neanderthal origin and 3.3% to Denisovan origin; 12.2% of fragments are of unknown origin. We find that Icelanders have more Denisovan-like fragments than expected through incomplete lineage sorting. This is best explained by Denisovan gene flow, either into ancestors of the introgressing Neanderthals or directly into humans. A within-individual, paired comparison of archaic fragments with syntenic non-archaic fragments revealed that, although the overall rate of mutation was similar in humans and Neanderthals during the 500 thousand years that their lineages were separate, there were differences in the relative frequencies of mutation types—perhaps due to different generation intervals for males and females. Finally, we assessed 271 phenotypes, report 5 associations driven by variants in archaic fragments and show that the majority of previously reported associations are better explained by non-archaic variants.
We sequenced the genome of a Neandertal from Chagyrskaya Cave in the Altai Mountains, Russia, to 27-fold genomic coverage. We show that this Neandertal was a female and that she was more related to Neandertals in western Eurasia [Prüfer et al., Science 358, 655-658 (2017); Hajdinjak et al., Nature 555, 652-656 (2018)] than to Neandertals who lived earlier in Denisova Cave [Prüfer et al., Nature 505, 43-49 (2014)], which is located about 100 km away. About 12.9% of the Chagyrskaya genome is spanned by homozygous regions that are between 2.5 and 10 centiMorgans (cM) long. This is consistent with the fact that Siberian Neandertals lived in relatively isolated populations of less than 60 individuals. In contrast, a Neandertal from Europe, a Denisovan from the Altai Mountains, and ancient modern humans seem to have lived in populations of larger sizes. The availability of three Neandertal genomes of high quality allows a view of genetic features that were unique to Neandertals and that are likely to have been at high frequency among them. We find that genes highly expressed in the striatum in the basal ganglia of the brain carry more amino-acid-changing substitutions than genes expressed elsewhere in the brain, suggesting that the striatum may have evolved unique functions in Neandertals.