Many or all present-day human genomes carry segments of DNA inherited from archaic humans due to admixture events occurring >25,000 years ago. Several methods have been published to detect such segments. These variously require phased haplotypes, an archaic reference sequence and/or an outgroup with little related archaic introgression. Inference from these approaches have been used to document purifying and positive selection of archaic segments and to understand their influence on the phenotype. However, the comparative accuracy of different methods at detecting archaic segments, and how well inferred segments overlap across methods, remains underexplored. Here, we used demographic simulations to evaluate accuracy in detecting archaic ancestry under three widely-used approaches: SPrime, IBDMix and HMmix, and a technique introduced here, cp-archaic, as well as applying all approaches to 1000 Genomes data. Our results reveal substantial variation in method performance, with recall and precision ranges differing significantly across approaches and parameter settings. cp-archaic achieved the highest accuracy overall (F1, the harmonic mean of precision and recall=0.92-0.94 across three different simulations). The choice of demographic simulation substantially impacted the distribution and characteristics of true archaic segments, with demographic scenarios creating both archaic ancestry ‘deserts’ and regions of very high archaic ancestry in the absence of selection. Notably, we also found low agreement between methods, with <22% overlap in individual archaic sites detected across all four approaches in real data. These findings highlight how conclusions about archaic introgression patterns, population differences, and genome-wide coverage depend critically on both methodological choice and underlying demographic assumptions. We recommend careful method selection and parameter optimisation, as well as caution when interpreting individual archaic segments, particularly in comparative studies across populations.
Abstract Characterized by the earliest use of pottery, the Jomon culture was a unique Neolithic culture that spread throughout the Japanese Archipelago. Previous archaeological evidence suggests that Jomon hunter-gatherers colonized the southernmost islands, the Ryukyu Archipelago, by approximately 7,000 years before present (YBP). However, genetic characteristics of the Ryukyu Jomon population and its contribution to the modern population have not been elucidated yet. In this study, we newly sequenced 273 modern and 25 ancient (6,700–900 YBP) whole genomes collected across the Ryukyu Archipelago. Our analysis demonstrated a genetic differentiation between the Hondo (Japanese mainland) and Ryukyu Jomon, dating back to ∼6,900 YBP. After the divergence from the Hondo Jomon, the Ryukyu Jomon experienced severe bottlenecks, with an effective population size of ∼2,000. Admixture between the Ryukyu Jomon and migrants from the historic Hondo population occurred ∼1,000 YBP, which corresponds to the widespread adoption of iron tools and agriculture in the Central Ryukyus. Different demographic histories between modern Hondo and Ryukyu populations resulted in different rates of Jomon ancestry in these populations. By providing a new perspective on the peopling of the Ryukyu Archipelago, this study significantly enhances our understanding of cultural transitions in the region.
Abstract At rare human genomic regions, DNA methylation states are established in the early embryo and maintained during cellular differentiation, yielding systemic (i.e. not tissue-specific) interindividual epigenetic variation. Previous screens for such correlated regions of systemic interindividual variation (CoRSIVs) were limited to White Americans. Here, we describe the first human CoRSIV screen including self-identified Black and White Americans. We integrate deep whole-genome bisulfite sequencing data for three tissues from each of ten Black and ten White donors in the NIH Genotype-Tissue Expression program. This approach identifies twice as many CoRSIVs among Black than White Americans. Establishment of CoRSIV methylation is sensitive to periconceptional environmental exposures including assisted reproduction, seasonal variation, and famine. CoRSIV-associated genes are enriched for GWAS variants linked to cancer and neurodevelopment. Although only 15% of Black CoRSIVs overlap with those among White individuals, both sets are associated with the same subfamilies of transposable elements. Within multiple cell lines, ranked enrichments of transcription factor binding to Black and White CoRSIVs are exquisitely coordinated and related to genome organization, indicating that CoRSIV methylation states established in the early embryo play an important role in guiding subsequent cellular differentiation.
BACKGROUND:Recessive dystrophic epidermolysis bullosa (RDEB) is a rare and severe blistering skin disorder caused by loss-of-function mutations in the type VII collagen gene (COL7A1). The COL7A1 c.6527insC mutation is curiously prevalent among individuals with RDEB and is found worldwide in Europe and the Americas. Previous research has suggested the possibility of a Sephardic Jewish origin of the mutation; however, individuals with RDEB are not known to have predominant Jewish ancestry. METHODS:In this study, a global cohort of individuals with RDEB with the c.6527insC founder mutation from Spain, France, Argentina, Chile, Colombia and the USA were investigated by autosomal genotyping, pairwise identical-by-descent matching and a local ancestry analysis. Age estimation analysis was performed to determine when Jewish founders introduced the c.6527insC mutation into Iberian and Native American populations (~900 CE and 1492 CE, respectively). RESULTS:Sephardic ancestry was identified at the haplotype spanning the c.6527insC mutation in 85% of the individuals, despite mixed ancestry elsewhere in the genome and no known recent Sephardic ancestry. Identical-by-descent matching between this RDEB subpopulation and a known crypto-Jewish community in Belmonte, Portugal was also ascertained, providing support for crypto-Jewish ancestry in this RDEB subpopulation. CONCLUSION:The identification of this unique RDEB subpopulation unified by the single most prevalent c.6527insC mutation holds great potential to facilitate promising new RDEB therapies using CRISPR Cas 9 gene and base editing. The identification of a single guide RNA allowing efficient and safe editing of this variant would represent a unique drug to treat a large cohort of patients with the same founder mutation.
Understanding genetic differences between populations is essential for avoiding confounding in genome-wide association studies and improving polygenic score (PGS) portability. We developed a statistical pipeline to infer fine-scale Ancestry Components and applied it to UK Biobank data. Ancestry Components identify population structure not captured by widely used principal components, improving stratification correction for geographically correlated traits. To estimate the similarity of genetic effect sizes between groups, we developed ANCHOR, which estimates changes in the predictive power of an existing PGS in distinct local ancestry segments. ANCHOR infers highly similar (estimated correlation 0.98 ± 0.07) effect sizes between UK Biobank participants of African and European ancestry for 47 of 53 quantitative phenotypes, suggesting that gene–environment and gene–gene interactions do not play major roles in poor cross-ancestry PGS transferability for these traits in the United Kingdom, and providing optimism that shared causal mutations operate similarly in different populations. This study introduces the concept of Ancestry Components and shows that they can offer improved population stratification correction for geographically correlated traits. By using ancestry-aware polygenic score construction in admixed individuals, the authors find that effect sizes are conserved across ancestry groups.
The history of human populations has been strongly shaped by admixture events, contributing to patterns of observed genetic diversity across populations. In this study, we introduce the Principal component Ancestry proportions using NNLS Estimation (PANE) method that leverages principal component analysis and non-negative least squares to assess the ancestral compositions of admixed individuals given a large set of populations. Our results show its ability to reliably estimate ancestry across several scenarios, even those with a significant proportion of missing genotypes, in a fraction of the time required when using other tools.
The North Eurasian forest and forest-steppe zones have sustained millennia of sociocultural connections among northern peoples, but much of their history is poorly understood. In particular, the genomic formation of populations that speak Uralic and Yeniseian languages today is unknown. Here, by generating genome-wide data for 180 ancient individuals spanning this region, we show that the Early-to-Mid-Holocene hunter-gatherers harboured a continuous gradient of ancestry from fully European-related in the Baltic, to fully East Asian-related in the Transbaikal. Contemporaneous groups in Northeast Siberia were off-gradient and descended from a population that was the primary source for Native Americans, which then mixed with populations of Inland East Asia and the Amur River Basin to produce two populations whose expansion coincided with the collapse of pre-Bronze Age population structure. Ancestry from the first population, Cis-Baikal Late Neolithic-Bronze Age (Cisbaikal_LNBA), is associated with Yeniseian-speaking groups and those that admixed with them, and ancestry from the second, Yakutia Late Neolithic-Bronze Age (Yakutia_LNBA), is associated with migrations of prehistoric Uralic speakers. We show that Yakutia_LNBA first dispersed westwards from the Lena River Basin around 4,000 years ago into the Altai-Sayan region and into West Siberian communities associated with Seima-Turbino metallurgy-a suite of advanced bronze casting techniques that expanded explosively from the Altai1. The 16 Seima-Turbino period individuals were diverse in their ancestry, also harbouring DNA from Indo-Iranian-associated pastoralists and from a range of hunter-gatherer groups. Thus, both cultural transmission and migration were key to the Seima-Turbino phenomenon, which was involved in the initial spread of early Uralic-speaking communities.
The Slavs are a major ethnolinguistic group of Europe, yet the process that led to their formation remains disputed. As of the sixth century CE, people supposedly belonging to the Slavs populated the space between the Avar Khaganate in the Carpathian Basin, the Merovingian Frankish Empire to the West and the Balkan Peninsula to the South. Proposed theories to explain those events are, however, conceptually incompatible, as some invoke major population movements while others stress the continuity of local populations. We report high-quality genomic data of 18 individuals from two nearby burial sites in South Moravia that span from the fifth to the tenth century CE, during which the region became the core of the ninth century Slavic principality. In contrast to existing data, the individuals reported here can be directly connected to an Early-Slavic-associated culture and include the earliest known inhumation associated with any such culture. The data indicates a strong genetic shift incompatible with local continuity between the fifth and seventh century, supporting the notion that the Slavic expansion in South Moravia was driven by population movement.
Background The six admixing baboon species offer a natural experiment to study negative selection on admixture and the nature of genetic incompatibility. In Tanzania, a secondary contact between olive and yellow baboons allows admixture despite 1.3 million years of divergence. An independent secondary contact occurred in Ethiopia, upon which olive baboons invaded and displaced an ancient Hamadryas-like population separated by 0.6 million years, mirroring the displacement of Neanderthals by modern humans. Results We analyze 156 high-coverage genomes sampled from seven olive and four yellow baboon populations in East Africa. Analyzing local ancestry across the whole genome, we find evidence of negative selection on minor parent ancestry in both Tanzanian yellow and olive baboon populations that reaches far beyond the hybrid zone. Across populations, we find that selection on minor parent ancestry is stronger on the X chromosome than on autosomes, most extremely in one yellow baboon population, which shows a seven-fold difference. The proportion of minor parent ancestry (MPA) is substantially higher on the X chromosome in Ethiopian olive and yellow baboon populations, which both displaced the populations now representing their minor parent ancestry, owing mainly to a few genomic regions with MPA at very high frequencies. We hypothesize that strong negative selection on MPA allowed these X chromosome regions to retain the original ancestry, as this was slowly displaced across the remaining genome. Conclusions Our findings provide deeper insights into admixture dynamics in primates, highlighting the persistence of selection against admixture across various levels of admixture, and underscoring the need to include chromosome X in admixture analyses.
BACKGROUND:The recent emergence of technologies that capture and analyse genetic variation patterns obtained from a person's DNA sample has led to numerous academic and commercial endeavours to infer individuals' ancestries. In theory, a person's genome contains a wealth of readily accessible information regarding their ancestors, despite only some of our ancestors contributing to the DNA we carry. This makes genetic tests an attractive alternative to the painstaking reconstruction of family trees or directly contacting long-lost relations, particularly when, unless there are notable individuals in the tree, historical and genealogical records tend to diminish in frequency with each generation. However, while powerful, there are limits to what genetic data can unearth, as well as important assumptions underlying these analyses. METHODS:This review describes some of the early history and latest advances in techniques and data used to infer ancestry using genetics, highlighting both the power and limitations of current studies. CONCLUSION:While genetics is a powerful means of exploring aspects of people's ancestry, a stronger focus on conveying uncertainty will allow both academics and non-academics to avoid the ever-present risks of over-interpretation.
Variable DNA methylation states during early embryonic development overlap with mechanisms of transposon silencing by Krüppel-associated box zinc finger proteins (KZFPs). We investigated the influence of genetic variation in KZFPs and their transposon targets and identified a variably methylated region (VMR) proximal to the human LY6S-AS1 gene linked to an intronic MER11C retrotransposon targeted by the primate-specific KZFP ZNF808. Mendelian randomisation analysis supported a causal link between VMR methylation and type 1 diabetes risk, and H1 stem cells with inactivated ZNF808 showed marked, transient upregulation of the LY6S-AS1 transcript during the early stages of pancreatic development. Loss of function genetic mutations in ZNF808 have previously been linked to pancreatic agenesis and neonatal diabetes in humans, and the VMR was previously associated with levels of insulin secretion in Gambian children. Together this evidence points to a link between KZFP-mediated DNA methylation of LY6S-AS1 , and pancreatic development and function at this locus. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was funded by UK Medical Research Council grants MR/ T032863/1 and MC\_PC\_21037. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethical approval for all human data analysed in this study was given by the joint Gambia Government/Medical Research Council Unit The Gambia Ethics Committee. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Population history-focused DNA and ancient DNA (aDNA) research in Africa has dramatically increased in the past decade, enabling increasingly fine-scale investigations into the continent's past. However, while international interest in human genomics research in Africa grows, major structural barriers limit the ability of African scholars to lead and engage in such research and impede local communities from partnering with researchers and benefitting from research outcomes. Because conversations about research on African people and their past are often held outside Africa and exclude African voices, an important step for African DNA and aDNA research is moving these conversations to the continent. In May 2023 we held the DNAirobi workshop in Nairobi, Kenya and here we synthesize what emerged most prominently in our discussions. We propose an ideal vision for population history-focused DNA and aDNA research in Africa in ten years' time and acknowledge that to realize this future, we need to chart a path connecting a series of "landmarks" that represent points of consensus in our discussions. These include effective communication across multiple audiences, reframed relationships and capacity building, and action toward structural changes that support science and beyond. We concluded there is no single path to creating an equitable and self-sustaining research ecosystem, but rather many possible routes linking these landmarks. Here we share our diverse perspectives as geneticists, anthropologists, archaeologists, museum curators, and educators to articulate challenges and opportunities for African DNA and aDNA research and share an initial map toward a more inclusive and equitable future.
First identified in isogenic mice, metastable epialleles (MEs) are loci where the extent of DNA methylation (DNAm) is variable between individuals but correlates across tissues derived from different germ layers within a given individual. This property, termed systemic interindividual variation (SIV), is attributed to stochastic methylation establishment before germ layer differentiation. Evidence suggests that some putative human MEs are sensitive to environmental exposures in early development. In this review we introduce key concepts pertaining to human MEs, describe methods used to identify MEs in humans, and review their genomic features. We also highlight studies linking DNAm at putative human MEs to early environmental exposures and postnatal (including disease) phenotypes.
The six admixing baboon species offer a natural experiment to study negative selection on admixture and potentially genetic incompatibility in species with divergence resembling that between anatomically modern and extinct human subspecies such as Neanderthals. We analyze 156 high-coverage genomes sampled from seven olive and four yellow baboon populations shaped by admixture in two different reticulation events. In Tanzania, olive and yellow baboons admix despite 1.3 million years of divergence. In Ethiopia, olive baboons have invaded and displaced an ancient Hamadryas-like population separated by 0.6 million years, thus mirroring the displacement of Neanderthals by modern humans. Analyzing local ancestry across whole genome, we find evidence of negative selection on minor parent ancestry in Tanzanian yellow and olive baboon populations reaching far behind the hybrid zone and in Ethiopian olive baboons. We find that evidence of selection against admixture is up to seven times stronger on the X chromosome. The proportion of minor parent ancestry is substantially higher on the X chromosome in Ethiopian olive and yellow baboon populations, which displaced the populations now representing their minor parent ancestry. This additional MPA is concentrated in a few genomic regions with high frequency. This indicates that this original ancestry was retained by negative selection on the invading ancestry, suggesting that these loci play a role in emerging reproductive barriers in the Papio as well as the Homo genus. ### Competing Interest Statement The authors have declared no competing interest.
Populations of the Eastern Highlands of Papua New Guinea (EHPNG, area 11,157 km2) lived in relative isolation from the rest of the world until the mid-20th century, and the region contains a wealth of linguistic and cultural diversity. Notably, several populations of EHPNG were devastated by an epidemic prion disease, kuru, which at its peak in the mid-twentieth century led to some villages being almost depleted of adult women. Until now, population genetic analyses to learn about genetic diversity, migration, admixture, and the impact of the kuru epidemic have been restricted to a small number of variants or samples. Here, we present a population genetic analysis of the region based on genome-wide genotype data of 943 individuals from 21 linguistic groups and 68 villages in EHPNG, including 34 villages in the South Fore linguistic group, the group most affected by kuru. We find a striking degree of genetic population structure in the relatively small region (average FST between linguistic groups 0.024). The genetic population structure correlates well with linguistic grouping, with some noticeable exceptions that reflect the clan system of community organization that has historically existed in EHPNG. We also detect the presence of migrant individuals within the EHPNG region and observe a significant excess of females among migrants compared to among non-migrants in areas of high kuru exposure (p = 0.0145, chi-squared test). This likely reflects the continued practice of patrilocality despite documented fears and strains placed on communities as a result of kuru and its associated skew in female incidence.
An understanding of genetic differences between populations is essential for avoiding confounding in genome-wide association studies (GWAS) and understanding the evolution of human traits. Polygenic risk scores constructed in one group perform poorly in highly genetically-differentiated populations, for reasons which remain controversial. We developed a statistical ancestry inference pipeline able to decompose ancestry both within and between countries, and applied it to the UK Biobank data. This identifies fine-scale patterns of genetic relatedness not captured by standard and widely used principal components (PCs), and allows fine-scale population stratification correction that removes both false positive and false negative associations for traits with geographic correlations. We also develop and apply ANCHOR, an approach leveraging segments of distinct ancestries within individuals to estimate similarity in underlying causal effect sizes between groups, using an existing PGS. Applying ANCHOR to >8000 people of mixed African and European ancestry, we demonstrate that estimated causal effect sizes are highly similar across these ancestries for 26 of 29 quantitative molecular and non-molecular phenotypes (mean correlation 0.98 +/-0.08), providing evidence that gene-environment and gene-gene interactions do not play major roles in the poor prediction of European-ancestry PRS scores in African populations for these traits, contradicting previous findings. Instead our results provide optimism that shared causal mutations operate similarly in different groups, focussing the challenge of improving GWAS “portability” between groups on joint fine-mapping.
Previous studies have highlighted how African genomes have been shaped by a complex series of historical events. Despite this, genome-wide data have only been obtained from a small proportion of present-day ethnolinguistic groups. By analyzing new autosomal genetic variation data of 1333 individuals from over 150 ethnic groups from Cameroon, Republic of the Congo, Ghana, Nigeria, and Sudan, we demonstrate a previously underappreciated fine-scale level of genetic structure within these countries, for example, correlating with historical polities in western Cameroon. By comparing genetic variation patterns among populations, we infer that many northern Cameroonian and Sudanese groups share genetic links with multiple geographically disparate populations, likely resulting from long-distance migrations. In Ghana and Nigeria, we infer signatures of intermixing dated to over 2000 years ago, corresponding to reports of environmental transformations possibly related to climate change. We also infer recent intermixing signals in multiple African populations, including Congolese, that likely relate to the expansions of Bantu language–speaking peoples.