Abstract Uterine disorders and menstrual abnormalities are prevalent reproductive conditions with significant clinical consequences. Recent genome-wide association studies have identified hundreds of variants contributing to different uterine disorders, while epidemiological evidence suggests that these disorders co-occur more frequently than expected, implying shared genetic mechanisms. However, the specific genetic variants and biological pathways underlying this shared architecture remain poorly characterized. Here, we conducted a uterus-centric, multi-trait genome-wide association analysis across ten uterine disorders in a European cohort to elucidate this shared genetic architecture. Further, we embedded this architecture in a functional and population genomics framework to investigate its plausible biological mechanisms and its evolutionary history in humans. We confirm strong positive genetic correlations between major uterine disorders and identify 31 independent susceptibility loci jointly affecting the genetic risk of multiple uterine pathologies, substantiating an intertwined biological basis. Populational analyses demonstrated that several of these susceptibility variants exhibit pronounced allele frequency differentiation across global populations suggestive of recent polygenic selection, including variants with well-supported regulatory functions at the ESR1-CCDC170 , WNT4 , SFR1, FOXO1, ITPR1, DMRT1 and CDKN2B loci. Notably, we show that most derived alleles acquired during recent human evolution increase risk across multiple uterine disorders and may evolve under antagonistic selection. These findings provide functional annotation and population-level prioritization of genetic variants influencing multiple uterine disorders. They also highlight how past evolutionary histories may contribute to population differences in uterine disease prevalence and pathogenesis.
P. vivax, the most geographically widespread human malaria parasite with millions of clinical cases per year, is however quasi absent in sub-Saharan Africa. Positive selection targeting the rs2814778 protective mutation, also known as the Duffy-null allele, may explain the absence (or quasi absence) of vivax in sub-Saharan Africa by a progressive purge of the pathogen due to a quasi-fixation of the Duffy-null allele and the resulting high rates of protected carriers in western, central and eastern populations. Yet, while positive selection has been clearly evidenced in admixed populations coexisting with vivax, the selection model currently admitted poorly explains the lack of the Duffy-null allele in Europe, or in Asia where the pathogen is mainly observed. In this article, several validated Deep Learning methods applied to high coverage sequence data obtained in 589 African individuals resolved this retention of the Duffy-null resistance to vivax in Africa. The CNN and GAN algorithms implemented in this study also predict a rise in frequency of the Duffy-null mutation due to selection 25-35 kya years ago in the western part of Africa, a geographical region and a time frame overlapping with the rise of another protective mutation, βS, the sickle-cell mutation protective at heterozygous state against the malaria caused by P. falciparum. In addition, the pattern of Duffy-null haplotypes highlights a quick spread of the Duffy-null allele in sub-Saharan Africa due to post-admixture selection events following the road of the recent Bantu expansion. Independent lines of evidence describing malaria as a life-threatening disease in West Africa from ~30 kya, together with a rise in frequency followed by recent disseminations of the Duffy-null resistance, open new perspectives about both the history of malaria as a major human disease and the history of the main protective mutations in Africa. ### Competing Interest Statement The authors have declared no competing interest.
Leveraging past allele frequencies has proven to be key for identifying the impact of natural selection across time. However, this approach suffers from imprecise estimations of the intensity (s) and timing (T) of selection, particularly when ancient samples are scarce in specific epochs. Here, we aimed to bypass the computation of allele frequencies across arbitrarily defined past epochs and refine the estimations of selection parameters by implementing convolutional neural networks (CNNs) algorithms that directly use ancient genotypes sampled across time. Using computer simulations, we first show that genotype-based CNNs consistently outperform an approximate Bayesian computation (ABC) approach based on past allele frequency trajectories, regardless of the selection model assumed and the number of available ancient genotypes. When applying this method to empirical data from modern and ancient Europeans, we replicated the reported increased number of selection events in post-Neolithic Europe, independently of the continental subregion studied. Furthermore, we substantially refined the ABC-based estimations of s and T for a set of positively and negatively selected variants, including iconic cases of positive selection and experimentally validated disease-risk variants. Our CNN predictions support a history of recent positive and negative selection targeting variants associated with host defence against pathogens, aligning with previous work that highlights the significant impact of infectious diseases, such as tuberculosis, in Europe. These findings collectively demonstrate that detecting the footprints of natural selection on ancient genomes is crucial for unravelling the history of severe human diseases.
Ancient genomics can directly detect human genetic adaptation to environmental cues. However, it remains unclear how pathogens have exerted selective pressures on human genome diversity across different epochs and affected present-day inflammatory disease risk. Here, we use an ancestry-aware approximate Bayesian computation framework to estimate the nature, strength, and time of onset of selection acting on 2,879 ancient and modern European genomes from the last 10,000 years. We found that the bulk of genetic adaptation occurred after the start of the Bronze Age, <4,500 years ago, and was enriched in genes relating to host-pathogen interactions. Furthermore, we detected directional selection acting on specific leukocytic lineages and experimentally demonstrated that the strongest negatively selected candidate variant in immunity genes, lipopolysaccharide-binding protein (LBP) D283G, is hypomorphic. Finally, our analyses suggest that the risk of inflammatory disorders has increased in post-Neolithic Europeans, possibly because of antagonistic pleiotropy following genetic adaptation to pathogens.
Admixture has been a pervasive phenomenon in human history, extensively shaping the patterns of population genetic diversity. There is increasing evidence to suggest that admixture can also facilitate genetic adaptation to local environments, i.e., admixed populations acquire beneficial mutations from source populations, a process that we refer to as “adaptive admixture.” However, the role of adaptive admixture in human evolution and the power to detect it remain poorly characterized. Here, we use extensive computer simulations to evaluate the power of several neutrality statistics to detect natural selection in the admixed population, assuming multiple admixture scenarios. We show that statistics based on admixture proportions, Fadm and LAD, show high power to detect mutations that are beneficial in the admixed population, whereas other statistics, including iHS and FST, falsely detect neutral mutations that have been selected in the source populations only. By combining Fadm and LAD into a single, powerful statistic, we scanned the genomes of 15 worldwide, admixed populations for signatures of adaptive admixture. We confirm that lactase persistence and resistance to malaria have been under adaptive admixture in West Africans and in Malagasy, North Africans, and South Asians, respectively. Our approach also uncovers other cases of adaptive admixture, including APOL1 in Fulani nomads and PKN2 in East Indonesians, involved in resistance to infection and metabolism, respectively. Collectively, our study provides evidence that adaptive admixture has occurred in human populations whose genetic history is characterized by periods of isolation and spatial expansions resulting in increased gene flow.
The Pacific region is of major importance for addressing questions regarding human dispersals, interactions with archaic hominins and natural selection processes(1). However, the demographic and adaptive history of Oceanian populations remains largely uncharacterized. Here we report high-coverage genomes of 317 individuals from 20 populations from the Pacific region. We find that the ancestors of Papuan-related ('Near Oceanian') groups underwent a strong bottleneck before the settlement of the region, and separated around 20,000-40,000 years ago. We infer that the East Asian ancestors of Pacific populations may have diverged from Taiwanese Indigenous peoples before the Neolithic expansion, which is thought to have started from Taiwan around 5,000 years ago(2-4). Additionally, this dispersal was not followed by an immediate, single admixture event with Near Oceanian populations, but involved recurrent episodes of genetic interactions. Our analyses reveal marked differences in the proportion and nature of Denisovan heritage among Pacific groups, suggesting that independent interbreeding with highly structured archaic populations occurred. Furthermore, whereas introgression of Neanderthal genetic information facilitated the adaptation of modern humans related to multiple phenotypes (for example, metabolism, pigmentation and neuronal development), Denisovan introgression was primarily beneficial for immune-related functions. Finally, we report evidence of selective sweeps and polygenic adaptation associated with pathogen exposure and lipid metabolism in the Pacific region, increasing our understanding of the mechanisms of biological adaptation to island environments.
During their dispersals over the last 100,000 years, modern humans have been exposed to a large variety of environments, resulting in genetic adaptation. While genome-wide scans for the footprints of positive Darwinian selection have increased knowledge of genes and functions potentially involved in human local adaptation, they have globally produced evidence of a limited contribution of selective sweeps in humans. Conversely, studies based on machine learning algorithms suggest that recent sweeps from standing variation are widespread in humans, an observation that has been recently questioned. Here, we sought to formally quantify the number of recent selective sweeps in humans, by leveraging approximate Bayesian computation and whole-genome sequence data. Our computer simulations revealed suitable ABC estimations, regardless of the frequency of the selected alleles at the onset of selection and the completion of sweeps. Under a model of recent selection from standing variation, we inferred that an average of 68 (from 56 to 79) and 140 (from 94 to 198) sweeps occurred over the last 100,000 years of human history, in African and Eurasian populations, respectively. The former estimation is compatible with human adaptation rates estimated since divergence with chimps, and reveals numbers of sweeps per generation per site in the range of values estimated in Drosophila. Our results confirm the rarity of selective sweeps in humans and show a low contribution of sweeps from standing variation to recent human adaptation.
Tuberculosis (TB), usually caused by Mycobacterium tuberculosis bacteria, is the first cause of death from an infectious disease at the worldwide scale, yet the mode and tempo of TB pressure on humans remain unknown. The recent discovery that homozygotes for the P1104A polymorphism of TYK2 are at higher risk to develop clinical forms of TB provided the first evidence of a common, monogenic predisposition to TB, offering a unique opportunity to inform on human co-evolution with a deadly pathogen. Here, we investigate the history of human exposure to TB by determining the evolutionary trajectory of the TYK2 P1104A variant in Europe, where TB is considered to be the deadliest documented infectious disease. Leveraging a large dataset of 1,013 ancient human genomes and using an approximate Bayesian computation approach, we find that the P1104A variant originated in the common ancestors of West Eurasians ∼30,000 years ago. Furthermore, we show that, following large-scale population movements of Anatolian Neolithic farmers and Eurasian steppe herders into Europe, P1104A has markedly fluctuated in frequency over the last 10,000 years of European history, with a dramatic decrease in frequency after the Bronze Age. Our analyses indicate that such a frequency drop is attributable to strong negative selection starting ∼2,000 years ago, with a relative fitness reduction on homozygotes of 20%, among the highest in the human genome. Together, our results provide genetic evidence that TB has imposed a heavy burden on European health over the last two millennia.
The Roma Diaspora-traditionally known as Gypsies-remains among the least explored population migratory events in historical times. It involved the migration of Roma ancestors out-of-India through the plateaus of Western Asia ultimately reaching Europe. The demographic effects of the Diaspora-bottlenecks, endogamy, and gene flow-might have left marked molecular traces in the Roma genomes. Here, we analyze the whole-genome sequence of 46 Roma individuals pertaining to four migrant groups in six European countries. Our analyses revealed a strong, early founder effect followed by a drastic reduction of ∼44% in effective population size. The Roma common ancestors split from the Punjabi population, from Northwest India, some generations before the Diaspora started, <2,000 years ago. The initial bottleneck and subsequent endogamy are revealed by the occurrence of extensive runs of homozygosity and identity-by-descent segments in all Roma populations. Furthermore, we provide evidence of gene flow from Armenian and Anatolian groups in present-day Roma, although the primary contribution to Roma gene pool comes from non-Roma Europeans, which accounts for >50% of their genomes. The linguistic and historical differentiation of Roma in migrant groups is confirmed by the differential proportion, but not a differential source, of European admixture in the Roma groups, which shows a westward cline. In the present study, we found that despite the strong admixture Roma had in their diaspora, the signature of the initial bottleneck and the subsequent endogamy is still present in Roma genomes.
Selective pressures imposed by pathogens have varied among human populations throughout their evolution, leading to marked inter-population differences at some genes mediating susceptibility to infectious and immune-related diseases. Here, we investigated the evolutionary history of a common polymorphism resulting in a Y-529 versus C-529 change in the cadherin related family member 3 (CDHR3) receptor which underlies variable susceptibility to rhinovirus-C infection and is associated with severe childhood asthma. The protective variant is the derived allele and is found at high frequency worldwide (69-95%). We detected genome-wide significant signatures of natural selection consistent with a rapid increase of the haplotypes carrying the allele, suggesting that non-neutral processes have acted on this locus across all human populations. However, the allele has not fixed in any population despite multiple lines of evidence suggesting that the mutation predates human migrations out of Africa. Using an approximate Bayesian computation method, we estimate the age of the mutation while explicitly accounting for past demography and positive or frequency-dependent balancing selection. Our analyses indicate a single emergence of the mutation in anatomically modern humans similar to 150 000 years ago and indicate that balancing selection has maintained the beneficial allele at high equilibrium frequencies worldwide. Apart from the well-known cases of the MHC and ABO genes, this study provides the first evidence that negative frequency-dependent selection plausibly acted on a human disease susceptibility locus, a form of balancing selection compatible with typical transmission dynamics of communicable respiratory viruses that might exploit CDHR3.
The hemoglobin beta(S) sickle mutation is a textbook case in which natural selection maintains a deleterious mutation at high frequency in the human population. Homozygous individuals for this mutation develop sickle-cell disease, whereas heterozygotes benefit from higher protection against severe malaria. Because the overdominant beta(S) allele should be purged almost immediately from the population in the absence of malaria, the study of the evolutionary history of this iconic mutation can provide important information about the history of human exposure to malaria. Here, we sought to increase our understanding of the origins and time depth of the beta(S) mutation in populations with different lifestyles and ecologies, and we analyzed the diversity of HBB in 479 individuals from 13 populations of African farmers and rainforest hunter-gatherers. Using an approximate Bayesian computation method, we estimated the age of the beta(S) allele while explicitly accounting for population subdivision, past demography, and balancing selection. When the effects of balancing selection are taken into account, our analyses indicate a single emergence of beta(S) in the ancestors of present-day agriculturalist populations similar to 22,000 years ago. Furthermore, we show that rainforest hunter-gatherers have more recently acquired the beta(S) mutation from the ancestors of agriculturalists through adaptive gene flow during the last similar to 6,000 years. Together, our results provide evidence for a more ancient exposure to malarial pressures among the ancestors of agriculturalists than previously appreciated, and they suggest that rainforest hunter-gatherers have been increasingly exposed to malaria during the last millennia.
Over the last 100,000 years, humans have spread across the globe and encountered a highly diverse set of environments to which they have had to adapt. Genome-wide scans of selection are powerful to detect selective sweeps. However, because of unknown fractions of undetected sweeps and false discoveries, the numbers of detected sweeps often poorly reflect actual numbers of selective sweeps in populations. The thousands of soft sweeps on standing variation recently evidenced in humans have also been interpreted as a majority of mis-classified neutral regions. In such a context, the extent of human adaptation remains little understood. We present a new rationale to estimate these actual numbers of sweeps expected over the last 100,000 years (denoted byX) from genome-wide population data, both considering hard sweeps and selective sweeps on standing variation. We implemented an approximate Bayesian computation framework and showed, based on computer simulations, that such a method can properly estimateX. We then jointly estimated the number of selective sweeps, their mean intensity and age in several 1000G African, European and Asian populations. Our estimations ofX, found weakly sensitive to demographic misspecifications, revealed very limited numbers of sweeps regardless the frequency of the selected alleles at the onset of selection and the completion of sweeps. We estimated ∼80 sweeps in average across fifteen 1000G populations when assuming incomplete sweeps only and ∼140 selective sweeps in non-African populations when incorporating complete sweeps in our simulations. The method proposed may help to address controversies on the number of selective sweeps in populations, guiding further genome-wide investigations of recent positive selection.
Selective pressures imposed by pathogens have varied among human populations throughout their evolution, leading to marked inter-population differences at some genes mediating susceptibility to infectious and immune-related diseases. A common polymorphism resulting in a C529 versus T529 change in the Cadherin-Related Family Member 3 ( CDHR3 ) receptor is associated with rhinovirus-C (RV-C) susceptibility and severe childhood asthma. Given the morbidity and mortality associated with RV-C dependent respiratory infections and asthma, we hypothesized that the protective variant has been under selection in the human population. Supporting this idea, a recent cross-species outbreak of RV-C among chimpanzees in Uganda, which carry the ancestral ‘risk’ allele at this position, resulted in a mortality rate of 8.9%. Using publicly available genomic data, we sought to determine the evolutionary history and role of selection acting on this infectious disease susceptibility locus. The protective variant is the derived allele and is found at high frequency worldwide, with the lowest relative frequency in African populations and highest in East Asian populations. There is minimal population structure among haplotypes, and we detect genomic signatures consistent with a rapid increase in frequency of the protective allele across all human populations. However, given strong evidence that the protective allele arose in anatomically modern humans prior to their migrations out of Africa and that the allele has not fixed in any population, the patterns observed here are not consistent with a classical selective sweep. We hypothesize that patterns may indicate frequency-dependent selection worldwide. Irrespective of the mode of selection, our analyses show the derived allele has been subject to selection in recent human evolution.
Elucidating population structure and levels of genetic diversity and recombination is necessary to understand the evolution and adaptation of species. Candida albicans is the second most frequent agent of human fungal infections worldwide, causing high-mortality rates. Here we present the genomic sequences of 182 C . albicans isolates collected worldwide, including commensal isolates, as well as ones responsible for superficial and invasive infections, constituting the largest dataset to date for this major fungal pathogen. Although, C . albicans shows a predominantly clonal population structure, we find evidence of gene flow between previously known and newly identified genetic clusters, supporting the occurrence of (para)sexuality in nature. A highly clonal lineage, which experimentally shows reduced fitness, has undergone pseudogenization in genes required for virulence and morphogenesis, which may explain its niche restriction. Candida albicans thus takes advantage of both clonality and gene flow to diversify.
Bantu languages are spoken by about 310 million Africans, yet the genetic history of Bantu-speaking populations remains largely unexplored. We generated genomic data for 1318 individuals from 35 populations in western central Africa, where Bantu languages originated. We found that early Bantu speakers first moved southward, through the equatorial rainforest, before spreading toward eastern and southern Africa. We also found that genetic adaptation of Bantu speakers was facilitated by admixture with local populations, particularly for the HLA and LCT loci. Finally, we identified a major contribution of western central African Bantu speakers to the ancestry of African Americans, whose genomes present no strong signals of natural selection. Together, these results highlight the contribution of Bantu-speaking peoples to the complex genetic history of Africans and African Americans.
Leprosy is a human infectious disease caused by Mycobacterium leprae. A strong host genetic contribution to leprosy susceptibility is well established. However, the modulation of the transcriptional response to infection and the mechanism(s) of disease control are poorly understood. To address this gap in knowledge of leprosy pathogenicity, we conducted a genome-wide search for expression quantitative trait loci (eQTL) that are associated with transcript variation before and after stimulation with M. leprae sonicate in whole blood cells. We show that M. leprae antigen stimulation mainly triggered the upregulation of immune related genes and that a substantial proportion of the differential gene expression is genetically controlled. Indeed, using stringent criteria, we identified 318 genes displaying cis-eQTL at an FDR of 0.01, including 66 genes displaying response-eQTL (reQTL), i.e. cis-eQTL that showed significant evidence for interaction with the M. leprae stimulus. Such reQTL correspond to regulatory variations that affect the interaction between human whole blood cells and M. leprae sonicate and, thus, likely between the human host and M. leprae bacilli. We found that reQTL were significantly enriched among binding sites of transcription factors that are activated in response to infection, and that they were enriched among single nucleotide polymorphisms (SNPs) associated with susceptibility to leprosy per se and Type-I Reaction, and seven of them have been targeted by recent positive selection. Our study suggested that natural selection shaped our genomic diversity to face pathogen exposure including M. leprae infection.
Humans differ in the outcome that follows exposure to life-threatening pathogens, yet the extent of population differences in immune responses and their genetic and evolutionary determinants remain undefined. Here, we characterized, using RNA sequencing, the transcriptional response of primary monocytes from Africans and Europeans to bacterial and viral stimuli-ligands activating Toll-like receptor pathways (TLR1/2, TLR4, and TLR7/8) and influenza virus-and mapped expression quantitative trait loci (eQTLs). We identify numerous cis-eQTLs that contribute to the marked differences in immune responses detected within and between populations and a strong trans-eQTL hotspot at TLR1 that decreases expression of pro-inflammatory genes in Europeans only. We find that immune-responsive regulatory variants are enriched in population-specific signals of natural selection and show that admixture with Neandertals introduced regulatory variants into European genomes, affecting preferentially responses to viral challenges. Together, our study uncovers evolutionarily important determinants of differences in host immune responsiveness between human populations.
Human genes governing innate immunity provide a valuable tool for the study of the selective pressure imposed by microorganisms on host genomes. A comprehensive, genome-wide study of how selective constraints and adaptations have driven the evolution of innate immunity genes is missing. Using full-genome sequence variation from the 1000 Genomes Project, we first show that innate immunity genes have globally evolved under stronger purifying selection than the remainder of protein-coding genes. We identify a gene set under the strongest selective constraints, mutations in which are likely to predispose individuals to life-threatening disease, as illustrated by STAT1 and TRAF3. We then evaluate the occurrence of local adaptation and detect 57 high-scoring signals of positive selection at innate immunity genes, variation in which has been associated with susceptibility to common infectious or autoimmune diseases. Furthermore, we show that most adaptations targeting coding variation have occurred in the last 6,000-13,000 years, the period at which populations shifted from hunting and gathering to farming. Finally, we show that innate immunity genes present higher Neandertal introgression than the remainder of the coding genome. Notably, among the genes presenting the highest Neandertal ancestry, we find the TLR6-TLR1-TLR10 cluster, which also contains functional adaptive variation in Europeans. This study identifies highly constrained genes that fulfill essential, non-redundant functions in host survival and reveals others that are more permissive to change containing variation acquired from archaic hominins or adaptive variants in specific populations improving our understanding of the relative biological importance of innate immunity pathways in natural conditions.