Speciation genomics researcher, author of the combinatorial theory of speciation and passionate birder
Genomes of high latitude killer whales (Orcinus orca) harbour signatures of post-glacial founding and expansion. Here, we investigate whether reduced efficacy of selection increased the predicted mutation burden in founder populations with low Ne, or whether recessive deleterious mutations exposed to selection in homozygous genotypes were purged. Comparing the accumulation of synonymous and non-synonymous mutations across pairs of globally sampled genomes reveals that high latitude North Atlantic genomes have accumulated proportionally fewer non-synonymous mutations than other populations. The genome of a 7.5-Kyr-old North Atlantic killer whale, inferred to be closely related to the direct ancestors of the present-day Icelandic and Norwegian populations, was used to calibrate the timing of the action of selection on non-synonymous mutations predominantly to the Holocene. Non-synonymous mutations purged in modern Norwegian killer whale genomes are found as globally shared standing variation in heterozygote genotypes more often than expected, suggesting associative- or pseudo-overdominance. Taken together, our findings are consistent with purging of recessive non-synonymous mutations exposed to selection in founder-associated homozygous genotypes.
Purifying selection is a critical factor in shaping genetic diversity. Current theoretical models mostly address scenarios of either very weak or strong selection, leaving a significant gap in our knowledge. The effects of purifying selection on patterns of genomic diversity remain poorly understood when selection against deleterious mutations is weak to moderate, particularly when recombination is limited or absent. In this study, we extend an existing approach, the fitness-class coalescent, to incorporate arbitrary levels of purifying selection in haploid populations. This model offers a comprehensive framework for exploring the influence of purifying selection in a wide range of demographic scenarios. Moreover, our research reveals potential sources of qualitative and quantitative biases in demographic inference, highlighting the significant risk of attributing genetic patterns to past demographic events rather than purifying selection. This work expands our understanding of the complex interplay between selection, drift, and population dynamics, and how purifying selection distorts demographic inference.
Most species have been through population bottlenecks and range expansions, and the impact of these events on patterns of diversity has been well studied. In particular, it has been shown that initially rare neutral variants could readily fix on the front of range expansions or during bottlenecks, giving genomic signatures looking like selective sweeps. Here we expand on previous work by considering the dynamics of genomic diversity in (functional) regions harboring deleterious variants during bottlenecks or during range expansions modeled as serial founder effects. We find that regions with very low levels of diversity (troughs) looking like selective sweeps can also readily form in these functional regions. Additionally, their properties depend on the dominance level of deleterious mutations. The number of troughs is larger and increases more rapidly in regions with co-dominant deleterious mutations than in regions with recessive mutations. Interestingly, we find that genetic diversity declines less rapidly in regions with partially recessive mutations than in regions with codominant ones or in regions with only neutral mutations. These features are generally enhanced in low recombination regions and for intermediate selection coefficients. If most deleterious mutations in a genome are partially recessive, it follows that functional low recombination regions should better preserve genetic diversity during range expansions than neutral regions of the genome.
The characterization of genes and biological functions underlying functional diversification and the formation of species is a major goal of evolutionary biology. In this study, we investigated the fast radiation of Microtus voles, one of the most speciose group of mammals, which shows strong genetic divergence despite few readily observable morphological differences. We produced an annotated reference genome for the common vole, Microtus arvalis, and resequenced the genomes of 10 different species and evolutionary lineages spanning the Microtus speciation continuum. Our full-genome sequences illustrate the recent and fast diversification of this group, and we identified genes in highly divergent genomic windows that have likely particular roles in their radiation. We found three biological functions enriched for highly divergent genes in most Microtus species and lineages: olfaction, immunity and metabolism. In particular, olfaction-related genes (mostly olfactory receptors and vomeronasal receptors) are fast evolving in all Microtus species indicating the exceptional importance of the olfactory system in the evolution of these rodents. Of note is e.g. the shared signature among vole species on Olfr1019 which has been associated with fear responses against predator odors in rodents. Our analyses provide a genome-wide basis for the further characterization of the ecological factors and processes of natural and sexual selection that have contributed to the fast radiation of Microtus voles.
Modern and ancient genomes are not necessarily drawn from homogeneous populations, as they may have been collected from different places and at different times. This heterogeneous sampling can be an issue for demographic inferences and results in biased demographic parameters and incorrect model choice if not properly considered. When explicitly accounted for, it can result in very complex models and high data dimensionality that are difficult to analyse. In this paper, we formally study the impact of such spatial and temporal sampling heterogeneity on demographic inference, and we introduce a way to circumvent this problem. To deal with structured samples without increasing the dimensionality of the site frequency spectrum (SFS), we introduce a new structured approach to the existing program fastsimcoal2. We assess the efficiency and relevance of this methodological update with simulated and modern human genomic data. We particularly focus on spatial and temporal heterogeneities to evidence the interest of this new SFS-based approach, which can be especially useful when handling scattered and ancient DNA samples, as in conservation genetics or archaeogenetics.
Admixture is a common biological phenomenon among populations of the same or different species. Identifying admixed tracts within individual genomes can provide valuable information to date admixture events, reconstruct ancestry-specific demographic histories, or detect adaptive introgression, genetic incompatibilities, as well as regions of the genomes affected by (associative-) overdominance. Although many local ancestry inference (LAI) methods have been developed in the last decade, their performance was accessed using large reference panels, which are rarely available for non-model organisms or ancient samples. Moreover, the demographic conditions for which LAI becomes unreliable have not been explicitly outlined. Here, we identify the demographic conditions for which local ancestries can be best estimated using very small reference panels. Furthermore, we compare the performance of two LAI methods (RFMix and MOSAIC) with the performance of a newly developed approach (simpLAI) that can be used even when reference populations consist of single individuals. Based on simulations of various demographic models, we also determine the limits of these LAI tools and propose post-painting filtering steps to reduce false-positive rates and improve the precision and accuracy of the inferred admixed tracts. Besides providing a guide for using LAI, our work shows that reasonable inferences can be obtained from a single diploid genome per reference under demographic conditions that are not uncommon among past human groups and non-model organisms.
Although some lineages of animals and plants have made impressive adaptive radiations when provided with ecological opportunity, the propensities to radiate vary profoundly among lineages for unknown reasons. In Africa's Lake Victoria region, one cichlid lineage radiated in every lake, with the largest radiation taking place in a lake less than 16,000 years old. We show that all of its ecological guilds evolved in situ. Cycles of lineage fusion through admixture and lineage fission through speciation characterize the history of the radiation. It was jump-started when several swamp-dwelling refugial populations, each of which were of older hybrid descent, met in the newly forming lake, where they fused into a single population, resuspending old admixture variation. Each population contributed a different set of ancient alleles from which a new adaptive radiation assembled in record time, involving additional fusion-fission cycles. We argue that repeated fusion-fission cycles in the history of a lineage make adaptive radiation fast and predictable.
A strong reduction in diversity around a specific locus is often interpreted as a recent rapid fixation of a positively selected allele, a phenomenon called a selective sweep. Rapid fixation of neutral variants can however lead to similar reduction in local diversity, especially when the population experiences changes in population size, e.g., bottlenecks or range expansions. The fact that demographic processes can lead to signals of nucleotide diversity very similar to signals of selective sweeps is at the core of an ongoing discussion about the roles of demography and natural selection in shaping patterns of neutral variation. Here we quantitatively investigate the shape of such neutral valleys of diversity under a simple model of a single population size change, and we compare it to signals of a selective sweep. We analytically describe the expected shape of such “neutral sweeps” and show that selective sweep valleys of diversity are, for the same fixation time, wider than neutral valleys. On the other hand, it is always possible to parametrize our model to find a neutral valley that has the same width as a given selected valley. We apply our framework to the case of a putative selective sweep signal around the gene Quetzalcoatl in D. melanogaster and show that the valley of diversity in the vicinity of this gene is compatible with a short bottleneck scenario without selection. Our findings provide further insight in how simple demographic models can create valleys of genetic diversity that may falsely be attributed to positive selection.
Range expansions have been common in the history of most species. Serial founder effects and subsequent population growth at expansion fronts typically lead to a loss of genomic diversity along the expansion axis. A frequent consequence is the phenomenon of "gene surfing," where variants located near the expanding front can reach high frequencies or even fix in newly colonized territories. Although gene surfing events have been characterized thoroughly for a specific locus, their effects on linked genomic regions and the overall patterns of genomic diversity have been little investigated. In this study, we simulated the evolution of whole genomes during several types of 1D and 2D range expansions differing by the extent of migration, founder events, and recombination rates. We focused on the characterization of local dips of diversity, or "troughs," taken as a proxy for surfing events. We find that, for a given recombination rate, once we consider the amount of diversity lost since the beginning of the expansion, it is possible to predict the initial evolution of trough density and their average width irrespective of the expansion condition. Furthermore, when recombination rates vary across the genome, we find that troughs are over-represented in regions of low recombination. Therefore, range expansions can leave local and global genomic signatures often interpreted as evidence of past selective events. Given the generality of our results, they could be used as a null model for species having gone through recent expansions, and thus be helpful to correctly interpret many evolutionary biology studies.
The field of population genomics has grown rapidly in response to the recent advent of affordable, large-scale sequencing technologies. As opposed to the situation during the majority of the 20th century, in which the development of theoretical and statistical population-genetic insights out-paced the generation of data to which they could be applied, genomic data are now being produced at a far greater rate than they can be meaningfully analyzed and interpreted. With this wealth of data has come a tendency to focus on fitting specific (and often rather idiosyncratic) models to data, at the expense of a careful exploration of the range of possible underlying evolutionary processes. For example, the approach of directly investigating models of adaptive evolution in each newly sequenced population or species often neglects the fact that a thorough characterization of ubiquitous non-adaptive processes is a prerequisite for accurate inference. We here describe the perils of these tendencies, present our consensus views on current best practices in population genomic data analysis, and highlight areas of statistical inference and theory that are in need of further attention. Thereby, we argue for the importance of defining a biologically relevant baseline model tuned to the details of each new analysis, of skepticism and scrutiny in interpreting model-fitting results, and of carefully defining addressable hypotheses and underlying uncertainties.
The precise genetic origins of the first Neolithic farming populations in Europe and Southwest Asia, as well as the processes and the timing of their differentiation, remain largely unknown. Demogenomic modeling of high-quality ancient genomes reveals that the early farmers of Anatolia and Europe emerged from a multiphase mixing of a Southwest Asian population with a strongly bottlenecked western hunter-gatherer population after the last glacial maximum. Moreover, the ancestors of the first farmers of Europe and Anatolia went through a period of extreme genetic drift during their westward range expansion, contributing highly to their genetic distinctiveness. This modeling elucidates the demographic processes at the root of the Neolithic transition and leads to a spatial interpretation of the population history of Southwest Asia and Europe during the late Pleistocene and early Holocene.
The Pacific region is of major importance for addressing questions regarding human dispersals, interactions with archaic hominins and natural selection processes(1). However, the demographic and adaptive history of Oceanian populations remains largely uncharacterized. Here we report high-coverage genomes of 317 individuals from 20 populations from the Pacific region. We find that the ancestors of Papuan-related ('Near Oceanian') groups underwent a strong bottleneck before the settlement of the region, and separated around 20,000-40,000 years ago. We infer that the East Asian ancestors of Pacific populations may have diverged from Taiwanese Indigenous peoples before the Neolithic expansion, which is thought to have started from Taiwan around 5,000 years ago(2-4). Additionally, this dispersal was not followed by an immediate, single admixture event with Near Oceanian populations, but involved recurrent episodes of genetic interactions. Our analyses reveal marked differences in the proportion and nature of Denisovan heritage among Pacific groups, suggesting that independent interbreeding with highly structured archaic populations occurred. Furthermore, whereas introgression of Neanderthal genetic information facilitated the adaptation of modern humans related to multiple phenotypes (for example, metabolism, pigmentation and neuronal development), Denisovan introgression was primarily beneficial for immune-related functions. Finally, we report evidence of selective sweeps and polygenic adaptation associated with pathogen exposure and lipid metabolism in the Pacific region, increasing our understanding of the mechanisms of biological adaptation to island environments.
Current procedures for inferring population history generally assume complete neutrality-that is, they neglect both direct selection and the effects of selection on linked sites. We here examine how the presence of direct purifying selection and background selection may bias demographic inference by evaluating two commonly-used methods (MSMC and fastsimcoal2), specifically studying how the underlying shape of the distribution of fitness effects and the fraction of directly selected sites interact with demographic parameter estimation. The results show that, even after masking functional genomic regions, background selection may cause the mis-inference of population growth under models of both constant population size and decline. This effect is amplified as the strength of purifying selection and the density of directly selected sites increases, as indicated by the distortion of the site frequency spectrum and levels of nucleotide diversity at linked neutral sites. We also show how simulated changes in background selection effects caused by population size changes can be predicted analytically. We propose a potential method for correcting for the mis-inference of population growth caused by selection. By treating the distribution of fitness effect as a nuisance parameter and averaging across all potential realizations, we demonstrate that even directly selected sites can be used to infer demographic histories with reasonable accuracy.
Formulating strategies for species conservation requires knowledge of evolutionary and genetic history. Tigers are among the most charismatic of endangered species and garner significant conservation attention. However, the evolutionary history and genomic variation of tigers remain poorly known. With 70% of the worlds wild tigers living in India, such knowledge is critical for tiger conservation. We re-sequenced 65 individual tiger genomes across their extant geographic range, representing most extant subspecies with a specific focus on tigers from India. As suggested by earlier studies, we found strong genetic differentiation between the putative tiger subspecies. Despite high total genomic diversity in India, individual tigers host longer runs of homozygosity, potentially suggesting recent inbreeding, possibly because of small and fragmented protected areas. Surprisingly, demographic models suggest recent divergence (within the last 10,000 years) between populations, and strong population bottlenecks. Amur tiger genomes revealed the strongest signals of selection mainly related to metabolic adaptation to cold, while Sumatran tigers show evidence of evolving under weak selection for genes involved in body size regulation. Depending on conservation objectives, our results support the isolation of Amur and Sumatran tigers, while geneflow between Malayan and South Asian tigers may be considered. Further, the impacts of ongoing connectivity loss on the health and persistence of tigers in India should be closely monitored.
Isobiotic mice, with an identical stable microbiota composition, potentially allow models of host-microbial mutualism to be studied over time and between different laboratories. To understand microbiota evolution in these models, we carried out a 6-year experiment in mice colonized with 12 representative taxa. Increased non-synonymous to synonymous mutation rates indicate positive selection in multiple taxa, particularly for genes annotated for nutrient acquisition or replication. Microbial sub-strains that evolved within a single taxon can stably coexist, consistent with niche partitioning of ecotypes in the complex intestinal environment. Dietary shifts trigger rapid transcriptional adaptation to macronutrient and micronutrient changes in individual taxa and alterations in taxa biomass. The proportions of different sub-strains are also rapidly altered after dietary shift. This indicates that microbial taxa within a mouse colony adapt to changes in the intestinal environment by long-term genomic positive selection and short-term effects of transcriptional reprogramming and adjustments in sub-strain proportions.
Matters Arising article 1 raised concerns about the interpretation of our findings reported in our recent publication on admixture-facilitated ecological speciation in Lake Constance stickleback 2 .After careful consideration of the criticism, including additional analyses testing the proposed alternative hypotheses, we can confirm our confidence in the inference of secondary contact between a West European and an East European stickleback lineage in the catchment of Lake Constance, and that this admixture facilitated the ecological divergence between lake and stream ecotypes within Lake Constance 2 .In particular, Berner 1 (i) questioned whether West and East European stickleback populations should be considered as divergent lineages, (ii) suggested that Lake Constance stickleback originated from the upper Danube instead of East Europe, (iii) questioned the suitability of our demographic modelling approach to reject an 'ecological vicariance' scenario, (iv) proposed that divergent selection within Lake Constance biased our inference of a secondary contact and admixture scenario, and (v) criticized our conclusion on admixture-facilitation of ecological speciation as premature.We address each of these concerns in this sequence.
In the last ten years, the next generation sequencing revolution has multiplied the amount of genetic data for many organisms by orders of magnitude. This has not only led to evolutionary biologists having more data available but also to new and different types of data: from a handful of allozyme markers in the 70s, we got dozens of restriction fragment length polymorphisms (RFLPs) in the 80s, hundreds of microsatellites in the 90s, thousands to hundreds of thousands of single nucleotide polymorphisms (SNPs) in the 2000s, a few full genomes in the 2010s, and thousands of full genomes in the 2020s. These data have provided information not only on the genetic diversity and evolution of the organisms studied but also on genome-wide patterns of selection, linkage disequilibrium, as well as recombination and mutation processes. Below, we will describe how these new genomic data can be used to infer the past demographic history of populations.
Genomes of high latitude killer whales harbour signatures of post-glacial founding and expansion. Here, we investigate whether reduced efficacy of selection increased mutation load in founder populations, or whether recessive deleterious mutations exposed to selection in homozygous genotypes were purged. Comparing the accumulation of synonymous and non-synonymous mutations across pairs of globally sampled genomes reveals that the most significant outliers are high latitude North Atlantic genomes, which have accumulated significantly fewer non-synonymous mutations than all other populations. Comparisons with the genome of a 7.5-Kyr-old North Atlantic killer whale, inferred to be closely related to the population directly ancestral to present-day Icelandic and Norwegian populations, calibrates the timing of the action of selection on non-synonymous mutations predominantly to the mid-late Holocene. Non-synonymous mutations purged in modern Norwegian killer whale genomes are found as globally shared standing variation in heterozygote genotypes more often than expected, suggesting overdominance. Taken together, our findings are consistent with purging of recessive non-synonymous mutations exposed to selection in founder-associated homozygous genotypes. ### Competing Interest Statement The authors have declared no competing interest.