Understanding how reproductive isolation (RI) evolves in the context of gene flow is central to explaining how new species arise. Theory predicts that the accumulation of barrier loci is shaped by the interplay between divergent selection, recombination, and genomic architecture, yet empirical tests spanning multiple stages of divergence at the intraspecific level remain rare. The pea aphid complex, comprising multiple sympatric biotypes specialised on different legume hosts and spanning a continuum of RI, provides an exceptional opportunity to examine how ecological adaptation drives reproductive isolation under ongoing gene flow. Using whole-genome sequencing data from 13 sympatric European biotypes, we constructed a dense, reference-genome anchored picture of the divergence continuum that controls for genetic background and recombination landscape. Joint patterns of genetic differentiation, absolute divergence and genetic diversity, as well as recombination and introgression rates, reveal a consistent signature of divergence with gene flow and allow robust identification of barrier loci. Across the divergence gradient, RI is highly polygenic, yet barrier loci are clustered and enriched in low-recombination regions and large chromosomal rearrangements, which promote strong linkage disequilibrium and coupling among barrier effects. While genome-wide FST and the number of barrier loci gradually accumulate as divergence increases, the dynamics of differentiation within barrier loci show stepwise increases consistent with theoretical "tipping points" in the accumulation of RI. Barrier loci are enriched in salivary effector, detoxification and chemosensory genes, highlighting their central role in host plant specialisation and RI. These candidate genes are dynamically recruited at different stages of divergence, and some of them are shared among independent biotype comparisons, suggesting generic functional determinants of host plant specialisation and reproductive isolation acquired via parallel evolution or introgression. Our study provides a comprehensive view of the genomic architecture and evolutionary dynamics underlying speciation with gene flow, illustrating how tightly linked, polygenic architectures can facilitate adaptation and diversification in natural populations. ### Competing Interest Statement The authors have declared no competing interest. ANR MECADAPT N° ANR-18-CE02-0012-03
Large-scale genomic resources can place genetic variation into an ecologically informed context. To advance our understanding of the population genetics of the fruit fly Drosophila melanogaster, we present an expanded release of the community-generated population genomics resource Drosophila Evolution over Space and Time (DEST 2.0; https://dest.bio/). This release includes 530 high-quality pooled libraries from flies collected across six continents over more than a decade (2009 to 2021), most at multiple time points per year; 211 of these libraries are sequenced and shared here for the first time. We used this enhanced resource to elucidate several aspects of the species' demographic history and identify novel signs of adaptation across spatial and temporal dimensions. For example, we showed that the spatial genetic structure of populations is stable over time, but that drift due to seasonal contractions of population size causes populations to diverge over time. We identified signals of adaptation that vary between continents in genomic regions associated with xenobiotic resistance, consistent with independent adaptation to common pesticides. Moreover, by analyzing samples collected during spring and fall across Europe, we provide new evidence for seasonal adaptation related to loci associated with pathogen response. Furthermore, we have also released an updated version of the DEST genome browser. This is a useful tool for studying spatiotemporal patterns of genetic variation in this classic model system.
Introduced over seventy years ago, F-statistics have been and remain central to population and evolutionary genetics. Among them, FST is one of the most commonly used descriptive statistics in empirical studies, notably to characterize the structure of genetic polymorphisms within and between populations, to shed light on the evolutionary history of populations, or to identify marker loci under differential selection for adaptive traits. However, the use of FST in simplified population models can overlook important hierarchical structures, such as geographic or temporal subdivisions, potentially leading to misleading interpretations and increasing false positives in genome scans for adaptive differentiation. Hierarchical F-statistics have been introduced to account for multiple pre-defined levels of population structure. Several estimators have also been proposed, including robust ones implemented in the popular R package hierfstat . Nevertheless, these were primarily designed for individual genotyping data and can be computationally intensive for large genomic datasets. In this study, we extend previous work by developing unbiased method-of-moments estimators for hierarchical F-statistics tailored for Pool-Seq data, a cost-effective alternative to individual genome sequencing. These Pool--Seq estimators have been developed in an Anova framework, using definitions based on identity-in-state probabilities. The new estimators have been implemented in an updated version of the R package poolfstat , together with estimators for sample allele count data derived from individual genotyping data. We validate and compare the performance of these estimators through extensive simulations under a hierarchical island model. Finally, we apply these estimators to real Pool-Seq data from Drosophila melanogaster populations, demonstrating their usefulness in revealing population structure and identifying loci with high differentiation within or between groups of subpopulations and associated with spatial or temporal genetic variation. ### Competing Interest Statement The authors have declared no competing interest.
Seed banking (or dormancy) is a widespread bet-hedging strategy, generating a form of population overlap, which decreases the magnitude of genetic drift. The methodological complexity of integrating this trait implies it is ignored when developing tools to detect selective sweeps. But, as dormancy lengthens the ancestral recombination graph (ARG), increasing times to fixation, it can change the genomic signatures of selection. To detect genes under positive selection in seed banking species it is important to 1) determine whether the efficacy of selection is affected, and 2) predict the patterns of nucleotide diversity at and around positively selected alleles. We present the first tree sequence-based simulation program integrating a weak seed bank to examine the dynamics and genomic footprints of beneficial alleles in a finite population. We find that seed banking does not affect the probability of fixation and confirm expectations of increased times to fixation. We also confirm earlier findings that, for strong selection, the times to fixation are not scaled by the inbreeding effective population size in the presence of seed banks, but are shorter than would be expected. As seed banking increases the effective recombination rate, footprints of sweeps appear narrower around the selected sites and due to the scaling of the ARG are detectable for longer periods of time. The developed simulation tool can be used to predict the footprints of selection and draw statistical inference of past evolutionary events in plants, invertebrates, or fungi with seed banks.
The reconstruction of geographic and demographic scenarios of dissemination for invasive pathogens of crops is a key step toward improving the management of emerging infectious diseases. Nowadays, the reconstruction of biological invasions typically uses the information of both genetic and historical information to test for different hypotheses of colonization. The Approximate Bayesian Computation framework and its recent Random Forest development (ABC-RF) have been successfully used in evolutionary biology to decipher multiple histories of biological invasions. Yet, for some organisms, typically plant pathogens, historical data may not be reliable notably because of the difficulty to identify the organism and the delay between the introduction and the first mention. We investigated the history of the invasion of Africa by the fungal pathogen of banana Pseudocercospora fijiensis, by testing the historical hypothesis against other plausible hypotheses. We analyzed the genetic structure of eight populations from six eastern and western African countries, using 20 microsatellite markers and tested competing scenarios of population foundation using the ABC-RF methodology. We do find evidence for an invasion front consistent with the historical hypothesis, but also for the existence of another front never mentioned in historical records. We question the historical introduction point of the disease on the continent. Crucially, our results illustrate that even if ABC-RF inferences may sometimes fail to infer a single, well-supported scenario of invasion, they can be helpful in rejecting unlikely scenarios, which can prove much useful to shed light on disease dissemination routes.
Resurrection studies are a useful tool to measure how phenotypic traits have changed in populations through time. If these trait modifications correlate with the environmental changes that occurred during the time period, it suggests that the phenotypic changes could be a response to selection. Selfing, through its reduction of effective size, could challenge the ability of a population to adapt to environmental changes. Here, we used a resurrection study to test for adaptation in a selfing population of Medicago truncatula, by comparing the genetic composition and flowering times across 22 generations. We found evidence for evolution toward earlier flowering times by about two days and a peculiar genetic structure, typical of highly selfing populations, where some multilocus genotypes (MLGs) are persistent through time. We used the change in frequency of the MLGs through time as a multilocus fitness measure and built a selection gradient that suggests evolution toward earlier flowering times. Yet, a simulation model revealed that the observed change in flowering time could be explained by drift alone, provided the effective size of the population is small enough (<150). These analyses suffer from the difficulty to estimate the effective size in a highly selfing population, where effective recombination is severely reduced.
Multipartite viruses have a segmented genome, with each segment encapsidated separately. In all multipartite virus species for which the question has been addressed, the distinct segments reproducibly accumulate at a specific and host-dependent relative frequency, defined as the 'genome formula'. Here, we test the hypothesis that the multipartite genome organization facilitates the regulation of gene expression via changes of the genome formula and thus via gene copy number variations. In a first experiment, the faba bean necrotic stunt virus (FBNSV), whose genome is composed of eight DNA segments each encoding a single gene, was inoculated into faba bean or alfalfa host plants, and the relative concentrations of the DNA segments and their corresponding messenger RNAs (mRNAs) were monitored. In each of the two host species, our analysis consistently showed that the genome formula variations modulate gene expression, the concentration of each genome segment linearly and positively correlating to that of its cognate mRNA but not of the others. In a second experiment, twenty parallel FBNSV lines were transferred from faba bean to alfalfa plants. Upon host switching, the transcription rate of some genome segments changes, but the genome formula is modified in a way that compensates for these changes and maintains a similar ratio between the various viral mRNAs. Interestingly, a deep-sequencing analysis of these twenty FBNSV lineages demonstrated that the host-related genome formula shift operates independently of DNA-segment sequence mutation. Together, our results indicate that nanoviruses are plastic genetic systems, able to transiently adjust gene expression at the population level in changing environments, by modulating the copy number but not the sequence of each of their genes.
Simulation‐based methods such as approximate Bayesian computation (ABC) are well‐adapted to the analysis of complex scenarios of populations and species genetic history. In this context, supervised machine learning (SML) methods provide attractive statistical solutions to conduct efficient inferences about scenario choice and parameter estimation. The Random Forest methodology (RF) is a powerful ensemble of SML algorithms used for classification or regression problems. Random Forest allows conducting inferences at a low computational cost, without preliminary selection of the relevant components of the ABC summary statistics, and bypassing the derivation of ABC tolerance levels. We have implemented a set of RF algorithms to process inferences using simulated data sets generated from an extended version of the population genetic simulator implemented in DIYABC v2.1.0. The resulting computer package, named DIYABC Random Forest v1.0, integrates two functionalities into a user‐friendly interface: the simulation under custom evolutionary scenarios of different types of molecular data (microsatellites, DNA sequences or SNPs) and RF treatments including statistical tools to evaluate the power and accuracy of inferences. We illustrate the functionalities of DIYABC Random Forest v1.0 for both scenario choice and parameter estimation through the analysis of pseudo‐observed and real data sets corresponding to pool‐sequencing and individual‐sequencing SNP data sets. Because of the properties inherent to the implemented RF methods and the large feature vector (including various summary statistics and their linear combinations) available for SNP data, DIYABC Random Forest v1.0 can efficiently contribute to the analysis of large SNP data sets to make inferences about complex population genetic histories.
Tracking genetic changes of populations through time allows a more direct study of the evolutionary processes acting on the population than a single contemporary sample.Several statistical methods have been developed to characterize the demography and selection from temporal population genetic data.However, these methods are usually developed under the assumption of outcrossing reproduction and might not be applicable when there is substantial selfing in the population.Here, we focus on a method to detect loci under selection based on a genome scan of temporal differentiation, adapting it to the particularities of selfing populations.Selfing reduces the effective recombination rate and can extend hitch-hiking effects to the whole genome, erasing any local signal of selection on a genome scan.Therefore, selfing is expected to reduce the power of the test.By means of simulations, we evaluate the performance of the method under scenarios of adaptation from new mutations or standing variation at different rates of selfing.We find that the detection of loci under selection in predominantly selfing populations remains challenging even with the adapted method.Still, selective sweeps from standing variation on predominantly selfing populations can leave some signal of selection around the selected site thanks to historical recombination before the sweep.Under this scenario, ancestral advantageous alleles at low frequency leave the strongest local signal, while new advantageous mutations leave no local footprint of the sweep.
By capturing various patterns of the structuring of genetic variation across populations, f-statistics have proved highly effective for the inference of demographic history. Such statistics are defined as covariances of SNP allele frequency differences among sets of populations without requiring haplotype information and are hence particularly relevant for the analysis of pooled sequencing (Pool-Seq) data. We here propose a reinterpretation of the F (and D) parameters in terms of probability of gene identity and derive from this unified definition unbiased estimators for both Pool-Seq data and standard allele count data obtained from individual genotypes. We implemented these estimators in a new version of the R package poolfstat, which now includes a wide range of inference methods: (i) three-population test of admixture; (ii) four-population test of treeness; (iii) F4-ratio estimation of admixture rates; and (iv) fitting, visualization and (semi-automatic) construction of admixture graphs. A comprehensive evaluation of the methods implemented in poolfstat on both simulated Pool-Seq (with various sequencing coverages and error rates) and allele count data confirmed the accuracy of these approaches, even for the most cost-effective Pool-Seq design involving relatively low sequencing coverages. We further analysed a real Pool-Seq data made of 14 populations of the invasive species Drosophila suzukii, which allowed refining both the demographic history of native populations and the invasion routes followed by this emblematic pest. Our new package poolfstat provides the community with a user-friendly and efficient all-in-one tool to unravel complex population genetic histories from large-size Pool-Seq or allele count SNP data.
The advent of high throughput sequencing and genotyping technologies enables the comparison of patterns of polymorphisms at a very large number of markers. While the characterization of genetic structure from individual sequencing data remains expensive for many nonmodel species, it has been shown that sequencing pools of individual DNAs (Pool-seq) represents an attractive and cost-effective alternative. However, analyzing sequence read counts from a DNA pool instead of individual genotypes raises statistical challenges in deriving correct estimates of genetic differentiation. In this article, we provide a method-of-moments estimator of FST for Pool-seq data, based on an analysis-of-variance framework. We show, by means of simulations, that this new estimator is unbiased and outperforms previously proposed estimators. We evaluate the robustness of our estimator to model misspecification, such as sequencing errors and uneven contributions of individual DNAs to the pools. Finally, by reanalyzing published Pool-seq data of different ecotypes of the prickly sculpin Cottus asper, we show how the use of an unbiased FST estimator may question the interpretation of population structure inferred from previous analyses.
Understanding the processes of adaptive divergence, which may ultimately lead to speciation, is a major question in evolutionary biology. Allochronic differentiation refers to a particular situation where gene flow is primarily impeded by temporal isolation between early and late reproducers. This process has been suggested to occur in a large array of organisms, even though it is still overlooked in the literature. We here focused on a well‐documented case of incipient allochronic speciation in the winter pine processionary moth Thaumetopoea pityocampa . This species typically reproduces in summer and larval development occurs throughout autumn and winter. A unique, phenologically shifted population (SP) was discovered in 1997 in Portugal. It was proved to be strongly differentiated from the sympatric “winter population” (WP), but its evolutionary history could only now be explored. We took advantage of the recent assembly of a draft genome and of the development of pan‐genomic RAD‐seq markers to decipher the demographic history of the differentiating populations and develop genome scans of adaptive differentiation. We showed that the SP diverged relatively recently, that is, few hundred years ago, and went through two successive bottlenecks followed by population size expansions, while the sympatric WP is currently experiencing a population decline. We identified outlier SNPs that were mapped onto the genome, but none were associated with the phenological shift or with subsequent adaptations. The strong genetic drift that occurred along the SP lineage certainly challenged our capacity to reveal functionally important loci.
Identifying the genomic bases of adaptation to novel environments is a long-term objective in evolutionary biology. Because genetic differentiation is expected to increase between locally adapted populations at the genes targeted by selection, scanning the genome for elevated levels of differentiation is a first step towards deciphering the genomic architecture underlying adaptive divergence. The pea aphid Acyrthosiphon pisum is a model of choice to address this question, as it forms a large complex of plant-specialized races and cryptic species, resulting from recent adaptive radiation. Here, we characterized genomewide polymorphisms in three pea aphid races specialized on alfalfa, clover and pea crops, respectively, which we sequenced in pools (poolseq). Using a model-based approach that explicitly accounts for selection, we identified 392 genomic hotspots of differentiation spanning 47.3Mb and 2,484 genes (respectively, 9.12% of the genome size and 8.10% of its genes). Most of these highly differentiated regions were located on the autosomes, and overall differentiation was weaker on the X chromosome. Within these hotspots, high levels of absolute divergence between races suggest that these regions experienced less gene flow than the rest of the genome, most likely by contributing to reproductive isolation. Moreover, population-specific analyses showed evidence of selection in every host race, depending on the hotspot considered. These hotspots were significantly enriched for candidate gene categories that control host-plant selection and use. These genes encode 48 salivary proteins, 14 gustatory receptors, 10 odorant receptors, five P450 cytochromes and one chemosensory protein, which represent promising candidates for the genetic basis of host-plant specialization and ecological isolation in the pea aphid complex. Altogether, our findings open new research directions towards functional studies, for validating the role of these genes on adaptive phenotypes.
Natural reservoirs of zoonotic pathogens generally seem to be capable of tolerating infections. Tolerance and its underlying mechanisms remain difficult to assess using experiments or wildlife surveys. High-throughput sequencing technologies give the opportunity to investigate the genetic bases of tolerance, and the variability of its mechanisms in natural populations. In particular, population genomics may provide preliminary insights into the genes shaping tolerance and potentially influencing epidemiological dynamics. Here, we addressed these questions in the bank vole Myodes glareolus, the specific asymptomatic reservoir host of Puumala hantavirus (PUUV), which causes nephropathia epidemica (NE) in humans. Despite the continuous spatial distribution of M. glareolus in Sweden, NE is endemic to the northern part of the country. Northern bank vole populations in Sweden might exhibit tolerance strategies as a result of coadaptation with PUUV. This may favor the circulation and maintenance of PUUV and lead to high spatial risk of NE in northern Sweden. We performed a genome-scan study to detect signatures of selection potentially correlated with spatial variations in tolerance to PUUV. We analyzed six bank vole populations from Sweden, sampled from northern NE-endemic to southern NE-free areas. We combined candidate gene analyses (Tlr4, Tlr7, and Mx2 genes) and high-throughput sequencing of restriction site-associated DNA (RAD) markers. Outlier loci showed high levels of genetic differentiation and significant associations with environmental data including variations in the regional number of NE human cases. Among the 108 outliers that matched to mouse protein-coding genes, 14 corresponded to immune-related genes. The main biological pathways found to be significantly enriched corresponded to immune processes and responses to hantavirus, including the regulation of cytokine productions, TLR cascades, and IL-7, VEGF, and JAK-STAT signaling. In the future, genome-scan replicates and functional experimentations should enable to assess the role of these biological pathways in M. glareolus tolerance to PUUV.
The relative female and male contributions to demography are of great importance to better understand the history and dynamics of populations. While earlier studies relied on uniparental markers to investigate sex-specific questions, the increasing amount of sequence data now enables us to take advantage of tens to hundreds of thousands of independent loci from autosomes and the X chromosome. Here, we develop a novel method to estimate effective sex ratios or ESR (defined as the female proportion of the effective population) from allele count data for each branch of a rooted tree topology that summarizes the history of the populations of interest. Our method relies on Kimura's time-dependent diffusion approximation for genetic drift, and is based on a hierarchical Bayesian model to integrate over the allele frequencies along the branches. We show via simulations that parameters are inferred robustly, even under scenarios that violate some of the model assumptions. Analyzing bovine SNP data, we infer a strongly female-biased ESR in both dairy and beef cattle, as expected from the underlying breeding scheme. Conversely, we observe a strongly male-biased ESR in early domestication times, consistent with an easier taming and management of cows, and/or introgression from wild auroch males, that would both cause a relative increase in male effective population size. In humans, analyzing a subsample of non-African populations, we find a male-biased ESR in Oceanians that may reflect complex marriage patterns in Aboriginal Australians. Because our approach relies on allele count data, it may be applied on a wide range of species.