We introduce ABLE (Approximate Blockwise Likelihood Estimation), a novel simulation-based composite likelihood method that uses the blockwise site frequency spectrum to jointly infer past demography and recombination. ABLE is explicitly designed for a wide variety of data from unphased diploid genomes to genome-wide multi-locus data (for example, RADSeq) and can also accommodate arbitrarily large samples. We use simulations to demonstrate the accuracy of this method to infer complex histories of divergence and gene flow and reanalyze whole genome data from two species of orangutan. ABLE is available for download at https://github.com/champost/ABLE .
Likelihood methods are being developed for inference of migration rates and past demographic changes from population genetic data. We survey an approach for such inference using sequential importance sampling techniques derived from coalescent and diffusion theory. The consistent application and assessment of this approach has required the re-implementation of methods often considered in the context of computer experiments methods, in particular of Kriging which is used as a smoothing technique to infer a likelihood surface from likelihoods estimated in various parameter points, as well as reconsideration of methods for sampling the parameter space appropriately for such inference. We illustrate the performance and application of the whole tool chain on simulated and actual data, and highlight desirable developments in terms of data types and biological scenarios.
Likelihood methods are being developed for inference of migration rates and past demographic changes from population genetic data. We survey an approach for such inference using sequential importance sampling techniques derived from coalescent and diffusion theory. The consistent application and assessment of this approach has required the re-implementation of methods often considered in the context of computer experiments methods, in particular of Kriging which is used as a smoothing technique to infer a likelihood surface from likelihoods estimated in various parameter points, as well as reconsideration of methods for sampling the parameter space appropriately for such inference. We illustrate the performance and application of the whole tool chain on simulated and actual data, and highlight desirable developments in terms of data types and biological scenarios.Résumé Diverses approches ont été développées pour l’inférence des taux de migration et des changements démo-graphiques passés à partir de la variation génétique des populations. Nous décrivons une de ces approches utilisant des techniques d’échantillonnage pondéré séquentiel, fondées sur la modélisation par approches de coalescence et de diffusion de l’évolution de ces polymorphismes. L’application et l’évaluation systématique de cette approche ont requis la ré-implémentation de méthodes souvent considérées pour l’analyse de fonctions simulées, en particulier le krigeage, ici utilisé pour inférer une surface de vraisemblance à partir de vraisemblances estimées en différents points de l’espace des paramètres, ainsi que des techniques d’échantillonage de ces points. Nous illustrons la performance et l’application de cette série de méthodes sur données simulées et réelles, et indiquons les améliorations souhaitables en termes de types de données et de scénarios biologiques.Mots-clés histoire démographique, processus de coalescence, importance sampling, genetic polymorphismAMS 2000 subject classifications 92D10, 62M05, 65C05
The origin and evolution of species’ ranges remains a central focus of historical biogeography and the advent of likelihood methods based on phylogenies has revolutionized the way in which range evolution has been studied. A decade ago, the first elements of what turned out to be a popular inference approach of ancestral ranges based on the processes of Dispersal, local Extinction and Cladogenesis ( DEC ) was proposed. The success of the DEC model lies in its use of a flexible statistical framework known as a Continuous Time Markov Chain and since, several conceptual and computational improvements have been proposed using this as a baseline approach. In the spirit of the original version of DEC , we introduce DEC eXtended ( DECX ) by accounting for rapid expansion and local extinction as possible anagenetic events on the phylogeny but without increasing model complexity ( i.e . in the number of free parameters). Classical vicariance as a cladogenetic event is also incorporated by making use of temporally flexible constraints on the connectivity between any two given areas in accordance with the movement of landmasses and dispersal opportunity over time. DECX is built upon a previous implementation in C/C++ and can analyze phylogenies on the order of several thousand tips in a few minutes. We test our model extensively on Pseudo Observed Datasets and on well-curated and recently published data from various island clades and a worldwide phylogeny of Amphibians (3309 species). We also propose the very first implementation of the DEC model that can specifically account for trees with fossil tips ( i.e . non-ultrametric) using the phylogeny of palpimanoid spiders as a case study. In this paper, we argue in favour of the proposed improvements, which have the advantage of being computationally efficient while toeing the line of increased biological realism.
According to fossil data, the wood mouse arrived in North Africa 7500 ya, while it was present in Europe since Early Pleistocene. Previous molecular studies suggested that its introduction in North Africa probably occurred via the Strait of Gibraltar more than 0.4 Mya ago. In this study, we widely sampled wood mice to get a better understanding of the geographic and demographic history of this species in North Africa and possibly to help resolving the discrepancy between genetic and palaeontological data. Specifically, we wanted to answer the following questions: (1) When and how did the wood mouse arrive in North Africa? and (2) What is its demographic and geographic history in North Africa since its colonization? We collected in the field 438 new individuals and used both mtDNA and six microsatellite markers to answer these questions. Our results confirm that North African wood mice have a south-western European origin and colonized the Maghreb through the Strait of Gibraltar probably during the Mesolithic or slightly after. They first colonized the Tingitana Peninsula and then expanded throughout North Africa. Our genetic data suggest that the ancestral population size comprised numerous individuals reinforcing the idea that wood mice did not colonize Morocco accidentally through rafting of a few individuals, but via recurrent/multiple anthropogenic translocations. No spatial structuring of the genetic variability was recorded in North Africa, from Morocco to Tunisia.
The wealth of information contained in genome-scale datasets has substantially encouraged the development of methods inferring population histories with unprecedented resolution. Methods based on the Site Frequency Spectrum (SFS) are computationally efficient but discard information about linkage disequilibrium, while methods making use of linkage and recombination are computationally more intensive and rely on approximations such as the Sequentially Markov Coalescent. Overcoming these limitations, we introduce a novel Composite Likelihood (CL) framework which allows for the joint inference of arbitrarily complex population histories and the average genome-wide recombination rate from multiple genomes. We build upon an existing analytic approach that partitions the genome into blocks of equal (and arbitrary) size and summarizes the polymorphism and linkage information as blockwise counts of SFS types (bSFS). This statistic is a richer summary than the SFS because it retains information on the variation in genealogies contained in short-range linkage blocks across the genome. Our method, ABLE (Approximate Blockwise Likelihood Estimation), approximates the CL of arbitrary population histories via Monte Carlo simulations form the coalescent with recombination and overcomes limitations arising from analytical likelihood calculations. ABLE is first assessed by comparing it to expected analytic results for small samples and no intra-block recombination. The power of this approach is further illustrated by using whole genome data from the two species of orangutan and comparing our inferences under a series of models involving divergence and various forms of continuous or pulsed admixture with previous analyses based on the SFS and the SMC. Finally, we explore the effects of sampling (different block lengths and number of individuals) and find that accurate inference of demography and recombination can be achieved with reasonable computational effort. Our approach is also notably adapted to unphased data and fragmented assemblies making it particularly suitable for model as well as non-model organisms.
AimUnderstanding the relative contribution of diversification rates (speciation and extinction) and dispersal in the formation of the latitudinal diversity gradient - the decrease in species richness with increasing latitude - is a main goal of biogeography. The mammalian order Carnivora, which comprises 286 species, displays the traditional latitudinal diversity gradient seen in almost all mammalian orders. Yet the processes driving high species richness in the tropics may be fundamentally different in this group from that in other mammalian groups. Indeed, a recent study suggested that in Carnivora, unlike in all other major mammalian orders, net diversification rates are not higher in the tropics than in temperate regions. Our goal was thus to understand the reasons why there are more species of Carnivora in the tropics.LocationWorld-wide.MethodsWe reconstructed the biogeographical history of Carnivora using a time-calibrated phylogeny of the clade comprising all terrestrial species and dispersal-extinction-cladogenesis models. We also analysed a fossil dataset of carnivoran genera to examine how the latitudinal distribution of Carnivora varied through time.ResultsOur biogeographical analyses suggest that Carnivora originated in the East Palaearctic (i.e. Central Asia, China) in the early Palaeogene. Multiple independent lineages dispersed to low latitudes following three main paths: toward Africa, toward India/Southeast Asia and toward South America via the Bering Strait. These dispersal events were probably associated with local extinctions at high latitudes. Fossil data corroborate a high-latitude origin of the group, followed by late dispersal events toward lower latitudes in the Neogene.Main conclusionsUnlike most other mammalian orders, which originated and diversified at low latitudes and dispersed out of the tropics', Carnivora originated at high latitudes, and subsequently dispersed southward. Our study provides an example of combining phylogenetic and fossil data to understand the generation and maintenance of global-scale geographical variations in species richness.
This study presents genetic evidence that whale sharks, Rhincodon typus, are comprised of at least two populations that rarely mix and is the first to document a population expansion. Relatively high genetic structure is found when comparing sharks from the Gulf of Mexico with sharks from the Indo-Pacific. If mixing occurs between the Indian and Atlantic Oceans, it is not sufficient to counter genetic drift. This suggests whale sharks are not all part of a single global metapopulation. The significant population expansion we found was indicated by both microsatellite and mitochondrial DNA. The expansion may have happened during the Holocene, when tropical species could expand their range due to sea-level rise, eliminating dispersal barriers and increasing plankton productivity. However, the historic trend of population increase may have reversed recently. Declines in genetic diversity are found for 6 consecutive years at Ningaloo Reef in Australia. The declines in genetic diversity being seen now in Australia may be due to commercial-scale harvesting of whale sharks and collision with boats in past decades in other countries in the Indo-Pacific. The study findings have implications for models of population connectivity for whale sharks and advocate for continued focus on effective protection of the world's largest fish at multiple spatial scales.
Understanding the demographic history of populations and species is a central issue in evolutionary biology and molecular ecology. In this work, we develop a maximum-likelihood method for the inference of past changes in population size from microsatellite allelic data. Our method is based on importance sampling of gene genealogies, extended for new mutation models, notably the generalized stepwise mutation model (GSM). Using simulations, we test its performance to detect and characterize past reductions in population size. First, we test the estimation precision and confidence intervals coverage properties under ideal conditions, then we compare the accuracy of the estimation with another available method (MSVAR) and we finally test its robustness to misspecification of the mutational model and population structure. We show that our method is very competitive compared with alternative ones. Moreover, our implementation of a GSM allows more accurate analysis of microsatellite data, as we show that the violations of a single step mutation assumption induce very high bias toward false contraction detection rates. However, our simulation tests also showed some limits, which most importantly are large computation times for strong disequilibrium scenarios and a strong influence of some form of unaccounted population structure. This inference method is available in the latest implementation of the MIGRAINE software package.
Neutral community theory postulates a fundamental quantity, θ, which reflects the species diversity on a regional scale. While the recent genealogical formulation of community dynamics has considerably enhanced quantitative neutral ecology, its inferential aspects have remained computationally prohibitive. Here, we make use of a generalized version of the original two-level hierarchical framework in order to define a novel estimator for θ, which proves to be computationally efficient and robust when tested on a wide range of simulated neutral communities. Estimating θ from field data is also illustrated using two tropical forest datasets consisting of spatially separated permanent field plots. Preliminary results also reveal that our inferred regional diversity parameter based on community dynamics may be linked to widely used ordination techniques in ecology. This paper essentially paves the way for future work dealing with the parameter inference of neutral communities with respect to their spatial scale and structure.
Neutral community models have shown that limited migration can have a pervasive influence on the taxonomic composition of local communities even when all individuals are assumed of equivalent ecological fitness. Notably, the spatially implicit neutral theory yields a single parameter I for the immigration-drift equilibrium in a local community. In the case of plants, seed dispersal is considered as a defining moment of the immigration process and has attracted empirical and theoretical work. In this paper, we consider a version of the immigration parameter I depending on dispersal limitation from the neighbourhood of a community. Seed dispersal distance is alternatively modelled using a distribution that decreases quickly in the tails (thin-tailed Gaussian kernel) and another that enhances the chance of dispersal events over very long distances (heavily fat-tailed Cauchy kernel). Our analysis highlights two contrasting situations, where I is either mainly sensitive to community size (related to ecological drift) under the heavily fat-tailed kernel or mainly sensitive to dispersal distance under the thin-tailed kernel. We review dispersal distances of rainforest trees from field studies and assess the consistency between published estimates of I based on spatially-implicit models and the predictions of the kernel-based model in tropical forest plots. Most estimates of I were derived from large plots (10–50 ha) and were too large to be accounted for by a Cauchy kernel. Conversely, a fraction of the estimates based on multiple smaller plots (1 ha) appeared too small to be consistent with reported ranges of dispersal distances in tropical forests. Very large estimates may reflect within-plot habitat heterogeneity or estimation problems, while the smallest estimates likely imply other factors inhibiting migration beyond dispersal limitation. Our study underscores the need for interpreting I as an integrative index of migration limitation which, besides the limited seed dispersal, possibly includes habitat filtering or fragmentation.
a for gene the of an improvement in of Bahlo on a the DNA be an Infinitely-many (ISM) . For loci (microsatellites) under a Step-wise Mutation model (SMM) , we the Importance Sampling proposal distributions of Iorio . implement the case of a ) of
Neutral models provide an alternative to niche-based assembly rules of ecological communities by assuming that communities' properties are shaped by the stochastic interplay between ecological drift, migration and speciation. The recent and ongoing interest about neutral assumptions has produced many developments on the theoretical side, with nevertheless limited echoes in terms of analyses of real-world data. The present review paper aims to help bridge the widening gap between modellers and field ecologists through two objectives. First, to provide a multi-criteria typology of the main neutral models, including those from population genetics that have not yet been transposed to ecology, by considering how the fundamental processes of ecological drift, speciation and migration are modelled and, specifically, how space is taken into account. Second, to review methods recently proposed to estimate models parameters from field data, a point that should be mastered to allow for broader applications.