Simulation-based inference (SBI) enables amortized Bayesian inference by first training a neural posterior estimator (NPE) on prior-simulator pairs, typically through low-dimensional summary statistics, which can then be cheaply reused for fast inference by querying it on new test observations. Because NPE is estimated under the training data distribution, it is susceptible to misspecification when observations deviate from the training distribution. Many robust SBI approaches address this by modifying NPE training or introducing error models, coupling robustness to the inference network and compromising amortization and modularity. We introduce minimum-distance summaries, a plug-in robust NPE method that adapts queried test-time summaries independently of the pretrained NPE. Leveraging the maximum mean discrepancy (MMD) as a distance between observed data and a summary-conditional predictive distribution, the adapted summary inherits strong robustness properties from the MMD. We demonstrate that the algorithm can be implemented efficiently with random Fourier feature approximations, yielding a lightweight, model-free test-time adaptation procedure. We provide theoretical guarantees for the robustness of our algorithm and empirically evaluate it on a range of synthetic and real-world tasks, demonstrating substantial robustness gains with minimal additional overhead.
Estimating a distribution given access to its unnormalized density is pivotal in Bayesian inference, where the posterior is generally known only up to an unknown normalizing constant. Variational inference and Markov chain Monte Carlo methods are the predominant tools for this task; however, both are often challenging to apply reliably, particularly when the posterior has complex geometry. Here, we introduce Soft Contrastive Variational Inference (SoftCVI), which allows a family of variational objectives to be derived through a contrastive estimation framework. The approach parameterizes a classifier in terms of a variational distribution, reframing the inference task as a contrastive estimation problem aiming to identify a single true posterior sample among a set of samples. Despite this framing, we do not require positive or negative samples, but rather learn by sampling the variational distribution and computing ground truth soft classification labels from the unnormalized posterior itself. The objectives have zero variance gradient when the variational approximation is exact, without the need for specialized gradient estimators. We empirically investigate the performance on a variety of Bayesian inference tasks, using both simple (e.g. normal) and expressive (normalizing flow) variational distributions. We find that SoftCVI can be used to form objectives which are stable to train and mass-covering, frequently outperforming inference with other variational approaches.
We study the problem of likelihood maximization when the likelihood function is intractable but model simulations are readily available. We propose a sequential, gradient-based optimization method that directly models the Fisher score based on a local score matching technique which uses simulations from a localized region around each parameter iterate. By employing a linear parameterization to the surrogate score model, our technique admits a closed-form, least-squares solution. This approach yields a fast, flexible, and efficient approximation to the Fisher score, effectively smoothing the likelihood objective and mitigating the challenges posed by complex likelihood landscapes. We provide theoretical guarantees for our score estimator, including bounds on the bias introduced by the smoothing. Empirical results on a range of synthetic and real-world problems demonstrate the superior performance of our method compared to existing benchmarks.
We introduce Sequential Neural Posterior Score Estimation (SNPSE), a score-based method for Bayesian inference in simulator-based models. Our method, inspired by the remarkable success of score-based methods in generative modelling, leverages conditional score-based diffusion models to generate samples from the posterior distribution of interest. The model is trained using an objective function which directly estimates the score of the posterior. We embed the model into a sequential training procedure, which guides simulations using the current approximation of the posterior at the observation of interest, thereby reducing the simulation cost. We also introduce several alternative sequential approaches, and discuss their relative merits. We then validate our method, as well as its amortised, non-sequential, variant on several numerical examples, demonstrating comparable or superior performance to existing state-of-the-art methods such as Sequential Neural Posterior Estimation (SNPE).
Environmental change is intensifying the biodiversity crisis and threatening species across the tree of life. Conservation genomics can help inform conservation actions and slow biodiversity loss. However, more training, appropriate use of novel genomic methods and communication with managers are needed. Here, we review practical guidance to improve applied conservation genomics. We share insights aimed at ensuring effectiveness of conservation actions around three themes: (1) improving pedagogy and training in conservation genomics including for online global audiences, (2) conducting rigorous population genomic analyses properly considering theory, marker types and data interpretation and (3) facilitating communication and collaboration between managers and researchers. We aim to update students and professionals and expand their conservation toolkit with genomic principles and recent approaches for conserving and managing biodiversity. The biodiversity crisis is a global problem and, as such, requires international involvement, training, collaboration and frequent reviews of the literature and workshops as we do here.
This paper asks the question: can genomic information be used to recover a species that is already on the pathway to extinction due to genetic swamping from a related and more numerous population? We show that a breeding strategy in a captive breeding program can use whole genome sequencing to identify and remove segments of DNA introgressed through hybridisation. The proposed policy uses a generalized measure of kinship or heterozygosity accounting for local ancestry, that is, whether a specific genetic location was inherited from the target of conservation. We then show that optimizing these measures would minimize undesired ancestry while also controlling kinship and/or heterozygosity, in a simulated breeding population. The process is applied to real data representing the hybridized Scottish wildcat breeding population, with the result that it should be possible to breed out domestic cat ancestry. The ability to reverse introgression is a powerful tool brought about through the combination of sequencing with computational advances in ancestry estimation. Since it works best when applied early in the process, important decisions need to be made about which genetically distinct populations should benefit from it and which should be left to reform into a single population.
Many machine learning problems can be seen as approximating a *target* distribution using a *particle* distribution by minimizing their statistical discrepancy. Wasserstein Gradient Flow can move particles along a path that minimizes the $f$-divergence between the target and particle distributions. To move particles, we need to calculate the corresponding velocity fields derived from a density ratio function between these two distributions. Previous works estimated such density ratio functions and then differentiated the estimated ratios. These approaches may suffer from overfitting, leading to a less accurate estimate of the velocity fields. Inspired by non-parametric curve fitting, we directly estimate these velocity fields using interpolation techniques. We prove that our estimators are consistent under mild conditions. We validate their effectiveness using novel applications on domain adaptation and missing data imputation. The code for reproducing our results can be found at https://github.com/anewgithubname/gradest2.
The European wildcat population in Scotland is considered critically endangered as a result of hybridization with introduced domestic cats,1,2 though the time frame over which this gene flow has taken place is unknown. Here, using genome data from modern, museum, and ancient samples, we reconstructed the trajectory and dated the decline of the local wildcat population from viable to severely hybridized. We demonstrate that although domestic cats have been present in Britain for over 2,000 years,3 the onset of hybridization was only within the last 70 years. Our analyses reveal that the domestic ancestry present in modern wildcats is markedly over-represented in many parts of the genome, including the major histocompatibility complex (MHC). We hypothesize that introgression provides wildcats with protection against diseases harbored and introduced by domestic cats, and that this selection contributes to maladaptive genetic swamping through linkage drag. Using the case of the Scottish wildcat, we demonstrate the importance of local ancestry estimates to both understand the impacts of hybridization in wild populations and support conservation efforts to mitigate the consequences of anthropogenic and environmental change.
Domestic cats were derived from the Near Eastern wildcat (Felis lybica), after which they dispersed with people into Europe. As they did so, it is possible that they interbred with the indigenous population of European wildcats (Felis silvestris). Gene flow between incoming domestic animals and closely related indigenous wild species has been previously demonstrated in other taxa, including pigs, sheep, goats, bees, chickens, and cattle. In the case of cats, a lack of nuclear, genome-wide data, particularly from Near Eastern wildcats, has made it difficult to either detect or quantify this possibility. To address these issues, we generated 75 ancient mitochondrial genomes, 14 ancient nuclear genomes, and 31 modern nuclear genomes from European and Near Eastern wildcats. Our results demonstrate that despite cohabitating for at least 2,000 years on the European mainland and in Britain, most modern domestic cats possessed less than 10% of their ancestry from European wildcats, and ancient European wildcats possessed little to no ancestry from domestic cats. The antiquity and strength of this reproductive isolation between introduced domestic cats and local wildcats was likely the result of behavioral and ecological differences. Intriguingly, this long-lasting reproductive isolation is currently being eroded in parts of the species’ distribution as a result of anthropogenic activities.
This paper asks the question: can genomic information recover a species that is already on the pathway to extinction due to genetic swamping from a related and more numerous population? We show that whole genome sequencing can be used to identify and remove hybrid segments of DNA, when used as part of the breeding policy in a captive breeding program. The proposed policy uses a generalised measure of kinship or heterozygosity accounting for local ancestry, that is, whether a specific genetic location was inherited from from the target of conservation. We then show that optimising these measures would minimise undesired ancestry whilst also controlling undesired kinship or heterozygosity respectively, in a simulated breeding population. The process is applied to real data representing the hybridized Scottish wildcat breeding population, with the result that it should be possible to breed out the domestic cat ancestry. The ability to reverse introgression is a powerful new tool brought about from both sequencing and computational advances in ancestry estimation. Since it works best when applied early in the process, important decisions need to be made about which genetically distinct populations should benefit from it and which should be left to reform into a single population.
Novel technologies for recovering DNA information from archaeological and historical specimens have made available an ever-increasing amount of temporally spaced genetic samples from natural populations. These genetic time series permit the direct assessment of patterns of temporal changes in allele frequencies and hold the promise of improving power for the inference of selection. Increased time resolution can further facilitate testing hypotheses regarding the drivers of past selection events such as the incidence of plant and animal domestication. However, studying past selection processes through ancient DNA (aDNA) still involves considerable obstacles such as postmortem damage, high fragmentation, low coverage, and small samples. To circumvent these challenges, we introduce a novel Bayesian framework for the inference of temporally variable selection based on genotype likelihoods instead of allele frequencies, thereby enabling us to model sample uncertainties resulting from the damage and fragmentation of aDNA molecules. Also, our approach permits the reconstruction of the underlying allele frequency trajectories of the population through time, which allows for a better understanding of the drivers of selection. We evaluate its performance through extensive simulations and demonstrate its utility with an application to the ancient horse samples genotyped at the loci for coat coloration. Our results reveal that incorporating sample uncertainties can further improve the inference of selection.
Many machine learning problems can be formulated as approximating a target distribution using a particle distribution by minimizing a statistical discrepancy. Wasserstein Gradient Flow can be employed to move particles along a path that minimizes the $f$-divergence between the \textit{target} and \textit{particle} distributions. To perform such movements we need to calculate the corresponding velocity fields which include a density ratio function between these two distributions. While previous works estimated the density ratio function first and then differentiated the estimated ratio, this approach may suffer from overfitting, which leads to a less accurate estimate. Inspired by non-parametric curve fitting, we directly estimate these velocity fields using interpolation. We prove that our method is asymptotically consistent under mild conditions. We validate the effectiveness using novel applications on domain adaptation and missing data imputation.
Preserving natural genetic diversity and ecological function of wild species is a central goal in conservation biology. As such, anthropogenic hybridization is considered a threat to wild populations, as it can lead to changes in the genetic makeup of wild species and even to the extinction of wild genomes. In European wildcats, the genetic and ecological impacts of gene flow from domestic cats are mostly unknown at the species scale. However, in small and isolated populations, it is known to include genetic swamping of wild genomes. In this context, it is crucial to better understand the dynamics of hybridization across the species range, to inform and implement management measures that maintain the genetic diversity and integrity of the European wildcat. In the present paper, we aim to provide an overview of the current scientific understanding of anthropogenic hybridization in European wildcats, to clarify important aspects regarding the evaluation of hybridization given the available methodologies, and to propose guidelines for management and research priorities.
Understanding the rate and extent to which populations can adapt to novel environments at their ecological margins is fundamental to predicting the persistence of biological communities during ongoing and rapid global change. Recent range expansion in response to climate change in the UK butterfly Aricia agestis is associated with the evolution of novel interactions with a larval food plant, and the loss of its ability to use an ancestral host species. Using ddRAD analysis of 61,210 variable SNPs from 261 females from throughout the UK range of this species, we identify genomic regions at multiple chromosomes that are associated with evolutionary responses, and their association with demographic history and ecological variation. Gene flow appears widespread throughout the range, despite the apparently fragmented nature of the habitats used by this species. Patterns of haplotype variation between selected and neutral genomic regions suggest that evolution associated with climate adaptation is polygenic, resulting from the independent spread of alleles throughout the established range of this species, rather than the colonization of pre-adapted genotypes from coastal populations. These data suggest that rapid responses to climate change do not depend on the availability of pre-adapted genotypes. Instead, the evolution of novel forms of biotic interaction in A. agestis has occurred during range expansion, through the assembly of novel genotypes from alleles from multiple localities.
Innovations in ancient DNA (aDNA) preparation and sequencing technologies have exponentially increased the quality and quantity of aDNA data extracted from ancient biological materials. The additional temporal component from the incoming aDNA data can provide improved power to address fundamental evolutionary questions like characterising selection processes that shape the phenotypes and genotypes of contemporary populations or species. However, utilising aDNA to study past selection processes still involves considerable hurdles like how to eliminate the confounding factor of genetic interactions in the inference of selection. To address this issue, we extend the approach of He et al. (2022) to infer temporally variable selection from the aDNA data in the form of genotype likelihoods with the flexibility of modelling linkage and epistasis in this work. Our posterior computation is carried out by a robust adaptive version of the particle marginal Metropolis-Hastings algorithm with a coerced acceptance rate. Our extension inherits the desirable features of He et al. (2022) such as modelling sample uncertainty resulting from the damage and fragmentation of aDNA molecules and reconstructing underlying gamete frequency trajectories of the population. We evaluate its performance through extensive simulations and show its utility with an application to the aDNA data from pigmentation loci in horses.
Domestic cats were derived from the Near Eastern wildcat (Felis lybica), after which they dispersed with people into Europe. As they did so, it is possible that they interbred with the indigenous population of European wildcats (Felis silvestris). Gene flow between incoming domestic animals and closely related indigenous wild species has been previously demonstrated in other taxa including pigs, sheep, goats, bees, chickens and cattle. In the case of cats, a lack of nuclear, genome-wide data, particularly from Near Eastern wildcats, has made this possibility difficult to either detect or quantify. To address these issues, we generated 75 ancient mitochondrial genomes, 14 ancient nuclear genomes and 31 modern nuclear genomes from European and Near Eastern wildcats. Our results demonstrate that despite cohabitating for at least 2,000 years on the European mainland and in Britain, most modern domestic cats possessed less than 10% of their ancestry from European wildcats, and ancient European wildcats possessed little to no ancestry from domestic cats. The antiquity and strength of this reproductive isolation between introduced domestic cats and local wildcats was likely the result of behavioural and ecological differences. Intriguingly, this long-lasting reproductive isolation is currently being eroded in parts of the species’ distribution as a result of anthropogenic activities.
Computer simulations have proven a valuable tool for understanding complex phenomena across the sciences. However, the utility of simulators for modelling and forecasting purposes is often restricted by low data quality, as well as practical limits to model fidelity. In order to circumvent these difficulties, we argue that modellers must treat simulators as idealistic representations of the true data generating process, and consequently should thoughtfully consider the risk of model misspecification. In this work we revisit neural posterior estimation (NPE), a class of algorithms that enable black-box parameter inference in simulation models, and consider the implication of a simulation-to-reality gap. While recent works have demonstrated reliable performance of these methods, the analyses have been performed using synthetic data generated by the simulator model itself, and have therefore only addressed the well-specified case. In this paper, we find that the presence of misspecification, in contrast, leads to unreliable inference when NPE is used naively. As a remedy we argue that principled scientific inquiry with simulators should incorporate a model criticism component, to facilitate interpretable identification of misspecification and a robust inference component, to fit 'wrong but useful' models. We propose robust neural posterior estimation (RNPE), an extension of NPE to simultaneously achieve both these aims, through explicitly modelling the discrepancies between simulations and the observed data. We assess the approach on a range of artificially misspecified examples, and find RNPE performs well across the tasks, whereas naively using NPE leads to misleading and erratic posteriors.
The field of population genomics has grown rapidly in response to the recent advent of affordable, large-scale sequencing technologies. As opposed to the situation during the majority of the 20th century, in which the development of theoretical and statistical population-genetic insights out-paced the generation of data to which they could be applied, genomic data are now being produced at a far greater rate than they can be meaningfully analyzed and interpreted. With this wealth of data has come a tendency to focus on fitting specific (and often rather idiosyncratic) models to data, at the expense of a careful exploration of the range of possible underlying evolutionary processes. For example, the approach of directly investigating models of adaptive evolution in each newly sequenced population or species often neglects the fact that a thorough characterization of ubiquitous non-adaptive processes is a prerequisite for accurate inference. We here describe the perils of these tendencies, present our consensus views on current best practices in population genomic data analysis, and highlight areas of statistical inference and theory that are in need of further attention. Thereby, we argue for the importance of defining a biologically relevant baseline model tuned to the details of each new analysis, of skepticism and scrutiny in interpreting model-fitting results, and of carefully defining addressable hypotheses and underlying uncertainties.
Understanding the rate and extent to which populations can adapt to novel environments at their ecological margins is fundamental to predicting the persistence of biological communities during ongoing and rapid global change. Recent range expansion in response to climate change in the UK butterfly Aricia agestis is associated with the evolution of novel interactions with a larval food plant, and the loss of its ability to use its ancestral larval host species. Using ddRAD analysis of 61210 variable SNPs from 261 females from throughout the UK range of this species, we identify genomic regions at multiple chromosomes that are associated with these evolutionary responses, and their association with demographic history and ecological variation. Gene flow appears widespread throughout the range, despite the apparently fragmented nature of the habitats used by this species. Patterns of haplotype variation between selected and neutral genomic regions suggest that evolution associated with climate adaptation is polygenic, resulting from the independent spread of existing alleles throughout the established range of this species, rather than the colonisation of pre-adapted genotypes from coastal populations. These data suggest that rapid responses to climate change do not depend on the availability of pre-adapted genotypes. Instead, the evolution of novel forms of biotic interaction in Aricia agestis has occurred during range expansion, through the assembly of novel genotypes from alleles from multiple localities.
In many scientific applications, we do not have explicit access to the likelihood function. However simulations of the process of interest, using different parameter settings, may give us access to the likelihood function implicitly. The methodology for approximating likelihoods and posterior distributions based on simulated observations can be described as simulation-based inference. In this paper, we propose a simulation-based inference algorithm in which we iteratively update particles to more closely resemble the posterior. Our approach utilises simulations to estimate a density ratio function at each iteration and then uses it to approximate the KL divergence between the particle density and the posterior density. By alternating between gradient descent and density ratio estimation, the approximated KL divergence is minimized. We benchmark the performance of our algorithm on a Gaussian mixture model and the M/G/1 queue process model and report promising results.