Environmental DNA (eDNA) analyses are powerful for describing marine biodiversity but must be optimized for their effective use in routine monitoring. To maximize eDNA detection probabilities of sparsely distributed populations, water samples are usually concentrated from larger volumes and filtered using fine-pore membranes, often a significant cost-time bottleneck in the workflow. This study aimed to streamline eDNA sampling by investigating plankton net versus bucket sampling, direct versus sequential filtration including self-preserving filters. Biodiversity was assessed using metabarcoding of the small ribosomal subunit (18S rRNA) and mitochondrial cytochrome c oxidase I (COI) genes. Multi-species detection probabilities were estimated for each workflow using a probabilistic occupancy modelling approach. Significant workflow-related differences in biodiversity metrics were reported. Highest amplicon sequence variant (ASV) richness was attained by the bucket sampling combined with self-preserving filters, comprising a large portion of micro-plankton. Less diversity but more metazoan taxa were captured in the net samples combined with 5 µm pore size filters. Pre-filtered 1.2 µm samples yielded few or no unique ASVs. The highest average (~32%) metazoan detection probabilities in the 5 µm pore size net samples confirmed the effectiveness of pre-concentrating plankton for biodiversity screening. These results contribute to streamlining eDNA sampling protocols for uptake and implementation in marine biodiversity research and surveillance.
Fitness landscapes map genotypes to organismal fitness. Their topography depends on how mutational effects interact–epistasis–and is important for understanding evolutionary processes such as speciation, the rate of adaptation, the advantage of recombination, and predictability versus stochasticity of evolution. The growing amount of empirical data has made it possible to better test landscape models empirically. We argue that this endeavor will benefit from the development and use of meaningful null models against which to compare more complex models. Here we develop statistical and computational methods for fitting fitness data from mutation combinatorial networks to three simple models: additive, multiplicative and stickbreaking. We employ a Bayesian framework for doing model selection. Using simulations, we demonstrate that our methods work and we explore their statistical performance: bias, error, and the power to discriminate among models. We then illustrate our approach and its flexibility by analyzing several previously published datasets. An R-package that implements our methods is available in the CRAN repository under the name Stickbreaker.
In Beisel et al. (2007), a likelihood framework, based on extreme value theory (EVT), was developed for determining the distribution of fitness effects for adaptive mutations. In this paper we extend this framework beyond the extreme distributions and develop a likelihood framework for testing whether or not extreme value theory applies. By making two simple adjustments to the Generalized Pareto Distribution (GPD) we introduce a new simple five parameter probability density function that incorporates nearly every common (continuous) probability model ever used. This means that all of the common models are nested. This has important implications in model selection beyond determining the distribution of fitness effects. However, we demonstrate the use of this distribution utilizing likelihood ratio testing to evaluate alternative distributions to the Gumbel and Weibull domains of attraction of fitness effects. We use a bootstrap strategy, utilizing importance sampling, to determine where in the parameter space will the test be most powerful in detecting deviations from these domains and at what sample size, with focus on small sample sizes (n<20). Our results indicate that the likelihood ratio test is most powerful in detecting deviation from the Gumbel domain when the shape parameters of the model are small while the test is more powerful in detecting deviations from the Weibull domain when these parameters are large. As expected, an increase in sample size improves the power of the test. This improvement is observed to occur quickly with sample size n≥10 in tests related to the Gumbel domain and n≥15 in the case of the Weibull domain. This manuscript is in tribute to the contributions of Dr. Paul Joyce to the areas of Population Genetics, Probability Theory and Mathematical Statistics. A Tribute section is provided at the end that includes Paul's original writing in the first iterations of this manuscript. The Introduction and Alternatives to the GPD sections were obtained from the last iteration that Paul and I have worked on, with some adjustments. I hope that the rest gives justice to his thoughts and possible conclusions that he might have wanted to portray.
Parallelism is important because it reveals how inherently stochastic adaptation is. Even as we come to better understand evolutionary forces, stochasticity limits how well we can predict evolutionary outcomes. Here we sought to quantify parallelism and some of its underlying causes by adapting a bacteriophage (ID11) with nine different first-step mutations, each with eight-fold replication, for 100 passages. This was followed by whole-genome sequencing five isolates from each endpoint. A large amount of variation arose-281 mutational events occurred representing 112 unique mutations. At least 41% of the mutations and 77% of the events were adaptive. Within wells, populations generally experienced complex interference dynamics. The genome locations and counts of mutations were highly uneven: mutations were concentrated in two regulatory elements and three genes and, while 103 of the 112 (92%) of the mutations were observed in ≤4 wells, a few mutations arose many times. 91% of the wells and 81% of the isolates had a mutation in the D-promoter. Parallelism was moderate compared to previous experiments with this system. On average, wells shared 27% of their mutations at the DNA level and 38% when the definition of parallel change is expanded to include the same regulatory feature or residue. About half of the parallelism came from D-promoter mutations. Background had a small but significant effect on parallelism. Similarly, an analyses of epistasis between mutations and their ancestral background was significant, but the result was mostly driven by four individual mutations. A second analysis of epistasis focused on de novo mutations revealed that no isolate ever had more than one D-promoter mutation and that 56 of the 65 isolates lacking a D-promoter mutation had a mutation in genes D and/or E. We assayed time to lysis in four of these mutually exclusive mutations (the two most frequent D-promoter and two in gene D) across four genetic backgrounds. In all cases lysis was delayed. We postulate that because host cells were generally rare (i.e., high multiplicity of infection conditions developed), selection favored phage that delayed lysis to better exploit their current host (i.e., 'love the one you're with'). Thus, the vast majority of wells (at least 64 of 68, or 94%) arrived at the same phenotypic solution, but through a variety of genetic changes. We conclude that answering questions about the range of possible adaptive trajectories, parallelism, and the predictability of evolution requires attention to the many biological levels where the process of adaptation plays out.
Parallelism is important because it reveals how inherently stochastic adaptation is.Even as we come to better understand evolutionary forces, stochasticity limits how well we can predict evolutionary outcomes.Here we sought to quantify parallelism and some of its underlying causes by adapting a bacteriophage (ID11) with nine different first-step mutations, each with eight-fold replication, for 100 passages.This was followed by wholegenome sequencing five isolates from each endpoint.A large amount of variation arose -281 mutational events occurred representing 112 unique mutations.At least 41% of the mutations and 77% of the events were adaptive.Within wells, populations generally experienced complex interference dynamics.The genome locations and counts of mutations were highly uneven: mutations were concentrated in two regulatory elements and three genes and, while 103 of the 112 (92%) of the mutations were observed in ≤ 4 wells, a few mutations arose many times.91% of the wells and 81% of the isolates had a mutation in the D-promoter.Parallelism was moderate compared to previous experiments with this system.On average, wells shared 27% of their mutations at the DNA level and 38% when the definition of parallel change is expanded to include the same regulatory feature or residue.About half of the parallelism came from D-promoter mutations.Background had a small but significant effect on parallelism.Similarly, an analyses of epistasis between mutations and their ancestral background was significant, but the result was mostly driven by four individual mutations.A second analysis of epistasis focused on de novo mutations revealed that no isolate ever had more than one D-promoter mutation and that 56 of the 65 isolates lacking a D-promoter mutation had a mutation in genes D and/or E. We assayed time to lysis in four of these mutually exclusive mutations (the two most frequent D-promoter and two in gene D) across four genetic backgrounds.In all cases lysis was delayed.We postulate that because host cells were generally rare (i.e., high multiplicity of infection conditions developed), selection favored phage that delayed lysis to better exploit their current host (i.e., 'love the one you're with').Thus, the vast majority of wells (at least 64 of 68, or 94%) arrived at the same phenotypic solution, but through a variety of genetic changes.We conclude that answering questions about the range of possible adaptive trajectories, parallelism, and the predictability of evolution
The spread of infectious diseases can be impacted by human behavior, and behavioral decisions often depend implicitly on a planning horizon-the time in the future over which options are weighed. We investigate the effects of planning horizons on epidemic dynamics. We developed an epidemiological agent-based model (along with an ODE analog) to explore the decision-making of self-interested individuals on adopting prophylactic behavior. The decision-making process incorporates prophylaxis efficacy and disease prevalence with the individuals' payoffs and planning horizon. Our results show that for short and long planning horizons individuals do not consider engaging in prophylactic behavior. In contrast, individuals adopt prophylactic behavior when considering intermediate planning horizons. Such adoption, however, is not always monotonically associated with the prevalence of the disease, depending on the perceived protection efficacy and the disease parameters. Adoption of prophylactic behavior reduces the epidemic peak size while prolonging the epidemic and potentially generates secondary waves of infection. These effects can be made stronger by increasing the behavioral decision frequency or distorting an individual's perceived risk of infection.
Researchers in evolutionary genetics recently have recognized an exciting opportunity in decomposing beneficial mutations into their proximal, mechanistic determinants. The application of methods and concepts from molecular biology and life history theory to studies of lytic bacteriophages (phages) has allowed them to understand how natural selection sees mutations influencing life history. This work motivated the research presented here, in which we explored whether, under consistent experimental conditions, small differences in the genome of bacteriophage φX174 could lead to altered life history phenotypes among a panel of eight genetically distinct clones. We assessed the clones’ phenotypes by applying a novel statistical framework to the results of a serially sampled parallel infection assay, in which we simultaneously inoculated each of a large number of replicate host volumes with ∼1 phage particle. We sequentially plated the volumes over the course of infection and counted the plaques that formed after incubation. These counts served as a proxy for the number of phage particles in a single volume as a function of time. From repeated assays, we inferred significant, genetically determined heterogeneity in lysis time and burst size, including lysis time variance. These findings are interesting in light of the genetic and phenotypic constraints on the single-protein lysis mechanism of φX174. We speculate briefly on the mechanisms underlying our results, and we discuss the potential importance of lysis time variance in viral evolution.
Parallelism is important because it reveals how inherently stochastic adaptation is.Even as we come to better understand evolutionary forces, stochasticity limits how well we can predict evolutionary outcomes.Here we sought to quantify parallelism and some of its underlying causes by adapting a bacteriophage (ID11) with nine different first-step mutations, each with eight-fold replication, for 100 passages.This was followed by wholegenome sequencing five isolates from each endpoint.A large amount of variation arose -281 mutational events occurred representing 112 unique mutations.At least 41% of the mutations and 77% of the events were adaptive.Within wells, populations generally experienced complex interference dynamics.The genome locations and counts of mutations were highly uneven: mutations were concentrated in two regulatory elements and three genes and, while 103 of the 112 (92%) of the mutations were observed in ≤ 4 wells, a few mutations arose many times.91% of the wells and 81% of the isolates had a mutation in the D-promoter.Parallelism was moderate compared to previous experiments with this system.On average, wells shared 27% of their mutations at the DNA level and 38% when the definition of parallel change is expanded to include the same regulatory feature or residue.About half of the parallelism came from D-promoter mutations.Background had a small but significant effect on parallelism.Similarly, an analyses of epistasis between mutations and their ancestral background was significant, but the result was mostly driven by four individual mutations.A second analysis of epistasis focused on de novo mutations revealed that no isolate ever had more than one D-promoter mutation and that 56 of the 65 isolates lacking a D-promoter mutation had a mutation in genes D and/or E. We assayed time to lysis in four of these mutually exclusive mutations (the two most frequent D-promoter and two in gene D) across four genetic backgrounds.In all cases lysis was delayed.We postulate that because host cells were generally rare (i.e., high multiplicity of infection conditions developed), selection favored phage that delayed lysis to better exploit their current host (i.e., 'love the one you're with').Thus, the vast majority of wells (at least 64 of 68, or 94%) arrived at the same phenotypic solution, but through a variety of genetic changes.We conclude that answering questions about the range of possible adaptive trajectories, parallelism, and the predictability of evolution
Experimental evolution is an important research method that allows for the study of evolutionary processes occurring in microorganisms. Here we present a novel approach to experimental evolution that is based on application of next generation sequencing. Under this approach population level sequencing is applied to an evolving population in which multiple first-step beneficial mutations occur concurrently. As a result, frequencies of multiple beneficial mutations are observed in each replicate of an experiment. For this new type of data we develop methods of statistical inference. In particular, we propose a method for imputing selection coefficients of first-step beneficial mutations. The imputed selection coefficient are then used for testing the distribution of first-step beneficial mutations and for estimation of mean selection coefficient. In the case when selection coefficients are uniformly distributed, collected data may also be used to estimate the total number of available first-step beneficial mutations.
From population genetics theory, elevating the mutation rate of a large population should progressively reduce average fitness. If the fitness decline is large enough, the population will go extinct in a process known as lethal mutagenesis. Lethal mutagenesis has been endorsed in the virology literature as a promising approach to viral treatment, and several in vitro studies have forced viral extinction with high doses of mutagenic drugs. Yet only one empirical study has tested the genetic models underlying lethal mutagenesis, and the theory failed on even a qualitative level. Here we provide a new level of analysis of lethal mutagenesis by developing and evaluating models specifically tailored to empirical systems that may be used to test the theory. We first quantify a bias in the estimation of a critical parameter and consider whether that bias underlies the previously observed lack of concordance between theory and experiment. We then consider a seemingly ideal protocol that avoids this bias-mutagenesis of virions-but find that it is hampered by other problems. Finally, results that reveal difficulties in the mere interpretation of mutations assayed from double-strand genomes are derived. Our analyses expose unanticipated complexities in testing the theory. Nevertheless, the previous failure of the theory to predict experimental outcomes appears to reside in evolutionary mechanisms neglected by the theory (e. g., beneficial mutations) rather than from a mismatch between the empirical setup and model assumptions. This interpretation raises the specter that naive attempts at lethal mutagenesis may augment adaptation rather than retard it.
BACKGROUND:Explanations for bacterial biofilm persistence during antibiotic treatment typically depend on non-genetic mechanisms, and rarely consider the contribution of evolutionary processes.RESULTS:Using Escherichia coli biofilms, we demonstrate that heritable variation for broad-spectrum antibiotic resistance can arise and accumulate rapidly during biofilm development, even in the absence of antibiotic selection.CONCLUSIONS:Our results demonstrate the rapid de novo evolution of heritable variation in antibiotic sensitivity and resistance during E. coli biofilm development. We suggest that evolutionary processes, whether genetic drift or natural selection, should be considered as a factor to explain the elevated tolerance to antibiotics typically observed in bacterial biofilms. This could be an under-appreciated mechanism that accounts why biofilm populations are, in general, highly resistant to antibiotic treatment.
Throughout the 1980s, Simon Tavaré made numerous significant contributions to population genetics theory. As genetic data, in particular DNA sequence, became more readily available, a need to connect population-genetic models to data became the central issue. The seminal work of Griffiths and Tavaré (1994a, 1994b, 1994c) was among the first to develop a likelihood method to estimate the population-genetic parameters using full DNA sequences. Now, we are in the genomics era where methods need to scale-up to handle massive data sets, and Tavaré has led the way to new approaches. However, performing statistical inference under non-neutral models has proved elusive. In tribute to Simon Tavaré, we present an article in spirit of his work that provides a computationally tractable method for simulating and analyzing data under a class of non-neutral population-genetic models. Computational methods for approximating likelihood functions and generating samples under a class of allele-frequency based non-neutral parent-independent mutation models were proposed by Donnelly, Nordborg, and Joyce (DNJ) (Donnelly et al., 2001). DNJ (2001) simulated samples of allele frequencies from non-neutral models using neutral models as auxiliary distribution in a rejection algorithm. However, patterns of allele frequencies produced by neutral models are dissimilar to patterns of allele frequencies produced by non-neutral models, making the rejection method inefficient. For example, in some cases the methods in DNJ (2001) require 109 rejections before a sample from the non-neutral model is accepted. Our method simulates samples directly from the distribution of non-neutral models, making simulation methods a practical tool to study the behavior of the likelihood and to perform inference on the strength of selection.
Mutations that confer a selective advantage to an organism are the raw material upon which natural selection acts. The number of such mutations that are available is a central quantity of interest for understanding the tempo and trajectory of adaptive evolution. While this quantity is typically unknown, it can be estimated with varying levels of accuracy based on data obtained experimentally. We propose a method for estimating the number of beneficial mutations that accounts for the evolutionary forces that generate the data. Our model-based parametric approach is compared to an adjusted nonparametric abundance-based coverage estimator. We show that, in general, our estimator performs better. When the number of mutations is small, however, the performances of the two estimators are similar.
In relating genotypes to fitness, models of adaptation need to both be computationally tractable and qualitatively match observed data. One reason that tractability is not a trivial problem comes from a combinatoric problem whereby no matter in what order a set of mutations occurs, it must yield the same fitness. We refer to this as the bookkeeping problem. Because of their commutative property, the simple additive and multiplicative models naturally solve the bookkeeping problem. However, the fitness trajectories and epistatic patterns they predict are inconsistent with the patterns commonly observed in experimental evolution. This motivates us to propose a new and equally simple model that we call stickbreaking. Under the stickbreaking model, the intrinsic fitness effects of mutations scale by the distance of the current background to a hypothesized boundary. We use simulations and theoretical analyses to explore the basic properties of the stickbreaking model such as fitness trajectories, the distribution of fitness achieved, and epistasis. Stickbreaking is compared to the additive and multiplicative models. We conclude that the stickbreaking model is qualitatively consistent with several commonly observed patterns of adaptive evolution.
Recent decades have seen a significant rise in studies in which evolution is observed and analysed directly-as it happens-under replicated, controlled conditions. Such 'experimental evolution' approaches offer a degree of resolution of evolutionary processes and their underlying genetics that is difficult or even impossible to achieve in more traditional comparative and retrospective analyses. In principle, experimental populations can be monitored for phenotypic and genetic changes with any desired level of replication and measurement precision, facilitating progress on fundamental and previously unresolved questions in evolutionary biology. Here, we summarize 10 invited papers in which experimental evolution is making significant progress on a variety of fundamental questions. We conclude by briefly considering future directions in this very active field of research, emphasizing the importance of quantitative tests of theories and the emerging role of genome-wide re-sequencing.
Bacterial biofilms are particularly resistant to a wide variety of antimicrobial compounds. Their persistence in the face of antibiotic therapies causes significant problems in the treatment of infectious diseases. Seldom have evolutionary processes like genetic drift and mutation been invoked to explain how resistance to antibiotics emerges in biofilms, and we lack a simple and tractable model for the genetic and phenotypic diversification that occurs in bacterial biofilms. Here, we introduce the ‘onion model’, a simple neutral evolutionary model for phenotypic diversification in biofilms. We explore its properties and show that the model produces patterns of diversity that are qualitatively similar to observed patterns of phenotypic diversity in biofilms. We suggest that models like our onion model, which explicitly invoke evolutionary process, are key to understanding biofilm resistance to bactericidal and bacteriostatic agents. Elevated phenotypic variance provides an insurance effect that increases the likelihood that some proportion of the population will be resistant to imposed selective agents and may thus enhance persistence of the biofilm. Accounting for evolutionary change in biofilms will improve our ability to understand and counter diseases that are caused by biofilm persistence.
Existing inference methods for estimating the strength of balancing selection in multi-locus genotypes rely on the assumption that there are no epistatic interactions between loci. Complex systems in which balancing selection is prevalent, such as sets of human immune system genes, are known to contain components that interact epistatically. Therefore, current methods may not produce reliable inference on the strength of selection at these loci. In this paper, we address this problem by presenting statistical methods that can account for epistatic interactions in making inference about balancing selection. A theoretical result due to Fearnhead (2006) is used to build a multi-locus Wright-Fisher model of balancing selection, allowing for epistatic interactions among loci. Antagonistic and synergistic types of interactions are examined. The joint posterior distribution of the selection and mutation parameters is sampled by Markov chain Monte Carlo methods, and the plausibility of models is assessed via Bayes factors. As a component of the inference process, an algorithm to generate multi-locus allele frequencies under balancing selection models with epistasis is also presented. Recent evidence on interactions among a set of human immune system genes is introduced as a motivating biological system for the epistatic model, and data on these genes are used to demonstrate the methods.
Evolutionary biologists since Darwin have been fascinated by differences in the rate of trait‐evolutionary change across lineages. Despite this continued interest, we still lack methods for identifying shifts in evolutionary rates on the growing tree of life while accommodating uncertainty in the evolutionary process. Here we introduce a Bayesian approach for identifying complex patterns in the evolution of continuous traits. The method (auteur) uses reversible‐jump Markov chain Monte Carlo sampling to more fully characterize the complexity of trait evolution, considering models that range in complexity from those with a single global rate to potentially ones in which each branch in the tree has its own independent rate. This newly introduced approach performs well in recovering simulated rate shifts and simulated rates for datasets nearing the size typical for comparative phylogenetic study (i.e., ≥64 tips). Analysis of two large empirical datasets of vertebrate body size reveal overwhelming support for multiple‐rate models of evolution, and we observe exceptionally high rates of body‐size evolution in a group of emydid turtles relative to their evolutionary background. auteur will facilitate identification of exceptional evolutionary dynamics, essential to the study of both adaptive radiation and stasis.
The coalescent has become perhaps the most widely-used population genetics model. By modeling the ancestry of a sample, rather than the evolution of the entire population from which the sample is drawn, it provides a computationally efficient framework for data simulation. Furthermore, from a theoretical perspective, it provides the under-pinnings for many useful analysis techniques. In this chapter, we introduce the coalescent, describe some of the problems that it has been used to address, discuss practical implications that follow from the insight it provides, and summarize some of the available software.
The relationship between mutation, protein stability and protein function plays a central role in molecular evolution. Mutations tend to be destabilizing, including those that would confer novel functions such as host-switching or antibiotic resistance. Elevated temperature may play an important role in preadapting a protein for such novel functions by selecting for stabilizing mutations. In this study, we test the stability change conferred by single mutations that arise in a G4-like bacteriophage adapting to elevated temperature. The vast majority of these mutations map to interfaces between viral coat proteins, suggesting they affect protein-protein interactions. We assess their effects by estimating thermodynamic stability using molecular dynamic simulations and measuring kinetic stability using experimental decay assays. The results indicate that most, though not all, of the observed mutations are stabilizing.