There are few, if any, algorithms in statistical phylogenetics which are used more heavily than Felsenstein's 1973 pruning method for computing the likelihood of a tree. We present LvD, (Likelihood via Decomposition), an alternative to Felsenstein's algorithm based on a different decomposition of the underlying phylogeny. It works for all standard nucleotide models. The new algorithm allows updates of the likelihood calculation in worst case O(log n) time with n taxa, as opposed to worst case O(n) time for existing methods. In practice this leads to appreciable improvements in likelihood calculations, the extent of speed-up depending on how balanced or unbalanced the trees are. We explore implications for parallel computing, and show that the approach allows likelihoods to be computed in O(log n) parallel time per site, compared to (worst case) O(n) time. We implemented and applied the algorithm to large numbers of simulated and empirical data sets and showed that these theoretical advances lead to a significant practical speed-up, although the extent of the improvement depends on how balanced the phylogenies already are.
While phylogenies have been essential in understanding how species evolve, they do not adequately describe some evolutionary processes. For instance, hybridization, a common phenomenon where interbreeding between 2 species leads to formation of a new species, must be depicted by a phylogenetic network, a structure that modifies a phylogenetic tree by allowing 2 branches to merge into 1, resulting in reticulation. However, existing methods for estimating networks become computationally expensive as the dataset size and/or topological complexity increase. The lack of methods for scalable inference hampers phylogenetic networks from being widely used in practice, despite accumulating evidence that hybridization occurs frequently in nature. Here, we propose a novel method, PhyNEST (Phylogenetic Network Estimation using SiTe patterns), that estimates binary, level-1 phylogenetic networks with a fixed, user-specified number of reticulations directly from sequence data. By using the composite likelihood as the basis for inference, PhyNEST is able to use the full genomic data in a computationally tractable manner, eliminating the need to summarize the data as a set of gene trees prior to network estimation. To search network space, PhyNEST implements both hill climbing and simulated annealing algorithms. PhyNEST assumes that the data are composed of coalescent independent sites that evolve according to the Jukes-Cantor substitution model and that the network has a constant effective population size. Simulation studies demonstrate that PhyNEST is often more accurate than 2 existing composite likelihood summary methods (SNaQand PhyloNet) and that it is robust to at least one form of model misspecification (assuming a less complex nucleotide substitution model than the true generating model). We applied PhyNEST to reconstruct the evolutionary relationships among Heliconius butterflies and Papionini primates, characterized by hybrid speciation and widespread introgression, respectively. PhyNEST is implemented in an open-source Julia package and is publicly available at https://github.com/sungsik-kong/PhyNEST.jl.
MOTIVATION:The multispecies coalescent model is now widely accepted as an effective model for incorporating variation in the evolutionary histories of individual genes into methods for phylogenetic inference from genome-scale data. However, because model-based analysis under the coalescent can be computationally expensive for large datasets, a variety of inferential frameworks and corresponding algorithms have been proposed for estimation of species-level phylogenies and associated parameters, including speciation times and effective population sizes. RESULTS:We consider the problem of estimating the timing of speciation events along a phylogeny in a coalescent framework. We propose a maximum a posteriori estimator based on composite likelihood (MAPCL) for inferring these speciation times under a model of DNA sequence evolution for which exact site-pattern probabilities can be computed under the assumption of a constant θ throughout the species tree. We demonstrate that the MAPCL estimates are statistically consistent and asymptotically normally distributed, and we show how this result can be used to estimate their asymptotic variance. We also provide a more computationally efficient estimator of the asymptotic variance based on the non-parametric bootstrap. We evaluate the performance of our method using simulation and by application to an empirical dataset for gibbons. AVAILABILITY AND IMPLEMENTATION:The method has been implemented in the PAUP* program, freely available at https://paup.phylosolutions.com for Macintosh, Windows and Linux operating systems. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
BEAGLE is a high-performance likelihood-calculation library for phylogenetic inference. The BEAGLE library defines a simple, but flexible, application programming interface (API), and includes a collection of efficient implementations for calculation under a variety of evolutionary models on different hardware devices. The library has been integrated into recent versions of popular phylogenetics software packages including BEAST and MrBayes and has been widely used across a diverse range of evolutionary studies. Here, we present BEAGLE 3 with new parallel implementations, increased performance for challenging data sets, improved scalability, and better usability. We have added new OpenCL and central processing unit-threaded implementations to the library, allowing the effective utilization of a wider range of modern hardware. Further, we have extended the API and library to support concurrent computation of independent partial likelihood arrays, for increased performance of nucleotide-model analyses with greater flexibility of data partitioning. For better scalability and usability, we have improved how phylogenetic software packages use BEAGLE in multi-GPU (graphics processing unit) and cluster environments, and introduced an automated method to select the fastest device given the data set, evolutionary model, and hardware. For application developers who wish to integrate the library, we also have developed an online tutorial. To evaluate the effect of the improvements, we ran a variety of benchmarks on state-of-the-art hardware. For a partitioned exemplar analysis, we observe run-time performance improvements as high as 5.9-fold over our previous GPU implementation. BEAGLE 3 is free, open-source software licensed under the Lesser GPL and available at https://beagle-dev.github.io.
Interactions between fungi and plants, including parasitism, mutualism, and saprotrophy, have been invoked as key to their respective macroevolutionary success. Here we evaluate the origins of plant-fungal symbioses and saprotrophy using a time-calibrated phylogenetic framework that reveals linked and drastic shifts in diversification rates of each kingdom. Fungal colonization of land was associated with at least two origins of terrestrial green algae and preceded embryophytes (as evidenced by losses of fungal flagellum, ca. 720 Ma), likely facilitating terrestriality through endomycorrhizal and possibly endophytic symbioses. The largest radiation of fungi (Leotiomyceta), the origin of arbuscular mycorrhizae, and the diversification of extant embryophytes occurred ca. 480 Ma. This was followed by the origin of extant lichens. Saprotrophic mushrooms diversified in the Late Paleozoic as forests of seed plants started to dominate the landscape. The subsequent diversification and explosive radiation of Agaricomycetes, and eventually of ectomycorrhizal mushrooms, were associated with the evolution of Pinaceae in the Mesozoic, and establishment of angiosperm-dominated biomes in the Cretaceous.
DNA sequence data from mitochondrial genomes and c. 1000 nuclear exons were analysed for a complete taxon sampling of manta and devilrays ( Mobulidae) to estimate a current molecular phylogeny for the family. The resulting inferences were combined with morphological information to adopt an integrated approach to resolving the taxonomic arrangement of the family. The members of the genus Manta were found to consistently nest within the Mobula species and consequently the genus Manta is placed into the synonymy of Mobula. Mobula eregoodootenkee, M. japanica and M. rochebrunei were each found to be junior synonyms of M. kuhlii, M. mobular and M. hypostoma, respectively. The mitochondrial and nuclear tree topologies were in agreement except for the placement of M. tarapacana which was basal to all other mobulids in the nuclear exon analysis, but as the sister group to the M. alfrediM. birostris- M. mobular clade in the mitochondrial genome analysis. Results from this study are used to a revise the taxonomy for the family Mobulidae. A single genus is now recognized ( where there were previously two) and eight nominal species ( where there were previously 11).
Despite decades of research and an enormity of resultant data, cancer remains a significant public health problem. New tools and fresh perspectives are needed to obtain fundamental insights, to develop better prognostic and predictive tools, and to identify improved therapeutic interventions. With increasingly common genome-scale data, one suite of algorithms and concepts with potential to shed light on cancer biology is phylogenetics, a scientific discipline used in diverse fields. From grouping subsets of cancer samples to tracing subclonal evolution during cancer progression and metastasis, the use of phylogenetics is a powerful systems biology approach. Well-developed phylogenetic applications provide fast, robust approaches to analyze high-dimensional, heterogeneous cancer data sets. This article is part of a Special Issue entitled: Evolutionary principles - heterogeneity in cancer?, edited by Dr. Robert A. Gatenby.
Phylogeographic analysis can be described as the study of the geological and climatological processes that have produced contemporary geographic distributions of populations and species. Here, we attempt to understand how the dynamic process of landscape change on Madagascar has shaped the distribution of a targeted clade of mouse lemurs (genus Microcebus) and, conversely, how phylogenetic and population genetic patterns in these small primates can reciprocally advance our understanding of Madagascar's prehuman environment. The degree to which human activity has impacted the natural plant communities of Madagascar is of critical and enduring interest. Today, the eastern rainforests are separated from the dry deciduous forests of the west by a large expanse of presumed anthropogenic grassland savanna, dominated by the Family Poaceae, that blankets most of the Central Highlands. Although there is firm consensus that anthropogenic activities have transformed the original vegetation through agricultural and pastoral practices, the degree to which closed-canopy forest extended from the east to the west remains debated. Phylogenetic and population genetic patterns in a five-species clade of mouse lemurs suggest that longitudinal dispersal across the island was readily achieved throughout the Pleistocene, apparently ending at ∼55 ka. By examining patterns of both inter- and intraspecific genetic diversity in mouse lemur species found in the eastern, western, and Central Highland zones, we conclude that the natural environment of the Central Highlands would have been mosaic, consisting of a matrix of wooded savanna that formed a transitional zone between the extremes of humid eastern and dry western forest types.
Phycas is open source, freely available Bayesian phylogenetics software written primarily in C++ but with a Python interface. Phycas specializes in Bayesian model selection for nucleotide sequence data, particularly the estimation of marginal likelihoods, central to computing Bayes Factors. Marginal likelihoods can be estimated using newer methods (Thermodynamic Integration and Generalized Steppingstone) that are more accurate than the widely used Harmonic Mean estimator. In addition, Phycas supports two posterior predictive approaches to model selection: Gelfand-Ghosh and Conditional Predictive Ordinates. The General Time Reversible family of substitution models, as well as a codon model, are available, and data can be partitioned with all parameters unlinked except tree topology and edge lengths. Phycas provides for analyses in which the prior on tree topologies allows polytomous trees as well as fully resolved trees, and provides for several choices for edge length priors, including a hierarchical model as well as the recently described compound Dirichlet prior, which helps avoid overly informative induced priors on tree length.
A fern from the French Pyrenees—×Cystocarpium roskamianum—is a recently formed intergeneric hybrid between parental lineages that diverged from each other approximately 60 million years ago (mya; 95% highest posterior density: 40.2–76.2 mya). This is an extraordinarily deep hybridization event, roughly akin to an elephant hybridizing with a manatee or a human with a lemur. In the context of other reported deep hybrids, this finding suggests that populations of ferns, and other plants with abiotically mediated fertilization, may evolve reproductive incompatibilities more slowly, perhaps because they lack many of the premating isolation mechanisms that characterize most other groups of organisms. This conclusion implies that major features of Earth's biodiversity—such as the relatively small number of species of ferns compared to those of angiosperms—may be, in part, an indirect by-product of this slower "speciation clock" rather than a direct consequence of adaptive innovations by the more diverse lineages.
Department of Ecology and Evolutionary Biology, University of Kansas, Lawrence, Kansas, USAPaul O. LewisDepartment of Ecology and Evolutionary Biology, University of Connecticut, Storrs, Connecticut, USADavid L. SwoffordDepartment of Biology and Institute for Genome Sciences and Policy, Duke University, Durham, North Carolina, USADavid BryantDepartment of Mathematics and Statistics, University of Otago, Dunedin, New ZealandCONTENTS5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 96 5.2 The generalized stepping-stone (GSS) method . . . . . . . . . . . . . . . . . . 96 5.3 Reference distribution for tree topology . . . . . . . . . . . . . . . . . . . . . . . . . 985.3.1 Tree topology reference distribution . . . . . . . . . . . . . . . . . . . . . 98 5.3.1.1 Tree simulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 99 5.3.1.2 Tree topology reference distribution . . . . . . . . . 1005.3.2 Edge length reference distribution . . . . . . . . . . . . . . . . . . . . . . 103 5.3.3 Comparison with CCD methods . . . . . . . . . . . . . . . . . . . . . . . . 1045.4 Example . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 5.4.1 Model details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 105 5.4.2 Brute-force approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 106 5.4.3 GSS performance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1085.5 Summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 5.6 Funding . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 111Algorithms, andThe marginal likelihood is central to Bayesian model selection. It is the normalizing constant in Bayes' formula, and the Bayes factor used to compare two models is a ratio of marginal likelihoods. The marginal likelihood is defined as the expected value of the likelihood with respect to the prior. Methods for accurately estimating the marginal likelihood in phylogenetics were developed only recently (Lartillot and Philippe, 2006a; Xie et al., 2011; Fan et al., 2011; Arima and Tardella, 2012). Two of these methods-thermodynamic integration (TI; Lartillot and Philippe, 2006a) and the stepping-stone method (SS; Xie et al., 2011)—allow marginal likelihood estimation when the tree topology varies, but the most efficient methods to date-generalized stepping-stone (GSS; Fan et al., 2011) and the inflated density ratio method (IDR; Arima and Tardella, 2012)—have thus far remained restricted to estimating marginal likelihoods for a fixed tree topology. This chapter is concerned with updating GSS to allow variable tree topology, and the chapter by Wu et al. (Chapter 6) is concerned with updating the IDR method to allow variable tree topology.
Here we provide a detailed comparative analysis across the candidate X-Inactivation Center (XIC) region and the XIST locus in the genomes of six primates and three mammalian outgroup species. Since lemurs and other strepsirrhine primates represent the sister lineage to all other primates, this analysis focuses on lemurs to reconstruct the ancestral primate sequences and to gain insight into the evolution of this region and the genes within it. This comparative evolutionary genomics approach reveals significant expansion in genomic size across the XIC region in higher primates, with minimal size alterations across the XIST locus itself. Reconstructed primate ancestral XIC sequences show that the most dramatic changes during the past 80 million years occurred between the ancestral primate and the lineage leading to Old World monkeys. In contrast, the XIST locus compared between human and the primate ancestor does not indicate any dramatic changes to exons or XIST-specific repeats; rather, evolution of this locus reflects small incremental changes in overall sequence identity and short repeat insertions. While this comparative analysis reinforces that the region around XIST has been subject to significant genomic change, even among primates, our data suggest that evolution of the XIST sequences themselves represents only small lineage-specific changes across the past 80 million years.
Phylogenetic inference is fundamental to our understanding of most aspects of the origin and evolution of life, and in recent years, there has been a concentration of interest in statistical approaches such as Bayesian inference and maximum likelihood estimation. Yet, for large data sets and realistic or interesting models of evolution, these approaches remain computationally demanding. High-throughput sequencing can yield data for thousands of taxa, but scaling to such problems using serial computing often necessitates the use of nonstatistical or approximate approaches. The recent emergence of graphics processing units (GPUs) provides an opportunity to leverage their excellent floating-point computational performance to accelerate statistical phylogenetic inference. A specialized library for phylogenetic calculation would allow existing software packages to make more effective use of available computer hardware, including GPUs. Adoption of a common library would also make it easier for other emerging computing architectures, such as field programmable gate arrays, to be used in the future. We present BEAGLE, an application programming interface (API) and library for high-performance statistical phylogenetic inference. The API provides a uniform interface for performing phylogenetic likelihood calculations on a variety of compute hardware platforms. The library includes a set of efficient implementations and can currently exploit hardware including GPUs using NVIDIA CUDA, central processing units (CPUs) with Streaming SIMD Extensions and related processor supplementary instruction sets, and multicore CPUs via OpenMP. To demonstrate the advantages of a common API, we have incorporated the library into several popular phylogenetic software packages. The BEAGLE library is free open source software licensed under the Lesser GPL and available from http://beagle-lib.googlecode.com. An example client program is available as public domain software.
This is an electronic version of an article published in Systematic Biology [Holder, Mark T., Paul O. Lewis, and David L. Swofford. The Akaike information criterion will not choose the no common mechanism model. Systematic Biology, 59(4):477{485, 2010. ] Systematic Biology is available online at informaworld http://dx.doi.org/10.1093/sysbio/syq028.
Because horizontal gene transfer can confound the recovery of the largely prokaryotic tree of life (ToL), most genome-based techniques seek to eliminate horizontal signal from ToL analyses, commonly by sieving out incongruent genes and data. This approach greatly limits the number of gene families analysed to a subset thought to be representative of vertical evolutionary history. However, formalized tests have not been performed to determine whether combining the massive amounts of information available in fully sequenced genomes can recover a reasonable ToL. Consequently, we used empirically defined gene homology definitions from a previous study that delineate xenologous gene families (gene families derived from a common transfer event) to generate a massively concatenated, combined-data ToL matrix derived from 323 404 translated open reading frames arranged into 12 381 gene homologue groups coded as amino acid data and 63 336, 64 105, 65 153, 66 922 and 67 109 gene homologue groups coded as gene presence/absence data for 166 fully sequenced genomes. This whole-genome gene presence/absence and amino acid sequence ToL data matrix is composed of 4867 184 characters (a combined data-type mega-matrix). Phylogenetic analysis of this mega-matrix yielded a fully resolved ToL that classifies all three commonly accepted domains of life as monophyletic and groups most taxa in traditionally recognized locations with high support. Most importantly, these results corroborate the existence of a common evolutionary history for these taxa present in both data types that is evident only when these data are analysed in combination.© The Willi Hennig Society 2010.
Methods for inferring evolutionary trees can be divided into two broad categories: those that operate on a matrix of discrete characters that assigns one or more attributes or character states to each taxon (i.e. sequence or gene-family member); and those that operate on a matrix of pairwise distances between taxa, with each distance representing an estimate of the amount of divergence between two taxa since they last shared a common ancestor (see Chapter 1). The most commonly employed discrete-character methods used in molecular phylogenetics are parsimony and maximum likelihood methods. For molecular data, the character-state matrix is typically an aligned set of DNA or protein sequences, in which the states are the nucleotides A, C, G, and T (i.e. DNA sequences) or symbols representing the 20 common amino acids (i.e. protein sequences); however, other forms of discrete data such as restriction-site presence/absence and gene-order information also may be used.
Inosine monophosphate dehydrogenase (IMPDH) catalyzes an essential step in the biosynthesis of guanine nucleotides. This reaction involves two different chemical transformations, an NAD-linked redox reaction and a hydrolase reaction, that utilize mutually exclusive protein conformations with distinct catalytic residues. How did Nature construct such a complicated catalyst? Here we employ a "Wang-Landau" metadynamics algorithm in hybrid quantum mechanical/molecular mechanical (QM/MM) simulations to investigate the mechanism of the hydrolase reaction. These simulations show that the lowest energy pathway utilizes Arg418 as the base that activates water, in remarkable agreement with previous experiments. Surprisingly, the simulations also reveal a second pathway for water activation involving a proton relay from Thr321 to Glu431. The energy barrier for the Thr321 pathway is similar to the barrier observed experimentally when Arg418 is removed by mutation. The Thr321 pathway dominates at low pH when Arg418 is protonated, which predicts that the substitution of Glu431 with Gln will shift the pH-rate profile to the right. This prediction is confirmed in subsequent experiments. Phylogenetic analysis suggests that the Thr321 pathway was present in the ancestral enzyme, but was lost when the eukaryotic lineage diverged. We propose that the primordial IMPDH utilized the Thr321 pathway exclusively, and that this mechanism became obsolete when the more sophisticated catalytic machinery of the Arg418 pathway was installed. Thus, our simulations provide an unanticipated window into the evolution of a complex enzyme.