Following the acceptance of plate tectonics theory in the latter half of the 20th century, vicariance became the dominant explanation for the distributions of many plant and animal groups. In recent years, however, molecular-clock analyses have challenged a number of well-accepted hypotheses of vicariance. As a widespread group of insects with a fossil record dating back 300 My, cockroaches provide an ideal model for testing hypotheses of vicariance through plate tectonics versus transoceanic dispersal. However, their evolutionary history remains poorly understood, in part due to unresolved relationships among the nine recognized families. Here, we present a phylogenetic estimate of all extant cockroach families, as well as a timescale for their evolution, based on the complete mitochondrial genomes of 119 cockroach species. Divergence dating analyses indicated that the last common ancestor of all extant cockroaches appeared ∼235 Ma, ∼95 My prior to the appearance of fossils that can be assigned to extant families, and before the breakup of Pangaea began. We reconstructed the geographic ranges of ancestral cockroaches and found tentative support for vicariance through plate tectonics within and between several major lineages. We also found evidence of transoceanic dispersal in lineages found across the Australian, Indo-Malayan, African, and Madagascan regions. Our analyses provide evidence that both vicariance and dispersal have played important roles in shaping the distribution and diversity of these insects.
Species are fundamental units in biological research and can be defined on the basis of various operational criteria. There has been growing use of molecular approaches for species delimitation. Among the most widely used methods, the generalized mixed Yule-coalescent (GMYC) and Poisson tree processes (PTP) were designed for the analysis of single-locus data but are often applied to concatenations of multilocus data. In contrast, the Bayesian multispecies coalescent approach in the software Bayesian Phylogenetics and Phylogeography (BPP) explicitly models the evolution of multilocus data. In this study, we compare the performance of GMYC, PTP, and BPP using synthetic data generated by simulation under various speciation scenarios. We show that in the absence of gene flow, the main factor influencing the performance of these methods is the ratio of population size to divergence time, while number of loci and sample size per species have smaller effects. Given appropriate priors and correct guide trees, BPP shows lower rates of species overestimation and underestimation, and is generally robust to various potential confounding factors except high levels of gene flow. The single-threshold GMYC and the best strategy that we identified in PTP generally perform well for scenarios involving more than a single putative species when gene flow is absent, but PTP outperforms GMYC when fewer species are involved. Both methods are more sensitive than BPP to the effects of gene flow and potential confounding factors. Case studies of bears and bees further validate some of the findings from our simulation study, and reveal the importance of using an informed starting point for molecular species delimitation. Our results highlight the key factors affecting the performance of molecular species delimitation, with potential benefits for using these methods within an integrative taxonomic framework.
Summary Statistical phylogenetic inference plays an important role in evolutionary biology. The accuracy of phylogenetic methods relies on having suitable models of the evolutionary process. Various tools allow comparisons of candidate phylogenetic models, but assessing the absolute performance of models remains a considerable challenge. We introduce PhyloMAd, a user-friendly application for assessing the adequacy of commonly used models of nucleotide substitution and among-lineage rate variation. Our software implements a fast, likelihood-based method of model assessment that is tractable for analyses of large multi-locus datasets. PhyloMAd provides a means of informing model improvement, or selecting data to enhance the evolutionary signal in phylogenomic analyses. Availability and implementation PhyloMAd, together with a manual, a tutorial and the source code, are freely available from the GitHub repository github.com/duchene/phylomad.
A fundamental challenge in resolving evolutionary relationships across the tree of life is to account for heterogeneity in the evolutionary signal across loci. Studies of marsupial mammals have demonstrated that this heterogeneity can be substantial, leaving considerable uncertainty in the evolutionary timescale and relationships within the group. Using simulations and a new phylogenomic data set comprising nucleotide sequences of 1550 loci from 18 of the 22 extant marsupial families, we demonstrate the power of a method for identifying clusters of loci that support different phylogenetic trees. We find two distinct clusters of loci, each providing an estimate of the species tree that matches previously proposed resolutions of the marsupial phylogeny. We also identify a well-supported placement for the enigmatic marsupial moles (Notoryctes) that contradicts previous molecular estimates but is consistent with morphological evidence. The pattern of gene-tree variation across tree-space is characterized by changes in information content, GC content, substitution-model adequacy, and signatures of purifying selection in the data. In a simulation study, we show that incomplete lineage sorting can explain the division of loci into the two tree-topology clusters, as found in our phylogenomic analysis of marsupials. We also demonstrate the potential benefits of minimizing uncertainty from phylogenetic conflict for molecular dating. Our analyses reveal that Australasian marsupials appeared in the early Paleocene, whereas the diversification of present-day families occurred primarily during the late Eocene and early Oligocene. Our methods provide an intuitive framework for improving the accuracy and precision of phylogenetic inference and molecular dating using genome-scale data.
The evolutionary timescale of angiosperms has long been a key question in biology. Molecular estimates of this timescale have shown considerable variation, being influenced by differences in taxon sampling, gene sampling, fossil calibrations, evolutionary models, and choices of priors. Here, we analyze a data set comprising 76 protein-coding genes from the chloroplast genomes of 195 taxa spanning 86 families, including novel genome sequences for 11 taxa, to evaluate the impact of models, priors, and gene sampling on Bayesian estimates of the angiosperm evolutionary timescale. Using a Bayesian relaxed molecular-clock method, with a core set of 35 minimum and two maximum fossil constraints, we estimated that crown angiosperms arose 221 (251-192) Ma during the Triassic. Based on a range of additional sensitivity and subsampling analyses, we found that our date estimates were generally robust to large changes in the parameters of the birth-death tree prior and of the model of rate variation across branches. We found an exception to this when we implemented fossil calibrations in the form of highly informative gamma priors rather than as uniform priors on node ages. Under all other calibration schemes, including trials of seven maximum age constraints, we consistently found that the earliest divergences of angiosperm clades substantially predate the oldest fossils that can be assigned unequivocally to their crown group. Overall, our results and experiments with genome-scale data suggest that reliable estimates of the angiosperm crown age will require increased taxon sampling, significant methodological changes, and new information from the fossil record. [Angiospermae, chloroplast, genome, molecular dating, Triassic.].
In statistical phylogenetic analyses of DNA sequences, models of evolutionary change commonly assume that base composition is stationary through time and across lineages. This assumption is violated by many data sets, but it is unclear whether the magnitude of these violations is sufficient to mislead phylogenetic inference. We investigated the impacts of compositional heterogeneity on phylogenetic estimates using a method for assessing model adequacy. Based on a detailed simulation study, we found that common frequentist criteria are highly conservative, such that the model is often rejected when the phylogenetic estimates do not show clear signs of bias. We propose new criteria and provide guidelines for their usage. We apply these criteria to genome-scale data from 40 birds and find that loci with severely non-homogeneous base composition are uncommon. Our results show the importance of using well-informed diagnostic statistics when testing model adequacy for phylogenomic analyses.
The higher termites (Termitidae) are keystone species and ecosystem engineers. They have exceptional biomass and play important roles in decomposition of dead plant matter, in soil manipulation, and as the primary food for many animals, especially in the tropics. Higher termites are most diverse in rainforests, with estimated origins in the late Eocene (similar to 54 Ma), postdating the breakup of Pangaea and Gondwana when most continents became separated. Since termites are poor fliers, their origin and spread across the globe requires alternative explanation. Here, we show that higher termites originated 42-54Ma in Africa and subsequently underwent at least 24 dispersal events between the continents in two main periods. Using phylogenetic analyses of mitochondrial genomes from 415 species, including all higher termite taxonomic and feeding groups, we inferred 10 dispersal events to South America and Asia 35-23Ma, coinciding with the sharp decrease in global temperature, sea level, and rainforest cover in the Oligocene. After global temperatures increased, 23-5Ma, there was only one more dispersal to South America but 11 to Asia and Australia, and one dispersal back to Africa. Most of these dispersal events were transoceanic and might have occurred via floating logs. The spread of higher termites across oceans was helped by the novel ecological opportunities brought about by environmental and ecosystem change, and led termites to become one of the few insect groups with specialized mammal predators. This has parallels with modern invasive species that have been able to thrive in human-impacted ecosystems.
MOTIVATION:In rapidly evolving pathogens, including viruses and some bacteria, genetic change can accumulate over short time-frames. Accordingly, their sampling times can be used to calibrate molecular clocks, allowing estimation of evolutionary rates. Methods for estimating rates from time-structured data vary in how they treat phylogenetic uncertainty and rate variation among lineages. We compiled 81 virus data sets and estimated nucleotide substitution rates using root-to-tip regression, least-squares dating and Bayesian inference.RESULTS:Although estimates from these three methods were often congruent, this largely relied on the choice of clock model. In particular, relaxed-clock models tended to produce higher rate estimates than methods that assume constant rates. Discrepancies in rate estimates were also associated with high among-lineage rate variation, and phylogenetic and temporal clustering. These results provide insights into the factors that affect the reliability of rate estimates from time-structured sequence data, emphasizing the importance of clock-model testing.CONTACT:sduchene@unimelb.edu.au or garzonsebastian@hotmail.comSupplementary information: Supplementary data are available at Bioinformatics online.
MOTIVATION Molecular-clock methods can be used to estimate evolutionary rates and timescales from DNA sequence data. However, different genes can display different patterns of rate variation across lineages, calling for the employment of multiple clock models. Selecting the optimal clock-partitioning scheme for a multigene dataset can be computationally demanding, but clustering methods provide a feasible alternative. We investigated the performance of different clustering methods using data from chloroplast genomes and data generated by simulation. RESULTS Our results show that mixture models provide a useful alternative to traditional partitioning algorithms. We found only a small number of distinct patterns of among-lineage rate variation among chloroplast genes, which were consistent across taxonomic scales. This suggests that the evolution of chloroplast genes has been governed by a small number of genomic pacemakers. Our study also demonstrates that clustering methods provide an efficient means of identifying clock-partitioning schemes for genome-scale datasets. AVAILABILITY AND IMPLEMENTATION The code and data sets used in this study are available online at https://github.com/sebastianduchene/pacemaker_clustering_methods CONTACT sebastian.duchene@sydney.edu.au SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
The order Passeriformes comprises the majority of extant avian species. Analyses of molecular data have provided important insights into the evolution of this diverse order. However, molecular estimates of the evolutionary and demographic timescales of passerine species have been hindered by a lack of reliable calibrations. This has led to a reliance on the application of standard substitution rates to mitochondrial DNA data, particularly rates estimated from analyses of the gene encoding cytochrome b (CYTB). To investigate patterns of rate variation across passerine lineages, we used a Bayesian phylogenetic approach to analyse the protein‐coding genes of 183 mitochondrial genomes. We found that the most commonly used mitochondrial marker, CYTB, has low variation in rates across passerine lineages. This lends support to its widespread use as a molecular clock in birds. However, we also found that the patterns of among‐lineage rate variation in CYTB are only weakly related to the evolutionary rate of the mitochondrial genome as a whole. Our analyses confirmed the presence of mutational saturation at third codon positions across the protein‐coding genes of the mitochondrial genome, reinforcing the view that these sites should be excluded in studies of deep passerine relationships. The results of our analyses have provided information that will be useful for molecular‐clock studies of passerine evolution.
Rates and timescales of viral evolution can be estimated using phylogenetic analyses of time-structured molecular sequences. This involves the use of molecular-clock methods, calibrated by the sampling times of the viral sequences. However, the spread of these sampling times is not always sufficient to allow the substitution rate to be estimated accurately. We conducted Bayesian phylogenetic analyses of simulated virus data to evaluate the performance of the date-randomization test, which is sometimes used to investigate whether time-structured data sets have temporal signal. An estimate of the substitution rate passes this test if its mean does not fall within the 95% credible intervals of rate estimates obtained using replicate data sets in which the sampling times have been randomized. We find that the test sometimes fails to detect rate estimates from data with no temporal signal. This error can be minimized by using a more conservative criterion, whereby the 95% credible interval of the estimate with correct sampling times should not overlap with those obtained with randomized sampling times. We also investigated the behavior of the test when the sampling times are not uniformly distributed throughout the tree, which sometimes occurs in empirical data sets. The test performs poorly in these circumstances, such that a modification to the randomization scheme is needed. Finally, we illustrate the behavior of the test in analyses of nucleotide sequences of cereal yellow dwarf virus. Our results validate the use of the date-randomization test and allow us to propose guidelines for interpretation of its results.
Molecular clock models are commonly used to estimate evolutionary rates and timescales from nucleotide sequences. The goal of these models is to account for rate variation among lineages, such that they are assumed to be adequate descriptions of the processes that generated the data. A common approach for selecting a clock model for a data set of interest is to examine a set of candidates and to select the model that provides the best statistical fit. However, this can lead to unreliable estimates if all the candidate models are actually inadequate. For this reason, a method of evaluating absolute model performance is critical. We describe a method that uses posterior predictive simulations to assess the adequacy of clock models. We test the power of this approach using simulated data and find that the method is sensitive to bias in the estimates of branch lengths, which tends to occur when using underparameterized clock models. We also compare the performance of the multinomial test statistic, originally developed to assess the adequacy of substitution models, but find that it has low power in identifying the adequacy of clock models. We illustrate the performance of our method using empirical data sets from coronaviruses, simian immunodeficiency virus, killer whales, and marine turtles. Our results indicate that methods of investigating model adequacy, including the one proposed here, should be routinely used in combination with traditional model selection in evolutionary studies. This will reveal whether a broader range of clock models to be considered in phylogenetic analysis.
Phylogenetic estimation of evolutionary timescales has become routine in biology, forming the basis of a wide range of evolutionary and ecological studies. However, there are various sources of bias that can affect these estimates. We investigated whether tree imbalance, a property that is commonly observed in phylogenetic trees, can lead to reduced accuracy or precision of phylogenetic timescale estimates. We analysed simulated data sets with calibrations at internal nodes and at the tips, taking into consideration different calibration schemes and levels of tree imbalance. We also investigated the effect of tree imbalance on two empirical data sets: mitogenomes from primates and serial samples of the African swine fever virus. In analyses calibrated using dated, heterochronous tips, we found that tree imbalance had a detrimental impact on precision and produced a bias in which the overall timescale was underestimated. A pronounced effect was observed in analyses with shallow calibrations. The greatest decreases in accuracy usually occurred in the age estimates for medium and deep nodes of the tree. In contrast, analyses calibrated at internal nodes did not display a reduction in estimation accuracy or precision due to tree imbalance. Our results suggest that molecular-clock analyses can be improved by increasing taxon sampling, with the specific aims of including deeper calibrations, breaking up long branches and reducing tree imbalance.
Evolutionary timescales can be estimated from genetic data using phylogenetic methods based on the molecular clock. To account for molecular rate variation among lineages, a number of relaxed-clock models have been developed. Some of these models assume that rates vary among lineages in an autocorrelated manner, so that closely related species share similar rates. In contrast, uncorrelated relaxed clocks allow all of the branch-specific rates to be drawn from a single distribution, without assuming any correlation between rates along neighbouring branches. There is uncertainty about which of these two classes of relaxed-clock models are more appropriate for biological data. We present an R package, NELSI, that allows the evolution of DNA sequences to be simulated according to a range of clock models. Using data generated by this package, we assessed the ability of two Bayesian phylogenetic methods to distinguish among different relaxed-clock models and to quantify rate variation among lineages. The results of our analyses show that rate autocorrelation is typically difficult to detect, even when there is complete taxon sampling. This provides a potential explanation for past failures to detect rate autocorrelation in a range of data sets.
We are writing in response to a recent critique by Emerson & Hickerson ( ), who challenge the evidence of a time‐dependent bias in molecular rate estimates. This bias takes the form of a negative relationship between inferred evolutionary rates and the ages of the calibrations on which these estimates are based. Here, we present a summary of the evidence obtained from a broad range of taxa that supports a time‐dependent bias in rate estimates, with a consideration of the potential causes of these observed trends. We also describe recent progress in improving the reliability of evolutionary rate estimation and respond to the concerns raised by Emerson & Hickerson ( ) about the validity of rates estimated from time‐structured sequence data. In doing so, we hope to dispel some misconceptions and to highlight several research directions that will improve our understanding of time‐dependent biases in rate estimates.
The evolution of hepatitis B virus (HBV), particularly its origins and evolutionary timescale, has been the subject of debate. Three major scenarios have been proposed, variously placing the origin of HBV in humans and great apes from some million years to only a few thousand years ago (ka). To compare these scenarios, we analyzed 105 full-length HBV genome sequences from all major genotypes sampled globally. We found a high correlation between the demographic histories of HBV and humans, as well as coincidence in the times of origin of specific subgenotypes with human migrations giving rise to their host indigenous populations. Together with phylogenetic evidence, this suggests that HBV has co-expanded with modern humans. Based on the co-expansion, we conducted a Bayesian dating analysis to estimate a precise evolutionary timescale for HBV. Five calibrations were used at the origins of F/H genotypes, D4, C3 and B6 from respective indigenous populations in the Pacific and Arctic and A5 from Haiti. The estimated time for the origin of HBV was 34.1ka (95% highest posterior density interval 27.6-41.3ka), coinciding with the dispersal of modern non-African humans. Our study, the first to use full-length HBV sequences, places a precise timescale on the HBV epidemic and also shows that the "branching paradox" of the more divergent genotypes F/H from Amerindians is due to an accelerated substitution rate, probably driven by positive selection. This may explain previously observed differences in the natural history of HBV between genotypes F1 and A2, B1, and D.
UNLABELLED:Genomic evolution is shaped by a dynamic combination of mutation, selection and genetic drift. These processes lead to evolutionary rate variation across loci and among lineages. In turn, interactions between these two forms of rate variation can produce residual effects, whereby the pattern of among-lineage rate heterogeneity varies across loci. The nature of rate variation is encapsulated in the pacemaker models of genome evolution, which differ in the degree of importance assigned to residual effects: none (Universal Pacemaker), some (Multiple Pacemaker) or total (Degenerate Multiple Pacemaker). Here we use a phylogenetic method to partition the rate variation across loci, allowing comparison of these pacemaker models. Our analysis of 431 genes from 29 mammalian taxa reveals that rate variation across these genes can be explained by 13 pacemakers, consistent with the Multiple Pacemaker model. We find no evidence that these pacemakers correspond to gene function. Our results have important consequences for understanding the factors driving genomic evolution and for molecular-clock analyses.AVAILABILITY AND IMPLEMENTATION:ClockstaR-G is freely available for download from github (https://github.com/sebastianduchene/clockstarg).
The plant pathogen Phytophthora infestans emerged in Europe in 1845, triggering the Irish potato famine and massive European potato crop losses that continued until effective fungicides were widely employed in the 20th century. Today the pathogen is ubiquitous, with more aggressive and virulent strains surfacing in recent decades. Recently, complete P. infestans mitogenome sequences from 19th-century herbarium specimens were shown to belong to a unique lineage (HERB-1) predicted to be rare or extinct in modern times. We report 44 additional P. infestans mitogenomes: four from 19th-century Europe, three from 1950s UK, and 37 from modern populations across the New World. We use phylogenetic analyses to identify the HERB-1 lineage in modern populations from both Mexico and South America, and to demonstrate distinct mitochondrial haplotypes were present in 19th-century Europe, with this lineage initially diversifying 75 years before the first reports of potato late blight.
The angiosperm genus Logania R.Br. (Loganiaceae) is endemic to the mainland of Australia. A recent genetic study challenged the monophyly of Logania, suggesting that its two sections, Logania sect. Logania and Logania sect. Stomandra, do not group together. Additionally, the genus has a disjunct distribution, with a gap at the Nullarbor Plain in southern Australia. Therefore, Logania is a favourable candidate to gain insight into phylogenetic relationships and how these might intersect with Earth-history events. Our phylogenetic analyses of DNA sequences of two chloroplast markers (petD and rps16) showed that Logania sect. Logania and L. sect. Stomandra were each resolved as monophyletic, but the genus (as currently circumscribed) was not. Based on our Bayesian estimates of divergence times, the disjunct distributions within Logania sect. Stomandra could have been caused by flooding of the Eucla Basin. However, this biogeographical process cannot account for the distribution of Logania sect. Logania, with long-distance dispersal and establishment seeming more likely.
The termite genus Coptotermes (Rhinotermitidae) is found in Asia, Africa, Central/South America and Australia, with greatest diversity in Asia. Some Coptotermes species are amongst the world's most damaging invasive termites, but the genus is also significant for containing the most sophisticated mound-building termites outside the family Termitidae. These mound-building Coptotermes occur only in Australia. Despite its economic and evolutionary significance, the biogeographic history of the genus has not been well investigated, nor has the evolution of the Australian mound-building species. We present here the first phylogeny of the Australian Coptotermes to include representatives from all described species. We combined our new data with previously generated data to estimate the first phylogeny to include representatives from all continents where the genus is found. We also present the first estimation of divergence dates during the evolution of the genus. We found the Australian Coptotermes to be monophyletic and most closely related to the Asian Coptotermes, with considerable genetic diversity in some Australian taxa possibly representing undescribed species. The Australian mound-building species did not form a monophyletic clade. Our ancestral state reconstruction analysis indicated that the ancestral Australian Coptotermes was likely to have been a tree nester, and that mound-building behaviour has arisen multiple times. The Australian Coptotermes were found to have diversified ∼13million years ago, which plausibly matches with the narrowing of the Arafura Sea allowing Asian taxa to cross into Australia. The first diverging Coptotermes group was found to be African, casting doubt on the previously raised hypothesis that the genus has an Asian origin.