Summary Haplotype networks are an intuitive method for visualising relationships between individual genotypes at the population level. Here, we present popart, an integrated software package that provides a comprehensive implementation of haplotype network methods, phylogeographic visualisation tools and standard statistical tests, together with publication‐ready figure production. popart also provides a platform for the implementation and distribution of new network‐based methods – we describe one such new method, integer neighbour‐joining. The software is open source and freely available for all major operating systems.
Simulation experiments are used widely throughout evolutionary biology and bioinformatics to compare models, promote methods, and test hypotheses. The biggest practical constraint on simulation experiments is the computational demand, particularly as the number of parameters increases. Given the extraordinary success of Monte Carlo methods for conducting inference in phylogenetics, and indeed throughout the sciences, we investigate ways in which Monte Carlo framework can be used to carry out simulation experiments more efficiently. The key idea is to sample parameter values for the experiments, rather than iterate through them exhaustively. Exhaustive analyses become completely infeasible when the number of parameters gets too large, whereas sampled approaches can fare better in higher dimensions. We illustrate the framework with applications to phylogenetics and genetic archaeology.
Predicting species’ chances of survival under climate change requires an understanding of their adaptive potential. Now research into hybridization—one mechanism that could facilitate adaptation—shows that species of the plant genus Pachycladon that survived the Last Glacial Maximum benefited from the transfer of genetic information through hybridization. Predicting survival and extinction scenarios for climate change requires an understanding of the present day ecological characteristics of species and future available habitats, but also the adaptive potential of species to cope with environmental change. Hybridization is one mechanism that could facilitate this. Here we report statistical evidence that the transfer of genetic information through hybridization is a feature of species from the plant genus Pachycladon that survived the Last Glacial Maximum in geographically separated alpine refugia in New Zealand’s South Island. We show that transferred glucosinolate hydrolysis genes also exhibit evidence of intra-locus recombination. Such gene exchange and recombination has the potential to alter the chemical defence in the offspring of hybridizing species. We use a mathematical model to show that when hybridization increases the adaptive potential of species, future biodiversity will be best protected by preserving closely related species that hybridize rather than by conserving distantly related species that are genetically isolated.
Simulations often involve the use of model parameters which are unknown or uncertain. For this reason, simulation experiments are often repeated for multiple combinations of parameter values, often iterating through parameter values lying on a fixed grid. However, the use of a discrete grid places limits on the dimension of the parameter space and creates the potential to miss important parameter combinations which fall in the gaps between grid points. Here we draw parallels with strategies for numerical integration and describe a Markov chain Monte-Carlo strategy for exploring parameter values. We illustrate the approach using examples from phylogenetics, archaeology, and epidemiology.
Congruence is a broadly applied notion in evolutionary biology used to justify multigene phylogeny or phylogenomics, as well as in studies of coevolution, lateral gene transfer, and as evidence for common descent. Existing methods for identifying incongruence or heterogeneity using character data were designed for data sets that are both small and expected to be rarely incongruent. At the same time, methods that assess incongruence using comparison of trees test a null hypothesis of uncorrelated tree structures, which may be inappropriate for phylogenomic studies. As such, they are ill-suited for the growing number of available genome sequences, most of which are from prokaryotes and viruses, either for phylogenomic analysis or for studies of the evolutionary forces and events that have shaped these genomes. Specifically, many existing methods scale poorly with large numbers of genes, cannot accommodate high levels of incongruence, and do not adequately model patterns of missing taxa for different markers. We propose the development of novel incongruence assessment methods suitable for the analysis of the molecular evolution of the vast majority of life and support the investigation of homogeneity of evolutionary process in cases where markers do not share identical tree structures.
Interest in congruence in phylogenetic data has largely focused on issues affecting multicellular organisms, and animals in particular, in which the level of incongruence is expected to be relatively low. In addition, assessment methods developed in the past have been designed for reasonably small numbers of loci and scale poorly for larger data sets. However, there are currently over a thousand complete genome sequences available and of interest to evolutionary biologists, and these sequences are predominantly from microbial organisms, whose molecular evolution is much less frequently tree-like than that of multicellular life forms. As such, the level of incongruence in these data is expected to be high. We present a congruence method that accommodates both very large numbers of genes and high degrees of incongruence. Our method uses clustering algorithms to identify subsets of genes based on similarity of phylogenetic signal. It involves only a single phylogenetic analysis per gene, and therefore, computation time scales nearly linearly with the number of genes in the data set. We show that our method performs very well with sets of sequence alignments simulated under a wide variety of conditions. In addition, we present an analysis of core genes of prokaryotes, often assumed to have been largely vertically inherited, in which we identify two highly incongruent classes of genes. This result is consistent with the complexity hypothesis.
The most conspicuous feature in previous phaeophycean phylogenies is a large polytomy known as the brown algal crown radiation (BACR). The BACR encompasses 10 out of the 17 currently recognized brown algal orders. A recent study has been able to resolve a few nodes of the BACR, suggesting that it may be a soft polytomy caused by a lack of signal in molecular markers. The present work aims to refine relationships within the BACR and investigate the nature and timeframe of the diversification in question using a dual approach. A multi-marker phylogeny of the brown algae was built from 10 mitochondrial, plastid and nuclear loci (>10,000 nt) of 72 phaeophycean taxa, resulting in trees with well-resolved inter-ordinal relationships within the BACR. Using Bayesian relaxed molecular clock analysis, it is shown that the BACR is likely to represent a gradual diversification spanning most of the Lower Cretaceous rather than a sudden radiation. Non-molecular characters classically used in ordinal delimitation were mapped on the molecular topology to study their evolutionary history.
Nearly all of eukaryotic diversity has been classified into 6 suprakingdom-level groups (supergroups) based on molecular and morphological/cell-biological evidence; these are Opisthokonta, Amoebozoa, Archaeplastida, Rhizaria, Chromalveolata, and Excavata. However, molecular phylogeny has not provided clear evidence that either Chromalveolata or Excavata is monophyletic, nor has it resolved the relationships among the supergroups. To establish the affinities of Excavata, which contains parasites of global importance and organisms regarded previously as primitive eukaryotes, we conducted a phylogenomic analysis of a dataset of 143 proteins and 48 taxa, including 19 excavates. Previous phylogenomic studies have not included all major subgroups of Excavata, and thus have not definitively addressed their interrelationships. The enigmatic flagellate Andalucia is sister to typical jakobids. Jakobids (including Andalucia ), Euglenozoa and Heterolobosea form a major clade that we name Discoba. Analyses of the complete dataset group Discoba with the mitochondrion-lacking excavates or “metamonads” (diplomonads, parabasalids, and Preaxostyla), but not with the final excavate group, Malawimonas . This separation likely results from a long-branch attraction artifact. Gradual removal of rapidly-evolving taxa from the dataset leads to moderate bootstrap support (69%) for the monophyly of all Excavata, and 90% support once all metamonads are removed. Most importantly, Excavata robustly emerges between unikonts (Amoebozoa + Opisthokonta) and “megagrouping” of Archaeplastida, Rhizaria, and chromalveolates. Our analyses indicate that Excavata forms a monophyletic suprakingdom-level group that is one of the 3 primary divisions within eukaryotes, along with unikonts and a megagroup of Archaeplastida, Rhizaria, and the chromalveolate lineages.
DNA flows between chromosomes and mobile elements, following rules that are poorly understood. This limited knowledge is partly explained by the limits of current approaches to study the structure and evolution of genetic diversity. Network analyses of 119,381 homologous DNA families, sampled from 111 cellular genomes and from 165,529 phage, plasmid, and environmental virome sequences, offer challenging insights. Our results support a disconnected yet highly structured network of genetic diversity, revealing the existence of multiple "genetic worlds." These divides define multiple isolated groups of DNA vehicles drawing on distinct gene pools. Mathematical studies of the centralities of these worlds' subnetworks demonstrate that plasmids, not viruses, were key vectors of genetic exchange between bacterial chromosomes, both recently and in the past. Furthermore, network methodology introduces new ways of quantifying current sampling of genetic diversity.
Several morphologically dissimilar ascomycete fungi including Schizosaccharomyces, Taphrina, Saitoella, Pneumocystis, and Neolecta have been grouped into the taxon Taphrinomycotina (Archiascomycota or Archiascomycotina), originally based on rRNA phylogeny. These analyses lack statistically significant support for the monophyly of this grouping, and although confirmed by more recent multigene analyses, this topology is contradicted by mitochondrial phylogenies. To resolve this inconsistency, we have assembled phylogenomic mitochondrial and nuclear data sets from four distantly related taphrinomycotina taxa: Schizosaccharomyces pombe, Pneumocystis carinii, Saitoella complicata, and Taphrina deformans. Our phylogenomic analyses based on nuclear data (113 proteins) conclusively support the monophyly of Taphrinomycotina, diverging as a sister group to Saccharomycotina + Pezizomycotina. However, despite the improved taxon sampling, Taphrinomycotina continue to be paraphyletic with the mitochondrial data set (13 proteins): Schizosaccharomyces species associate with budding yeasts (Saccharomycotina) and the other Taphrinomycotina group as a sister group to Saccharomycotina + Pezizomycotina. Yet, as Schizosaccharomyces and Saccharomycotina species are fast evolving, the mitochondrial phylogeny may be influenced by a long-branch attraction (LBA) artifact. After removal of fast-evolving sequence positions from the mitochondrial data set, we recover the monophyly of Taphrinomycotina. Our combined results suggest that Taphrinomycotina is a legitimate taxon, that this group of species diverges as a sister group to Saccharomycotina + Pezizomycotina, and that phylogenetic positioning of yeasts and fission yeasts with mitochondrial data is plagued by a strong LBA artifact.
Phylogenomic analyses of large sets of genes or proteins have the potential to revolutionize our understanding of the tree of life. However, problems arise because estimated phylogenies from individual loci often differ because of different histories, systematic bias, or stochastic error. We have developed Concaterpillar, a hierarchical clustering method based on likelihood-ratio testing that identifies congruent loci for phylogenomic analysis. Concaterpillar also includes a test for shared relative evolutionary rates between genes indicating whether they should be analyzed separately or by concatenation. In simulation studies, the performance of this method is excellent when a multiple comparison correction is applied. We analyzed a phylogenomic data set of 60 translational protein sequences from the major supergroups of eukaryotes and identified three congruent subsets of proteins. Analysis of the largest set indicates improved congruence relative to the full data set and produced a phylogeny with stronger support for five eukaryote supergroups including the Opisthokonts, the Plantae, the stramenopiles + Apicomplexa (chromalveolates), the Amoebozoa, and the Excavata. In contrast, the phylogeny of the second largest set indicates a close relationship between stramenopiles and red algae, to the exclusion of alveolates, suggesting gene transfer from the red algal secondary symbiont to the ancestral stramenopile host nucleus during the origin of their chloroplast. Investigating phylogenomic data sets for conflicting signals has the potential to both improve phylogenetic accuracy and inform our understanding of genome evolution.
Background: Fornicata is a relatively recently established group of protists that includes the diplokaryotic diplomonads (which have two similar nuclei per cell), and the monokaryotic enteromonads, retortamonads and Carpediemonas, with the more typical one nucleus per cell. The monophyly of the group was confirmed by molecular phylogenetic studies, but neither the internal phylogeny nor its position on the eukaryotic tree has been clearly resolved.Results: Here we have introduced data for three genes (SSU rRNA, alpha-tubulin and HSP90) with a wide taxonomic sampling of Fornicata, including ten isolates of enteromonads, representing the genera Trimitus and Enteromonas, and a new undescribed enteromonad genus. The diplomonad sequences formed two main clades in individual gene and combined gene analyses, with Giardia (and Octomitus) on one side of the basal divergence and Spironucleus, Hexamita and Trepomonas on the other. Contrary to earlier evolutionary scenarios, none of the studied enteromonads appeared basal to diplokaryotic diplomonads. Instead, the enteromonad isolates were all robustly situated within the second of the two diplomonad clades. Furthermore, our analyses suggested that enteromonads do not constitute a monophyletic group, and enteromonad monophyly was statistically rejected in 'approximately unbiased' tests of the combined gene data.Conclusion: We suggest that all higher taxa intended to unite multiple enteromonad genera be abandoned, that Trimitus and Enteromonas be considered as part of Hexamitinae, and that the term 'enteromonads' be used in a strictly utilitarian sense. Our result suggests either that the diplokaryotic condition characteristic of diplomonads arose several times independently, or that the monokaryotic cell of enteromonads originated several times independently by secondary reduction from the diplokaryotic state. Both scenarios are evolutionarily complex. More comparative data on the similarity of the genomes of the two nuclei of diplomonads will be necessary to resolve which evolutionary scenario is more probable.
It has recently been proposed that a well-resolved Tree of Life can be achieved through concatenation of shared genes. There are, however, several difficulties with such an approach, especially in the prokaryotic part of this tree. We tackled some of them using a new combination of maximum likelihood-based methods, developed in order to practice as safe and careful concatenations as possible. First, we used the application concaterpillar on carefully aligned core genes. This application uses a hierarchical likelihood-ratio test framework to assess both the topological congruence between gene phylogenies (i.e., whether different genes share the same evolutionary history) and branch-length congruence (i.e., whether genes that share the same history share the same pattern of relative evolutionary rates). We thus tested if these core genes can be concatenated or should be instead categorized into different incongruent sets. Second, we developed a heat map approach studying the evolution of the phylogenetic support for different bipartitions, when the number of sites of different phylogenetic quality in the concatenation increases. These heatmaps allow us to follow which phylogenetic signals increase or decrease as the concatenation progresses and to detect emerging artifactual groupings, that is, groups that are more and more supported when more and more homoplasic sites are thrown in the analysis. We showed that, as far as 7 major prokaryotic lineages are concerned, only 22 core genes can be said to be congruent and can be safely concatenated. This number is even smaller than the number of genes retained to reconstruct a "Tree of One Per Cent." Furthermore, the concatenation of these 22 markers leads to an unresolved tree as the only groupings in the concatenation tree seem to reflect emerging artifacts. Using concatenated core genes as a valid framework to classify uncharacterized environmental sequences can thus be misleading.
Background Homing endonuclease genes (HEGs) are superfluous, but are capable of invading populations that mix alleles by biasing their inheritance patterns through gene conversion. One model suggests that their long-term persistence is achieved through recurrent invasion. This circumvents evolutionary degeneration, but requires reasonable rates of transfer between species to maintain purifying selection. Although HEGs are found in a variety of microbes, we found the previous discovery of this type of selfish genetic element in the mitochondria of a sea anemone surprising. Methods/Principal Findings We surveyed 29 species of Cnidaria for the presence of the COXI HEG. Statistical analyses provided evidence for HEG invasion. We also found that 96 individuals of Metridium senile, from five different locations in the UK, had identical HEG sequences. This lack of sequence divergence illustrates the stable nature of Anthozoan mitochondria. Our data suggests this HEG conforms to the recurrent invasion model of evolution. Conclusions Ordinarily such low rates of HEG transfer would likely be insufficient to enable major invasion. However, the slow rate of Anthozoan mitochondrial change lengthens greatly the time to HEG degeneration: this significantly extends the periodicity of the HEG life-cycle. We suggest that a combination of very low substitution rates and rare transfers facilitated metazoan HEG invasion.
Here, we address a much-debated topic: is there or is there not an organismal tree of gamma-proteobacteria that can be unambiguously inferred from a core of shared genes? We apply several recently developed analytical methods to this problem, for the first time. Our heat map analyses of P values and of bootstrap bipartitions show the presence of conflicting phylogenetic signals among these core genes. Our synthesis reconstruction suggests that at least 10% of these genes have been laterally transferred during the divergence of the gamma-proteobacteria, and that for most of the rest, there is too little phylogenetic signal to permit firm conclusions about the mode of inheritance. Although there is clearly a central tendency in this data set (it is far from random), lateral gene transfers cannot be ruled out. Instead of an organismal tree, we propose that these core genes could be used to define a more subtle and partially reticulated pattern of relationships.
BACKGROUND:Since Darwin's Origin of Species, reconstructing the Tree of Life has been a goal of evolutionists, and tree-thinking has become a major concept of evolutionary biology. Practically, building the Tree of Life has proven to be tedious. Too few morphological characters are useful for conducting conclusive phylogenetic analyses at the highest taxonomic level. Consequently, molecular sequences (genes, proteins, and genomes) likely constitute the only useful characters for constructing a phylogeny of all life. For this reason, tree-makers expect a lot from gene comparisons. The simultaneous study of the largest number of molecular markers possible is sometimes considered to be one of the best solutions in reconstructing the genealogy of organisms. This conclusion is a direct consequence of tree-thinking: if gene inheritance conforms to a tree-like model of evolution, sampling more of these molecules will provide enough phylogenetic signal to build the Tree of Life. The selection of congruent markers is thus a fundamental step in simultaneous analysis of many genes.RESULTS:Heat map analyses were used to investigate the congruence of orthologues in four datasets (archaeal, bacterial, eukaryotic and alpha-proteobacterial). We conclude that we simply cannot determine if a large portion of the genes have a common history. In addition, none of these datasets can be considered free of lateral gene transfer.CONCLUSION:Our phylogenetic analyses do not support tree-thinking. These results have important conceptual and practical implications. We argue that representations other than a tree should be investigated in this case because a non-critical concatenation of markers could be highly misleading.
To generate data for comparative analyses of zygomycete mitochondrial gene expression, we sequenced mtDNAs of three distantly related zygomycetes, Rhizopus oryzae, Mortierella verticillata and Smittium culisetae. They all contain the standard fungal mitochondrial gene set, plus rnpB, the gene encoding the RNA subunit of the mitochondrial RNase P (mtP-RNA) and rps3, encoding ribosomal protein S3 (the latter lacking in R.oryzae). The mtP-RNAs of R.oryzae and of additional zygomycete relatives have the most eubacteria-like RNA structures among fungi. Precise mapping of the 5′ and 3′ termini of the R.oryzae and M.verticillata mtP-RNAs confirms their expression and processing at the exact sites predicted by secondary structure modeling. The 3′ RNA processing of zygomycete mitochondrial mRNAs, SSU-rRNA and mtP-RNA occurs at the C-rich sequence motifs similar to those identified in fission yeast and basidiomycete mtDNAs. The C-rich motifs are included in the mature transcripts, and are likely generated by exonucleolytic trimming of RNA 3′ termini. Zygomycete mtDNAs feature a variety of insertion elements: (i) mtDNAs of R.oryzae and M.verticillata were subject to invasions by double hairpin elements; (ii) genes of all three species contain numerous mobile group I introns, including one that is closest to an intron that invaded angiosperm mtDNAs; and (iii) at least one additional case of a mobile element, characterized by a homing endonuclease insertion between partially duplicated genes [Paquin,B., Laforest,M.J., Forget,L., Roewer,I., Wang,Z., Longcore,J. and Lang,B.F. (1997) Curr. Genet., 31, 380–395]. The combined mtDNA-encoded proteins contain insufficient phylogenetic signal to demonstrate monophyly of zygomycetes.
Lateral gene transfer (LGT) is often seen as a form of noise, obscuring the phylogenetic signal with which we might hope to reconstruct the evolution of a group of organisms, or indeed the history of all life (the Tree of Life). Such reconstruction might still be possible if the subset of genes conserved among all genomes in a group (or common to all genomes) comprise a core that is relatively refractory to LGT. Several papers designed to test this notion have recently appeared, and here we re-analyze one, which claims that the core of single-copy orthologs shared by all sequenced genomes of the gammaproteobacteria is essentially free of LGT. This conclusion is unfortunately premature, and it is very hard to determine what fraction of this core has been affected by LIST. We discuss other difficulties with the core concept and suggest that, although the core idea must remain part of our understanding of phylogenetic relationships, it should not be the sole basis for defining such relationships, because these are not exclusively tree-like. We suggest instead a more complex but more natural framework for classification, which we call the Synthesis of Life.