There is a long-standing consensus that the animal phyla closest to our own phylum of Chordata are the Echinodermata and Hemichordata. These three phyla constitute the major clade of Deuterostomia. Recent analyses have questioned the support for the monophyly of Deuterostomia, however, showing that the branch leading to deuterostomes is very short and may be influenced by systematic error. Here we use a site-by-site approach to explore multiple sources of error. Under conditions that promote long-branch attraction (LBA), especially branch-length heterogeneity and sites constrained in their amino acid composition, we find that deuterostome monophyly is strongly supported. When we make efforts to mitigate these sources of error, we cannot distinguish between monophyletic and paraphyletic Deuterostomia. Our findings have implications for the interpretation of putative deuterostome fossils, for the reconstruction of a bilaterian ancestor and, more generally, for how datasets for deep-time phylogenetic analyses are assembled and analyzed. ### Competing Interest Statement The authors have declared no competing interest.
The evolutionary origins of Bilateria remain enigmatic. One of the more enduring proposals highlights similarities between a cnidarian-like planula larva and simple acoel-like flatworms. This idea is based in part on the view of the Xenacoelomorpha as an outgroup to all other bilaterians which are themselves designated the Nephrozoa (protostomes and deuterostomes). Genome data can provide important comparative data and help understand the evolution and biology of enigmatic species better. Here, we assemble and analyze the genome of the simple, marine xenacoelomorph Xenoturbella bocki, a key species for our understanding of early bilaterian evolution. Our highly contiguous genome assembly of X. bocki has a size of ~111 Mbp in 18 chromosome-like scaffolds, with repeat content and intron, exon, and intergenic space comparable to other bilaterian invertebrates. We find X. bocki to have a similar number of genes to other bilaterians and to have retained ancestral metazoan synteny. Key bilaterian signaling pathways are also largely complete and most bilaterian miRNAs are present. Overall, we conclude that X. bocki has a complex genome typical of bilaterians, which does not reflect the apparent simplicity of its body plan that has been so important to proposals that the Xenacoelomorpha are the simple sister group of the rest of the Bilateria.
The CODEML program in the PAML package has been widely used to analyze protein-coding gene sequences to estimate the synonymous and nonsynonymous rates (dS and dN) and to detect positive Darwinian selection driving protein evolution. For users not familiar with molecular evolutionary analysis, the program is known to have a steep learning curve. Here, we provide a step-by-step protocol to illustrate the commonly used tests available in the program, including the branch models, the site models, and the branch-site models, which can be used to detect positive selection driving adaptive protein evolution affecting particular lineages of the species phylogeny, affecting a subset of amino acid residues in the protein, and affecting a subset of sites along prespecified lineages, respectively. A data set of the myxovirus (Mx) genes from ten mammal and two bird species is used as an example. We discuss a new feature in CODEML that allows users to perform positive selection tests for multiple genes for the same set of taxa, as is common in modern genome-sequencing projects. The PAML package is distributed at https://github.com/abacus-gene/paml under the GNU license, with support provided at its discussion site (https://groups.google.com/g/pamlsoftware). Data files used in this protocol are available at https://github.com/abacus-gene/paml-tutorial.
Mountains play a key role in forming biodiversity by acting both as barriers to gene flow among populations and as corridors for the migration of populations adapted to the conditions prevailing at high elevations. The Anatolian and the Zagros Mountains are located in the Alpine-Himalayan belt. The formation of these mountains has influenced the distribution and isolation of the animal population since the late Cenozoic. Apathya is a genus of lacertid lizards distributed along these mountains with two species, i.e., Apathya cappadocica and Apathya yassujica. The taxonomy status of lineages within the genus is complicated. In this study, we tried to collect extensive samples from throughout the distribution range, especially within the Zagros Mountains. Also, we used five genetic markers, two mitochondrial (COI and Cyt b) and three nuclear (C-mos, NKTR, and MCIR), to resolve the phylogenetic relationships within the genus and explain several possible scenarios that shaped multiple genetic structures. The combination of results in the current study indicated eight well-support monophyletic lineages that separated to two main groups; group 1 including A. c. cappadocica, A. c. muhtari and A. c. wolteri, group 2 contains four regional clades Turkey, Urmia, Baneh and Ilam, and finally a single clade belonging to the species A. yassujica. In contrast to previous studies, Apathya cappadocica urmiana was divided into four clades and three clades were recognized within Iranian boundaries. The clades have dispersed from Anatolia to adjacent regions in the south of Anatolia and the western Zagros Mountains. According to the evidence generated in this study this clade is paraphyletic. Based on our assumption, orogeny activities and also climate fluctuations in Middle Miocene and Pleistocene have influenced to formation of lineages. In this study we revisit the taxonomy of the genus and demonstrate that the species diversity was substantially underestimated. Our findings suggest that each of the eight clades corresponding to subspecies and distinct geographic regions deserve to be promoted to species level.
Uropeltidae is a clade of small fossorial snakes (ca. 64 extant species) endemic to peninsular India and Sri Lanka. Uropeltid taxonomy has been confusing, and the status of some species has not been revised for over a century. Attempts to revise uropeltid systematics and undertake evolutionary studies have been hampered by incompletely sampled and incompletely resolved phylogenies. To address this issue, we take advantage of historical museum collections, including type specimens, and apply genome-wide shotgun (GWS) sequencing, along with recent field sampling (using Sanger sequencing) to establish a near-complete multilocus species-level phylogeny (ca. 87% complete at species level). This results in a phylogeny that supports the monophyly of all genera (if Brachyophidium is considered a junior synonym of Teretrurus), and provides a firm platform for future taxonomic revision. Sri Lankan uropeltids are probably monophyletic, indicating a single colonisation event of this island from Indian ancestors. However, the position of Rhinophis goweri (endemic to Eastern Ghats, southern India) is unclear and warrants further investigation, and evidence that it may nest within the Sri Lankan radiation indicates a possible recolonisation event. DNA sequence data and morphology suggest that currently recognised uropeltid species diversity is substantially underestimated. Our study highlights the benefits of integrating museum collections in molecular genetic analyses and their role in understanding the systematics and evolutionary history of understudied organismal groups.
Inference of deep phylogenies has almost exclusively used protein rather than DNA sequences based on the perception that protein sequences are less prone to homoplasy and saturation or to issues of compositional heterogeneity than DNA sequences. Here, we analyze a model of codon evolution under an idealized genetic code and demonstrate that those perceptions may be misconceptions. We conduct a simulation study to assess the utility of protein versus DNA sequences for inferring deep phylogenies, with protein-coding data generated under models of heterogeneous substitution processes across sites in the sequence and among lineages on the tree, and then analyzed using nucleotide, amino acid, and codon models. Analysis of DNA sequences under nucleotide-substitution models (possibly with the third codon positions excluded) recovered the correct tree at least as often as analysis of the corresponding protein sequences under modern amino acid models. We also applied the different data-analysis strategies to an empirical dataset to infer the metazoan phylogeny. Our results from both simulated and real data suggest that DNA sequences may be as useful as proteins for inferring deep phylogenies and should not be excluded from such analyses. Analysis of DNA data under nucleotide models has a major computational advantage over protein-data analysis, potentially making it feasible to use advanced models that account for among-site and among-lineage heterogeneity in the nucleotide-substitution process in inference of deep phylogenies.
The evolutionary origins of Bilateria remain enigmatic. One of the more enduring proposals highlights similarities between a cnidarian-like planula larva and simple acoel-like flatworms. This idea is based in part on the view of the Xenacoelomorpha as an outgroup to all other bilaterians which are themselves designated the Nephrozoa (protostomes and deuterostomes). Genome data can help to elucidate phylogenetic relationships and provide important comparative data. Here we assemble and analyse the genome of the simple, marine xenacoelomorph Xenoturbella bocki , a key species for our understanding of early bilaterian and deuterostome evolution. Our highly contiguous genome assembly of X. bocki has a size of ∼111 Mbp in 18 chromosome like scaffolds, with repeat content and intron, exon and intergenic space comparable to other bilaterian invertebrates. We find X. bocki to have a similar number of genes to other bilaterians and to have retained ancestral metazoan synteny. Key bilaterian signalling pathways are also largely complete and most bilaterian miRNAs are present. We conclude that X. bocki has a complex genome typical of bilaterians, in contrast to the apparent simplicity of its body plan. Overall, our data do not provide evidence supporting the idea that Xenacoelomorpha are a primitively simple outgroup to other bilaterians and gene presence/absence data support a relationship with Ambulacraria.
The multispecies coalescent (MSC) model accommodates both species divergences and within-species coalescent and provides a natural framework for phylogenetic analysis of genomic data when the gene trees vary across the genome. The MSC model implemented in the program bpp assumes a molecular clock and the Jukes-Cantor model, and is suitable for analyzing genomic data from closely related species. Here we extend our implementation to more general substitution models and relaxed clocks to allow the rate to vary among species. The MSC-with-relaxed-clock model allows the estimation of species divergence times and ancestral population sizes using genomic sequences sampled from contemporary species when the strict clock assumption is violated, and provides a simulation framework for evaluating species tree estimation methods. We conducted simulations and analyzed two real datasets to evaluate the utility of the new models. We confirm that the clock-JC model is adequate for inference of shallow trees with closely related species, but it is important to account for clock violation for distant species. Our simulation suggests that there is valuable phylogenetic information in the gene-tree branch lengths even if the molecular clock assumption is seriously violated, and the relaxed-clock models implemented in bpp are able to extract such information. Our Markov chain Monte Carlo algorithms suffer from mixing problems when used for species tree estimation under the relaxed clock and we discuss possible improvements. We conclude that the new models are currently most effective for estimating population parameters such as species divergence times when the species tree is fixed.
A wide range of data types can be used to delimit species and various computer-based tools dedicated to this task are now available. Although these formalized approaches have significantly contributed to increase the objectivity of species delimitation (SD) under different assumptions, they are not routinely used by alpha-taxonomists. One obvious shortcoming is the lack of interoperability among the various independently developed SD programs. Given the frequent incongruences between species partitions inferred by different SD approaches, researchers applying these methods often seek to compare these alternative species partitions to evaluate the robustness of the species boundaries. This procedure is excessively time consuming at present, and the lack of a standard format for species partitions is a major obstacle. Here, we propose a standardized format, SPART, to enable compatibility between different SD tools exporting or importing partitions. This format reports the partitions and describes, for each of them, the assignment of individuals to the "inferred species". The syntax also allows support values to be optionally reported, as well as original trees and the full command lines used in the respective SD analyses. Two variants of this format are proposed, overall using the same terminology but presenting the data either optimized for human readability (matricial SPART) or in a format in which each partition forms a separate block (SPART.XML). ABGD, DELINEATE, GMYC, PTP and TR2 have already been adapted to output SPART files and a new version of LIMES has been developed to import, export, merge and split them.
The effort to reconstruct the tree of life was revolutionized by the use of sequences of proteins and nucleic acids. Phylogenetic trees are now routinely inferred using hundreds of thousands of amino acid or nucleotide characters. It thus seems surprising that many aspects of the tree of life are still controversial; conflicting results between large scale phylogenomic studies show that errors remain common despite large datasets. These errors often result from systematic biases in the way sequences evolve. While the resulting systematic errors are well understood, it requires careful efforts to reduce their effects.
The availability of complete sets of genes from many organisms makes it possible to identify genes unique to (or lost from) certain clades. This information is used to reconstruct phylogenetic trees; identify genes involved in the evolution of clade specific novelties; and for phylostratigraphy—identifying ages of genes in a given species. These investigations rely on accurately predicted orthologs. Here we use simulation to produce sets of orthologs that experience no gains or losses. We show that errors in identifying orthologs increase with higher rates of evolution. We use the predicted sets of orthologs, with errors, to reconstruct phylogenetic trees; to count gains and losses; and for phylostratigraphy. Our simulated data, containing information only from errors in orthology prediction, closely recapitulate findings from empirical data. We suggest published downstream analyses must be informed to a large extent by errors in orthology prediction that mimic expected patterns of gene evolution.
The bilaterally symmetric animals (Bilateria) are considered to comprise two monophyletic groups, Protostomia (Ecdysozoa and the Lophotrochozoa) and Deuterostomia (Chordata and the Xenambulacraria). Recent molecular phylogenetic studies have not consistently supported deuterostome monophyly. Here, we compare support for Protostomia and Deuterostomia using multiple, independent phylogenomic datasets. As expected, Protostomia is always strongly supported, especially by longer and higher-quality genes. Support for Deuterostomia, however, is always equivocal and barely higher than support for paraphyletic alternatives. Conditions that cause tree reconstruction errors-inadequate models, short internal branches, faster evolving genes, and unequal branch lengths-coincide with support for monophyletic deuterostomes. Simulation experiments show that support for Deuterostomia could be explained by systematic error. The branch between bilaterian and deuterostome common ancestors is, at best, very short, supporting the idea that the bilaterian ancestor may have been deuterostome-like. Our findings have important implications for the understanding of early animal evolution.
Knowing phylogenetic relationships among species is fundamental for many studies in biology. An accurate phylogenetic tree underpins our understanding of the major transitions in evolution, such as the emergence of new body plans or metabolism, and is key to inferring the origin of new genes, detecting molecular adaptation, understanding morphological character evolution and reconstructing demographic changes in recently diverged species. Although data are ever more plentiful and powerful analysis methods are available, there remain many challenges to reliable tree building. Here, we discuss the major steps of phylogenetic analysis, including identification of orthologous genes or proteins, multiple sequence alignment, and choice of substitution models and inference methodologies. Understanding the different sources of errors and the strategies to mitigate them is essential for assembling an accurate tree of life. Understanding evolutionary relationships between species requires the generation of accurate phylogenetic trees. In this Review, Kapli, Yang and Telford discuss the principles, steps and computational tools for phylogenetic tree building. They describe the impact of burgeoning genomic datasets as well as the diverse sources of errors and how they can be mitigated.
The evolutionary relationships of two animal phyla, Ctenophora and Xenacoelomorpha, have proved highly contentious. Ctenophora have been proposed as the most distant relatives of all other animals (Ctenophora-first rather than the traditional Porifera-first). Xenacoelomorpha may be primitively simple relatives of all other bilaterally symmetrical animals (Nephrozoa) or simplified relatives of echinoderms and hemichordates (Xenambulacraria). In both cases, one of the alternative topologies must be a result of errors in tree reconstruction. Here, using empirical data and simulations, we show that the Ctenophora-first and Nephrozoa topologies (but not Porifera-first and Ambulacraria topologies) are strongly supported by analyses affected by systematic errors. Accommodating this finding suggests that empirical studies supporting Ctenophora-first and Nephrozoa trees are likely to be explained by systematic error. This would imply that the alternative Porifera-first and Xenambulacraria topologies, which are supported by analyses designed to minimize systematic error, are the most credible current alternatives.
Introductory paragraph The availability of complete sets of genes from many organisms makes it possible to identify genes unique to (or lost from) certain clades. This information is used to reconstruct phylogenetic trees; to identify genes involved in the evolution of clade specific novelties; and for phylostratigraphy - identifying ages of genes in a given species. These investigations rely on accurately predicted orthologs. Here we use simulation to produce sets of orthologs which experience no gains or losses. We show that errors in identifying orthologs increase with higher rates of evolution. We use the predicted sets of orthologs, with errors, to reconstruct phylogenetic trees; to count gains and losses; and for phylostratigraphy. Our simulated data, containing information only from errors in orthology prediction, closely recapitulate findings from empirical data. We suggest published downstream analyses must be informed to a large extent by errors in orthology prediction which mimic expected patterns of gene evolution.
An amendment to this paper has been published and can be accessed via the original article.
Background:The classification of hepatitis viruses still predominantly relies on ad hoc criteria, i.e., phenotypic traits and arbitrary genetic distance thresholds.Given the subjectivity of such practices coupled with the constant sequencing of samples and discovery of new strains, this manual approach to virus classification becomes cumbersome and impossible to generalize. Methods:Using two well-studied hepatitis virus datasets, HBV and HCV, we assess if computational methods for molecular species delimitation that are typically applied to barcoding biodiversity studies can also be successfully deployed for hepatitis virus classification.For comparison, we also used ABGD, a tool that in contrast to other distance methods attempts to automatically identify the barcoding gap using pairwise genetic distances for a set of aligned input sequences. Results -Discussion:We find that, the mPTP species delimitation tool identified even without adapting its default parameters taxonomic clusters that either correspond to the currently acknowledged genotypes or to known subdivision of genotypes (subtypes or subgenotypes).In the cases where the delimited cluster corresponded to subtype or subgenotype, there were previous concerns that their status may be underestimated.The clusters obtained from the ABGD analysis differed depending on the parameters used.However, under certain values the results were very similar to the taxonomy and mPTP which indicates the usefulness of distance based methods in virus taxonomy under appropriate parameter settings .The overlap of predicted clusters with taxonomically acknowledged genotypes implies that virus classification can be successfully automated.
BACKGROUND:The classification of hepatitis viruses still predominantly relies on ad hoc criteria, i.e., phenotypic traits and arbitrary genetic distance thresholds. Given the subjectivity of such practices coupled with the constant sequencing of samples and discovery of new strains, this manual approach to virus classification becomes cumbersome and impossible to generalize.METHODS:Using two well-studied hepatitis virus datasets, HBV and HCV, we assess if computational methods for molecular species delimitation that are typically applied to barcoding biodiversity studies can also be successfully deployed for hepatitis virus classification. For comparison, we also used ABGD, a tool that in contrast to other distance methods attempts to automatically identify the barcoding gap using pairwise genetic distances for a set of aligned input sequences.RESULTS—DISCUSSION:We found that the mPTP species delimitation tool identified even without adapting its default parameters taxonomic clusters that either correspond to the currently acknowledged genotypes or to known subdivision of genotypes (subtypes or subgenotypes). In the cases where the delimited cluster corresponded to subtype or subgenotype, there were previous concerns that their status may be underestimated. The clusters obtained from the ABGD analysis differed depending on the parameters used. However, under certain values the results were very similar to the taxonomy and mPTP which indicates the usefulness of distance based methods in virus taxonomy under appropriate parameter settings. The overlap of predicted clusters with taxonomically acknowledged genotypes implies that virus classification can be successfully automated.
Polyneoptera represents one of the major lineages of winged insects, comprising around 40,000 extant species in 10 traditional orders, including grasshoppers, roaches, and stoneflies. Many important aspects of polyneopteran evolution, such as their phylogenetic relationships, changes in their external appearance, their habitat preferences, and social behavior, are unresolved and are a major enigma in entomology. These ambiguities also have direct consequences for our understanding of the evolution of winged insects in general; for example, with respect to the ancestral habitats of adults and juveniles. We addressed these issues with a large-scale phylogenomic analysis and used the reconstructed phylogenetic relationships to trace the evolution of 112 characters associated with the external appearance and the lifestyle of winged insects. Our inferences suggest that the last common ancestors of Polyneoptera and of the winged insects were terrestrial throughout their lives, implying that wings did not evolve in an aquatic environment. The appearance of the first polyneopteran insect was mainly characterized by ancestral traits such as long segmented abdominal appendages and biting mouthparts held below the head capsule. This ancestor lived in association with the ground, which led to various specializations including hardened forewings and unique tarsal attachment structures. However, within Polyneoptera, several groups switched separately to a life on plants. In contrast to a previous hypothesis, we found that social behavior was not part of the polyneopteran ground plan. In other traits, such as the biting mouthparts, Polyneoptera shows a high degree of evolutionary conservatism unique among the major lineages of winged insects.
Mesalina are small lacertid lizards occurring in the Saharo‐Sindian deserts from North Africa to the east of the Iranian plateau. Earlier phylogenetic studies indicated that there are several species complexes within the genus and that thorough taxonomic revisions are needed. In this study, we aim at resolving the phylogeny and taxonomy of the M. brevirostris species complex distributed from the Middle East to the Arabian/Persian Gulf region and Pakistan. We sequenced three mitochondrial and three nuclear gene fragments, and in combination with species delimitation and species‐tree estimation, we infer a time‐calibrated phylogeny of the complex. The results of the genetic analyses support the presence of four clearly delimited species in the complex that diverged approximately between the middle Pliocene and the Pliocene/Pleistocene boundary. Species distribution models of the four species show that the areas of suitable habitat are geographically well delineated and nearly allopatric, and that most of the species have rather divergent environmental niches. Morphological characters also confirm the differences between the species, although sometimes minute. As a result of all these lines of evidence, we revise the taxonomy of the Mesalina brevirostris species complex. We designate a lectotype for Mesalina brevirostris Blanford, 1874; resurrect the available name Eremias bernoullii Schenkel, 1901 from the synonymy of M. brevirostris; elevate M. brevirostris microlepis (Angel, 1936) to species status; and describe Mesalina saudiarabica, a new species from Saudi Arabia.