Finding the correct position of new sequences within an established phylogenetic tree is an increasingly relevant problem in evolutionary bioinformatics and metagenomics. Recently, alignment-free approaches for this task have been proposed. One such approach is based on the concept of phylogenetically-informative k-mers or phylo- k-mers for short. In practice, phylo- k-mers are inferred from a set of related reference sequences and are equipped with scores expressing the probability of their appearance in different locations within the input reference phylogeny. Computing phylo- k-mers, however, represents a computational bottleneck to their applicability in real-world problems such as the phylogenetic analysis of metabarcoding reads and the detection of novel recombinant viruses. Here we consider the problem of phylo- k-mer computation: how can we efficiently find all k-mers whose probability lies above a given threshold for a given tree node? We describe and analyze algorithms for this problem, relying on branch-and-bound and divide-and-conquer techniques. We exploit the redundancy of adjacent windows of the alignment to save on computation. Besides computational complexity analyses, we provide an empirical evaluation of the relative performance of their implementations on simulated and real-world data. The divide-and-conquer algorithms are found to surpass the branch-and-bound approach, especially when many phylo- k-mers are found.
MOTIVATION:Phylogenetic placement enables phylogenetic analysis of massive collections of newly sequenced DNA, when de novo tree inference is too unreliable or inefficient. Assuming that a high-quality reference tree is available, the idea is to seek the correct placement of the new sequences in that tree. Recently, alignment-free approaches to phylogenetic placement have emerged, both to circumvent the need to align the new sequences and to avoid the calculations that typically follow the alignment step. A promising approach is based on the inference of k-mers that can be potentially related to the reference sequences, also called phylo-k-mers. However, its usage is limited by the time and memory-consuming stage of reference data preprocessing and the large numbers of k-mers to consider.RESULTS:We suggest a filtering method for selecting informative phylo-k-mers based on mutual information, which can significantly improve the efficiency of placement, at the cost of a small loss in placement accuracy. This method is implemented in IPK, a new tool for computing phylo-k-mers that significantly outperforms the software previously available. We also present EPIK, a new software for phylogenetic placement, supporting filtered phylo-k-mer databases. Our experiments on real-world data show that EPIK is the fastest phylogenetic placement tool available, when placing hundreds of thousands and millions of queries while still providing accurate placements.AVAILABILITY AND IMPLEMENTATION:IPK and EPIK are freely available at https://github.com/phylo42/IPK and https://github.com/phylo42/EPIK. Both are implemented in C++ and Python and supported on Linux and MacOS.
To help address the underrepresentation of arthropods and Asian biodiversity from climate-change assessments, we carried out year-long, weekly sampling campaigns with Malaise traps at different elevations and latitudes in Gaoligongshan National Park in southwestern China. From these 623 samples, we barcoded 10,524 beetles and compared scenarios of climate-change-induced biodiversity loss, by designating seasonal, elevational, and latitudinal subsets of beetles as communities that plausibly could go extinct as a group, which we call "loss sets". The availability of a published mitochondrial-genome-based phylogeny of the Coleoptera allowed us to compare the loss of species diversity with and without accounting for phylogenetic relatedness. We hypothesised that phylogenetic relatedness would mitigate extinction, since the extinction of any loss set would result in the disappearance of all its species but only part of its evolutionary history, which is still extant in the remaining loss sets. We found different patterns of community clustering by season and latitude, depending on whether phylogenetic information was incorporated. However, accounting for phylogeny only slightly mitigated the amount of biodiversity loss under climate change scenarios, against our expectations: there is no phylogenetic "escape clause" for biodiversity conservation. We achieve the same results whether phylogenetic information was derived from the mitogenome phylogeny or from a de novo barcode-gene tree. We encourage interested researchers to use this data set to study lineage-specific community assembly patterns in conjunction with life-history traits and environmental covariates.
Accurate determination of the evolutionary relationships between genes is a foundational challenge in biology. Homology-evolutionary relatedness-is in many cases readily determined based on sequence similarity analysis. By contrast, whether or not two genes directly descended from a common ancestor by a speciation event (orthologs) or duplication event (paralogs) is more challenging, yet provides critical information on the history of a gene. Since 2009, this task has been the focus of the Quest for Orthologs (QFO) Consortium. The sixth QFO meeting took place in Okazaki, Japan in conjunction with the 67th National Institute for Basic Biology conference. Here, we report recent advances, applications, and oncoming challenges that were discussed during the conference. Steady progress has been made toward standardization and scalability of new and existing tools. A feature of the conference was the presentation of a panel of accessible tools for phylogenetic profiling and several developments to bring orthology beyond the gene unit-from domains to networks. This meeting brought into light several challenges to come: leveraging orthology computations to get the most of the incoming avalanche of genomic data, integrating orthology from domain to biological network levels, building better gene models, and adapting orthology approaches to the broad evolutionary and genomic diversity recognized in different forms of life and viruses.
MOTIVATION:Novel recombinant viruses may have important medical and evolutionary significance, as they sometimes display new traits not present in the parental strains. This is particularly concerning when the new viruses combine fragments coming from phylogenetically distinct viral types. Here, we consider the task of screening large collections of sequences for such novel recombinants. A number of methods already exist for this task. However, these methods rely on complex models and heavy computations that are not always practical for a quick scan of a large number of sequences. RESULTS:We have developed SHERPAS, a new program to detect novel recombinants and provide a first estimate of their parental composition. Our approach is based on the precomputation of a large database of 'phylogenetically-informed k-mers', an idea recently introduced in the context of phylogenetic placement in metagenomics. Our experiments show that SHERPAS is hundreds to thousands of times faster than existing software, and enables the analysis of thousands of whole genomes, or long-sequencing reads, within minutes or seconds, and with limited loss of accuracy. AVAILABILITY AND IMPLEMENTATION:The source code is freely available for download at https://github.com/phylo42/sherpas. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Abstract Multiple sequence alignment is an essential preliminary step for a large number of different algorithms in bioinformatics. With the decrease of the sequencing cost, the need to process more and more data increases, making alignment-based approaches computationally expensive in practice. This led to the emergence of many alignment-free algorithms and the widespread adoption of k-mer based approaches. We describe phylogenetically-informed k-mers, or phylo k-mers , a concept re-cently introduced in the context of phylogenetic placement in metagenomics [1], and successfully applied for viral recombination detection [2]. A phylo k-mer is a k-mer that is present with a non-negligible probability in unknown relatives of the sequences contained in an alignment. While the calculation of these probabilities is computationally heavy and requires the reference alignment as an input, it has to be done only once per alignment. Once calculated, phylo k-mers can be applied in alignment-free algorithms that require a massive input of new query sequences: in RAPPAS [1] for phylogenetic placement, and SHERPAS [2] for viral recombination detection. We discuss methods of calculation, or construction of phylo k-mers, and present xpas , the phylo k-mer construction library. It allows for fast and memory-efficient construction of databases of phylo k-mers. Those databases are used by SHERPAS and the new version of RAPPAS, which is currently under development.
High-throughput DNA methods hold great promise for phylogenetic analysis of lineages that are difficult to study with conventional molecular and morphological approaches. The mites (Acari), and in particular the highly diverse soildwelling lineages, are among the least known branches of the metazoan Tree-of-Life. We extracted numerous minute mites from soils in an area of mixed forest and grassland in southern Iberia. Selected specimens representing the full morphological diversity were shotgun sequenced in bulk, followed by genome assembly of short reads from the mixture, which produced >100 mitochondrial genomes representing diverse acarine lineages. Phylogenetic analyses in combination with taxonomically limited mitogenomes available publicly resulted in plausible trees defining basal relationships of the Acari. Several critical nodes were supported by ancestral-state reconstructions of mitochondrial gene rearrangements. Molecular calibration placed the minimum age for the common ancestor of the superorder Acariformes, which includes most soil-dwelling mites, to the Cambrian-Ordovician (likely within 455-552 Ma), whereas the origin of the superorder Parasitiformes was placed later in the Carboniferous-Permian. Most family-level taxa within the Acariformes were dated to the Jurassic and Triassic. The ancient origin of Acariformes and the early diversification of major extant lineages linked to the soil are consistent with a pioneering role for mites in building the earliest terrestrial ecosystems.
MOTIVATION:Phylogenetic placement (PP) is a process of taxonomic identification for which several tools are now available. However, it remains difficult to assess which tool is more adapted to particular genomic data or a particular reference taxonomy. We developed Placement Evaluation WOrkflows (PEWO), the first benchmarking tool dedicated to PP assessment. Its automated workflows can evaluate PP at many levels, from parameter optimization for a particular tool, to the selection of the most appropriate genetic marker when PP-based species identifications are targeted. Our goal is that PEWO will become a community effort and a standard support for future developments and applications of PP.AVAILABILITY AND IMPLEMENTATION:https://github.com/phylo42/PEWO.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
MOTIVATION:Taxonomic classification is at the core of environmental DNA analysis. When a phylogenetic tree can be built as a prior hypothesis to such classification, phylogenetic placement (PP) provides the most informative type of classification because each query sequence is assigned to its putative origin in the tree. This is useful whenever precision is sought (e.g. in diagnostics). However, likelihood-based PP algorithms struggle to scale with the ever-increasing throughput of DNA sequencing.RESULTS:We have developed RAPPAS (Rapid Alignment-free Phylogenetic Placement via Ancestral Sequences) which uses an alignment-free approach, removing the hurdle of query sequence alignment as a preliminary step to PP. Our approach relies on the precomputation of a database of k-mers that may be present with non-negligible probability in relatives of the reference sequences. The placement is performed by inspecting the stored phylogenetic origins of the k-mers in the query, and their probabilities. The database can be reused for the analysis of several different metagenomes. Experiments show that the first implementation of RAPPAS is already faster than competing likelihood-based PP algorithms, while keeping similar accuracy for short reads. RAPPAS scales PP for the era of routine metagenomic diagnostics.AVAILABILITY AND IMPLEMENTATION:Program and sources freely available for download at https://github.com/blinard-BIOINFO/RAPPAS.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Gene families evolve by the processes of speciation (creating orthologs), gene duplication (paralogs), and horizontal gene transfer (xenologs), in addition to sequence divergence and gene loss. Orthologs in particular play an essential role in comparative genomics and phylogenomic analyses. With the continued sequencing of organisms across the tree of life, the data are available to reconstruct the unique evolutionary histories of tens of thousands of gene families. Accurate reconstruction of these histories, however, is a challenging computational problem, and the focus of the Quest for Orthologs Consortium. We review the recent advances and outstanding challenges in this field, as revealed at a symposium and meeting held at the University of Southern California in 2017. Key advances have been made both at the level of orthology algorithm development and with respect to coordination across the community of algorithm developers and orthology end-users. Applications spanned a broad range, including gene function prediction, phylostratigraphy, genome evolution, and phylogenomics. The meetings highlighted the increasing use of meta-analyses integrating results from multiple different algorithms, and discussed ongoing challenges in orthology inference as well as the next steps toward improvement and integration of orthology resources.
High-throughput DNA methods hold great promise for the study of the hyperdiverse arthropod fauna of the soil. We used the mitochondrial metagenomic approach to generate 39 mitochondrial genomes from adult and larval specimens of Coleoptera collected from soil samples. The mitogenomes correspond to species from the families Carabidae (6), Chrysomelidae (1), Curculionidae (9), Dermestidae (1), Elateridae (1), Latridiidae (1), Scarabaeidae (3), Silvanidae (1), Staphylinidae (12), and Tenebrionidae (4). All the mitogenomes followed the putative ancestral gene order for Coleoptera. We provide the first available mitogenome for 30 genera of Coleoptera, including endogean representatives of the genera Torneuma, Coiffaitiella, Otiorhynchus, Oligotyphlopsis, and Typhlocharis.
OrthoInspector is one of the leading software suites for orthology relations inference. In this paper, we describe a major redesign of the OrthoInspector online resource along with a significant increase in the number of species: 4753 organisms are now covered across the three domains of life, making OrthoInspector the most exhaustive orthology resource to date in terms of covered species (excluding viruses). The new website integrates original data exploration and visualization tools in an ergonomic interface. Distributions of protein orthologs are represented by heatmaps summarizing their evolutionary histories, and proteins with similar profiles can be directly accessed. Two novel tools have been implemented for comparative genomics: a phylogenetic profile search that can be used to find proteins with a specific presence-absence profile and investigate their functions and, inversely, a GO profiling tool aimed at deciphering evolutionary histories of molecular functions, processes or cell components. In addition to the re-designed website, the OrthoInspector resource now provides a REST interface for programmatic access. OrthoInspector 3.0 is available at http://lbgi.fr/orthoinspectorv3.
A phylogenetic tree at the species level is still far off for highly diverse insect orders, including the Coleoptera, but the taxonomic breadth of public sequence databases is growing. In addition, new types of data may contribute to increasing taxon coverage, such as metagenomic shotgun sequencing for assembly of mitogenomes from bulk specimen samples. The current study explores the application of these techniques for large-scale efforts to build the tree of Coleoptera. We used shotgun data from 17 different ecological and taxonomic datasets (5 unpublished) to assemble a total of 1942 mitogenome contigs of >3000 bp. These sequences were combined into a single dataset together with all mitochondrial data available at GenBank, in addition to nuclear markers widely used in molecular phylogenetics. The resulting matrix of nearly 16000 species with two or more loci produced trees (RAxML) showing overall congruence with the Linnaean taxonomy at hierarchical levels from suborders to genera. We tested the role of full-length mitogenomes in stabilizing the tree from GenBank data, as mitogenomes might link terminals with non-overlapping gene representation. However, the mitogenome data were only partly useful in this respect, presumably because of the purely automated approach to assembly and gene delimitation, but improvements in future may be possible by using multiple assemblers and manual curation. In conclusion, the combination of data mining and metagenomic sequencing of bulk samples provided the largest phylogenetic tree of Coleoptera to date, which represents a summary of existing phylogenetic knowledge and a defensible tree of great utility, in particular for studies at the intra-familial level, despite some shortcomings for resolving basal nodes.
We introduce a multi-factorial, multi-level approach to build and explore evolutionary scenarios of complex protein networks. EvoKEN combines a unique formalism for integrating multiple types of data associated with network molecular components and knowledge extraction techniques for detecting cohesive/anomalous evolutionary processes. We analyzed known human pathway maps and identified perturbations or specializations at the local topology level that reveal important evolutionary and functional aspects of these cellular systems.
Field‐collected specimens of invertebrates are regularly killed and preserved in ethanol, prior to DNA extraction from the specimens, while the ethanol fraction is usually discarded. However, DNA may be released from the specimens into the ethanol, which can potentially be exploited to study species diversity in the sample without the need for DNA extraction from tissue. We used shallow shotgun sequencing of the total DNA to characterize the preservative ethanol from two pools of insects (from a freshwater habitat and terrestrial habitat) to evaluate the efficiency of DNA transfer from the specimens to the ethanol. In parallel, the specimens themselves were subjected to bulk DNA extraction and shotgun sequencing, followed by assembly of mitochondrial genomes for 39 of 40 species in the two pools. Shotgun sequencing from the ethanol fraction and read‐matching to the mitogenomes detected ~40% of the arthropod species in the ethanol, confirming the transfer of DNA whose quantity was correlated to the biomass of specimens. The comparison of diversity profiles of microbiota in specimen and ethanol samples showed that ‘closed association’ (internal tissue) bacterial species tend to be more abundant in DNA extracted from the specimens, while ‘open association’ symbionts were enriched in the preservative fluid. The vomiting reflex of many insects also ensures that gut content is released into the ethanol, which provides easy access to DNA from prey items. Shotgun sequencing of DNA from preservative ethanol provides novel opportunities for characterizing the functional or ecological components of an ecosystem and their trophic interactions.
Achieving high accuracy in orthology inference is essential for many comparative, evolutionary and functional genomic analyses, yet the true evolutionary history of genes is generally unknown and orthologs are used for very different applications across phyla, requiring different precision-recall trade-offs. As a result, it is difficult to assess the performance of orthology inference methods. Here, we present a community effort to establish standards and an automated web-based service to facilitate orthology benchmarking. Using this service, we characterize 15 well-established inference methods and resources on a battery of 20 different benchmarks. Standardized benchmarking provides a way for users to identify the most effective methods for the problem at hand, sets a minimum requirement for new tools and resources, and guides the development of more accurate orthology inference methods.
Abstract The complete mitochondrial genome of the recently discovered beetle family Iberobaeniidae is described and compared with known coleopteran mitogenomes. The mitochondrial sequence was obtained by shotgun metagenomic sequencing using the Illumina Miseq technology and resulted in an average coverage of 130 × and a minimum coverage of 35×. The mitochondrial genome of Iberobaeniidae includes 13 protein-coding genes, 2 rRNAs, 22 tRNAs genes, and 1 putative control region, and showed a unique rearrangement of protein-coding genes. This is the first rearrangement affecting the relative position of protein-coding and ribosomal genes reported for the order Coleoptera.
MiSeq Illumina reads (TruSeq library, 250 bp paired-end or single-end, 500 cycles, insert size 600-900 bp, v2 chemistry) of the DNA gut content of five pooled Cycloneda sanguinea (Coleoptera: Coccinellidae), one Harmonia axyridis (Coleoptera: Coccinellidae), six pooled Hippodamia convergens (Coleoptera: Coccinellidae) and ten pooled Doru luteipes (Dermaptera: Forficulidae).
MiSeq Illumina reads (TruSeq library, 250 bp paired-end, 500 cycles, insert size 600-900 bp, v2 chemistry) of the DNA gut content of six pooled ladybird beetle Hippodamia convergens (Coleoptera: Coccinellidae).
Characterizing trophic networks is fundamental to many questions in ecology, but this typically requires painstaking efforts, especially to identify the diet of small generalist predators. Several attempts have been devoted to develop suitable molecular tools to determine predatory trophic interactions through gut content analysis, and the challenge has been to achieve simultaneously high taxonomic breadth and resolution. General and practical methods are still needed, preferably independent of PCR amplification of barcodes, to recover a broader range of interactions. Here we applied shotgun-sequencing of the DNA from arthropod predator gut contents, extracted from four common coccinellid and dermapteran predators co-occurring in an agroecosystem in Brazil. By matching unassembled reads against six DNA reference databases obtained from public databases and newly assembled mitogenomes, and filtering for high overlap length and identity, we identified prey and other foreign DNA in the predator guts. Good taxonomic breadth and resolution was achieved (93% of prey identified to species or genus), but with low recovery of matching reads. Two to nine trophic interactions were found for these predators, some of which were only inferred by the presence of parasitoids and components of the microbiome known to be associated with aphid prey. Intraguild predation was also found, including among closely related ladybird species. Uncertainty arises from the lack of comprehensive reference databases and reliance on low numbers of matching reads accentuating the risk of false positives. We discuss caveats and some future prospects that could improve the use of direct DNA shotgun-sequencing to characterize arthropod trophic networks.
Paul D. Thomas合作论文数Artificial Intelligence Center2