ABSTRACT The Leray-XT primer pair has been widely used to amplify the mitochondrial cytochrome c oxidase subunit I (COI) gene from animals. In some marine metabarcoding studies, protists have also been amplified and sequenced using these primers. Here, we ask if the Leray-XT COI primer pair is suitable for observing ciliates and radiolarians, which are numerically and ecologically important components of marine protistan communities. We show that while there are sufficient COI reference sequences for ciliates in NCBI for taxonomic assignments, there are currently only two COI reference sequences for radiolarians. Using in-silico analyses, we additionally show that while the reverse primer Leray-XT primer can bind and potentially amplify both ciliates and radiolarians, the forward primer cannot bind to either taxon. These results show that the Leray-XT primer pair is not suitable for observing ciliates and radiolarians, although it may be useful for observing other marine protistan taxa.
Short-branch Microsporidia were previously shown to form a basal grade within the expanded Microsporidia clade and to branch near the classical, long-branch Microsporidia. Although they share simpler versions of some morphological characteristics, they do not show accelerated evolutionary rates, making them ideal candidates to study the evolutionary trajectories that have led to long-branch microsporidian unique characteristics. However, most sequences assigned to the short-branch Microsporidia are undescribed, novel environmental lineages for which the identification requires knowledge of where they can be found. To direct future isolation, we used the EukBank database of the global UniEuk initiative that contains the majority of the publicly available environmental V4 SSU rRNA gene sequences of protists. The curated OTU table and corresponding metadata were used to evaluate the occurrence of short-branch Microsporidia across freshwater, hypersaline, marine benthic, marine pelagic, and terrestrial environments. Presence-absence analyses infer that short-branch Microsporidia are most abundant in freshwater and terrestrial environments, and alpha- and beta-diversity measures indicate that focusing our sampling effort on these two environments would cover a large part of their overall diversity. These results can be used to coordinate future isolation and sampling campaigns to better understand the enigmatic evolution of microsporidians' unique characteristics.
Marine Stramenopiles (MAST) were first described two decades ago through ribosomal RNA gene (rRNA gene) sequences from marine surveys of microbial eukaryotes. MAST comprise several independent lineages at the base of the Stramenopiles. Despite their prevalence in the ocean, the majority of MAST diversity remains uncultured. Previous studies, mainly in marine environments, have explored MAST's cell morphology, distribution, trophic strategies, and genomics using culturing-independent methods. In comparison, less is known about their presence outside marine habitats. Here, we analyse the extensive EukBank dataset to assess the extent to which MAST can be considered marine protists. Additionally, by incorporating newly available rRNA gene sequences, we update Stramenopiles phylogeny, identifying three novel MAST lineages. Our results indicate that MAST are primarily marine with notable exceptions within MAST-2 and MAST-12, where certain subclades are prevalent in freshwater and soil habitats. In the marine water column, only a few MAST species, particularly within clades -1, -3, -4, and -7, dominate and exhibit clear latitudinal distribution patterns. Overall, the massive sequencing dataset analysed in our study confirms and partially expands the previously described diversity of MASTs groups and underscores the predominantly marine nature of most of these uncultured lineages.
Satellite remote sensing is a powerful tool to monitor the global dynamics of marine plankton. Previous research has focused on developing models to predict the size or taxonomic groups of phytoplankton. Here, we present an approach to identify community types from a global plankton network that includes phytoplankton and heterotrophic protists and to predict their biogeography using global satellite observations. Six plankton community types were identified from a co-occurrence network inferred using a novel rDNA 18 S V4 planetary-scale eukaryotic metabarcoding dataset. Machine learning techniques were then applied to construct a model that predicted these community types from satellite data. The model showed an overall 67% accuracy in the prediction of the community types. The prediction using 17 satellite-derived parameters showed better performance than that using only temperature and/or the concentration of chlorophyll a . The constructed model predicted the global spatiotemporal distribution of community types over 19 years. The predicted distributions exhibited strong seasonal changes in community types in the subarctic–subtropical boundary regions, which were consistent with previous field observations. The model also identified the long-term trends in the distribution of community types, which suggested responses to ocean warming.
In this study, we used a predator-enabled metagenomics strategy to sample the virome of a remote and difficult-to-access densely forested African tropical region. Specifically, we focused our study on the use of army ants of the genus Dorylus that are obligate collective foragers and group predators that attack and overwhelm a broad array of animal prey. Using 209 army ant samples collected from 29 colonies and the virion-associated nucleic acid-based metagenomics approach, we showed that a broad diversity of bacterial, plant, invertebrate and vertebrate viral sequences were accumulated by army ants: including sequences from 157 different viral genera in 56 viral families. This suggests that using predators and scavengers such as army ants to sample broad swathes of tropical forest viromes can shed light on the composition and the structure of viral populations of these complex and inaccessible ecosystems.
Taxonomic assignment of operational taxonomic units (OTUs) is an important bioinformatics step in analyzing environmental sequencing data. Pairwise alignment and phylogenetic-placement methods represent two alternative approaches to taxonomic assignments, but their results can differ. Here we used available colpodean ciliate OTUs from forest soils to compare the taxonomic assignments of VSEARCH (which performs pairwise alignments) and EPA-ng (which performs phylogenetic placements). We showed that when there are differences in taxonomic assignments between pairwise alignments and phylogenetic placements at the subtaxon level, there is a low pairwise similarity of the OTUs to the reference database. We then showcase how the output of EPA-ng can be further evaluated using GAPPA to assess the taxonomic assignments when there exist multiple equally likely placements of an OTU, by taking into account the sum over the likelihood weights of the OTU placements within a subtaxon, and the branch distances between equally likely placement locations. We also inferred the evolutionary and ecological characteristics of the colpodean OTUs using their placements within subtaxa. This study demonstrates how to fully analyze the output of EPA-ng, by using GAPPA in conjunction with knowledge of the taxonomic diversity of the clade of interest.
The Arctic Ocean (AO) is being rapidly transformed by global warming, but its biodiversity remains understudied for many planktonic organisms, in particular for unicellular eukaryotes that play pivotal roles in marine food webs and biogeochemical cycles. The aim of this study was to characterize the biogeographic ranges of species that comprise the contemporary pool of unicellular eukaryotes in the AO as a first step toward understanding mechanisms that structure these communities and identifying potential target species for monitoring. Leveraging the Tara Oceans DNA metabarcoding data, we mapped the global distributions of operational taxonomic units (OTUs) found on Arctic shelves into five biogeographic categories, identified biogeographic indicators, and inferred the degree to which AO communities of unicellular eukaryotes share members with assemblages from lower latitudes. Arctic/Polar indicator OTUs, as well as some globally ubiquitous OTUs, dominated the detection and abundance of DNA reads in the Arctic samples. OTUs detected only in Arctic samples (Arctic-exclusives) showed restricted distribution with relatively low abundances, accounting for 10–16% of the total Arctic OTU pool. OTUs with high abundances in tropical and/or temperate latitudes (non-Polar indicators) were also found in the AO but mainly at its periphery. We observed a large change in community taxonomic composition across the Atlantic-Arctic continuum, supporting the idea that advection and environmental filtering are important processes that shape plankton assemblages in the AO. Altogether, this study highlights the connectivity between the AO and other oceans, and provides a framework for monitoring and assessing future changes in this vulnerable ecosystem.
In every liter of seawater there are between 10 and 100 billion life forms, mostly invisible, called marine plankton or marine microbiome, which form the largest and most dynamic ecosystem on our planet, at the heart of global ecological and economic processes. While physical and chemical parameters of planktonic ecosystems are fairly well measured and modeled at the planetary scale, biological data are still scarce due to the extreme cost and relative inflexibility of the classical vessels and instruments used to explore marine biodiversity. Here we introduce ‘Plankton Planet’, an initiative whose goal is to engage the curiosity and creativity of researchers, makers, and mariners to ( i ) co-develop a new generation of cost-effective (frugal) universal scientific instrumentation to measure the genetic and morphological diversity of marine microbiomes in context, ( ii ) organize their systematic deployment through coastal or open ocean communities of sea-users/farers, to generate uniform plankton data across global and long-term spatio-temporal scales, and ( iii ) setup tools to flow the data without embargo into public and explorable databases. As proof-of-concept, we show how 20 crews of sailors were able to sample plankton biomass from the world surface ocean in a single year, generating the first seatizen-based, planetary dataset of marine plankton biodiversity based on DNA barcodes. The quality of this dataset is comparable to that generated by Tara Oceans and is not biased by the multiplication of samplers. The data unveil significant genetic novelty and can be used to explore the taxonomic and ecological diversity of plankton at both regional and global scales. This pilot project paves the way for construction of a miniaturized, modular, evolvable, affordable and open-source citizen field-platform that will allow systematic assessment of the eco/morpho/genetic variation of aquatic ecosystems and microbiomes across the dimensions of the Earth system.
Major seasonal community reorganizations and associated biomass variations are landmarks of plankton ecology. However, the processes determining marine species and community turnover rates have not been fully elucidated so far. Here, we analyse patterns of planktonic protist community succession in temperate latitudes, based on quantitative taxonomic data from both microscopy counts and ribosomal DNA metabarcoding from plankton samples collected biweekly over 8 years (2009-2016) at the SOMLIT-Astan station (Roscoff, Western English Channel). Considering the temporal structure of community dynamics (creating temporal correlation), we elucidated the recurrent seasonal pattern of the dominant species and OTUs (rDNA-derived taxa) that drive annual plankton successions. The use of morphological and molecular analyses in combination allowed us to assess absolute species abundance while improving taxonomic resolution, and revealed a greater diversity. Overall, our results underpinned a protist community characterised by a seasonal structure, which is supported by the dominant OTUs. We detected that some were partly benthic as a result of the intense tidal mixing typical of the French coasts in the English Channel. While the occurrence of these microorganisms is driven by the physical and biogeochemical conditions of the environment, internal community processes, such as the complex network of biotic interactions, also play a key role in shaping protist communities.
Metabarcoding of microbial eukaryotes (collectively known as protists ) has developed tremendously in the last decade, almost uniquely relying on the 18S rRNA gene. As microbial eukaryotes are extremely diverse, many primers and primer pairs have been developed. To cover a relevant and representative fraction of the protist community in a given study system, a wise primer choice is needed as no primer pair can target all protists equally well. As such, a smart primer choice is very difficult even for experts and there are very few on-line resources available to list existing primers. We built a database listing 179 primers and 76 primer pairs that have been used for eukaryotic 18S rRNA metabarcoding. In silico performance of primer pairs was tested against two sequence databases: PR 2 for eukaryotes and a subset of Silva for prokaryotes. This allowed to determine the taxonomic specificity of primer pairs, the location of mismatches as well as amplicon size. We developed a R-based web application that allows to browse the database, visualize the taxonomic distribution of the amplified sequences with the number of mismatches, and to test any user-defined primer set ( https://app.pr2-primers.org ). This tool will provide the basis for guided primer choices that will help a wide range of ecologists to implement protists as part of their investigations.
MOTIVATION:Previously we presented swarm, an open-source amplicon clustering programme that produces fine-scale molecular operational taxonomic units (OTUs) that are free of arbitrary global clustering thresholds. Here, we present swarm v3 to address issues of contemporary datasets that are growing towards tera-byte sizes. RESULTS:When compared with previous swarm versions, swarm v3 has modernized C++ source code, reduced memory footprint by up to 50%, optimized CPU-usage and multithreading (more than 7 times faster with default parameters), and it has been extensively tested for its robustness and logic. AVAILABILITY AND IMPLEMENTATION:Source code and binaries are available at https://github.com/torognes/swarm. SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
One promising avenue for reconciling the goals of crop production and ecosystem preservation consists in the manipulation of beneficial biotic interactions, such as between insects and microbes. Insect gut microbiota can affect host fitness by contributing to development, host immunity, nutrition, or behavior. However, the determinants of gut microbiota composition and structure, including host phylogeny and host ecology, remain poorly known. Here, we used a well-studied community of eight sympatric fruit fly species to test the contributions of fly phylogeny, fly specialization, and fly sampling environment on the composition and structure of bacterial gut microbiota. Comprising both specialists and generalists, these species belong to five genera from to two tribes of the Tephritidae family. For each fly species, one field and one laboratory samples were studied. Bacterial inventories to the genus level were produced using 16S metabarcoding with the Oxford Nanopore Technology. Sample bacterial compositions were analyzed with recent network-based clustering techniques. Whereas gut microbiota were dominated by the Enterobacteriaceae family in all samples, microbial profiles varied across samples, mainly in relation to fly identity and sampling environment. Alpha diversity varied across samples and was higher in the Dacinae tribe than in the Ceratitinae tribe. Network analyses allowed grouping samples according to their microbial profiles. The resulting groups were very congruent with fly phylogeny, with a significant modulation of sampling environment, and with a very low impact of fly specialization. Such a strong imprint of host phylogeny in sympatric fly species, some of which share much of their host plants, suggests important control of fruit flies on their gut microbiota through vertical transmission and/or intense filtering of environmental bacteria.
EukRibo is a manually curated, public reference database of small-subunit ribosomal RNA gene (18S rDNA) sequences of eukaryotes, specifically aimed at taxonomic annotation of high-throughput metabarcoding datasets. Unlike other reference databases of ribosomal genes, it is not meant to exhaustively capture all publicly available 18S rDNA sequences from the INSDC repositories, but to represent a subset of highly trustable sequences covering the whole known diversity of eukaryotes. EukRibo strives to include only sequences with verified, up-to-date taxonomic identifications, with a strong focus on protists, and relatively low genetic redundancy, to keep the database compact yet comprehensive. Environmental clone sequences representing previously identified novel diversity are accepted as reference sequences only if they have a precise lineage designation, useful for taxonomic annotation. EukRibo is part of a suite of public resources generated by the UniEuk project, which all follow a common taxonomic framework for maximal interoperability. The high level of taxonomic accuracy of EukRibo allows higher confidence in the taxonomic annotation of environmental metabarcodes, and should facilitate identification of new eukaryotic diversity at various taxonomic levels. The database is currently in version 2, and all versions are permanently stored and made available via the FAIR open platform Zenodo. It is our hope that EukRibo will help ongoing curation efforts of other 18S rDNA reference databases, and we welcome suggestions of corrections and new features to be included in subsequent versions.
Taxonomic assignment of OTUs is an important bioinformatics step in analyzing environmental sequencing data. Pairwise-alignment and phylogenetic-placement methods represent two alternative approaches to taxonomic assignments, but their results can differ. Here we used available colpodean ciliate OTUs from forest soils to compare the taxonomic assignments of VSEARCH (which performs pairwise alignments) and EPA-ng (which performs phylogenetic placements). We showed that when there are differences in taxonomic assignments between pairwise alignments and phylogenetic placements at the subtaxon level, there is a low pairwise similarity of the OTUs to the reference database. We then showcase how the output of EPA-ng can be further evaluated using GAPPA to assess the taxonomic assignments when there exist multiple equally likely placements of an OTU, by taking into account the sum over the likelihood weights of the OUT placements within a subtaxon, and the branch distances between equally likely placement locations. We also inferred evolutionary and ecological characteristics of the colpodean OTUs using their placements within subtaxa. This study demonstrates how to fully analyse the output of EPA-ng, by using GAPPA in conjunction with knowledge of the taxonomic diversity of the clade of interest.
This repository contains the rDNA 18S V9 OTU table, its rarefied version, and the related contextual data from Plankton Planet Pilot Project. In the file P2_TO_18SV9_otu_table.tsv.gz, each OTU, one per row, is described by the following fields: amplicon = identifier of the representative (most abundant) sequence of the swarm; total = total number of reads in the entire dataset; cloud = number of unique sequences constituting the OTU; length = length of the representative sequence; spread = number of samples in which the OTU has been found; quality = minimum expected error observed for the representative sequence, divided by sequence length; sequence = nucleic acid sequence of the representative sequence; identity = percentage of identity of the representative sequence to the closest reference sequence from PR2_V9 (https://doi.org/10.5281/zenodo.3768951); references = best hit reference sequence(s); taxonomy = taxonomic path assigned to the representative sequence; taxogroup = high-taxonomic level assignation of the representative barcode; chloroplast = yes: presence of permanent chloroplast / no: absence of permanent chloroplast / NA: undetermined; symb_small = parasite: the species is a parasite / commensal: the species is a commensal / mutualist: the species is a mutualist symbiont, most often a microalgal taxa involved in photosymbiosis / no: the species is not involved in a symbiosis as small partner / NA: undetermined; symbiont_host = photo: the host species relies on a mutualistic microalgal photosymbiont to survive (obligatory photosymbiosis) / photo_falc: same as photo, but facultative relationship / photo_klep: the host species maintains chloroplasts from microalgal prey(s) to survive / photo_klep_falc: same as photo_klep, but facultative / Nfix = the host species must interact with a mutualistic symbiont providing N2 fixation to survive / Nfix_falc = same as Nfix, but facultative / no: the species is not involved in any mutualistic symbioses; NA: undetermined; silicification = yes: the species has a silicified skeleton / no: it does not / NA: undetermined; calcification = yes: the species has a calcified skeleton / no: it does not / NA: undetermined; strontification = yes: the species has a skeleton made of strontium / no: it does not / NA: undetermined; PPXXX = number of reads in each of the 214 Plankton Planet samples; TARA_XXXXXXXXXX = number of reads in each of the 386 Tara Oceans samples. The file P2_TO_18SV9_otu_table_raref_313539.tsv.gz contains the same fields but with number of reads (total and per sample) obtained after random subsampling (313,539 reads per sample). In the file P2_TO_18SV9_context.tsv.gz, each sample is described by the following fields: sample = identifier of the sample; lower_size_fraction = lower limit of the size fraction in µm; upper_size_fraction = lower limit of the size fraction in µm; event_date = date (year-month-day); event_latitude = geographic position (latitude in DD); event_longitude = geographic position (longitude in DD); depth = depth in meters; temperature = sea water temperature in °C
Satellite remote sensing from space is a powerful way to monitor the global dynamics of marine plankton. Previous research has focused on developing models to predict the size or taxonomic groups of phytoplankton. Here we present an approach to identify representative communities from a global plankton network that included both zooplankton and phytoplankton and using global satellite observations to predict their biogeography. Six representative plankton communities were identified from a global co-occurrence network inferred using a novel rDNA 18S V4 planetary-scale eukaryotic metabarcoding dataset. Machine learning techniques were then applied to train a model that predicted these representative communities from satellite data. The model showed an overall 67% accuracy in the prediction of the representative communities. The prediction based on 17 satellite-derived parameters showed better performance than based only on temperature and/or the concentration of chlorophyll a . The trained model allowed to predict the global spatiotemporal distribution of communities over 19-years. Our model exhibited strong seasonal changes in the community compositions in the subarctic-subtropical boundary regions, which were consistent with previous field observations. This network-oriented approach can easily be extended to more comprehensive models including prokaryotes as well as viruses.
Nucleic acid based studies of marine biodiversity often focus on Kingdom-level diversity. Such approaches often largely miss diversity of less studied groups, likely to harbour many unknown lineages which are likely playing significant ecological roles. Among these elusive groups are Endomyxa (Rhizaria), a ubiquitous, but understudied lineage comprising parasites (e.g. Ascetosporea, Phytomyxea), free-living amoebae (e.g. Vampyrellida, Gromia, Filoreta ), flagellates (e.g. Tremula, Aquavolon ), and unknown environmental lineages. Using Endomyxa-biased primers targeting the hypervariable region V4 of the 18S rRNA gene, we explored the diversity of Endomyxa in marine samples from European coastal sites and compared it to that found in pan-eukaryote V4 libraries of the same samples. In total 458 endomyxan OTUs were identified, of which 38% were only detected by the specific primers. Most are distinct from published sequences, and include novel diversity within known clades, and putative novel lineages. The data revealed variations in endomyxan assemblages related to habitat (benthic vs. pelagic, sampling site), mode of nutrition (parasitic vs. free-living) and nucleic acid type (DNA vs. RNA). Overall, the vast majority of endomyxan diversity occurs in sediments (including Vampyrellida, Reticulosida, and the environmental “Novel Clade 12”) where they form diverse and active communities including many uncharacterised lineages.
Background: Efficiently managing large, heterogeneous data in a structured yet flexible way is a challenge to research laboratories working with genomic data. Specifically regarding both shotgun- and metabarcoding-based metagenomics, while online reference databases and user-friendly tools exist for running various types of analyses (e.g., Qiime, Mothur, Megan, IMG/VR, Anvi'o, Qiita, MetaVir), scientists lack comprehensive software for easily building scalable, searchable, online data repositories on which they can rely during their ongoing research. Results: metaXplor is a scalable, distributable, fully web-interfaced application for managing, sharing, and exploring metagenomic data. Being based on a flexible NoSQL data model, it has few constraints regarding dataset contents and thus proves useful for handling outputs from both shotgun and metabarcoding techniques. By supporting incremental data feeding and providing means to combine filters on all imported fields, it allows for exhaustive content browsing, as well as rapid narrowing to find specific records. The application also features various interactive data visualization tools, ways to query contents by BLASTing external sequences, and an integrated pipeline to enrich assignments with phylogenetic placements. The project home page provides the URL of a live instance allowing users to test the system on public data. Conclusion: metaXplor allows efficient management and exploration of metagenomic data. Its availability as a set of Docker containers, making it easy to deploy on academic servers, on the cloud, or even on personal computers, will facilitate its adoption.
Two 915-MHz microwave treatments (2 kW-8 min and 4 kW-4 min) were applied to soil and their effects were monitored just after treatments (T0) and 26 days (T26) in soil microcosms. Densities of culturable bacteria, fluorescent pseudomonads and nematodes, hydrolysis activity and soil DNA content declined by over 50% immediately after both microwave treatments (T0), excluding the total fungal 18S rRNA (−13 or −17%) and bacterial 16S rRNA copies (non-significant). A rapid shift in bacterial community composition occurred from T0 towards a large increase in the relative abundance of Firmicutes (+1650%) and a concomitant decrease in various phyla (e.g. Acidobacteria, Actinobacteria, Bacteroidetes, Proteobacteria) from −85 to −61%. At T26 and for both treatments, fluorescein diacetate hydrolysis, density of culturable bacteria, 18S rRNA gene numbers, Simpson diversity, relative abundances of Bacteroidetes and Proteobacteria regained levels similar to controls. Alpha, Delta and Gamma classes Proteobacteria also end up reaching a similar level. In contrast, fluorescent pseudomonad density, nematode diversity and abundance, soil DNA content and relative abundances of some phylum (Acidobacteria, Actinobacteria, Chloroflexi) did not increase in such proportions. Most of the soil biological properties have not been permanently impacted. Nevertheless, the recovery kinetics highly differed according properties, with resilience indices at T26 varying from −87 to +99%.