The rapid rate of virus discovery renders manual curation by taxonomy experts increasingly impractical, creating a need for reliable software that can reproducibly assign viral contigs to taxa at all fifteen ranks of the virus taxonomy. We led an open community challenge for the computational taxonomic classification of viruses and assembled a dataset of virus sequences combining expert-curated and metagenomic sequences. Seventeen teams contributed a total of thirty-four automated, fully reproducible classification pipelines. Most tools correctly assigned viruses belonging to established species, genera, or families, but viruses that are unclassified at those lower ranks remain challenging. This study provides datasets, open-source software, novel approaches, and recommendations to benchmark computational taxonomic classification of viruses, and support organizing the many viruses discovered in big omics data.
Although virus ecogenomics has expanded access to and understanding of the virosphere, existing classification tools lack taxonomic resolution and are unable to scale to modern discovery-based datasets or classify previously unknown sequence space. Here we develop vConTACT3-a machine learning-based tool that improves scalability and accuracy of virus taxonomy. By optimizing gene-sharing thresholds and leveraging adaptive, realm-specific cut-offs, vConTACT3 expands classification to both eukaryote and prokaryote viruses for four of the six officially recognized realms, and establishes accurate hierarchical taxonomy from genus to order. Specifically, vConTACT3 achieves >95% agreement with official taxonomy for 35,545 and 13,524 public prokaryotic and eukaryotic virus genomes, respectively, to surpass vConTACT2 across most realms, while still uniquely classifying previously uncharacterized taxa, and doing so even faster. vConTACT3 application provides taxonomy assignments for tens of thousands of unclassified taxa rapidly, automatically and systematically; evaluates virus sequence space to reveal support for fewer taxonomic ranks than currently available and identifies taxonomically challenging areas across the virosphere.
Viruses are key players in diverse ecosystems, but studying their impacts is technically and taxonomically challenging. Taxonomic complexities derive from undersampling, diverse DNA and RNA genomes with multiple evolutionary origins, and lack of a universal barcode gene. While virus ecogenomics has expanded access to and understanding of the virosphere, available classification tools poorly scale to modern discovery-based datasets, lack taxonomic resolution, and/or are unable to classify novel sequence space. Here we develop, benchmark, and release vConTACT3, a machine learning-based tool that improves scalability and accuracy, adds extensive user-requested features, expands classification to both eukaryote and prokaryote viruses for 4/6 officially recognized realms, and establishes accurate hierarchical taxonomy from genus to order. Application to 48,069 public virus genomes provided new taxonomy assignments for thousands of taxa, revealed support for fewer taxonomic ranks than currently available, and systematically identified taxonomically problematic areas across the virosphere. ### Competing Interest Statement The authors have declared no competing interest. U.S. National Science Foundation, https://ror.org/021nxhr62, DBI-2149506, DBI-2022070 Ohio Supercomputer Center, https://ror.org/01apna436 Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, EXC 2051 – Project-ID 390713860 European Research Council, https://ror.org/0472cxd90, Consolidator grant 865694 Alexander von Humboldt Foundation Biotechnology and Biological Sciences Research Council, BBS/E/F/000PR13631, BBS/E/F/000PR13633, BB/X011011/1, BBS/E/F/000PR13634, BBS/E/F/000PR13635, BBS/E/F/000PR13636
High-fidelity (HF) long-read sequencing enables accurate profiling of microorganisms and pathogens at single-molecule resolution. However, current Oxford Nanopore Technologies (ONT)-a revolutionary platform offering real-time, portable sequencing at relatively low instrumental cost-suffer from severe read-length bias, limited accuracy (often <Q20), and low throughput. Here, Circular- and Linear-Amplicon-Mediated Error Correction (CLAE) is introduced, a biochemical and computational approach that addresses these limitations by integrating hairpin ligation, pre-circling, single-stranded DNA linearization, and targeted nickase-based debranching. CLAE significantly enhances rolling-circle amplification (RCA) efficiency for long DNA templates, markedly improving Nanopore sequencing yield and accuracy. CLAE achieves Q30-level accuracy in up to 27% of RCA reads, throughput exceeding 800 Mb per 100 pores, and an N50 of ≈15 Kb (bacterial genome). Moreover, its bidirectional subreads and high throughput substantially boost accuracy without compromising read length. CLAE is validated by resolving SARS-CoV-2 quasi-species from community wastewater and recovering novel, full-length RNA virus genomes from marine samples. CLAE enables precise variant detection in complex samples and corrects short-read misassemblies, significantly broadening ONT's utility in metaviromics, epidemiology, and environmental surveillance. Thus, CLAE establishes a versatile, field-compatible platform for high-fidelity viral genome sequencing in targeted and agnostic contexts.
Global atmospheric methane concentrations are rapidly rising and becoming isotopically more depleted, implying an unresolved microbial contribution. Rising Arctic temperatures are variably altering soil methane cycling, causing consequential uncertainty in the atmospheric methane budget. We demonstrated in an Arctic wetland that below-ground microbiota and methane-cycling features parallelled above-ground plant communities. To upscale emissions, we applied machine learning to remote sensing data to identify habitats, which were assigned average emissions. To upscale dynamically, we incorporated climate data, remotely-sensed water table variation, and habitat classes into a temporally-resolved biogeochemical model, to predict methane flux and isotope dynamics. This accurately estimated more depleted 13C-methane than previously used for Arctic habitats in global source partitioning. Remote-sensing of these rapidly changing inaccessible landscapes can thus help constrain the role of the Arctic in ongoing changes in global methane emissions. ### Competing Interest Statement The authors have declared no competing interest.
Climate change thaws permafrost, which releases greenhouse gases partly from dormant microorganisms awakening and metabolizing organic matter. Though DNA viruses that infect these soil microbes have been studied, little is known on soil RNA viruses, which typically infect microeukaryotes. Here we identify and characterize 2,651 RNA viruses from a 4-year time series of bulk soil metatranscriptomes derived from the climatically fragile Stordalen Mire ecosystem - a long-studied permafrost peatland. RNA virus diversity was structured by habitat (palsa, bog, and fen), and these patterns correlated with pH and carbon dioxide and methane emissions. Further, host prediction, virus-encoded metabolite-transforming and information-processing functions suggested roles in ecosystem-scale carbon fluxes and contribute to greenhouse gases emissions. Together, these RNA virus ecogenomic data in permafrost provide essential baseline information for integration into predictive models to support hypothesis testing. ### Competing Interest Statement The authors have declared no competing interest.
Prokaryotic microbes have impacted marine biogeochemical cycles for billions of years. Viruses also impact these cycles, through lysis, horizontal gene transfer, and encoding and expressing genes that contribute to metabolic reprogramming of prokaryotic cells. While this impact is difficult to quantify in nature, we hypothesized that it can be examined by surveying virus-encoded auxiliary metabolic genes (AMGs) and assessing their ecological context. We systematically developed a global ocean AMG catalog by integrating previously described and newly identified AMGs and then placed this catalog into ecological and metabolic contexts relevant to ocean biogeochemistry. From 7.6 terabases of Tara Oceans paired prokaryote- and virus-enriched metagenomic sequence data, we increased known ocean virus populations to 579,904 (up 16
In environmental research, cross-disciplinary analyses enable the discovery of novel insights that may not otherwise be evident. Doing these analyses efficiently requires integration of heterogeneous data into a common data structure; however, this type of data integration represents a major challenge, especially for large, multi-institutional projects. Not only should the sharing of individual datasets follow FAIR principles (Findable, Accessible, Interoperable, Reusable), but the ideal data management system should also include a central multidisciplinary data organization framework. The EMERGE Database (EMERGE-DB; https://emerge-db.asc.ohio-state.edu/) is the central data hub of the EMERGE Biology Integration Institute (NSF award # 2022070), which investigates the changing dynamics of a thawing permafrost ecosystem in Stordalen Mire, northern Sweden. The EMERGE-DB accomplishes the essential tasks of data management (i.e., data storage and sharing), while also offering more advanced functionality to facilitate interdisciplinary collaboration. Data and standardized metadata—including both sample and file metadata—are integrated within a Neo4j graph database, which allows combined datasets from different source files to be obtained via efficient custom queries. A front-end web portal provides access to this data for both the public and for EMERGE project members (who can access non-public data via login), with different pages providing different “views” of the database for different common use cases. Although data are still deposited to external community repositories (e.g. Zenodo, NCBI databases) to ensure cost-effective long-term accessibility, these depositions are tracked within the EMERGE-DB’s standardized metadata system, with all internally- and externally-stored datasets displayed within a centralized page on the web portal. Although this data integration and sharing framework is customized for the EMERGE project’s needs, many of its guiding principles—such as the centralized web access point for all datasets, and general file formatting standards to streamline the detailed integration of sample metadata—are broadly applicable as “best practices” that other projects can apply in their own data management systems.
Soil microorganisms are pivotal in the global carbon cycle, but the viruses that affect them and their impact on ecosystems are less understood. In this study, we explored the diversity, dynamics, and ecology of soil viruses through 379 metagenomes collected annually from 2010 to 2017. These samples spanned the seasonally thawed active layer of a permafrost thaw gradient, which included palsa, bog, and fen habitats. We identified 5051 virus operational taxonomic units (vOTUs), doubling the known viruses for this site. These vOTUs were largely ephemeral within habitats, suggesting a turnover at the vOTU level from year to year. While the diversity varied by thaw stage and depth-related patterns were specific to each habitat, the virus communities did not significantly change over time. The abundance ratios of virus to host at the phylum level did not show consistent trends across the thaw gradient, depth, or time. To assess potential ecosystem impacts, we predicted hosts in silico and found viruses linked to microbial lineages involved in the carbon cycle, such as methanotrophy and methanogenesis. This included the identification of viruses of Candidatus Methanoflorens, a significant global methane contributor. We also detected a variety of potential auxiliary metabolic genes, including 24 carbon-degrading glycoside hydrolases, six of which are uniquely terrestrial. In conclusion, these long-term observations enhance our understanding of soil viruses in the context of climate-relevant processes and provide opportunities to explore their role in terrestrial carbon cycling.
Our knowledge of viral sequence space has exploded with advancing sequencing technologies and large-scale sampling and analytical efforts. Though archaea are important and abundant prokaryotes in many systems, our knowledge of archaeal viruses outside of extreme environments is limited. This largely stems from the lack of a robust, high-throughput, and systematic way to distinguish between bacterial and archaeal viruses in datasets of curated viruses. Here we upgrade our prior text-based tool (MArVD) via training and testing a random forest machine learning algorithm against a newly curated dataset of archaeal viruses. After optimization, MArVD2 presented a significant improvement over its predecessor in terms of scalability, usability, and flexibility, and will allow user-defined custom training datasets as archaeal virus discovery progresses. Benchmarking showed that a model trained with viral sequences from the hypersaline, marine, and hot spring environments correctly classified 85% of the archaeal viruses with a false detection rate below 2% using a random forest prediction threshold of 80% in a separate benchmarking dataset from the same habitats.
Whereas DNA viruses are known to be abundant, diverse, and commonly key ecosystem players, RNA viruses are insufficiently studied outside disease settings. In this study, we analyzed ≈28 terabases of Global Ocean RNA sequences to expand Earth’s RNA virus catalogs and their taxonomy, investigate their evolutionary origins, and assess their marine biogeography from pole to pole. Using new approaches to optimize discovery and classification, we identified RNA viruses that necessitate substantive revisions of taxonomy (doubling phyla and adding >50% new classes) and evolutionary understanding. “Species”-rank abundance determination revealed that viruses of the new phyla “ Taraviricota ,” a missing link in early RNA virus evolution, and “ Arctiviricota ” are widespread and dominant in the oceans. These efforts provide foundational knowledge critical to integrating RNA viruses into ecological and epidemiological models.
DNA viruses are increasingly recognized as influencing marine microbes and microbe-mediated biogeochemical cycling. However, little is known about global marine RNA virus diversity, ecology, and ecosystem roles. In this study, we uncover patterns and predictors of marine RNA virus community- and "species"-level diversity and contextualize their ecological impacts from pole to pole. Our analyses revealed four ecological zones, latitudinal and depth diversity patterns, and environmental correlates for RNA viruses. Our findings only partially parallel those of cosampled plankton and show unexpectedly high polar ecological interactions. The influence of RNA viruses on ecosystems appears to be large, as predicted hosts are ecologically important. Moreover, the occurrence of auxiliary metabolic genes indicates that RNA viruses cause reprogramming of diverse host metabolisms, including photosynthesis and carbon cycling, and that RNA virus abundances predict ocean carbon export.
Microbes and their viruses are hidden engines driving Earth’s ecosystems from the oceans and soils to humans and bioreactors. Though gene marker approaches can now be complemented by genome-resolved studies of inter-(macrodiversity) and intra-(microdiversity) population variation, analytical tools to do so remain scattered or under-developed. Here, we introduce MetaPop, an open-source bioinformatic pipeline that provides a single interface to analyze and visualize microbial and viral community metagenomes at both the macro- and microdiversity levels. Macrodiversity estimates include population abundances and α- and β-diversity. Microdiversity calculations include identification of single nucleotide polymorphisms, novel codon-constrained linkage of SNPs, nucleotide diversity (π and θ), and selective pressures (pN/pS and Tajima’s D) within and fixation indices (FST) between populations. MetaPop will also identify genes with distinct codon usage. Following rigorous validation, we applied MetaPop to the gut viromes of autistic children that underwent fecal microbiota transfers and their neurotypical peers. The macrodiversity results confirmed our prior findings for viral populations (microbial shotgun metagenomes were not available) that diversity did not significantly differ between autistic and neurotypical children. However, by also quantifying microdiversity, MetaPop revealed lower average viral nucleotide diversity (π) in autistic children. Analysis of the percentage of genomes detected under positive selection was also lower among autistic children, suggesting that higher viral π in neurotypical children may be beneficial because it allows populations to better “bet hedge” in changing environments. Further, comparisons of microdiversity pre- and post-FMT in autistic children revealed that the delivery FMT method (oral versus rectal) may influence viral activity and engraftment of microdiverse viral populations, with children who received their FMT rectally having higher microdiversity post-FMT. Overall, these results show that analyses at the macro level alone can miss important biological differences. These findings suggest that standardized population and genetic variation analyses will be invaluable for maximizing biological inference, and MetaPop provides a convenient tool package to explore the dual impact of macro- and microdiversity across microbial communities.
Microbes drive myriad ecosystem processes, but under strong influence from viruses. Because studying viruses in complex systems requires different tools than those for microbes, they remain underexplored. To combat this, we previously aggregated double-stranded DNA (dsDNA) virus analysis capabilities and resources into ‘iVirus’ on the CyVerse collaborative cyberinfrastructure. Here we substantially expand iVirus’s functionality and accessibility, to iVirus 2.0, as follows. First, core iVirus apps were integrated into the Department of Energy’s Systems Biology KnowledgeBase (KBase) to provide an additional analytical platform. Second, at CyVerse, 20 software tools (apps) were upgraded or added as new tools and capabilities. Third, nearly 20-fold more sequence reads were aggregated to capture new data and environments. Finally, documentation, as “live” protocols, was updated to maximize user interaction with and contribution to infrastructure development. Together, iVirus 2.0 serves as a uniquely central and accessible analytical platform for studying how viruses, particularly dsDNA viruses, impact diverse microbial ecosystems.
Downloading reads from the Ocean Sampling Day (2014) using fastq-dump, a tool from the SRA toolkit implemented in Cyverse.
Background:Viruses influence global patterns of microbial diversity and nutrient cycles.Though viral metagenomics (viromics), specifically targeting dsDNA viruses, has been critical for revealing viral roles across diverse ecosystems, its analyses differ in many ways from those used for microbes.To date, viromics benchmarking has covered read pre-processing, assembly, relative abundance, read mapping thresholds and diversity estimation, but other steps would benefit from benchmarking and standardization.Here we use in silico-generated datasets and an extensive literature survey to evaluate and highlight how dataset composition (i.e.viromes vs bulk metagenomes) and assembly fragmentation impact (i) viral contig identification tool, (ii) virus taxonomic classification, and (iii) identification and curation of auxiliary metabolic genes (AMGs). Results:The in silico benchmarking of five commonly used virus identification tools show that genecontent-based tools consistently performed well for long (≥3 kbp) contigs, while k-mer-and blast-based tools were uniquely able to detect viruses from short (<3 kbp) contigs.Notably, however, the performance increase of k-mer-and blast-based tools for short contigs was obtained at the cost of increased false positives (sometimes up to ~5% for virome and ~75% bulk samples), particularly when eukaryotic or mobile genetic element sequences were included in the test datasets.For viral classification, variously sized genome fragments were assessed using gene-sharing network analytics to quantify drop-offs in taxonomic assignments, which revealed correct assignations ranging from ~95% (whole genomes) down to ~80% (3 kbp sized genome fragments).A similar trend was also observed for other viral classification tools such as VPF-class, ViPTree and VIRIDIC, suggesting that caution is warranted when classifying short genome fragments and not full genomes.Finally, we highlight how fragmented assemblies can lead to erroneous identification of AMGs and outline a best-practices workflow to curate candidate AMGs in viral genomes assembled from metagenomes.Conclusion:Together, these benchmarking experiments and annotation guidelines should aid researchers seeking to best detect, classify, and characterize the myriad viruses 'hidden' in diverse sequence datasets.
A collection of protocols designed to guide the user in processing a viral metagenome from raw sequence data to assembly, and subsequent analysis. The user usesactual reads fromOcean Sampling Day (2014)and processes them entirely within Cyverse, a NSF-supported cyberinfrastructure.
Viruses, despite their great abundance and significance in biological systems, remain largely mysterious. Indeed, the vast majority of the perhaps hundreds of millions of viral species on the planet remain undiscovered. Additionally, many viruses deposited in central databases like GenBank and RefSeq are littered with genes annotated as “hypothetical protein” or the equivalent. Cenote-Taker2, a virus discovery and annotation tool available on command line and with a graphical user interface with free high-performance computation access, utilizes highly sensitive models of hallmark virus genes to discover familiar or divergent viral sequences from user-input contigs. Additionally, Cenote-Taker2 uses a flexible set of modules to automatically annotate the sequence features of contigs, providing more gene information than comparable tools. The outputs include readable and interactive genome maps, virome summary tables, and files that can be directly submitted to GenBank. We expect Cenote-Taker2 to facilitate virus discovery, annotation, and expansion of the known virome.