A virome survey of banana plantations and their surrounding plants was carried out at nation-wide level in Malawi using virion associated nucleic acids (VANA) high throughput sequencing (HTS) on pooled samples and appropriate alien controls. In total, 366 plants were sequenced, and 23 plant virus species were detected, three species on banana (275 plants) and 20 species in surrounding plants (91 plants). Two putative novel virus species; ginger tymo-like virus and pepper derived totivirus were detected and confirmed by RT-PCR on ginger and pepper. Nine known virus species and detected a host plant was identified for two of them. No viral exchange between banana and surrounding plants was observed. Results from the VANA protocol, applied to pooled banana samples, were compared with previous targeted PCR results obtained from individual banana samples. HTS test detected better BanMMV than IC-(RT)-PCR on individual samples (better inclusivity) but detected with much lower sensitivity BBTV and BSV species, often with less than 10 reads per sample. Detection of novel and known viruses and new host plants calls for strengthened sanitory and phytosanitory measures within and beyond banana production systems. Our research confirms that HTS sensitivity depends on sampling, pooling protocol and targeted virus species.
Saffron (Crocus sativus L.) is a vegetatively propagated crop of high economic and cultural value, potentially affected by viral infections that may impact its productivity. Despite Iran’s dominance in global saffron production, knowledge of its virome remains limited. In this study, we conducted the first nationwide virome survey of saffron in Iran employing a high-throughput sequencing (HTS) approach on pooled samples obtained from eleven provinces in Iran and one location in Afghanistan. Members of three virus families were detected—Potyviridae (Potyvirus), Solemoviridae (Polerovirus), and Geminiviridae (Mastrevirus)—as well as one satellite from the family Alphasatellitidae (Clecrusatellite). A novel Potyvirus, tentatively named saffron Iran virus (SaIRV) and detected in three provinces, shares less than 68% nucleotide identity with known Potyvirus species, thus meeting the ICTV criteria for designation as a new species. Genetic diversity analyses revealed substantial intrapopulation SNP variation but no clear geographical clustering. Among the two wild Crocus species sampled, only Crocus speciosus harbored turnip mosaic virus. Virome network and phylogenetic analyses confirmed widespread viral circulation likely driven by corm-mediated propagation. Our findings highlight the need for targeted certification programs and biological characterization of key viruses to mitigate potential impacts on saffron yield and quality.
Background High-throughput sequencing (HTS) technologies completed by the bioinformatic analysis of the generated data are becoming an important detection technique for virus diagnostics. They have the potential to replace or complement the current PCR-based methods thanks to their improved inclusivity and analytical sensitivity, as well as their overall good repeatability and reproducibility. Cross-contamination is a well-known phenomenon in molecular diagnostics and corresponds to the exchange of genetic material between samples. Cross-contamination management was a key drawback during the development of PCR-based detection and is now adequately monitored in routine diagnostics. HTS technologies are facing similar difficulties due to their very high analytical sensitivity. As a single viral read could be detected in millions of sequencing reads, it is mandatory to fix a detection threshold that will be informed by estimated cross-contamination. Cross-contamination monitoring should therefore be a priority when detecting viruses by HTS technologies. Results We present Cont-ID, a bioinformatic tool designed to check for cross-contamination by analysing the relative abundance of virus sequencing reads identified in sequence metagenomic datasets and their duplication between samples. It can be applied when the samples in a sequencing batch have been processed in parallel in the laboratory and with at least one specific external control called Alien control. Using 273 real datasets, including 68 virus species from different hosts (fruit tree, plant, human) and several library preparation protocols (Ribodepleted total RNA, small RNA and double-stranded RNA), we demonstrated that Cont-ID classifies with high accuracy (91%) viral species detection into (true) infection or (cross) contamination. This classification raises confidence in the detection and facilitates the downstream interpretation and confirmation of the results by prioritising the virus detections that should be confirmed. Conclusions Cross-contamination between samples when detecting viruses using HTS (Illumina technology) can be monitored and highlighted by Cont-ID (provided an alien control is present). Cont-ID is based on a flexible methodology relying on the output of bioinformatics analyses of the sequencing reads and considering the contamination pattern specific to each batch of samples. The Cont-ID method is adaptable so that each laboratory can optimise it before its validation and routine use.
A cockspur coral tree ( Erythrina crista-galli , Fabaceae) from the collection of woody ornamentals of the Botanical Garden (Faculty of Science) in Zagreb showed conspicuous virus-like symptoms. The leaf tissue was analyzed by high throughput sequencing (HTS), revealing the presence of two viruses: prune dwarf virus (PDV, genus Ilarvirus ) and an unknown virus belonging to the genus Capillovirus (family Betaflexiviridae ). The complete sequence of PDV RNA3 of 2,129 nucleotides (nts) coding for the coat and movement proteins was obtained. The complete coding region spanning 6,483 nts was obtained for the capillovirus. It contained all expected open reading frames, and its maximum nucleotide identity was 42% with apple stem grooving capillovirus sequence (GenBank accession number LC143387). As this is far below the threshold for species delineation within the genus Capillovirus , we propose the name Erythrina capillovirus (ErCV) for this putatively new virus and a new virus species Capillovirus ErCV in the family Betaflexiviridae . RT-PCR confirms the presence of ErCV and PDV in symptomatic E. crista-galli , an unusual and exotic fabaceous host with no viruses recorded yet. Asymptomatic Erythrina plants were tested negative for the two viruses in RT-PCR, which, together with their presence in the symptomatic plant, suggests the capillovirus and/or PDV might play a role in eliciting observed symptoms and deserve further investigations.
The advances in high-throughput sequencing (HTS) technologies and bioinformatic tools have provided new opportunities for virus and viroid discovery and diagnostics. Hence, new sequences of viral origin are being discovered and published at a previously unseen rate. Therefore, a collective effort was undertaken to write and propose a framework for prioritizing the biological characterization steps needed after discovering a new plant virus to evaluate its impact at different levels. Even though the proposed approach was widely used, a revision of these guidelines was prepared to consider virus discovery and characterization trends and integrate novel approaches and tools recently published or under development. This updated framework is more adapted to the current rate of virus discovery and provides an improved prioritization for filling knowledge and data gaps. It consists of four distinct steps adapted to include a multi-stakeholder feedback loop. Key improvements include better prioritization and organization of the various steps, earlier data sharing among researchers and involved stakeholders, public database screening, and exploitation of genomic information to predict biological properties.
High-throughput sequencing (HTS) technologies have brought tremendous improvements in the ability to detect plant viruses and have great potential for application in virus routine diagnostics. The performance criteria of an HTS test need therefore to be estimated and compared with traditional virus indexing tests before it can be used in routine diagnostics. In this study, 78 Musa accessions previously indexed for viruses by molecular tests and/or electron microscopy were tested individually or in pools using an HTS protocol based on total RNA sequencing. The analytical sensitivity of HTS and RT-PCR was also compared by independent testing on serial dilutions of RNA extracts. In total, 136 libraries were sequenced in five batches, and the sequences were analyzed for virus detection. The external alien control, a wheat sample infected by barley yellow dwarf virus, monitored the contamination burden and determined an adaptative detection threshold. Overall, the HTS test displayed a better analytical sensitivity than the RT-PCR and a better inclusivity than the classical indexing protocol, as distant isolates and new viral species were only detected by the HTS test. The repeatability and reproducibility of virus detection were both 100%, although differences in number of sequencing reads per virus were observed between replicates. The diagnostic sensitivity was very high, but false positive results were observed. Finally, the results also underlined the need for expert judgement in the interpretation of the results. In conclusion, the HTS test with an alien control and completed by expert evaluation fulfilled the criteria of the virus indexing protocol for Musa germplasm. [Formula: see text] Copyright © 2023 The Author(s). This is an open access article distributed under the CC BY-NC-ND 4.0 International license .
Recent developments in high-throughput sequencing (HTS) technologies and bioinformatics have drastically changed research in virology, especially for virus discovery. Indeed, proper monitoring of the viral population requires information on the different isolates circulating in the studied area. For this purpose, HTS has greatly facilitated the sequencing of new genomes of detected viruses and their comparison. However, bioinformatics analyses allowing reconstruction of genome sequences and detection of single nucleotide polymorphisms (SNPs) can potentially create bias and has not been widely addressed so far. Therefore, more knowledge is required on the limitations of predicting SNPs based on HTS-generated sequence samples. To address this issue, we compared the ability of 14 plant virology laboratories, each employing a different bioinformatics pipeline, to detect 21 variants of pepino mosaic virus (PepMV) in three samples through large-scale performance testing (PT) using three artificially designed datasets. To evaluate the impact of bioinformatics analyses, they were divided into three key steps: reads pre-processing, virus-isolate identification, and variant calling. Each step was evaluated independently through an original, PT design including discussion and validation between participants at each step. Overall, this work underlines key parameters influencing SNPs detection and proposes recommendations for reliable variant calling for plant viruses. The identification of the closest reference, mapping parameters and manual validation of the detection were recognized as the most impactful analysis steps for the success of the SNPs detections. Strategies to improve the prediction of SNPs are also discussed.
Cassava Brown Streak Disease (CBSD), which is caused by cassava brown streak virus (CBSV) and Ugandan cassava brown streak virus (UCBSV), represents one of the most devastating threats to cassava production in Africa, including in Rwanda where a dramatic epidemic in 2014 dropped cassava yield from 3.3 million to 900,000 tonnes (1). Studying viral genetic diversity at the genome level is essential in disease management, as it can provide valuable information on the origin and dynamics of epidemic events. To fill the current lack of genome-based diversity studies of UCBSV, we performed a nationwide survey of cassava ipomovirus genomic sequences in Rwanda by high-throughput sequencing (HTS) of pools of plants sampled from 130 cassava fields in thirteen cassava-producing districts, spanning seven agro-ecological zones with contrasting climatic conditions and different cassava cultivars. HTS allowed the assembly of a nearly complete consensus genome of UCBSV in twelve districts. The phylogenetic analysis revealed high homology between UCBSV genome sequences, with a maximum of 0.8 per cent divergence between genomes at the nucleotide level. An in-depth investigation based on Single Nucleotide Polymorphisms (SNPs) was conducted to explore the genome diversity beyond the consensus sequences. First, to ensure the validity of the result, a panel of SNPs was confirmed by independent reverse transcription polymerase chain reaction (RT-PCR) and Sanger sequencing. Furthermore, the combination of fixation index (FST) calculation and Principal Component Analysis (PCA) based on SNP patterns identified three different UCBSV haplotypes geographically clustered. The haplotype 2 (H2) was restricted to the central regions, where the NAROCAS 1 cultivar is predominantly farmed. RT-PCR and Sanger sequencing of individual NAROCAS1 plants confirmed their association with H2. Haplotype 1 was widely spread, with a 100 per cent occurrence in the Eastern region, while Haplotype 3 was only found in the Western region. These haplotypes’ associations with specific cultivars or regions would need further confirmation. Our results prove that a much more complex picture of genetic diversity can be deciphered beyond the consensus sequences, with practical implications on virus epidemiology, evolution, and disease management. Our methodology proposes a high-resolution analysis of genome diversity beyond the consensus between and within samples. It can be used at various scales, from individual plants to pooled samples of virus-infected plants. Our findings also showed how subtle genetic differences could be informative on the potential impact of agricultural practices, as the presence and frequency of a virus haplotype could be correlated with the dissemination and adoption of improved cultivars.
High-throughput sequencing (HTS) technologies have the potential to become one of the most significant advances in molecular diagnostics. Their use by researchers to detect and characterize plant pathogens and pests has been growing steadily for more than a decade and they are now envisioned as a routine diagnostic test to be deployed by plant pest diagnostics laboratories. Nevertheless, HTS technologies and downstream bioinformatics analysis of the generated datasets represent a complex process including many steps whose reliability must be ensured. The aim of the present guidelines is to provide recommendations for researchers and diagnosticians aiming to reliably use HTS technologies to detect plant pathogens and pests. These guidelines are generic and do not depend on the sequencing technology or platform. They cover all the adoption processes of HTS technologies from test selection to test validation as well as their routine implementation. A special emphasis is given to key elements to be considered: undertaking a risk analysis, designing sample panels for validation, using proper controls, evaluating performance criteria, confirming and interpreting results. These guidelines cover any HTS test used for the detection and identification of any plant pest (viroid, virus, bacteria, phytoplasma, fungi and fungus-like protists, nematodes, arthropods, plants) from any type of matrix. Overall, their adoption by diagnosticians and researchers should greatly improve the reliability of pathogens and pest diagnostics and foster the use of HTS technologies in plant health.
High-throughput sequencing (HTS) is a powerful tool that enables the simultaneous detection and potential identification of any organisms present in a sample. The growing interest in the application of HTS technologies for routine diagnostics in plant health laboratories is triggering the development of guidelines on how to prepare laboratories for performing HTS testing. This paper describes general and technical recommendations to guide laboratories through the complex process of preparing a laboratory for HTS tests within existing quality assurance systems. From nucleic acid extractions to data analysis and interpretation, all of the steps are covered to ensure reliable and reproducible results. These guidelines are relevant for the detection and identification of any plant pest (e.g. arthropods, bacteria, fungi, nematodes, invasive plants or weeds, protozoa, viroids, viruses), and from any type of matrix (e.g. pure microbial culture, plant tissue, soil, water), regardless of the HTS technology (e.g. amplicon sequencing, shotgun sequencing) and of the application (e.g. surveillance programme, phytosanitary certification, quarantine, import control). These guidelines are written in general terms to facilitate the adoption of HTS technologies in plant pest routine diagnostics and enable broader application in all plant health fields, including research. A glossary of relevant terms is provided among the Supplementary Material.
High-throughput sequencing (HTS) technologies have become indispensable tools assisting plant virus diagnostics and research thanks to their ability to detect any plant virus in a sample without prior knowledge. As HTS technologies are heavily relying on bioinformatics analysis of the huge amount of generated sequences, it is of utmost importance that researchers can rely on efficient and reliable bioinformatic tools and can understand the principles, advantages, and disadvantages of the tools used. Here, we present a critical overview of the steps involved in HTS as employed for plant virus detection and virome characterization. We start from sample preparation and nucleic acid extraction as appropriate to the chosen HTS strategy, which is followed by basic data analysis requirements, an extensive overview of the in-depth data processing options, and taxonomic classification of viral sequences detected. By presenting the bioinformatic tools and a detailed overview of the consecutive steps that can be used to implement a well-structured HTS data analysis in an easy and accessible way, this paper is targeted at both beginners and expert scientists engaging in HTS plant virome projects.
Background Plant viruses cause half of the emerging plant diseases and pose a great threat to agricultural crops worldwide. INEXTVIR Marie-Curie Training Network proposes the use of High-Throughput Sequencing (HTS) technologies for studying at an unprecedented depth the virome of selected agricultural crops and ecosystems across Europe. Our specific aims are: a) detecting the virome present in selected agricultural crops across Europe, b) understanding the biological and ecological impacts of the virome, c) improving virus detection capabilities in plant health diagnostic and certification settings through the development and validation of HTS methods, d) assessing the agronomic and socio-economic impact of virome and translate it into practical decision tools. Description Within the INEXTVIR project three bioinformatics PhDs are responsible for the development of methods and software tools for virome data analysis. We will evaluate the performance of existing pipelines for the analysis of HTS data for plant virus, while also developing new methods. Furthermore, we will propose an approach to simplify and automate bioinformatics pipelines to be user-friendlier by integrating them within a laboratory data management system and workflow tools. Moreover, we plan to develop a machine learning approach for decision support helping to select the most efficient pipeline based on the type of HTS dataset available and the goal of the analysis. Conclusions The simplified bioinformatics pipelines for virome HTS data analysis will make their use easier by researchers with little bioinformatic background. While working on that we are also contributing to a unified project-wide management of HTS data.
Large-scale genome sequencing and the increasingly massive use of high-throughput approaches produce a vast amount of new information that completely transforms our understanding of thousands of microbial species. However, despite the development of powerful bioinformatics approaches, full interpretation of the content of these genomes remains a difficult task. Launched in 2005, the MicroScope platform (https://www.genoscope.cns.fr/agc/ microscope) has been under continuous development and provides analysis for prokaryotic genome projects together with metabolic network reconstruction and post-genomic experiments allowing users to improve the understanding of gene functions. Here we present new improvements of the MicroScope user interface for genome selection, navigation and expert gene annotation. Automatic functional annotation procedures of the platform have also been updated and we added several new tools for the functional annotation of genes and genomic regions. We finally focus on new tools and pipeline developed to perform comparative analyses on hundreds of genomes based on pangenome graphs. To date, MicroScope contains data for >11 800 micro-bial genomes, part of which are manually curated and maintained by microbiologists (>4500 personal accounts in September 2019). The platform enables collaborative work in a rich comparative genomic context and improves community-based curation efforts.
The overwhelming list of new bacterial genomes becoming available on a daily basis makes accurate genome annotation an essential step that ultimately determines the relevance of thousands of genomes stored in public databanks. The MicroScope platform (http://www.genoscope.cns.fr/agc/microscope) is an integrative resource that supports systematic and efficient revision of microbial genome annotation, data management and comparative analysis. Starting from the results of our syntactic, functional and relational annotation pipelines, MicroScope provides an integrated environment for the expert annotation and comparative analysis of prokaryotic genomes. It combines tools and graphical interfaces to analyze genomes and to perform the manual curation of gene function in a comparative genomics and metabolic context. In this article, we describe the free-of-charge MicroScope services for the annotation and analysis of microbial (meta)genomes, transcriptomic and resequencing data. Then, the functionalities of the platform are presented in a way providing practical guidance and help to the nonspecialists in bioinformatics. Newly integrated analysis tools (i.e. prediction of virulence and resistance genes in bacterial genomes) and original method recently developed (the pan-genome graph representation) are also described. Integrated environments such as MicroScope clearly contribute, through the user community, to help maintaining accurate resources.
The annotation of genomes from NGS platforms needs to be automated and fully integrated. However, maintaining consistency and accuracy in genome annotation is a challenging problem because millions of protein database entries are not assigned reliable functions. This shortcoming limits the knowledge that can be extracted from genomes and metabolic models. Launched in 2005, the MicroScope platform (http://www.genoscope.cns.fr/agc/microscope) is an integrative resource that supports systematic and efficient revision of microbial genome annotation, data management and comparative analysis. Effective comparative analysis requires a consistent and complete view of biological data, and therefore, support for reviewing the quality of functional annotation is critical. MicroScope allows users to analyze microbial (meta)genomes together with post-genomic experiment results if any (i.e. transcriptomics, re-sequencing of evolved strains, mutant collections, phenotype data). It combines tools and graphical interfaces to analyze genomes and to perform the expert curation of gene functions in a comparative context. Starting with a short overview of the MicroScope system, this paper focuses on some major improvements of the Web interface, mainly for the submission of genomic data and on original tools and pipelines that have been developed and integrated in the platform: computation of pan-genomes and prediction of biosynthetic gene clusters. Today the resource contains data for more than 6000 microbial genomes, and among the 2700 personal accounts (65% of which are now from foreign countries), 14% of the users are performing expert annotations, on at least a weekly basis, contributing to improve the quality of microbial genome annotations.