Searching publicly archived sequence data for emerging aquatic animal pathogens is a powerful but challenging approach for increasing our understanding of newly identified or poorly characterized organisms. However, searching for target sequences within the sequence read archive (SRA) database requires significant time, data storage, and computing power, limiting its accessibility. Utilizing a new database, Logan, we undertook a meta-analysis of SRA data sets to investigate the presence of an emerging virus, Macrobrachium rosenbergii golda virus (MrGV). MrGV was first characterized in M. rosenbergii larvae in 2020, associated with repeated mass mortalities in Bangladesh hatcheries. MrGV has since been detected in two separate reports from the Jiangsu Province of central, coastal China, and during a larval mortality event in India. Here, we discovered that MrGV is present in two additional provinces in southern China, Thailand, and India. We also found molecular evidence to confirm, as previously suspected, the circulation of the virus within Southern Asian populations of M. rosenbergii as far back as 2011, and that, based on relative abundance, MrGV is mostly associated with larvae. Overall, the identification of MrGV sequences in data sets that are largely unpublished within the scientific literature has provided novel insights into the pathogen's biology, including the prevalence of MrGV globally and the life stages of prawns that should be screened to prevent the spread of the virus. This work illustrates how mining public sequencing data, supported by databases like Logan and standardized metadata submissions, can support cost-effective epidemiological studies of pathogens and strengthen One Health approaches to global disease monitoring.IMPORTANCESearching for target sequences within the sequence read archive (SRA) database requires significant time, data storage, and computing power, limiting its accessibility. This study demonstrates how the Logan database, constructed from an SRA-wide genome assembly, can be utilized to rapidly and efficiently find target sequences within the SRA database, expanding the use of these publicly available data sets outside of their original intended purposes. Here, we searched for an emerging virus, Macrobrachium rosenbergii golda virus, in prawns to reveal insights into its geographic distribution, host range, and relative abundance, without the need for additional sampling. We demonstrate how, with careful application of this approach, alongside improvements in metadata quality and accessibility, sequencing data sets can be used to uncover critical insights into pathogen biology. This type of data mining could add otherwise unknown data to epidemiological studies of emerging, re-emerging, and rare pathogens globally, allowing the determination of the spread of agents within and between populations.
Macrobrachium rosenbergii golda virus (MrGV) was first characterised in Macrobrachium rosenbergii larvae in 2020, associated with mass mortalities in multiple Southern Bangladesh prawn hatcheries. MrGV has since been detected in two metatranscriptomic datasets from M. rosenbergii in China and in relation to a larval mortality event in India. The major objective of this study was to further characterise the geographical spread of the virus by mining the NCBI Sequence Read Archive (SRA) database. Utilising a new database, Logan, generated from assembling each SRA dataset, and a custom Snakemake pipeline, we discovered MrGV sequence data in M. rosenbergii SRA datasets from China, Thailand, and India, and determined that presence and relative abundance of MrGV is mostly associated with the larval life stage of M. rosenbergii . These results provide insights into the current prevalence of MrGV globally, suggest the life stages of prawn that should be screened to prevent spread of the virus, and demonstrate how the Logan database can be used to inform epidemiological studies. ### Competing Interest Statement The authors have declared no competing interest. Cefas Seedcorn project DP1000 Defra project AHPFFX The Genomics for Animal and Plant Disease Consortium (GAP-DC) provided by Defra and UK Research and Innovation (UKRI)
Water and Pacific Oyster samples were collected from an estuary in Southwest England. Primary identification of bacteria suggested isolates (n = 10) were Vibrio alginolyticus; however, phylogenetic analysis using whole genome sequencing Illumina and Nanopore data showed they were Vibrio diabolicus.
Background Patagonian toothfish ( Dissostichus eleginoides ) is an economically and ecologically important fish species in the family Nototheniidae. Juveniles occupy progressively deeper waters as they mature and grow, and adults have been caught as deep as 2500 m, living on or in just above the southern shelves and slopes around the sub-Antarctic islands of the Southern Ocean. As apex predators, they are a key part of the food web, feeding on a variety of prey, including krill, squid, and other fish. Despite its importance, genomic sequence data, which could be used for more accurate dating of the divergence between Patagonian and Antarctic toothfish, or establish whether it shares adaptations to temperature with fish living in more polar or equatorial climes, has so far been limited. Results A high-quality D. eleginoides genome was generated using a combination of Illumina, PacBio and Omni-C sequencing technologies. To aid the genome annotation, the transcriptome derived from a variety of toothfish tissues was also generated using both short and long read sequencing methods. The final genome assembly was 797.8 Mb with a N50 scaffold length of 3.5 Mb. Approximately 31.7% of the genome consisted of repetitive elements. A total of 35,543 putative protein-coding regions were identified, of which 50% have been functionally annotated. Transcriptomics analysis showed that approximately 64% of the predicted genes (22,617 genes) were found to be expressed in the tissues sampled. Comparative genomics analysis revealed that the anti-freeze glycoprotein (AFGP) locus of D. eleginoides does not contain any AFGP proteins compared to the same locus in the Antarctic toothfish ( Dissostichus mawsoni ). This is in agreement with previously published results looking at hybridization signals and confirms that Patagonian toothfish do not possess AFGP coding sequences in their genome. Conclusions We have assembled and annotated the Patagonian toothfish genome, which will provide a valuable genetic resource for ecological and evolutionary studies on this and other closely related species.
Norovirus is one of the largest causes of gastroenteritis worldwide, and Hepatitis E virus (HEV) is an emerging pathogen that has become the most dominant cause of acute viral hepatitis in recent years. The presence of norovirus and HEV has been reported within wastewater in many countries previously. Here we used amplicon deep sequencing (metabarcoding) to identify norovirus and HEV strains in wastewater samples from England collected in 2019 and 2020. For HEV, we sequenced a fragment of the RNA-dependent RNA polymerase (RdRp) gene targeting genotype three strains. For norovirus, we sequenced the 5′ portion of the major capsid protein gene (VP1) of genogroup II strains. Sequencing of the wastewater samples revealed eight different genotypes of norovirus GII (GII.2, GII.3, GII.4, GII.6, GII.7, GII.9, GII.13 and GII.17). Genotypes GII.3 and GII.4 were the most commonly found. The HEV metabarcoding assay was able to identify HEV genotype 3 strains in some samples with a very low viral concentration determined by RT-qPCR. Analysis showed that most HEV strains found in influent wastewater were typed as G3c and G3e and were likely to have originated from humans or swine. However, the small size of the HEV nested PCR amplicon could cause issues with typing, and so this method is more appropriate for samples with high CTs where methods targeting longer genomic regions are unlikely to be successful. This is the first report of HEV RNA in wastewater in England. This study demonstrates the utility of wastewater sequencing and the need for wider surveillance of norovirus and HEV within host species and environments.
While both virulent and putatively avirulent Yersinia ruckeri strains exist in aquaculture environments, the relationship between the distribution of virulence-associated factors and de facto pathogenicity in fish remains poorly understood. Pan-genome analysis of 18 complete genomes, representing established virulent and putatively avirulent lineages of Y. ruckeri, revealed the presence of a number of accessory genetic determinants. Further investigation of 68 draft genome assemblies revealed that the distribution of certain putative virulence factors correlated well with virulence and host-specificity. The inverse-autotransporter invasin locus yrIlm was, however, the only gene present in all virulent strains, while absent in lineages regarded as avirulent. Strains known to be associated with significant mortalities in salmonid aquaculture display a combination of serotype O1-LPS and yrIlm, with the well-documented highly virulent lineages, represented by MLVA clonal complexes 1 and 2, displaying duplication of the yrIlm locus. Duplication of the yrIlm locus was further found to have evolved over time in clonal complex 1, where some modern, highly virulent isolates display up to three copies.
Bacteria from the family Vibrionaceae have been implicated in mass mortalities of farmed Pacific oysters (Magallana gigas) in multiple countries, leading to substantial impairment of growth in the sector. In Ireland there has been concern that Vibrio have been involved in serious summer outbreaks. There is evidence that Vibrio aestuarianus is increasingly becoming the main pathogen of concern for the Pacific oyster industry in Ireland. While bacteria belonging to the Vibrio splendidus clade are also detected frequently in mortality episodes, their role in the outbreaks of summer mortality is not well understood. To identify and characterize strains involved in these outbreaks, 43 Vibrio isolates were recovered from Pacific oyster summer mass mortality episodes in Ireland from 2008 to 2015 and these were whole-genome sequenced. Among these, 25 were found to be V. aestuarianus (implicated in disease) and 18 were members of the V. splendidus species complex (role in disease undetermined). Two distinct clades of V. aestuarianus - clade A and clade B - were found that had previously been described as circulating within French oyster culture. The high degree of similarity between the Irish and French V. aestuarianus isolates points to translocation of the pathogen between Europe's two major oyster-producing countries, probably via trade in spat and other age classes. V. splendidus isolates were more diverse, but the data reveal a single clone of this species that has spread across oyster farms in Ireland. This underscores that Vibrio could be transmitted readily across oyster farms. The presence of V. aestuarianus clades A and B in not only France but also Ireland adds weight to growing concern that this pathogen is spreading and impacting Pacific oyster production within Europe.
Infectious disease causes significant mortality in wild and farmed systems, threatening biodiversity, conservation and animal welfare, as well as food security. To mitigate impacts and inform policy, tools such as mathematical models and computer simulations are valuable for predicting the potential spread and impact of disease. This paper describes the development of the Aquaculture Disease Network Model, AquaNet-Mod, and demonstrates its application to evaluating disease epidemics and the efficacy of control, using a Viral Haemorrhagic Septicaemia (VHS) case study. AquaNet-Mod is a data-driven, stochastic, state-transition model. Disease spread can occur via four different mechanisms, i) live fish movement, ii) river based, iii) short distance mechanical and iv) distance independent mechanical. Sites transit between three disease states: susceptible, clinically infected and subclinically infected. Disease spread can be interrupted by the application of disease mitigation measures and controls such as contact tracing, culling, fallowing and surveillance. Results from a VHS case study highlight the potential for VHS to spread to 96% of sites over a 10 year time horizon if no disease controls are applied. Epidemiological impact is significantly reduced when live fish movement restrictions are placed on the most connected sites and further still, when disease controls, representative of current disease control policy in England and Wales, are applied. The importance of specific disease control measures, particularly contact tracing and disease detection rate, are also highlighted. The merit of this model for evaluation of disease spread and the efficacy of controls, in the context of policy, along with potential for further application and development of the model, for example to include economic parameters, is discussed.
The armoured dinoflagellate Alexandrium can be found throughout many of the world’s temperate and tropical marine environments. The genus has been studied extensively since approximately half of its members produce a family of potent neurotoxins, collectively called saxitoxin. These compounds represent a significant threat to animal and environmental health. Moreover, the consumption of bivalve molluscs contaminated with saxitoxin poses a threat to human health. The identification of Alexandrium cells collected from sea water samples using light microscopy can provide early warnings of a toxic event, giving harvesters and competent authorities time to implement measures that safeguard consumers. However, this method cannot reliably resolve Alexandrium to a species level and, therefore, is unable to differentiate between toxic and non-toxic variants. The assay outlined in this study uses a quick recombinase polymerase amplification and nanopore sequencing method to first target and amplify a 500 bp fragment of the ribosomal RNA large subunit and then sequence the amplicon so that individual species from the Alexandrium genus can be resolved. The analytical sensitivity and specificity of the assay was assessed using seawater samples spiked with different Alexandrium species. When using a 0.22 µm membrane to capture and resuspend cells, the assay was consistently able to identify a single cell of A. minutum in 50 mL of seawater. Phylogenetic analysis showed the assay could identify the A. catenella, A. minutum, A. tamutum, A. tamarense, A. pacificum, and A. ostenfeldii species from environmental samples, with just the alignment of the reads being sufficient to provide accurate, real-time species identification. By using sequencing data to qualify when the toxic A. catenella species was present, it was possible to improve the correlation between cell counts and shellfish toxicity from r = 0.386 to r = 0.769 (p ≤ 0.05). Furthermore, a McNemar’s paired test performed on qualitative data highlighted no statistical differences between samples confirmed positive or negative for toxic species of Alexandrium by both phylogenetic analysis and real time alignment with the presence or absence of toxins in shellfish. The assay was designed to be deployed in the field for the purposes of in situ testing, which required the development of custom tools and state-of-the-art automation. The assay is rapid and resilient to matrix inhibition, making it suitable as a potential alternative detection method or a complementary one, especially when applying regulatory controls.
The World Health Organization considers antimicrobial resistance as one of the most pressing global issues which poses a fundamental threat to human health, development, and security. Due to demographic and environmental factors, the marine environment of the Gulf Cooperation Council (GCC) region may be particularly susceptible to the threat of antimicrobial resistance. However, there is currently little information on the presence of AMR in the GCC marine environment to inform the design of appropriate targeted surveillance activities. The objective of this study was to develop, implement and conduct a rapid regional baseline monitoring survey of the presence of AMR in the GCC marine environment, through the analysis of seawater collected from high-risk areas across four GCC states: (Bahrain, Oman, Kuwait, and the United Arab Emirates). 560 Escherichia coli strains were analysed as part of this monitoring programme between December 2018 and May 2019. Multi-drug resistance (resistance to three or more structural classes of antimicrobials) was observed in 32.5% of tested isolates. High levels of reduced susceptibility to ampicillin (29.6%), nalidixic acid (27.9%), tetracycline (27.5%), sulfamethoxazole (22.5%) and trimethoprim (22.5%) were observed. Reduced susceptibility to the high priority critically important antimicrobials: azithromycin (9.3%), ceftazidime (12.7%), cefotaxime (12.7%), ciprofloxacin (44.6%), gentamicin (2.7%) and tigecycline (0.5%), was also noted. A subset of 173 isolates was whole genome sequenced, and high carriage rates of qnrS1 (60/173) and blaCTX-M-15 (45/173) were observed, correlating with reduced susceptibility to the fluoroquinolones and third generation cephalosporins, respectively. This study is important because of the resistance patterns observed, the demonstrated utility in applying genomic-based approaches to routine microbiological monitoring, and the overall establishment of a transnational AMR surveillance framework focussed on coastal and marine environments.
Eukaryote symbionts of animals are major drivers of ecosystems not only because of their diversity and host interactions from variable pathogenicity but also through different key roles such as commensalism and to different types of interdependence. However, molecular investigations of metazoan eukaryomes require minimising coamplification of homologous host genes. In this study we (1) identified a previously published "antimetazoan" reverse primer to theoretically enable amplification of a wider range of microeukaryotic symbionts, including more evolutionarily divergent sequence types, (2) evaluated in silico several antimetazoan primer combinations, and (3) optimised the application of the best performing primer pair for high throughput sequencing (HTS) by comparing one-step and two-step PCR amplification approaches, testing different annealing temperatures and evaluating the taxonomic profiles produced by HTS and data analysis. The primer combination 574*F - UNonMet_DB tested in silico showed the largest diversity of nonmetazoan sequence types in the SILVA database and was also the shortest available primer combination for broadly-targeting antimetazoan amplification across the 18S rRNA gene V4 region. We demonstrate that the one-step PCR approach used for library preparation produces significantly lower proportions of metazoan reads, and a more comprehensive coverage of host-associated microeukaryote reads than the two-step approach. Using higher PCR annealing temperatures further increased the proportion of nonmetazoan reads in all sample types tested. The resulting V4 region amplicons were taxonomically informative even when only the forward read is analysed. This region also revealed a diversity of known and putatively parasitic lineages and a wider diversity of host-associated eukaryotes.
The Severn Estuary is a large macrotidal estuary which includes an extensive mudflat with microphytobenthos (MPB) playing a key role in the ecosystem. This study evaluated the impact of chlorination at two different dosing levels (0.05 and 0.5 mg/l as total residual oxidants, TRO, representative of potential concentrations in the mixing zone and within the cooling water systems of a power station) on a MPB community representative of the Severn Estuary. Biomass and diversity were not negatively impacted while physiology was partially affected at the beginning of the experiment, and it recovered towards the end of the experiment. Further investigations for diversity are needed to consolidate our findings. In conclusion our results show that MPB is resilient to chlorination up to a concentration of 0.5 mg/l which is much higher (>10 times) than what might be expected near the chlorinated discharges for most coastal power stations.
Diseases of bivalve molluscs caused by paramyxid parasites of the genus Marteilia have been linked to mass mortalities and the collapse of commercially important shellfish populations. Until recently, no Marteilia spp. have been detected in common cockle (Cerastoderma edule) populations in the British Isles. Molecular screening of cockles from ten sites on the Welsh coast indicates that a Marteilia parasite is widespread in Welsh C. edule populations, including major fisheries. Phylogenetic analysis of ribosomal DNA (rDNA) gene sequences from this parasite indicates that it is a closely related but different species to Marteilia cochillia, a parasite linked to mass mortality of C. edule fisheries in Spain, and that both are related to Marteilia octospora, for which we provide new rDNA sequence data. Preliminary light and transmission electron microscope (TEM) observations support this conclusion, indicating that the parasite from Wales is located primarily within areas of inflammation in the gills and the connective tissue of the digestive gland, whereas M. cochillia is found mainly within the epithelium of the digestive gland. The impact of infection by the new species, here described as Marteilia cocosarum n. sp., upon Welsh fisheries is currently unknown.
Disease poses a significant threat to aquaculture. While there are a number of factors contributing to pathogen transmission risk, movement of live fish is considered the most important. Understanding live fish movement patterns for different aquaculture sectors is therefore crucial to predicting disease occurrence and necessary for the development of effective, risk-based biosecurity, surveillance and containment policies. However, despite this, our understanding of live movement patterns of key aquaculture species, namely salmonids and cyprinids, within England and Wales remains limited. In this study, networks reflecting live fish movements associated with the cyprinid and salmonid sectors in England and Wales were constructed. The structure, composition and key attributes of each network were examined and compared to provide insight into the nature of trading patterns and connectedness, as well as highlight sites at a high risk of spreading disease. Connectivity at both site and catchment level was considered to facilitate understanding at different resolutions, providing further insight into disease outbreaks, with industry wide implications. The study highlighted that connectivity through live fish movements was extensive for both industries. The salmonid and cyprinid networks comprised 2533 and 3645 nodes, with a network density of 5.81 × 10-4 and 4.2 × 10-4, respectively. The maximum network reach of 2392 in the salmonid network was higher, both in absolute terms and as a proportion of the overall network, compared to maximum network reach of 2085 in the cyprinid network. However, in contrast, the number of sites in the cyprinid network with a network reach greater than one was 513, compared to 171 in the salmonid network. Patterns of connectivity indicated potential for more frequent yet smaller scale disease outbreaks in the cyprinid industry and less frequent but larger scale outbreaks in the salmonid industry. Further, high connectivity between river catchments within both networks was shown, posing challenges for zoning at the catchment level for the purpose of disease management. In addition to providing insight into pathogen transmission and epidemic potential within the salmonid and cyprinid networks, the study highlights the utility of network analysis, and the value of accessible, accurate live fish movement data in this context. The application of outputs from this study, and network analysis methodology, to inform future disease surveillance and control policies, both within England and Wales and more broadly, is discussed.
The Viral Hemorrhagic Septicemia Virus (VHSV) is an OIE notifiable pathogen widespread in the Northern Hemisphere that encompasses four genotypes and nine subtypes. In Europe, subtype Ia impairs predominantly the rainbow trout industry causing severe rates of mortality, while other VHSV genotypes and subtypes affect a number of marine and freshwater species, both farmed and wild. VHSV has repeatedly proved to be able to jump to rainbow trout from the marine reservoir, causing mortality episodes. The molecular mechanisms regulating VHSV virulence and host tropism are not fully understood, mainly due to the scarce availability of complete genome sequences and information on the virulence phenotype. With the scope of identifying in silico molecular markers for VHSV virulence, we generated an extensive dataset of 55 viral genomes and related mortality data obtained from rainbow trout experimental challenges. Using statistical association analyses that combined genetic and mortality data, we found 38 single amino acid polymorphisms scattered throughout the complete coding regions of the viral genome that were putatively involved in virulence of VHSV in trout. Specific amino acid signatures were recognized as being associated with either low or high virulence phenotypes. The phylogenetic analysis of VHSV coding regions supported the evolution toward greater virulence in rainbow trout within subtype Ia, and identified several other subtypes which may be prone to be virulent for this species. This study sheds light on the molecular basis for VHSV virulence, and provides an extensive list of putative virulence markers for their subsequent validation.
Intracellular microcolonies of bacteria (IMC), in some cases developing large extracellular cysts (bacterial aggregates), infecting primarily gill and digestive gland, have been historically reported in a wide diversity of economically important mollusk species worldwide, sometimes associated with severe lesions and mass mortality events. As an effort to characterize those organisms, traditionally named as Rickettsia or Chlamydia-like organisms, 1950 specimens comprising 22 mollusk species were collected over 10 countries and after histology examination, a selection of 99 samples involving 20 species were subjected to 16S rRNA gene amplicon sequencing. Phylogenetic analysis showed Endozoicomonadaceae sequences in all the mollusk species analyzed. Geographical differences in the distribution of Operational Taxonomic Units (OTUs) and a particular OTU associated with pathology in king scallop (OTU_2) were observed. The presence of Endozoicomonadaceae sequences in the IMC was visually confirmed by in situ hybridization (ISH) in eight selected samples. Sequencing data also indicated other symbiotic bacteria. Subsequent phylogenetic analysis of those OTUs revealed a novel microbial diversity associated with molluskan IMC infection distributed among different taxa, including the phylum Spirochetes, the families Anaplasmataceae and Simkaniaceae, the genera Mycoplasma and Francisella, and sulfur-oxidizing endosymbionts. Sequences like Francisella halioticida/philomiragia and Candidatus Brownia rhizoecola were also obtained, however, in the absence of ISH studies, the association between those organisms and the IMCs were not confirmed. The sequences identified in this study will allow for further molecular characterization of the microbial community associated with IMC infection in marine mollusks and their correlation with severity of the lesions to clarify their role as endosymbionts, commensals or true pathogens.
We analyse the network structure of the British salmonid aquaculture industry from the perspective of infectious disease control. We combine for the first time live fish transport (or movement) data covering England and Wales with data covering Scotland and include network layers representing potential transmission by rivers, sea water and local transmission via human or animal vectors in the immediate vicinity of each farm or fishery site. We find that 7.2% of all live fish transports cross the England-Scotland border and network analysis shows that 87% of English and Welsh nodes and 72% of Scottish nodes are reachable from cross-border connections via live fish transports alone. Consequently, from a disease-control perspective, the contact structures of England and Wales and of Scotland should not be considered in isolation. We also show that large epidemics require the live fish movement network and so control strategies targeting movements can be very effective. While there is relatively low risk of widespread epidemics on the live fish transport network alone, the potential risk is substantially amplified by the combined interaction of multiple network layers.
Background Next generation sequencing (NGS) is becoming widely used among diagnostics and research laboratories, and nowadays it is applied to a variety of disciplines, including veterinary virology. The NGS workflow comprises several steps, namely sample processing, library preparation, sequencing and primary/secondary/tertiary bioinformatics (BI) analyses. The latter is constituted by a complex process extremely difficult to standardize, due to the variety of tools and metrics available. Thus, it is of the utmost importance to assess the comparability of results obtained through different methods and in different laboratories. To achieve this goal, we have organized a proficiency test focused on the bioinformatics components for the generation of complete genome sequences of salmonid rhabdoviruses. Methods Three partners, that performed virus sequencing using different commercial library preparation kits and NGS platforms, gathered together and shared with each other 75 raw datasets which were analyzed separately by the participants to produce a consensus sequence according to their own bioinformatics pipeline. Results were then compared to highlight discrepancies, and a subset of inconsistencies were investigated more in detail. Results In total, we observed 526 discrepancies, of which 39.5% were located at genome termini, 14.1% at intergenic regions and 46.4% at coding regions. Among these, 10 SNPs and 99 indels caused changes in the protein products. Overall reproducibility was 99.94%. Based on the analysis of a subset of inconsistencies investigated more in-depth, manual curation appeared the most critical step affecting sequence comparability, suggesting that the harmonization of this phase is crucial to obtain comparable results. The analysis of a calibrator sample allowed assessing BI accuracy, being 99.983%. Conclusions We demonstrated the applicability and the usefulness of BI proficiency testing to assure the quality of NGS data, and recommend a wider implementation of such exercises to guarantee sequence data uniformity among different virology laboratories.