Harmful algal blooms have important implications for the health, functioning, and services of aquatic ecosystems. Our ability to detect and monitor these events is often challenged by the lack of rapid and cost-effective methods to identify bloom-forming organisms and their potential for toxin production. Here, we developed and applied a combination of DNA barcoding and Next Generation Sequencing (NGS) for the rapid assessment of phytoplankton community composition with a focus on two important indicators of ecosystem health: toxigenic bloom-forming cyanobacteria and impaired planktonic biodiversity. To develop this molecular toolset for identification of cyanobacterial and algal species present in HABs (harmful algal blooms), hereafter called HAB-ID, we achieved three goals: creating a validated reference database, optimizing molecular protocols, and developing original bioinformatics pipeline tailored to uncertainty of algal taxonomy. The BOLD (Barcode of Life Data System) 16S reference database from cultures of 211 cyanobacterial and algal strains representing 102 species with particular focus on bloom and toxin producing taxa was constructed with Sanger sequencing and further refined using Single Molecule Real Time Sequencing (SMRT-sequencing). Using the new reference database of 16S rDNA sequences and constructed mock communities of mixed strains for protocol validation, we developed new NGS primer sets which can recover 16S from both cyanobacteria and eukaryotic algal chloroplasts. We also developed DNA extraction protocols for cultured algal strains and environmental samples, which match commercial kit performance and offer a cost-efficient solution for large scale ecological assessments of harmful blooms while giving benefits of reproducibility and increased accessibility. Our innovative bioinformatics pipeline was designed to handle low taxonomic resolution for problematic genera of cyanobacteria such as the Anabaena-Aphanizomenon-Dolichospermum complex, two clusters of Anabaena (I and II), Planktothrix and Microcystis. This newly developed HAB-ID toolset was further validated by applying it to assess cyanobacterial and algal composition in field samples from waterbodies with recurrent HABs events.
This paper presents Sequenoscope a bioinformatics pipeline for analyzing Oxford Nanopore Technologies (ONT) adaptive sampling sequencing data. Sequenoscope features three main modules: filter_ONT for filtering raw reads and creating a FASTQ file with a subset of reads for further analyses, analyze for generating sequencing and read mapping statistics against the provided reference taxon sequences, and plot for interactive data summarization, comparison, and visualization between two datasets/test conditions. Here we demonstrate the ability of the pipeline to analyze ONT adaptive sampling sequence data and provide examples of the outputs users can expect using data we generated. Adaptive sampling was performed on two ZymoBIOMICS Microbial Community DNA Standards (Cat# D6311 and D6306) with targeted depletions of Listeria monocytogenes. By comparing the test and control experimental data in FASTQ file from the sequencing runs, Sequenoscope showed that depletion of L. monocytogenes was successful by providing users with parameters to compare such as taxon coverage, read length, and types of pore-level decisions made during sequencing. Although Sequenoscope was designed for ONT adaptive sampling data analysis, it can also be used to compare any two experimental conditions and supports short-read data from other sequencing platforms.
Dissemination of antimicrobial resistance (AMR) is a growing global public health burden. The aim of this study was to characterize AMR plasmid transmissions within a tertiary care hospital and identify relevant AMR plasmid transmission pathways. During an 18-month observation period, 540 clinical gram-negative multidrug-resistant bacterial (MDRB) isolates were collected during routine hospital surveillance and subjected to Pacific Biosciences long-read whole genome sequencing. Potential clonal transmissions were determined based on core genome multilocus sequence typing (cgMLST), and plasmid transmissions were detected using a novel real-time applicable tool for plasmid transmission detection. Potential transmissions were validated using epidemiological data. Among the 471 eligible MDRB isolates, we detected 1,539 plasmids; 84.41% of these were circularized. We identified 38 potential clonal transmissions in 24 clusters based on cgMLST and 121 potential plasmid transmissions in 24 clusters containing genetically related AMR plasmids. Among the latter clusters, 10 contained different multilocus sequence types (involving 2-38 isolates, median: 3 isolates), and nine contained multiple species (2-18 isolates, median: 4). Epidemiological data confirmed 19 clonal transmissions (in seven clusters) and an additional 12 plasmid transmissions (within eight plasmid clusters). Among these, we identified seven cases of intra-host and five patient-to-patient plasmid transmissions. We demonstrate that intra-host and patient-to-patient transmissions of AMR plasmids can be identified by combining long-read sequencing with real-time applicable tools during routine molecular surveillance. In addition, our study highlights that more than a decade of bacterial genomic surveillance missed at least one-third of all AMR transmission events due to plasmids. IMPORTANCE:Antimicrobial resistance (AMR) poses a significant threat to human health. Most AMR determinants are encoded extra-chromosomally on plasmids. Although current infection control strategies primarily focus on clonal transmission of multidrug-resistant bacteria, until today, AMR plasmid transmission routes are neither understood nor analyzed in the hospital setting. In our study, we simultaneously determined both clonal, that is, based on chromosomes, and AMR plasmid transmissions during routine molecular surveillance by combining long-read sequencing with a novel real-time applicable software tool and validated all potential transmission events with epidemiological data. Our analysis determined not only the yet unknown plasmid transmissions within healthcare facilities or within the community but also resulted, in addition to the clonal transmissions, in at least a third more transmissions due to AMR plasmids.
Ingestion of antibiotic-resistant bacteria following antibiotic treatments may lead to the transfer of antimicrobial resistance genes (ARGs) within a disturbed gut microbiota. However, it remains unclear whether and how microbes present in food matrices influence ARG transfer. Thus, a previously established mouse model, which demonstrated the conjugative transfer of a multi-drug resistance plasmid (pIncA/C) from Salmonella Heidelberg (donor) to Salmonella Typhimurium (recipient), was used to assess the effects of food-borne microbes derived from fresh carrots on pIncA/C transfer. Mice were pre-treated with ampicillin, streptomycin, sulfamethazine, or left untreated as a control to facilitate bacterial colonization. Contrary to previous findings where high-density colonization of the donor and recipient bacteria occurred in the absence of food-borne microbes, the presence of these microbes resulted in a low abundance of S. Typhimurium and no detection of S. Typhimurium transconjugants in the fecal samples from any of the mice. However, in mice pre-treated with streptomycin, a significant reduction in microbial species richness allowed for the significant enrichment of Enterobacteriaceae and pIncA/C transfer to bacteria from the genera Escherichia, Enterobacter, Citrobacter, and Proteus. These findings suggest that food-borne microbes may enhance ARG dissemination by influencing the population dynamics of bacterial hosts within a pre-disturbed gut microbiome.
Shiga toxin-producing Escherichia coli (STEC) can cause severe clinical disease in humans, particularly in young children. Recent advances have led to greater availability of sequencing technologies. We sought to use whole genome sequencing data to identify the presence or absence of known virulence factors in all clinical isolates submitted to our laboratory from Southern Alberta dated 2020–2022 and correlate these virulence factors with clinical outcomes obtained through chart review. Overall, the majority of HUS and hospitalizations were seen in patients with O157:H7 serotypes, and HUS cases were primarily in young children. The frequency of virulence factors differed between O157:H7 and non-O157 serotypes. Within the O157:H7 cases, certain virulence factors, including espP, espX1, and katP, were more frequent in HUS cases. The number of samples was too low to determine statistical significance.
Plasmids are the primary vector for horizontal transfer of antimicrobial resistance (AMR) within bacterial populations. We applied the MOB-suite, a toolset for reconstructing and typing plasmids, to 150 767 publicly available Salmonella whole-genome sequencing samples covering 1204 distinct serovars to produce a large-scale population survey of plasmids based on the MOB-suite plasmid nomenclature. Reconstruction yielded 183 017 plasmids representing 1044 primary MOB-clusters and 830 potentially novel MOB-clusters. Replicon and relaxase typing were able to type 83.4 and 58 % of plasmids, respectively, compared to 99.9 % for MOB-clusters. Within this work, we developed an approach to characterize the horizonal transfer of MOB-clusters and AMR genes across different serotypes, as well as the diversity of MOB-cluster associations with AMR genes. Aggregating conjugative mobility predictions provided by the MOB-suite and their corresponding serovar entropy demonstrated that non-mobilizable plasmids were associated with fewer serotypes compared to mobilizable or conjugative MOB-clusters. The host-range predictions for MOB-clusters also showed differences between the mobility classes, with mobilizable MOB-clusters accounting for 88.3 % of the multi-phyla (broad-host-range) predictions compared to 3 and 8.6 % for conjugative and non-mobilizable, respectively. A total of 296 (22 %) of identified MOB-clusters were associated with at least one resistance gene, indicating that the majority of Salmonella plasmids are not involved in AMR dissemination. Shannon entropy analysis of horizontal transfer of AMR genes across serovars and MOB-clusters demonstrated that horizonal transfer of genes is higher between serovars compared to transfer between different MOB-clusters. In addition to the population structure characterization based on primary MOB-clusters, we characterized a multi-plasmid outbreak responsible for the global dissemination of bla CMY-2 across different serotypes using higher resolution MOB-suite secondary cluster codes. The plasmid characterization approach developed here can be applied to different organisms to identify plasmids and genes which pose high risks for horizontal transfer.
Enterococcus faecium is a ubiquitous opportunistic pathogen that is exhibiting increasing levels of antimicrobial resistance (AMR). Many of the genes that confer resistance and pathogenic functions are localized on mobile genetic elements (MGEs), which facilitate their transfer between lineages. Here, features including resistance determinants, virulence factors and MGEs were profiled in a set of 1273 E. faecium genomes from two disparate geographic locations (in the UK and Canada) from a range of agricultural, clinical and associated habitats. Neither lineages of E. faecium , type A and B, nor MGEs are constrained by geographic proximity, but our results show evidence of a strong association of many profiled genes and MGEs with habitat. Many features were associated with a group of clinical and municipal wastewater genomes that are likely forming a new human-associated ecotype within type A. The evolutionary dynamics of E. faecium make it a highly versatile emerging pathogen, and its ability to acquire, transmit and lose features presents a high risk for the emergence of new pathogenic variants and novel resistance combinations. This study provides a workflow for MGE-centric surveillance of AMR in Enterococcus that can be adapted to other pathogens.
Hierarchical genotyping approaches can provide insights into the source, geography and temporal distribution of bacterial pathogens. Multiple hierarchical SNP genotyping schemes have previously been developed so that new isolates can rapidly be placed within pre-computed population structures, without the need to rebuild phylogenetic trees for the entire dataset. This classification approach has, however, seen limited uptake in routine public health settings due to analytical complexity and the lack of standardized tools that provide clear and easy ways to interpret results. The BioHansel tool was developed to provide an organism-agnostic tool for hierarchical SNP-based genotyping. The tool identifies split k-mers that distinguish predefined lineages in whole genome sequencing (WGS) data using SNP-based genotyping schemes. BioHansel uses the Aho-Corasick algorithm to type isolates from assembled genomes or raw read sequence data in a matter of seconds, with limited computational resources. This makes BioHansel ideal for use by public health agencies that rely on WGS methods for surveillance of bacterial pathogens. Genotyping results are evaluated using a quality assurance module which identifies problematic samples, such as low-quality or contaminated datasets. Using existing hierarchical SNP schemes for Mycobacterium tuberculosis and Salmonella Typhi, we compare the genotyping results obtained with the k-mer-based tools BioHansel and SKA, with those of the organism-specific tools TBProfiler and genotyphi, which use gold-standard reference-mapping approaches. We show that the genotyping results are fully concordant across these different methods, and that the k-mer-based tools are significantly faster. We also test the ability of the BioHansel quality assurance module to detect intra-lineage contamination and demonstrate that it is effective, even in populations with low genetic diversity. We demonstrate the scalability of the tool using a dataset of ~8100 S. Typhi public genomes and provide the aggregated results of geographical distributions as part of the tool's output. BioHansel is an open source Python 3 application available on PyPI and Conda repositories and as a Galaxy tool from the public Galaxy Toolshed. In a public health context, BioHansel enables rapid and high-resolution classification of bacterial pathogens with low genetic diversity.
Escherichia coli is a priority foodborne pathogen of public health concern and phenotypic serotyping provides critical information for surveillance and outbreak detection activities. Public health and food safety laboratories are increasingly adopting whole-genome sequencing (WGS) for characterizing pathogens, but it is imperative to maintain serotype designations in order to minimize disruptions to existing public health workflows. Multiple in silico tools have been developed for predicting serotypes from WGS data, including SRST2, SerotypeFinder and EToKi EBEis, but these tools were not designed with the specific requirements of diagnostic laboratories, which include: speciation, input data flexibility (fasta/fastq), quality control information and easily interpretable results. To address these specific requirements, we developed ECTyper (https://github.com/phac-nml/ ecoli_serotyping) for performing both speciation within Escherichia and Shigella, and in silico serotype prediction. We compared the serotype prediction performance of each tool on a newly sequenced panel of 185 isolates with confirmed phenotypic serotype information. We found that all tools were highly concordant, with 92-97 % for O-antigens and 98-100 % for H-antigens, and ECTyper having the highest rate of concordance. We extended the benchmarking to a large panel of 6954 publicly available E. coli genomes to assess the performance of the tools on a more diverse dataset. On the public data, there was a considerable drop in concordance, with 75-91 % for O-antigens and 62-90 % for H-antigens, and ECTyper and SerotypeFinder being the most concordant. This study highlights that in silico predictions show high concordance with phenotypic serotyping results, but there are notable differences in tool performance. ECTyper provides highly accurate and sensitive in silico serotype predictions, in addition to speciation, and is designed to be easily incorporated into bioinformatic workflows.
Laboratory-based wastewater surveillance for SARS-CoV-2, the causative agent of the ongoing COVD-19 pandemic, can be conducted using RT-qPCR-based screening of municipal wastewater samples. Although it provides rapid viral detection and can inform SARS-CoV-2 abundance in wastewater, this approach lacks the resolution required for viral genotyping and does not support tracking of viral genome evolution. The recent emergence of several variants of concern, a result of mutations across the genome including the accrual of important mutations within the viral spike glycoprotein, has highlighted the need for a method capable of detecting the cohort of mutations associated with these and newly emerging genotypes. Here we provide an innovative methodology for the recovery of a near-complete SARS-CoV-2 sequence from a wastewater sample collected from across Canadian municipalities including one that experienced a significant outbreak attributable to the SARS-CoV-2 B.1.1.7 variant of concern. Our results demonstrate that a combined interrogation of genome consensus-level sequences and alternative alleles enables the identification of a SARS-CoV-2 variant of concern and the detection of a new allele within a viral accessory gene that may be representative of a recently evolved B.1.1.7 sublineage.
Ingestion of food- or waterborne antibiotic-resistant bacteria may lead to dissemination of antibiotic resistance genes (ARGs) in the gut microbiota. The gut microbiota often suffers from various disturbances. It is not clear whether and how disturbed microbiota may affect ARG mobility under antibiotic treatments. For proof of concept, in the presence or absence of streptomycin pre-treatment, mice were inoculated orally with a beta-lactam-susceptible Salmonella enterica serovar Heidelberg clinical isolate (recipient) and a beta-lactam resistant Escherichia coli O80:H26 isolate (donor) carrying a bla(CMY-2) gene on an IncI2 plasmid. Immediately following inoculation, mice were treated with or without ampicillin in drinking water for 7 days. Faeces were sampled, donor, recipient and transconjugant were enumerated, bla(CMY-2) abundance was determined by quantitative PCR, faecal microbial community composition was determined by 16S rRNA amplicon sequencing and cecal samples were observed histologically for evidence of inflammation. In faeces of mice that received streptomycin pre-treatment, the donor abundance remained high, and the abundance of S. Heidelberg transconjugant and the relative abundance of Enterobacteriaceae increased significantly during the ampicillin treatment. Co-blooming of the donor, transconjugant and commensal Enterobacteriaceae in the inflamed intestine promoted significantly (P<0.05) higher and possibly wider dissemination of the bla CMY-2 gene in the gut microbiota of mice that received the combination of streptomycin pre-treatment and ampicillin treatment (Str-Amp) compared to the other mice. Following cessation of the ampicillin treatment, faecal shedding of S. Heidelberg transconjugant persisted much longer from mice in the Str-Amp group compared to the other mice. In addition, only mice in the Str-Amp group shed a commensal E. coli O2:H6 transconjugant, which carries three copies of the bla(CMY-2) gene, one on the IncI2 plasmid and two on the chromosome. The findings highlight the significance of pre-existing gut microbiota for ARG dissemination and persistence during and following antibiotic treatments of infectious diseases.
Cryptosporidium is a protozoan parasite that is transmitted to both humans and animals through zoonotic or anthroponotic means. When a host is infected with this parasite, it causes a gastrointestinal disease known as cryptosporidiosis. To understand the transmission dynamics of Cryptosporidium, the small subunit (SSU or 18S) rRNA and gp60 genes are commonly studied through PCR analysis and conventional Sanger sequencing. However, analyzing sequence chromatograms manually is both time consuming and prone to human error, especially in the presence of poorly resolved, heterozygous peaks and the absence of a validated database. For this study, we developed a Cryptosporidium genotyping tool, called CryptoGenotyper, which has the capability to read raw Sanger sequencing data for the two common Cryptosporidium gene targets (SSU rRNA and gp60) and classify the sequence data into standard nomenclature. The CryptoGenotyper has the capacity to perform quality control and properly classify sequences using a high quality, manually curated reference database, saving users' time and removing bias during data analysis. The incorporated heterozygous base calling algorithms for the SSU rRNA gene target resolves double peaks, therefore recovering data previously classified as inconclusive. The CryptoGenotyper successfully genotyped 99.3% (428/431) and 95.1% (154/162) of SSU rRNA chromatograms containing single and mixed sequences, respectively, and correctly subtyped 95.6% (947/991) of gp60 chromatograms without manual intervention. This new, user-friendly tool can provide both fast and reproducible analyses of Sanger sequencing data for the two most common Cryptosporidium gene targets.
Ingestion of food- or waterborne antibiotic-resistant bacteria may lead to the dissemination of antibiotic-resistance genes in the gut microbiota and the development of antibiotic-resistant bacterial infection, a significant threat to animal and public health. Food or water may be contaminated with multiple resistant bacteria, but animal models on gene transfer were mainly based on single-strain infections. In this study, we investigated the mobility of β-lactam resistance following infection with single- versus multi-strain of resistant bacteria under ampicillin treatment. We characterized three bacterial strains isolated from food-animal production systems, Escherichia coli O80:H26 and Salmonella enterica serovars Bredeney and Heidelberg. Each strain carries at least one conjugative plasmid that encodes a β-lactamase. We orally infected mice with each or all three bacterial strain(s) in the presence or absence of ampicillin treatment. We assessed plasmid transfer from the three donor bacteria to an introduced E. coli CV601gfp recipient in the mouse gut, and evaluated the impacts of the bacterial infection on gut microbiota and gut health. In the absence of ampicillin treatment, none of the donor or recipient bacteria established in the normal gut microbiota and plasmid transfer was not detected. In contrast, the ampicillin treatment disrupted the gut microbiota and enabled S. Bredeney and Heidelberg to colonize and transfer their plasmids to the E. coli CV601gfp recipient. E. coli O80:H26 on its own failed to colonize the mouse gut. However, during co-infection with the two Salmonella strains, E. coli O80:H26 colonized and transferred its plasmid to the E. coli CV601gfp recipient and a residential E. coli O2:H6 strain. The co-infection significantly increased plasmid transfer frequency, enhanced Proteobacteria expansion and resulted in inflammation in the mouse gut. Our findings suggest that single-strain infection models for evaluating in vivo gene transfer may underrepresent the consequences of multi-strain infections following the consumption of heavily contaminated food or water.
Supplemental figure and tables from the manuscript: Universal whole-sequence based plasmid typing and utility to prediction of host-range and epidemiological surveillance
bioRxiv - the preprint server for biology, operated by Cold Spring Harbor Laboratory, a research and educational institution.
The reliable taxonomic identification of organisms through DNA sequence data requires a well parameterized library of curated reference sequences. However, it is estimated that just 15% of described animal species are represented in public sequence repositories. To begin to address this deficiency, we provide DNA barcodes for 1,500,003 animal specimens collected from 23 terrestrial and aquatic ecozones at sites across Canada, a nation that comprises 7% of the planet's land surface. In total, 14 phyla, 43 classes, 163 orders, 1123 families, 6186 genera, and 64,264 Barcode Index Numbers (BINs; a proxy for species) are represented. Species-level taxonomy was available for 38% of the specimens, but higher proportions were assigned to a genus (69.5%) and a family (99.9%). Voucher specimens and DNA extracts are archived at the Centre for Biodiversity Genomics where they are available for further research. The corresponding sequence and taxonomic data can be accessed through the Barcode of Life Data System, GenBank, the Global Biodiversity Information Facility, and the Global Genome Biodiversity Network Data Portal.
The reliable taxonomic identification of organisms through DNA sequence data requires a well parameterized library of curated reference sequences. However, it is estimated that just 15% of described animal species are represented in public sequence repositories. To begin to address this deficiency, we provide DNA barcodes for 1,500,003 animal specimens collected from 23 terrestrial and aquatic ecozones at sites across Canada, a nation that comprises 7% of the planet’s land surface. In total, 14 phyla, 43 classes, 163 orders, 1123 families, 6186 genera, and 64,264 Barcode Index Numbers (BINs; a proxy for species) are represented. Species-level taxonomy was available for 38% of the specimens, but higher proportions were assigned to a genus (69.5%) and a family (99.9%). Voucher specimens and DNA extracts are archived at the Centre for Biodiversity Genomics where they are available for further research. The corresponding sequence and taxonomic data can be accessed through the Barcode of Life Data System, GenBank, the Global Biodiversity Information Facility, and the Global Genome Biodiversity Network Data Portal. ![Figure][1] [1]: pending:yes
Environmental DNA (eDNA) is an effective approach for detecting vertebrates and plants, especially in aquatic ecosystems, but prior studies have largely examined eDNA in cool temperate settings. By contrast, this study employs eDNA to survey the fish fauna in tropical Lake Bacalar (Mexico) with the additional goal of assessing the possible presence of invasive fishes, such as Amazon sailfin catfish and tilapia. Sediment and water samples were collected from eight stations in Lake Bacalar on three occasions over a 4-month interval. Each sample was stored in the presence or absence of lysis buffer to compare eDNA recovery. Short fragments (184-187 bp) of the cytochrome c oxidase I (COI) gene were amplified using fusion primers and then sequenced on Ion Torrent PGM or S5 before their source species were determined using a custom reference sequence database constructed on BOLD. In total, eDNA sequences were recovered from 75 species of vertebrates including 47 fishes, 15 birds, 7 mammals, 5 reptiles, and 1 amphibian. Although all species are known from this region, six fish species represent new records for the study area, while two require verification. Sequences for five species (2 birds, 2 mammals, 1 reptile) were only detected from sediments, while sequences from 52 species were only recovered from water. Because DNA from the Amazon sailfin catfish was not detected, we used a mock eDNA experiment to confirm our methods would enable its detection. In summary, we developed protocols that recovered eDNA from tropical oligotrophic aquatic ecosystems and confirmed their effectiveness in detecting fishes and diverse species of vertebrates.
Harmful algal blooms have important implications for the health, functioning and services of aquatic ecosystems. Our ability to detect and monitor these events is often challenged by the lack of rapid and cost-effective methods to identify bloom-forming organisms and their potential for toxin production, Here, we developed and applied a combination of DNA barcoding and Next Generation Sequencing (NGS) for the rapid assessment of phytoplankton community composition with focus on two important indicators of ecosystem health: toxigenic bloom-forming cyanobacteria and impaired planktonic biodiversity. To develop this molecular toolset for identification of cyanobacterial and algal species present in HABs (Harmful Algal Blooms), hereafter called HAB-ID, we optimized NGS protocols, applied a newly developed bioinformatics pipeline and constructed a BOLD (Barcode of Life Data System) 16S reference database from cultures of 203 cyanobacterial and algal strains representing 101 species with particular focus on bloom and toxin producing taxa. Using the new reference database of 16S rDNA sequences and constructed mock communities of mixed strains for protocol validation we developed new NGS primer set which can recover 16S from both cyanobacteria and eukaryotic algal chloroplasts. We also developed DNA extraction protocols for cultured algal strains and environmental samples, which match commercial kit performance and offer a cost-efficient solution for large scale ecological assessments of harmful blooms while giving benefits of reproducibility and increased accessibility. Our bioinformatics pipeline was designed to handle low taxonomic resolution for problematic genera of cyanobacteria such as the Anabaena-Aphanizomenon-Dolichospermum species complex, two clusters of Anabaena (I and II), Planktothrix and Microcystis. This newly developed HAB-ID toolset was further validated by applying it to assess cyanobacterial and algal composition in field samples from waterbodies with recurrent HABs events.