In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.
Background: Antimicrobial resistance (AMR) poses significant risks to human and animal health, while the environment can contribute to its spread. National AMR surveillance programs are pivotal for assessing AMR prevalence, trends, and intervention outcomes; however, integrating advanced surveillance tools can be difficult. This pilot study, conducted by FAO ECTAD Indonesia and DGLAHS, the Indonesian Ministry of Agriculture, evaluated the costs and benefits of integrating the Nanopore MinION, Illumina MiSeq, and Sensititre system into a culture-based slaughterhouse-river surveillance system. Methods: Water samples were collected from six chicken slaughterhouses and adjacent rivers (pre- and post-treatment effluent, upstream, and downstream). Culture-based ESBL and general E. coli concentrations were estimated via the WHO Tricycle Protocol, while isolates (n = 42) were sequenced (MinION, MiSeq) and antimicrobial susceptibility testing conducted (Sensititre). Results: The Tricycle Protocol results provided estimates of effluent and river concentrations of ESBL and general E. coli identifying ESBL-to-general E. coli ratios of 13.8% and 6.2%, respectively. Compared to hybrid sequencing assemblies, MinION had a higher concordance than MiSeq for ARG identification (98%), virulence genes (96%), and locations for both (predominately plasmids). Furthermore, MinION concordance with Sensititre AST was 91%. Conclusions: Cost-benefit comparisons suggest sequencing can complement culture-based methods but is dependent on the value placed on the additional information gained.
Case-based infectious disease surveillance is fundamental to public health, but is resource-intensive, logistically complex, and prone to sampling bias. Wastewater testing and sequencing have increasingly been used for population-scale monitoring of pathogen dynamics, including in low-resource settings. Broader adoption of wastewater genomic surveillance, however, is limited by a lack of flexibility across sequencing platforms and approaches, and adaptability to additional pathogens. Here, we describe “Freyja 2”, an integrated bioinformatics tool enabling robust real-time inference of pathogen lineage prevalence and growth dynamics from wastewater and other complex samples. In Freyja 2, we develop new methods for estimating lineage prevalence and growth rates, and demonstrate robustness across common sequencing platforms and to low genomic coverage. By incorporating global pathogen data streams, we extend Freyja 2 to support multi-pathogen surveillance. We demonstrate tracking of multiple recent or ongoing public health emergencies, including COVID-19, mpox, and H5N1 influenza, revealing unreported diversity and lineage co-circulation.
The capacity for pathogen genomics in public health expanded rapidly during the coronavirus disease 2019 (COVID- 19) pan-demic, but many public health laboratories did not have the infrastructure in place to handle the vast amount of severe acute respiratory syndrome coronavirus 2 (SARS-CoV- 2) sequence data generated. The California Department of Public Health, in partnership with Theiagen Genomics, was an early adopter of cloud -based resources for bioinformatics and genomic epi-demiology, resulting in the creation of a SARS- CoV- 2 genomic surveillance system that combined the efforts of more than 40 sequencing laboratories across government, academia and industry to form California COVIDNet, California's SARS- CoV- 2 Whole-Genome Sequencing Initiative. Open-source bioinformatics workflows, ongoing training sessions for the public health workforce, and automated data transfer to visualization tools all contributed to the success of California COVIDNet. While chal-lenges remain for public health genomic surveillance worldwide, California COVIDNet serves as a framework for a scaled and successful bioinformatics infrastructure that has expanded beyond SARS- CoV- 2 to other pathogens of public health importance,
The 2022 multicountry mpox outbreak concurrent with the ongoing Coronavirus Disease 2019 (COVID-19) pandemic further highlighted the need for genomic surveillance and rapid pathogen whole-genome sequencing. While metagenomic sequencing approaches have been used to sequence many of the early mpox infections, these methods are resource intensive and require samples with high viral DNA concentrations. Given the atypical clinical presentation of cases associated with the outbreak and uncertainty regarding viral load across both the course of infection and anatomical body sites, there was an urgent need for a more sensitive and broadly applicable sequencing approach. Highly multiplexed amplicon-based sequencing (PrimalSeq) was initially developed for sequencing of Zika virus, and later adapted as the main sequencing approach for Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2). Here, we used PrimalScheme to develop a primer scheme for human monkeypox virus that can be used with many sequencing and bioinformatics pipelines implemented in public health laboratories during the COVID-19 pandemic. We sequenced clinical specimens that tested presumptively positive for human monkeypox virus with amplicon-based and metagenomic sequencing approaches. We found notably higher genome coverage across the virus genome, with minimal amplicon drop-outs, in using the amplicon-based sequencing approach, particularly in higher PCR cycle threshold (Ct) (lower DNA titer) samples. Further testing demonstrated that Ct value correlated with the number of sequencing reads and influenced the percent genome coverage. To maximize genome coverage when resources are limited, we recommend selecting samples with a PCR Ct below 31 Ct and generating 1 million sequencing reads per sample. To support national and international public health genomic surveillance efforts, we sent out primer pool aliquots to 10 laboratories across the United States, United Kingdom, Brazil, and Portugal. These public health laboratories successfully implemented the human monkeypox virus primer scheme in various amplicon sequencing workflows and with different sample types across a range of Ct values. Thus, we show that amplicon-based sequencing can provide a rapidly deployable, cost-effective, and flexible approach to pathogen whole-genome sequencing in response to newly emerging pathogens. Importantly, through the implementation of our primer scheme into existing SARS-CoV-2 workflows and across a range of sample types and sequencing platforms, we further demonstrate the potential of this approach for rapid outbreak response.
The sharing of genome sequences in online data repositories allows for large scale analyses of specific genes or gene families. This can result in the detection of novel gene subtypes as well as the development of improved detection methods. Here, we used publicly available WGS data to detect a novel Stx subtype, Stx2n in two clinical E. coli strains isolated in the USA. During this process, additional Stx2 subtypes were detected; six Stx2j, one Stx2m strain, and one Stx2o, were all analyzed for variability from the originally described subtypes. Complete genome sequences were assembled from short- or long-read sequencing and analyzed for serotype, and ST types. The WGS data from Stx2n- and Stx2o-producing STEC strains were further analyzed for virulence genes pro-phage analysis and phage insertion sites. Nucleotide and amino acid maximum parsimony trees showed expected clustering of the previously described subtypes and a clear separation of the novel Stx2n subtype. WGS data were used to design OMNI PCR primers for the detection of all known stx1 (283 bp amplicon), stx2 (400 bp amplicon), intimin encoded by eae (221 bp amplicon), and stx2f (438 bp amplicon) subtypes. These primers were tested in three different laboratories, using standard reference strains. An analysis of the complete genome sequence showed variability in serogroup, virulence genes, and ST type, and Stx2 pro-phages showed variability in size, gene composition, and phage insertion sites. The strains with Stx2j, Stx2m, Stx2n, and Stx2o showed toxicity to Vero cells. Stx2j carrying strain, 2012C-4221, was induced when grown with sub-inhibitory concentrations of ciprofloxacin, and toxicity was detected. Taken together, these data highlight the need to reinforce genomic surveillance to identify the emergence of potential new Stx2 or Stx1 variants. The importance of this surveillance has a paramount impact on public health. Per our description in this study, we suggest that 2017C-4317 be designated as the Stx2n type-strain.
We have adopted an open bioinformatics ecosystem to address the challenges of bioinformatics implementation in public health laboratories (PHLs). Bioinformatics implementation for public health requires practitioners to undertake standardized bioinformatic analyses and generate reproducible, validated and auditable results. It is essential that data storage and analysis are scalable, portable and secure, and that implementation of bioinformatics fits within the operational constraints of the laboratory. We address these requirements using Terra, a web-based data analysis platform with a graphical user interface connecting users to bioinformatics analyses without the use of code. We have developed bioinformatics workflows for use with Terra that specifically meet the needs of public health practitioners. These Theiagen workflows perform genome assembly, quality control, and characterization, as well as construction of phylogeny for insights into genomic epidemiology. Additonally, these workflows use open-source containerized software and the WDL workflow language to ensure standardization and interoperability with other bioinformatics solutions, whilst being adaptable by the user. They are all open source and publicly available in Dockstore with the version-controlled code available in public GitHub repositories. They have been written to generate outputs in standardized file formats to allow for further downstream analysis and visualization with separate genomic epidemiology software. Testament to this solution meeting the requirements for bioinformatic implementation in public health, Theiagen workflows have collectively been used for over 5 million sample analyses in the last 2 years by over 90 public health laboratories in at least 40 different countries. Continued adoption of technological innovations and development of further workflows will ensure that this ecosystem continues to benefit PHLs.
In the United States, reports of Salmonella enterica carrying mcr-1 remain rare in humans, but when observed, the infection is often associated with travel. Here, we report 14 mcr-1-positive Salmonella enterica isolates from patients in the United States that reported travel to the Dominican Republic within the 12 months before illness.
We have developed and implemented an undergraduate microbiology course in which students isolate, characterize, and perform whole genome assembly and analysis of Salmonella enterica from stream sediments and poultry litter. In the development of the course and over three semesters, successive teams of undergraduate students collected field samples and performed enrichment and isolation techniques specific for the detection of S. enterica. Eighty-eight strains were confirmed using standard microbiological methods and PCR of the invA gene. The isolates' genomes were Illumina-sequenced by the Center for Food Safety and Applied Nutrition at the FDA and the Virginia state Division of Consolidated Laboratory Services as part of the GenomeTrakr program. Students used GalaxyTrakr and other web- and non-web-based platforms and tools to perform quality control on raw and assembled sequence data, assemble, and annotate genomes, identify antimicrobial resistance and virulence genes, putative plasmids, and other mobile genetic elements. Strains with putative plasmid-borne antimicrobial resistance genes were further sequenced by students in our research lab using the Oxford Nanopore MinIONTM platform. Strains of Salmonella that were isolated include human infectious serotypes such as Typhimurium and Infantis. Over 31 of the isolates possessed antibiotic resistance genes, some of which were located on large, multidrug resistance plasmids. Plasmid pHJ-38, identified in a Typhimurium isolate, is an apparently self-transmissible 183 kb IncA/C2 plasmid that possesses multiple antimicrobial resistance and heavy-metal resistance genes. Plasmid pFHS-02, identified in an Infantis isolate, is an apparently self-transmissible 303 kb IncF1B plasmid that also possesses numerous heavy-metal and antimicrobial resistance genes. Using direct and indirect measures to assess student outcomes, results indicate that course participation contributed to cognitive gains in relevant content knowledge and research skills such as field sampling, molecular techniques, and computational analysis. Furthermore, participants self-reported a deeper interest in scientific research and careers as well as psychosocial outcomes (e.g., sense of belonging and self-efficacy) commonly associated with student success and persistence in STEM. Overall, this course provided a powerful combination of field, wet lab, and computational biology experiences for students, while also providing data potentially useful in pathogen surveillance, epidemiological tracking, and for the further study of environmental reservoirs of S. enterica.
Laboratories that run Whole Genome Sequencing (WGS) produce a tremendous amount of data, up to 10 gigabytes for some common instruments. There is a need to standardize the quality assurance and quality control process (QA/QC). Therefore we have created SneakerNet to automate the QA/QC for WGS.
Salmonella enterica subsp. enterica serovar Corvallis is commonly reported in avian populations and avian by-products. We report the draft genome sequence of a multidrug-resistant S. Corvallis strain (NPHL 15376). To our knowledge, this is the first reported case of this serovar isolated from human blood in the United States.
4. CONCLUSIONS • Transmissible plasmids affect environmental ecosystems via exchange and recombination of antibiotic resistance genes. • Exchange occurs between native bacterial populations and introduced fecal pathogens selected for resistance in farm animals. • Public Health Relevance: Native bacteria in aquatic and soil habitats may act as incubators and sites for recombination of genes that are subsequently transferred to human pathogens. Plasmid 1-1 Plasmid 1-20