Shiga toxins in Shiga toxin-producing Escherichia coli (STEC) infections are responsible for bloody diarrhea and serious complications such as hemolytic uremic syndrome. Two types of toxins have been identified: Shiga toxin type 1 (Stx1) and the immunologically distinct Shiga toxin type 2 (Stx2). Numerous STEC that express toxin variants within those two major groups have been characterized, some of which confer unique biological properties. These variants are grouped within the Stx1 or Stx2 types and are often assigned subtypes to indicate they are not identical in sequence or phenotype. Because serious outcomes of infection are associated with certain Stx subtypes, there is a need to assign Stx sequences to the proper subtype. Here, we report a comprehensive analysis of known Stx subtypes and describe a scheme and algorithm to classify the Stx toxins and Stx operon sequences by phylogenetic sequence-based relatedness of the holotoxin conforming to historical type designations. We used this analysis to develop the free and open-source StxTyper software and database that implements this typing algorithm; StxTyper is also integrated into AMRFinderPlus 4.0 at the National Center for Biotechnology Information (NCBI). We validated and compared the results to PCR assays on a set of isolates and current state of the art surveillance methods used in the Danish public health system, and we summarize StxTyper results for over 111,000 publicly available E. coli genomes. We further propose a procedure to coordinate naming and identification for newly discovered and characterized Stx subtypes.
Acinetobacter baumannii is a clinically important, Gram-negative pathogen responsible for a wide variety of nosocomial and community-acquired infections. Antibiotic resistance is a serious concern, as the organism has a wide variety of intrinsic resistance mechanisms, including chromosomal class C (blaADC) and D (blaOXA-51 family) β-lactamases, and the ability to readily acquire additional β-lactamases. Surveillance studies can reveal the diversity and distribution of β-lactamase alleles, but are difficult and expensive to conduct. Herein, we describe an approach using publicly available data derived from whole genome sequences, to explore the diversity and distribution of β-lactamase alleles across 28,330 isolates. The most common intrinsic alleles at the time of writing were blaADC-73, blaADC-30, blaADC-222, blaADC-33, and blaOXA-66, and the most common acquired allele was blaOXA-23. Interestingly, only 63.0% of assigned blaADC alleles were encountered and the 10 most common blaADC and intrinsic blaOXA alleles represented approximately 75% of their respective gene totals while dozens were extremely infrequent. Differences were observed over time and geography. Surprisingly, more distinct unassigned (i.e., lacking a blaADC or blaOXA number) alleles were encountered than distinct, assigned alleles. Understanding the diversity and distribution of β-lactamase alleles helps to prioritize variants for further research, selects targets for drug development, and may aid in selecting therapies for a given infection.
ABSTRACT Pseudomonas aeruginosa is a clinically important Gram-negative pathogen responsible for a wide variety of serious nosocomial and community-acquired infections. Antibiotic resistance is a major concern, as this organism has a wide variety of resistance mechanisms, including chromosomal class C ( bla PDC ) and D ( bla OXA-50 family) β-lactamases, efflux pumps, porin channels, and the ability to readily acquire additional β-lactamases. Surveillance studies can reveal the diversity and distribution of β-lactamase alleles but are difficult and expensive to conduct. Herein, we apply a novel approach, using publicly available data derived from whole genome sequences, to explore the diversity and distribution of β-lactamase alleles across 30,452 P . aeruginosa isolates. The most common alleles were bla PDC-3 , bla PDC-5 , bla PDC-8 , bla OXA-488 , bla OXA-50 , and bla OXA-486 . Interestingly, only 43.6% of assigned bla PDC alleles were encountered, and the 10 most common bla PDC and intrinsic bla OXA alleles represent approximately 75% of their respective total alleles, while many other assigned alleles were extremely uncommon. As anticipated, differences were observed over time and geography. Surprisingly, more distinct unassigned alleles were encountered than distinct assigned alleles. Understanding the diversity and distribution of β-lactamase alleles helps to prioritize variants for further research, select targets for drug development, and may aid in selecting therapies for a given infection.
The National Center for Biotechnology Information (NCBI) provides online information resources for biology, including the GenBank® nucleic acid sequence database and the PubMed® database of citations and abstracts published in life science journals. NCBI provides search and retrieval operations for most of these data from 35 distinct databases. The E-utilities serve as the programming interface for most of these databases. New resources include the Comparative Genome Resource (CGR) and the BLAST ClusteredNR database. Resources receiving significant updates in the past year include PubMed, PMC, Bookshelf, IgBLAST, GDV, RefSeq, NCBI Virus, GenBank type assemblies, iCn3D, ClinVar, GTR, dbGaP, ALFA, ClinicalTrials.gov, Pathogen Detection, antimicrobial resistance resources, and PubChem. These resources can be accessed through the NCBI home page at https://www.ncbi.nlm.nih.gov.
Reference sequences and annotations serve as the foundation for many lines of research today, from organism and sequence identification to providing a core description of the genes, transcripts and proteins found in an organism's genome. Interpretation of data including transcriptomics, proteomics, sequence variation and comparative analyses based on reference gene annotations informs our understanding of gene function and possible disease mechanisms, leading to new biomedical discoveries. The Reference Sequence (RefSeq) resource created at the National Center for Biotechnology Information (NCBI) leverages both automatic processes and expert curation to create a robust set of reference sequences of genomic, transcript and protein data spanning the tree of life. RefSeq continues to refine its annotation and quality control processes and utilize better quality genomes resulting from advances in sequencing technologies as well as RNA-Seq data to produce highquality annotated genomes, ortholog predictions across more organisms and other products that are easily accessible through multiple NCBI resources. This report summarizes the current status of the eukaryotic, prokaryotic and viral RefSeq resources, with a focus on eukaryotic annotation, the increase in taxonomic representation and the effect it will have on comparative genomics.
Fast, efficient public health actions require well-organized and coordinated systems that can supply timely and accurate knowledge. Public databases of pathogen genomic data, such as the International Nucleotide Sequence Database Collaboration (INSDC), have become essential tools for efficient public health decisions. However, these international resources began primarily for academic purposes, rather than for surveillance or interventions. Now, queries need to access not only the whole genomes of multiple pathogens but also make connections using robust contextual metadata to identify issues of public health relevance. Databases that over time developed a patchwork of submission formats and requirements need to be consistently organized and coordinated internationally to allow effective searches.To help resolve these issues, we propose a common pathogen data structure called the Pathogen Data Object Model (DOM) that will formalize the minimum pieces of sequence data and contextual data necessary for general public health uses, while recognizing that submitters will likely withhold a wide range of non-public contextual data. Further, we propose contributors use the Pathogen DOM for all pathogen submissions (bacterial, viral, fungal, and parasites), which will simplify data submissions and provide a consistent and transparent data structure for downstream data analyses. We also highlight how improved submission tools can support the Pathogen DOM, offering users additional easy-to-use methods to ensure this structure is followed.
The sharing of genome sequences in online data repositories allows for large scale analyses of specific genes or gene families. This can result in the detection of novel gene subtypes as well as the development of improved detection methods. Here, we used publicly available WGS data to detect a novel Stx subtype, Stx2n in two clinical E. coli strains isolated in the USA. During this process, additional Stx2 subtypes were detected; six Stx2j, one Stx2m strain, and one Stx2o, were all analyzed for variability from the originally described subtypes. Complete genome sequences were assembled from short- or long-read sequencing and analyzed for serotype, and ST types. The WGS data from Stx2n- and Stx2o-producing STEC strains were further analyzed for virulence genes pro-phage analysis and phage insertion sites. Nucleotide and amino acid maximum parsimony trees showed expected clustering of the previously described subtypes and a clear separation of the novel Stx2n subtype. WGS data were used to design OMNI PCR primers for the detection of all known stx1 (283 bp amplicon), stx2 (400 bp amplicon), intimin encoded by eae (221 bp amplicon), and stx2f (438 bp amplicon) subtypes. These primers were tested in three different laboratories, using standard reference strains. An analysis of the complete genome sequence showed variability in serogroup, virulence genes, and ST type, and Stx2 pro-phages showed variability in size, gene composition, and phage insertion sites. The strains with Stx2j, Stx2m, Stx2n, and Stx2o showed toxicity to Vero cells. Stx2j carrying strain, 2012C-4221, was induced when grown with sub-inhibitory concentrations of ciprofloxacin, and toxicity was detected. Taken together, these data highlight the need to reinforce genomic surveillance to identify the emergence of potential new Stx2 or Stx1 variants. The importance of this surveillance has a paramount impact on public health. Per our description in this study, we suggest that 2017C-4317 be designated as the Stx2n type-strain.
<ns4:p>Virulence is a complex mix of microbial traits and host susceptibility that could ultimately lead to disease. The increased prevalence of multidrug resistant infections complicates treatment options, augmenting the need for developing robust computational methods and pipelines that enable researchers and clinicians to rapidly identify the underlying mechanism(s) of virulence in any given sample/isolate. Consequently, the National Center for Biotechnology and Information at the National Institutes of Health hosted an in-person hackathon in Bethesda, Maryland during July 2019 to assist with developing cloud-based methods to reduce reliance on local computational infrastructure. Groups of attendees were assigned tasks that are relevant to identifying relevant tools, constructing pipelines capable of identifying microbial virulence factors, and managing the associated data and metadata. Specifically, the assigned tasks consisted of the following: data indexing, metabolic functions, virulence factors, antimicrobial resistance, mobile elements in enterococci, and metatranscriptomics. The cloud-based framework established by this hackathon can be augmented and built upon by the research community to aid in the rapid identification of microbial virulence factors.</ns4:p>
Antimicrobial resistance (AMR) is a significant public health threat. Low- cost whole- genome sequencing, which is often used in surveillance programmes, provides an opportunity to assess AMR gene content in these genomes using in silico approaches. A variety of bioinformatic tools have been developed to identify these genomic elements. Most of those tools rely on reference databases of nucleotide or protein sequences and collections of models and rules for analysis. While the tools are critical for the identification of AMR genes, the databases themselves also provide significant utility for researchers, for applications ranging from sequence analysis to information about AMR phenotypes. Additionally, these databases can be evaluated by domain experts and others to ensure their accuracy. Here we describe how we curate the genes, point mutations and blast rules, and hidden Markov models used in NCBI???s AMRFinderPlus, along with the quality- control steps we take to ensure database quality. We also describe the web interfaces that display the full structure of the database and their newly developed cross- browser relationships. Then, using the Reference Gene Catalog as an example, we detail how the databases, rules and models are made publicly available, as well as how to access the software. In addition, as part of the Pathogen Detection system, we have analysed over 1 million publicly available genomes using AMRFinderPlus and its databases. We discuss how the computed analyses generated by those tools can be accessed through a web interface. Finally, we conclude with NCBI???s plans to make these databases accessible over the long- term.
This multiagency report developed by the Interagency Collaboration for Genomics for Food and Feed Safety provides an overview of the use of and transition to whole genome sequencing (WGS) technology for detection and characterization of pathogens transmitted commonly by food and for identification of their sources. We describe foodbome pathogen analysis, investigation, and harmonization efforts among the following federal agencies: National Institutes of Health; Department of Health and Human Services, Centers for Disease Control and Prevention (CDC) and U.S. Food and Drug Administration (FDA); and the U.S. Department of Agriculture, Food Safety and Inspection Service, Agricultural Research Service, and Animal and Plant Health Inspection Service. We describe single nucleotide polymorphism, core-genome, and whole genome tnultilocus sequence typing data analysis methods as used in the PulseNet (CDC) and GenomeTrakr (FDA) networks, underscoring the complementary nature of the results for linking genetically related foodbome pathogens during outbreak investigations while allowing flexibility to meet the specific needs of Interagency Collaboration partners. We highlight how we apply WGS to pathogen characterization (virulence and antimicrobial resistance profiles) and source attribution efforts and increase transparency by making the sequences and other data publicly available through the National Center for Biotechnology Information. We also highlight the impact of current trends in the use of culture-independent diagnostic tests for human diagnostic testing on analytical approaches related to food safety and what is next for the use of WGS in the area of food safety.
Antimicrobial resistance (AMR) is a significant public health threat. With the rise of affordable whole genome sequencing, in silico approaches to assessing AMR gene content can be used to detect known resistance mechanisms and potentially identify novel mechanisms. To enable accurate assessment of AMR gene content, as part of a multi-agency collaboration, NCBI developed a comprehensive AMR gene database, the Bacterial Antimicrobial Resistance Reference Gene Database and the AMR gene detection tool AMRFinder. Here, we describe the expansion of the Reference Gene Database, now called the Reference Gene Catalog, to include putative acid, biocide, metal, stress resistance genes, in addition to virulence genes and species-specific point mutations. Genes and point mutations are classified by broad functions, as well as more detailed functions. As we have expanded both the functional repertoire of identified genes and functionality, NCBI released a new version of AMRFinder, known as AMRFinderPlus. This new tool allows users the option to utilize only the core set of AMR elements, or include stress response and virulence genes, too. AMRFinderPlus can detect acquired genes and point mutations in both protein and nucleotide sequence. In addition, the evidence used to identify the gene has been expanded to include whether nucleotide or protein sequence was used, its location in the contig, and presence of an internal stop codon. These database improvements and functional expansions will enable increased precision in identifying AMR genes, linking AMR genotypes and phenotypes, and determining possible relationships between AMR, virulence, and stress response.
aNational Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA bFood and Drug Administration, Center for Veterinary Medicine, Office of Research, Laurel, Maryland, USA cUSDA Food Safety and Inspection Service, Office of Public Health Science, Eastern Laboratory, Athens, Georgia, USA dEnteric Diseases Laboratory Branch, Centers for Disease Control and Prevention, Atlanta, Georgia, USA
Antimicrobial resistance (AMR) is a major public health problem that requires publicly available tools for rapid analysis. To identify acquired AMR genes in whole genome sequences, the National Center for Biotechnology Information (NCBI) has produced a high-quality, curated, AMR gene reference database consisting of up-to-date protein and gene nomenclature, a set of hidden Markov models (HMMs), and a curated protein family hierarchy. Currently, the Bacterial Antimicrobial Resistance Reference Gene Database contains 4,579 antimicrobial resistance gene proteins and more than 560 HMMs. Here, we describe AMRFinder, a tool that uses this reference dataset to identify AMR genes. To assess the predictive ability of AMRFinder, we measured the consistency between predicted AMR genotypes from AMRFinder against resistance phenotypes of 6,242 isolates from the National Antimicrobial Resistance Monitoring System (NARMS). This included 5,425 Salmonella enterica , 770 Campylobacter spp., and 47 Escherichia coli phenotypically tested against various antimicrobial agents. Of 87,679 susceptibility tests performed, 98.4% were consistent with predictions. To assess the accuracy of AMRFinder, we compared its gene symbol output with that of a 2017 version of ResFinder, another publicly available resistance gene database. Most gene calls were identical, but there were 1,229 gene symbol differences between them, with differences due to both algorithmic differences and database composition. AMRFinder missed 16 loci that Resfinder found, while Resfinder missed 1,147 loci AMRFinder identified. Two missing drug classes from the 2017 version of ResFinder contributed 81% of missed loci. Based on these results, AMRFinder appears to be a highly accurate AMR gene detection system. Importance Antimicrobial resistance is a major public health problem. Traditionally, antimicrobial resistance has been identified using phenotypic assays. With the advent of genome sequencing, we now can identify resistance genes and deduce if an isolate could be resistant to antibiotics. We describe a database of 4,579 acquired antimicrobial resistance genes, the largest publicly available, and a software tool to identify genes in bacterial genomes, AMRFinder. Unlike other tools, AMRFinder uses a gene hierarchy to prevent overpredicting what the correct gene call should be, enabling more accurate assessment. To assess these resources, we determined the resistance gene content of over 6,200 bacterial isolates from the National Antimicrobial Resistance Monitoring System that have been assayed using traditional methods and that also have had their genomes sequenced. We also compared our gene assessments to those of a popularly used tool. We found that AMRFinder has a high overall consistency between genotypes and phenotypes.
Antimicrobial resistance (AMR) is a major public health problem that requires publicly available tools for rapid analysis. To identify AMR genes in whole-genome sequences, the National Center for Biotechnology Information (NCBI) has produced AMRFinder, a tool that identifies AMR genes using a high-quality curated AMR gene reference database. The Bacterial Antimicrobial Resistance Reference Gene Database consists of up-to-date gene nomenclature, a set of hidden Markov models (HMMs), and a curated protein family hierarchy. Currently, it contains 4,579 antimicrobial resistance proteins and more than 560 HMMs. Here, we describe AMRFinder and its associated database. To assess the predictive ability of AMRFinder, we measured the consistency between predicted AMR genotypes from AMRFinder and resistance phenotypes of 6,242 isolates from the National Antimicrobial Resistance Monitoring System (NARMS). This included 5,425 Salmonella enterica, 770 Campylobacter spp., and 47 Escherichia coli isolates phenotypically tested against various antimicrobial agents. Of 87,679 susceptibility tests performed, 98.4% were consistent with predictions. To assess the accuracy of AMRFinder, we compared its gene symbol output with that of a 2017 version of ResFinder, another publicly available resistance gene detection system. Most gene calls were identical, but there were 1,229 gene symbol differences (8.8%) between them, with differences due to both algorithmic differences and database composition. AMRFinder missed 16 loci that ResFinder found, while ResFinder missed 216 loci that AMRFinder identified. Based on these results, AMRFinder appears to be a highly accurate AMR gene detection system.
FDA proactively invests in tools to support innovation of emerging technologies, such as infectious disease next generation sequencing (ID-NGS). Here, we introduce FDA-ARGOS quality-controlled reference genomes as a public database for diagnostic purposes and demonstrate its utility on the example of two use cases. We provide quality control metrics for the FDA-ARGOS genomic database resource and outline the need for genome quality gap filling in the public domain. In the first use case, we show more accurate microbial identification of Enterococcus avium from metagenomic samples with FDA-ARGOS reference genomes compared to non-curated GenBank genomes. In the second use case, we demonstrate the utility of FDA-ARGOS reference genomes for Ebola virus target sequence comparison as part of a composite validation strategy for ID-NGS diagnostic tests. The use of FDA-ARGOS as an in silico target sequence comparator tool combined with representative clinical testing could reduce the burden for completing ID-NGS clinical trials.
The initial report of the mcr-1 (mobile colistin resistance) gene has led to many reports of mcr-1 variants and other mcr genes from different bacterial species originating from human, animal and environmental samples in different geographical locations. Resistance gene nomenclature is complex and unfortunately problems such as different names being used for the same gene/protein or the same name being used for different genes/proteins are not uncommon. Registries exist for some families, such as bla (β-lactamase) genes, but there is as yet no agreed nomenclature scheme for mcr genes. The National Center for Biotechnology Information (NCBI) recently took over assigning bla allele numbers from the longstanding Lahey β-lactamase website and has agreed to do the same for mcr genes. Here, we propose a nomenclature scheme that we hope will be acceptable to researchers in this area and that will reduce future confusion.
Infectious disease next generation sequencing (ID-NGS) diagnostics are on the cusp of revolutionizing the clinical market. To facilitate this transition, FDA proactively invested in tools to support innovation of emerging technologies. FDA and collaborators established a publicly available database, FDA dAtabase for Regulatory-Grade micrObial Sequences (FDA-ARGOS), as a tool to fill reference database gaps with quality-controlled genomes. This manuscript discusses quality control metrics for the proposed FDA-ARGOS genomic resource and outlines the need for quality-controlled genome gap filling in the public domain. Here, we also present three case studies showcasing potential applications for FDA-ARGOS in infectious disease diagnostics, specifically: assay design, reference database and in silico sequence comparison in combination with representative microbial organism wet lab testing; a novel composite validation strategy for ID-NGS diagnostics. The use of FDA-ARGOS as an in silico comparator tool could reduce the burden for completing ID-NGS clinical trials. In addition, use cases identifying Enterococcus avium and Ebola virus (Zaire ebolavirus variant Makona) demonstrate the utility of FDA-ARGOS as a reference database for independent performance validation of new tests and for documenting how one would use this database as an in silico sequence target comparator tool for ID-NGS validation, respectively.
Background As next generation sequence technology has advanced, there have been parallel advances in genome-scale analysis programs for determining evolutionary relationships as proxies for epidemiological relationship in public health. Most new programs skip traditional steps of ortholog determination and multi-gene alignment, instead identifying variants across a set of genomes, then summarizing results in a matrix of single-nucleotide polymorphisms or alleles for standard phylogenetic analysis. However, public health authorities need to document the performance of these methods with appropriate and comprehensive datasets so they can be validated for specific purposes, e.g., outbreak surveillance. Here we propose a set of benchmark datasets to be used for comparison and validation of phylogenomic pipelines. Methods We identified four well-documented foodborne pathogen events in which the epidemiology was concordant with routine phylogenomic analyses (reference-based SNP and wgMLST approaches). These are ideal benchmark datasets, as the trees, WGS data, and epidemiological data for each are all in agreement. We have placed these sequence data, sample metadata, and “known” phylogenetic trees in publicly-accessible databases and developed a standard descriptive spreadsheet format describing each dataset. To facilitate easy downloading of these benchmarks, we developed an automated script that uses the standard descriptive spreadsheet format. Results Our “outbreak” benchmark datasets represent the four major foodborne bacterial pathogens (Listeria monocytogenes, Salmonella enterica, Escherichia coli, and Campylobacter jejuni) and one simulated dataset where the “known tree” can be accurately called the “true tree”. The downloading script and associated table files are available on GitHub: https://github.com/WGS-standards-and-analysis/datasets. Discussion These five benchmark datasets will help standardize comparison of current and future phylogenomic pipelines, and facilitate important cross-institutional collaborations. Our work is part of a global effort to provide collaborative infrastructure for sequence data and analytic tools—we welcome additional benchmark datasets in our recommended format, and, if relevant, we will add these on our GitHub site. Together, these datasets, dataset format, and the underlying GitHub infrastructure present a recommended path for worldwide standardization of phylogenomic pipelines.
The National Center for Biotechnology Information (NCBI) provides a large suite of online resources for biological information and data, including the GenBank((R)) nucleic acid sequence database and the PubMed database of citations and abstracts for published life science journals. The Entrez system provides search and retrieval operations for most of these data from 37 distinct databases. The E-utilities serve as the programming interface for the Entrez system. Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. New resources released in the past year include iCn3D, MutaBind, and the Antimicrobial Resistance Gene Reference Database; and resources that were updated in the past year include My Bibliography, SciENcv, the Pathogen Detection Project, Assembly, Genome, the Genome Data Viewer, BLAST and PubChem. All of these resources can be accessed through the NCBI home page at www.ncbi.nlm.nih.gov.