The role of Real-Time PCR assays for surveillance and rapid screening for pathogens is garnering more and more attention because of its versatility and ease of adoption. The goal of this study was to design, test, and evaluate Real-Time TaqMan PCR assays for the detection of botulinum neurotoxin (bont/A-G) genes from currently recognized BoNT subtypes. Assays were computationally designed and then laboratory tested for sensitivity and specificity using DNA preparations containing bont genes from 82 target toxin subtypes, including nine bivalent toxin types; 31 strains representing other clostridial species; and an extensive panel that consisted of DNA from a diverse set of prokaryotic (bacterial) and eukaryotic (fungal, protozoan, plant, and animal) species. In addition to laboratory testing, the assays were computationally evaluated using in silico analysis for their ability to detect bont gene sequences from recently identified toxin subtypes. Seventeen specific assays (two for each of the bont/C, bont/D, bont/E, and bont/G subtypes and three for each of the bont/A, bont/B, and bont/F subtypes) were designed and evaluated for their ability to detect bont genes encoding multiple subtypes from all seven serotypes. These assays could provide an additional tool for the detection of botulinum neurotoxins in clinical, environmental and food samples that can complement other existing methods used in clinical diagnostics, regulatory, public health, and research laboratories.
To enable personalized cancer treatment, machine learning models have been developed to predict drug response as a function of tumor and drug features. However, most algorithm development efforts have relied on cross-validation within a single study to assess model accuracy. While an essential first step, cross-validation within a biological data set typically provides an overly optimistic estimate of the prediction performance on independent test sets. To provide a more rigorous assessment of model generalizability between different studies, we use machine learning to analyze five publicly available cell line-based data sets: National Cancer Institute 60, ancer Therapeutics Response Portal (CTRP), Genomics of Drug Sensitivity in Cancer, Cancer Cell Line Encyclopedia and Genentech Cell Line Screening Initiative (gCSI). Based on observed experimental variability across studies, we explore estimates of prediction upper bounds. We report performance results of a variety of machine learning models, with a multitasking deep neural network achieving the best cross-study generalizability. By multiple measures, models trained on CTRP yield the most accurate predictions on the remaining testing data, and gCSI is the most predictable among the cell line data sets included in this study. With these experiments and further simulations on partial data, two lessons emerge: (1) differences in viability assays can limit model generalizability across studies and (2) drug diversity, more than tumor diversity, is crucial for raising model generalizability in preclinical screening.
Viral pathogens can rapidly evolve, adapt to novel hosts, and evade human immunity. The early detection of emerging viral pathogens through biosurveillance coupled with rapid and accurate diagnostics are required to mitigate global pandemics. However, RNA viruses can mutate rapidly, hampering biosurveillance and diagnostic efforts. Here, we present a novel computational approach called FEVER (Fast Evaluation of Viral Emerging Risks) to design assays that simultaneously accomplish: 1) broad-coverage biosurveillance of an entire group of viruses, 2) accurate diagnosis of an outbreak strain, and 3) mutation typing to detect variants of public health importance. We demonstrate the application of FEVER to generate assays to simultaneously 1) detect sarbecoviruses for biosurveillance; 2) diagnose infections specifically caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2); and 3) perform rapid mutation typing of the D614G SARS-CoV-2 spike variant associated with increased pathogen transmissibility. These FEVER assays had a high in silico recall (predicted positive) up to 99.7% of 525,708 SARS-CoV-2 sequences analyzed and displayed sensitivities and specificities as high as 92.4% and 100% respectively when validated in 100 clinical samples. The D614G SARS-CoV-2 spike mutation PCR test was able to identify the single nucleotide identity at position 23,403 in the viral genome of 96.6% SARS-CoV-2 positive samples without the need for sequencing. This study demonstrates the utility of FEVER to design assays for biosurveillance, diagnostics, and mutation typing to rapidly detect, track, and mitigate future outbreaks and pandemics caused by emerging viruses.
waves) in the analyzed datasets and automatically identify their optimal number. The features are extracted by identifying counties that have similarities between the county-level societal variables and the COVID-19 parameters. These demonstration analyses will facilitate the ongoing pandemic simulations and predictions performed by Los Alamos other institutions, as well as lay the groundwork for future work.
There is an urgent need for rapid, accurate, and ultrasensitive diagnostics to identify respiratory viruses. Amplification-free detection of nucleic acids provides a promising technique for rapid diagnosis of viral infections. The objective of this study was to design, characterize, and optimize mutation-resistant molecular beacon probes for the detection of influenza versus SARS-CoV-2 viruses. High-coverage probe cocktails were computationally designed using our "fast evaluation of viral emerging risks" (FEVER) pipeline. Thermodynamic compatibility of the probe cocktails was ensured by incorporating design constraints, such as length and GC content. The resulting assays comprised three influenza probes (one for influenza A and two for influenza B viruses) and four coronavirus probes (two for SARS-CoV-2 and two for other SARS-like viruses). The signal-to-noise ratios of the probes were optimized using a PCR thermocycler with varying concentrations of target RNA, buffer concentrations and annealing temperatures (from 25°C to 95°C). Probes were tested against targets containing one or two mismatches representing the most common natural variants. The influenza probes were able to bind to influenza target sequences with up to two mismatches, which accounts for 100% of known influenza strains. The optimal hybridization conditions for all seven probes included a 1:4 target to probe ratio at room temperature with 3 mM magnesium chloride. The limit of detection, without amplification, was 2.44 x 109 copies/µL. To detect influenza and SARS-CoV-2 in patient samples without amplification, we will explore other ultrasensitive detection platforms in future experiments. Here we have successfully developed a pipeline for the characterization and optimization of probe-target hybridization using our high-coverage influenza and coronavirus probes. Optimizing hybridization conditions is the first step towards developing an ultrasensitive amplification-free method for detecting viral nucleic acids.
Detection methods that do not require nucleic acid amplification are advantageous for viral diagnostics due to their rapid results. These platforms could provide information for both accurate diagnoses and pandemic surveillance. Influenza virus is prone to pandemic-inducing genetic mutations, so there is a need to apply these detection platforms to influenza diagnostics. Here, we analyzed the Fast Evaluation of Viral Emerging Risks (FEVER) pipeline on ultrasensitive detection platforms, including a waveguide-based optical biosensor and a flow cytometry bead-based assay. The pipeline was also evaluated in silico for sequence coverage in comparison to the U.S. Centers for Disease Control and Prevention's (CDC) influenza A and B diagnostic assays. The influenza FEVER probe design had a higher tolerance for mismatched bases than the CDC's probes, and the FEVER probes altogether had a higher detection rate for influenza isolate sequences from GenBank. When formatted for use as molecular beacons, the FEVER probes detected influenza RNA as low as 50 nM on the waveguide-based optical biosensor and 1 nM on the flow cytometer. In addition to molecular beacons, which have an inherently high background signal we also developed an exonuclease selection method that could detect 500 pM of RNA. The combination of high-coverage probes developed using the FEVER pipeline coupled with ultrasensitive optical biosensors is a promising approach for future influenza diagnostic and biosurveillance applications.
Abstract Summary Polymerase chain reaction-based assays are the current gold standard for detecting and diagnosing SARS-CoV-2. However, as SARS-CoV-2 mutates, we need to constantly assess whether existing PCR-based assays will continue to detect all known viral strains. To enable the continuous monitoring of SARS-CoV-2 assays, we have developed a web-based assay validation algorithm that checks existing PCR-based assays against the ever-expanding genome databases for SARS-CoV-2 using both thermodynamic and edit-distance metrics. The assay-screening results are displayed as a heatmap, showing the number of mismatches between each detection and each SARS-CoV-2 genome sequence. Using a mismatch threshold to define detection failure, assay performance is summarized with the true-positive rate (recall) to simplify assay comparisons. Availability and implementation The assay evaluation website and supporting software are Open Source and freely available at https://covid19.edgebioinformatics.org/#/assayValidation, https://github.com/jgans/thermonucleotide BLAST and https://github.com/LANL-Bioinformatics/assay_validation. Supplementary information Supplementary data are available at Bioinformatics online.
Microbial biomass is increasingly used to predict respiration in soil organic carbon (SOC) models. Its increased use combined with the difficulty of accurately measuring this variable points a need to directly assess the importance of microbial biomass abundance for carbon (C) cycling. To test the hypothesis that the initial microbial biomass abundance (i.e. biomass abundance on new plant litter) is a strong driver of plant litter C cycling, we manipulated biomass abundance by 10 and 100-fold dilution and composition using 12 source communities on sterile pine litter and measured respiration in microcosms for 30 days. In the first two days of microbial growth on fresh litter, a 100-fold difference in initial biomass abundance caused an average difference in respiration of nearly 300%, but the effect rapidly declined to less than 30% in 10 days and to 14% in 30 days. Parallel simulations with a soil carbon model, SOMIC 1.0, also predicted a 14% difference over 30 days, consistent with the experimental results. Model simulations predicted convergence of cumulative CO 2 to within 10% in three months and within 4% in three years. Rapid microbial growth likely attenuates the effects of large initial differences in biomass abundance. In contrast, the persistence of source community as an explanatory factor in driving differences in respiration across microcosms supports the importance of microbial composition in C cycling. Overall, the results suggest that the initial abundance of microbial biomass on litter is a weak driver of C flux from litter decomposition over long timescales (months to years) when litter communities have equal nutrient availability. By extension, slight variation in the timing of microbial dispersal to fresh litter is likely to be a minor factor in long-term C flux. Importance Microbial biomass is one of the most common microbial parameters used in land carbon (C) cycle models, however, it is notoriously difficult to measure accurately. To understand the consequences of mismeasurement, as well as the broader importance of microbial biomass abundance as a direct driver of ecological phenomena, greater quantitative understanding of the role of microbial biomass abundance in environmental processes is needed. Using microcosms, we manipulated the initial biomass of numerous microbial communities across a 100-fold range and measured effects on CO 2 production during plant litter decomposition. We found that the effects of initial biomass abundance on CO 2 production was largely attenuated within a week, while the effects of community type remained significant over the course of the experiment. Overall, our results suggest that initial microbial biomass abundance in litter decomposition within an ecosystem is a weak driver of long-term C cycling dynamics.
We describe the use of in silico approaches to improve the process of molecular assay development and reduce time and cost by utilizing available databases of whole genome pathogen sequences combined with modern bioinformatics and physical modeling tools. Well-characterized assays are needed for accurately detecting pathogens in environmental and patient samples and also for evaluation of the efficacy of a medical countermeasure that may be administered to patients. The polymerase chain reaction (PCR) remains the gold standard for pathogen detection due to the simplicity of its instrumentation, low cost of reagents, and outstanding limit of detection (LOD), sensitivity, and specificity. However, creation of such PCR assays often involves iterations of design, preliminary testing, and thorough validation with clinical isolates and testing in relevant matrices, which can be time consuming, costly, and result in suboptimal assays. Since formal validation (e.g., for Emergency Use Authorization [EUA] or Food and Drug Administration [FDA] licensure) of an infectious disease assay can be very expensive and can require extensive time of development, having a well-designed assay up front is a critical first step. Yet, many assays described in the literature utilized limited design capabilities and many initially promising assays fail the validation process, resulting in increased costs and timelines for successful product development. While the computational approaches outlined in this document by no means obviate the need for wet lab testing, they can reduce the amount of effort wasted on empirical optimization and iterative redesigns and also guide validation studies. The proposed computational approaches also result in higher performing assays with better sensitivity, specificity, and lower LOD and reduce the possibility of assay failure due to signature erosion. To provide clarity, an extensive glossary of defined terms is provided. Received: 19 March 2020; Accepted: 19 March 2020 VC AOAC INTERNATIONAL 2020. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited. 882 Journal of AOAC INTERNATIONAL, 103(4), 2020, 882–899 doi: 10.1093/jaoacint/qsaa045 Technical Communication 1 Background and Rationale Nucleic acid-based assays, such as real-time PCR, are the mainstay of clinical diagnostics and biosurveillance. A typical PCR assay design begins with computational (“in silico”) identification of a unique region (signature) that can support the binding of primer and probe sequences for target-specific amplification as a means of detecting the presence of the target organism. This step is followed by wet lab testing of the primers and probes using genomic deoxyribonucleic acid (DNA) or reverse transcribed ribonucleic acid (RNA) and performance-optimization of selected assays. In addition, extensive testing of the assay in the intended clinical matrix is required to evaluate assay parameters, such as LOD, sensitivity (probability of detection), and specificity (see glossary for definitions). The sensitivity and specificity of the assay are experimentally determined using a set of target (inclusivity) strains, near-neighbor (exclusivity) strains, and matrix-relevant (background) organisms. Assay performance also needs to be measured in assay-specific matrices (i.e., blood, stool, water, soil, etc.). Often, assays are computationally designed using a set of available genomic/gene sequences at that time and then experimentally validated for signature presence in all available samples of the target organism (inclusivity panel) and validated for signature absence in many other samples that do not contain the target (exclusivity panel and matrix panel). In an ideal scenario, a laboratory routinely engaged in assay development could complete this process within 6 to 12 months. Detection assays are typically designed using all sequences available at that time. Many of the biodefense assays were designed and tested at least a decade ago when available sequences were limited. Thanks to recent advances in modern sequencing technologies, there is a sharp increase in the availability of whole genome sequences (Figure 1). Hence, these older assays have the potential to fail if evaluated against currently available sequences. Moreover, publicly available sequence databases (e.g., GenBank) typically contain only a small fraction of naturally occurring sequence diversity. As a result, detection assays are vulnerable to “overfitting”: correctly differentiating known (i.e., sequenced) targets and nontargets but failing to detect novel target variants or falsely detecting novel nontargets. Knowledge of the true genetic diversity is limited for some biodefense agents and their near neighbors, as often only several geographical and temporal representatives are fully characterized while other geographic locations have been ignored or significantly undersampled and hence are under-represented. In addition, while some agents, such as the bacterium Bacillus anthracis, are monomorphic (i.e., highly conserved), other agents, especially RNA viruses, are very diverse [e.g., Lymphocytic choriomeningitis virus (LCMV), Lassa virus, and Crimean-Congo hemorrhagic fever virus (CCHFV)]. In general, detection assays targeting highly conserved targets tend to fail due to unsequenced near-neighbor cross-reactivity, while assays targeting diverse targets tend to fail due to false negatives against unsequenced target variants. While the recent revolution in next-generation sequencing technologies combined with decreasing sequencing costs has increased knowledge of population genomic structure, the capability for laboratory-based evaluation of newly sequenced strains has not kept pace. In this scenario, replacing or redesigning older assays to incorporate new knowledge of the target genomic landscape is critical. However, wet lab testing may not be feasible due to limitations on the timely availability of samples/strains. This problem is further exacerbated by policy decisions, such as the 2015 Department of Defense (DoD) moratorium that decreased access to live/inactivated biodefense pathogens for various applications, including assay development and validation (1). 1.1 Additional Considerations with the Status Quo Testing of Assays Against Inclusivity/Exclusivity Panels The AOAC Stakeholder Panel on Agent Detection Assays (SPADA) inclusivity/exclusivity panels for the biodefenserelevant bacterial pathogens, such as Bacillus anthracis, Yersinia pestis, Brucella suis, Burkholderia mallei, Burkholderia pseudomallei, and Francisella tularensis, comprise a total of approximately 100 strains. These strains are used to validate the inclusivity/exclusivity criteria for the respective detection assays. Most of the inclusivity strains, and some exclusivity strains, are considered Biosafety Level 3 (BSL3) agents and, as a result, are limited to laboratories that are registered and certified for such work. Moreover, extensive laboratory testing adds cost and time to the assay development effort. Many whole genome sequences of these bacterial strains are available now (2–6), which allows the in silico evaluation of assays. An example set of assays developed prior to the nextgeneration sequencing revolution with representative analyses is illustrated in Figure 2. As expected, the majority of the evaluated assay signatures had perfect sequence matches to the target inclusivity genome sequences, and much less (0 to 40%) sequence identity to the exclusivity panel genome sequences. However, for all the assays evaluated, there was no “perfect” assay (i.e., no false positives and no false negatives). Some assays were computationally predicted to have both false negatives (e.g., Bacillus anthracis assay 1 against strain 10 in the inclusivity panel) and false positives (e.g., Bacillus anthracis assay 1 against strain 8 in the exclusivity panel). Many of these predicted assay failures correspond to expected deviations based on the genotypes of these strains. There are other assays that simply fail the inclusivity and/or exclusivity criteria (e.g., Bacillus anthracis assay 7 or Yersinia pestis assay 15) and are therefore not reliable diagnostics due to low specificity. However, given the high conservation of the assay signatures to the target strains in the inclusivity panel and their low conservation in the exclusivity panel, the “brute force” testing of all available strains is not cost Figure 1. Availability of whole genome sequences for representative bacteria. Black bar represents the assay design time frame. SantaLucia et al.: Journal of AOAC INTERNATIONAL Vol. 103, No. 4, 2020 | 883
Progress in modern biology is being driven, in part, by the large amounts of freely available data in public resources such as the International Nucleotide Sequence Database Collaboration (INSDC), the world's primary database of biological sequence (and related) information. INSDC and similar databases have dramatically increased the pace of fundamental biological discovery and enabled a host of innovative therapeutic, diagnostic, and forensic applications. However, as high-value, openly shared resources with a high degree of assumed trust, these repositories share compelling similarities to the early days of the Internet. Consequently, as public biological databases continue to increase in size and importance, we expect that they will face the same threats as undefended cyberspace. There is a unique opportunity, before a significant breach and loss of trust occurs, to ensure they evolve with quality and security as a design philosophy rather than costly "retrofitted" mitigations. This Perspective surveys some potential quality assurance and security weaknesses in existing open genomic and proteomic repositories, describes methods to mitigate the likelihood of both intentional and unintentional errors, and offers recommendations for risk mitigation based on lessons learned from cybersecurity.
In terrestrial ecosystems, stochasticity in the assembly of surface litter decomposer communities is widely believed to shape community composition. The compositional variation may drive functional variation, resulting in different patterns of carbon flow from decomposing litter. Is this important for climate feedbacks? The importance depends foremost on 1) the magnitude of functional variation and 2) the likelihood of substantial functional variation to occur within ecosystems. Here, we examined the likelihood of microbial driven variation in surface litter carbon (C) flow as a function of geographic scale. Our null hypothesis is that stochastic assembly of decomposer communities creates only minor functional variation (e.g. a few percent difference in carbon flow). Consequently, substantial functional variation among decomposer communities is likely to be found only among communities over large geographic scales (e.g. >100km), where climate and ecosystem gradients can create persistent functional differences between distant microbial communities. We performed a test of this hypothesis with a collection of over 400 soil samples from locations representing varied geographic scales (meters to 1000km). We suspended the soil microbial communities suspended in water, transferred aliquots to laboratory microcosms with sterile plant litter, and measured carbon flow (CO2 and DOC) arising from 45 days of decomposition. We present the likelihood of substantial functional variation among communities as a function of the original distance between the communities, ranging from <1cm (replicate microcosms derived from the same gram of soil) to >1000km. This work was supported by grants 2015SFAF260 and 2019SFAF255 from the OBER Genomic Sciences program.
DNA-based monitoring of pathogens in aerosol samples requires extraction methods that provide high recovery of DNA. To identify a suitable method, we evaluated six DNA extraction methods for recovery of target-specific DNA from samples with four bacterial agents at low abundance (< 10,000 genome copies per detection assay). These methods differed in rigor of cell disruption, approach for DNA capture, and extent of DNA purification. The six methods varied 1000-fold in the recovery of DNA from spores or cells of surrogates of Bacillus anthracis, Yersinia pestis, Burkholderia pseudomallei, and Francisella tularensis, each at about 10(5) CFU per sample. A custom method using paramagnetic Dynabeads for DNA capture greatly outperformed the other five methods. The cDynabead method provided about 80% recovery of target-specific DNA. The cDynabead method and a filtration method were further evaluated for DNA recovery from bacterial agents spiked on filters (ca. 10(5) CFU of each agent per filter quadrant) that were subsequently used to collect background outdoor air particulates for 24-h. The filtration method generally failed to recover detectable quantities of target DNA from the spiked filters, suggesting at least a 100-fold loss of target DNA during extraction, whereas the custom cDynabead method consistently yielded DNA sufficient for target detection.
Background: We are transforming the field of infectious disease diagnostics with the development of the Sample Prep for Infectious Disease Recognition With EDGE Bioinformatics (SPIDR-WEB). SPIDR-WEB is a sample-to-result biotechnology platform that enables efficient use of next generation sequencing (NGS) for pathogen detection in clinical samples. NGS has become a powerful tool for detection and characterization of both known and emerging pathogens. The main advantage of NGS is its non-biased approach that identifies all organisms in a sample. This is in contrast to traditional molecular assays that force us to look for a set of specific pathogens. In most clinical samples, the relative abundance of pathogen nucleic acids (DNA or RNA) is vanishingly small. Therefore, vast amounts of sequence data must be generated and analyzed to identify rare pathogen sequences. SPIDR-WEB is a sample-to-result process that relies on efficient laboratory and in silico steps. Methods & Materials: Clinical samples mostly comprise non-informative host RNAs or abundant housekeeping gene transcripts. SPIDR-WEB incorporates removal of non-informative RNAs (RNR), thereby enriching all other RNAs, including those from pathogens. This step enables either higher sensitivity and specificity, or less expensive and faster sequencing. Our custom EDGE bioinformatics data analysis platform provides rapid read classification at all taxonomic levels, and reliably detects all organisms present in a sample. EDGE is an efficient process, as it uses databases with pre-computed signatures, instead of aligning sequencing reads to the entire Genbank. In addition to RNR and EDGE, SPIDR-WEB includes robust, inexpensive and rapid sample lysis, RNA extraction, and library preparation steps. Results: We will describe SPIDR-WEB technology and show clinically-relevant results obtained from human blood, stool, respiratory, and other sample types. Conclusion: We are implementing SPIDR-WEB in both research and clinical settings to support a multitude of applications, such as discovery of novel mechanisms and biomarkers, study host-pathogen interactions, improve vaccines and therapeutics, and complement current diagnostic tools and help improve their utility.
Environmental biosurveillance and microbial ecology studies use PCR-based assays to detect and quantify microbial taxa and gene sequences within a complex background of microorganisms. However, the fragmentary nature and growing quantity of DNA-sequence data make group-specific assay design challenging. We solved this problem by developing a software platform that enables PCR-assay design at an unprecedented scale. As a demonstration, we developed quantitative PCR assays for a globally widespread, ecologically important bacterial group in soil, Acidobacteria Group 1. A total of 33 684 Acidobacteria 16S rRNA gene sequences were used for assay design. Following 1 week of computation on a 376-core cluster, 83 assays were obtained. We validated the specificity of the top three assays, collectively predicted to detect 42% of the Acidobacteria Group 1 sequences, by PCR amplification and sequencing of DNA from soil. Based on previous analyses of 16S rRNA gene sequencing, Acidobacteria Group 1 species were expected to decrease in response to elevated atmospheric CO2. Quantitative PCR results, using the Acidobacteria Group 1-specific PCR assays, confirmed the expected decrease and provided higher statistical confidence than the 16S rRNA gene-sequencing data. These results demonstrate a powerful capacity to address previously intractable assay design challenges.
Extensive use of antibiotics in both public health and animal husbandry has resulted in rapid emergence of antibiotic resistance in almost all human pathogens, including biothreat pathogens. Antibiotic resistance has thus become a major concern for both public health and national security. We developed multiplexed assays for rapid, simultaneous pathogen detection and characterization of ciprofloxacin and doxycycline resistance in Bacillus anthracis, Yersinia pestis, and Francisella tularensis. These assays are SNP-based and use Multiplexed Oligonucleotide Ligation-PCR (MOL-PCR). The MOL-PCR assay chemistry and MOLigo probe design process are presented. A web-based tool – MOLigoDesigner (http://MOLigoDesigner.lanl.gov) was developed to facilitate the probe design. All probes were experimentally validated individually and in multiplexed assays, and minimal sets of multiplexed MOLigo probes were identified for simultaneous pathogen detection and antibiotic resistance characterization.
Human dental plaque is a complex microbial community containing an estimated 700 to 19,000 species/phylotypes. Despite numerous studies analysing species richness in healthy and diseased human subjects, the true genomic composition of the human dental plaque microbiota remains unknown. Here we report a metagenomic analysis of a healthy human plaque sample using a combination of second-generation sequencing platforms. A total of 860 million base pairs of non-human sequences were generated. Various analysis tools revealed the presence of 12 well-characterized phyla, members of the TM-7 and BRC1 clade, and sequences that could not be classified. Both pathogens and opportunistic pathogens were identified, supporting the ecological plaque hypothesis for oral diseases. Mapping the metagenomic reads to sequenced reference genomes demonstrated that 4% of the reads could be assigned to the sequenced species. Preliminary annotation identified genes belonging to all known functional categories. Interestingly, although 73% of the total assembled contig sequences were predicted to code for proteins, only 51% of them could be assigned a functional role. Furthermore, ~2.8% of the total predicted genes coded for proteins involved in resistance to antibiotics and toxic compounds, suggesting that the oral cavity is an important reservoir for antimicrobial resistance.
BACKGROUND:New and improved antimicrobial countermeasures are urgently needed to counteract increased resistance to existing antimicrobial treatments and to combat currently untreatable or new emerging infectious diseases. We demonstrate that computational comparative genomics, together with experimental screening, can identify potential generic (i.e., conserved across multiple pathogen species) and novel virulence-associated genes that may serve as targets for broad-spectrum countermeasures.RESULTS:Using phylogenetic profiles of protein clusters from completed microbial genome sequences, we identified seventeen protein candidates that are common to diverse human pathogens and absent or uncommon in non-pathogens. Mutants of 13 of these candidates were successfully generated in Yersinia pseudotuberculosis and the potential role of the proteins in virulence was assayed in an animal model. Six candidate proteins are suggested to be involved in the virulence of Y. pseudotuberculosis, none of which have previously been implicated in the virulence of Y. pseudotuberculosis and three have no record of involvement in the virulence of any bacteria.CONCLUSION:This work demonstrates a strategy for the identification of potential virulence factors that are conserved across a number of human pathogenic bacterial species, confirming the usefulness of this tool.