Chemicals involved in plutonium uranium reduction extraction (PUREX) can be released from nuclear reprocessing facilities and accumulate in the environment. We exposed chemically diverse soils to a range of concentrations of key chemicals used in the PUREX process. The responses of soil microbial communities are dependent on soil type, and tributyl phosphate exposure generates the most reproducible changes in microbial communities. We reconstructed the genomes of key bacteria and find several phosphotriesterase genes found only in Rhizobiaceae. The abundance of phosphotriesterase genes is significantly higher in samples exposed to tributyl phosphate. These phosphotriesterase genes may be involved in breakdown of tributyl phosphate, and a means of accessing phosphate for these bacteria.
Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R2 of 0.604 [0.593-0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156-5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.
As genomic and related data continue to expand, research biologists are often hampered by the computational hurdles required to analyze their data. The National Institute of Allergy and Infectious Diseases (NIAID) established the Bioinformatics Resource Centers (BRC) to assist researchers with their analysis of genome sequence and other omics-related data. Recently, the PAThosystems Resource Integration Center (PATRIC), the Influenza Research Database (IRD), and the Virus Pathogen Database and Analysis Resource (ViPR) BRCs merged to form the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) at https://www.bv-brc.org/ . The combined BV-BRC leverages the functionality of the original resources for bacterial and viral research communities with a unified data model, enhanced web-based visualization and analysis tools, and bioinformatics services. Here we demonstrate how antimicrobial resistance data can be analyzed in the new resource.
In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scientific exploration, bringing significant advancements across sectors from drug development to renewable energy. To answer this call, we present DeepSpeed4Science initiative (deepspeed4science.ai) which aims to build unique capabilities through AI system technology innovations to help domain experts to unlock today's biggest science mysteries. By leveraging DeepSpeed's current technology pillars (training, inference and compression) as base technology enablers, DeepSpeed4Science will create a new set of AI system technologies tailored for accelerating scientific discoveries by addressing their unique complexity beyond the common technical approaches used for accelerating generic large language models (LLMs). In this paper, we showcase the early progress we made with DeepSpeed4Science in addressing two of the critical system challenges in structural biology research.
In this study, we built machine learning classifiers for predicting the presence or absence of the variable genes occurring in 10-90% of all publicly available high-quality Escherichia coli genomes. The BV-BRC genus-specific protein families were used to define orthologs across the set of genomes, and a single binary classifier was built for predicting the presence or absence of each family in each genome. Each model was built using the nucleotide k-mers from a set of 100 conserved genes as features. The resulting set of 3,259 XGBoost classifiers had a per-genome average macro F1 score of 0.944 [0.943-0.945, 95% CI]. We show that the F1 scores are stable across MLSTs, and that the trend can be recapitulated through sampling with a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including “hypothetical proteins”, were easily predicted (F1 = 0.902 [0.898-0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions, including transposition- (F1 = 0.895 [0.882-0.907, 95% CI]), phage- (F1 = 0.872 [0.868-0.876, 95% CI]), and plasmid-related (F1 = 0.824 [0.814-0.834, 95% CI]) functions had slightly lower F1 scores, but were still accurate. Finally, we applied the models to a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources and observed an average per-genome F1 score of 0.880 [0.876-0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data. Importance Having the ability to predict the protein-encoding gene content of a genome is important for a variety of bioinformatic tasks, including assessing genome quality, binning genomes from shotgun metagenomic assemblies, and assessing risk due to the presence of antimicrobial resistance (AMR) and other virulence genes. In this study, we built a series of binary classifiers for predicting the presence or absence of variable genes occurring in 10-90% of all publicly available E. coli genomes. Overall, the results show that a large portion of the E. coli variable gene content can be predicted with high accuracy, including genes with functions relating to horizontal gene transfer.
Having the ability to predict the protein-encoding gene content of an incomplete genome or metagenome-assembled genome is important for a variety of bioinformatic tasks. In this study, as a proof of concept, we built machine learning classifiers for predicting variable gene content in Escherichia coli genomes using only the nucleotide k-mers from a set of 100 conserved genes as features. Protein families were used to define orthologs, and a single classifier was built for predicting the presence or absence of each protein family occurring in 10%-90% of all E. coli genomes. The resulting set of 3,259 extreme gradient boosting classifiers had a per-genome average macro F1 score of 0.944 [0.943-0.945, 95% CI]. We show that the F1 scores are stable across multi-locus sequence types and that the trend can be recapitulated by sampling a smaller number of core genes or diverse input genomes. Surprisingly, the presence or absence of poorly annotated proteins, including "hypothetical proteins" was accurately predicted (F1 = 0.902 [0.898-0.906, 95% CI]). Models for proteins with horizontal gene transfer-related functions had slightly lower F1 scores but were still accurate (F1s = 0.895, 0.872, 0.824, and 0.841 for transposon, phage, plasmid, and antimicrobial resistance-related functions, respectively). Finally, using a holdout set of 419 diverse E. coli genomes that were isolated from freshwater environmental sources, we observed an average per-genome F1 score of 0.880 [0.876-0.883, 95% CI], demonstrating the extensibility of the models. Overall, this study provides a framework for predicting variable gene content using a limited amount of input sequence data. IMPORTANCE Having the ability to predict the protein-encoding gene content of a genome is important for assessing genome quality, binning genomes from shotgun metagenomic assemblies, and assessing risk due to the presence of antimicrobial resistance and other virulence genes. In this study, we built a set of binary classifiers for predicting the presence or absence of variable genes occurring in 10%-90% of all publicly available E. coli genomes. Overall, the results show that a large portion of the E. coli variable gene content can be predicted with high accuracy, including genes with functions relating to horizontal gene transfer. This study offers a strategy for predicting gene content using limited input sequence data.
The National Institute of Allergy and Infectious Diseases (NIAID) established the Bioinformatics Resource Center (BRC) program to assist researchers with analyzing the growing body of genome sequence and other omics-related data. In this report, we describe the merger of the PAThosystems Resource Integration Center (PATRIC), the Influenza Research Database (IRD) and the Virus Pathogen Database and Analysis Resource (ViPR) BRCs to form the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) https://www.bv-brc.org/. The combined BV-BRC leverages the functionality of the bacterial and viral resources to provide a unified data model, enhanced web-based visualization and analysis tools, bioinformatics services, and a powerful suite of command line tools that benefit the bacterial and viral research communities.
High-throughput genome sequencing technologies enable the investigation of complex genetic interactions, including the horizontal gene transfer of plasmids and bacteriophages. However, identifying these elements from assembled reads remains challenging due to genome sequence plasticity and the difficulty in assembling complete sequences. In this study, we developed a classifier, using random forest, to identify whether sequences originated from bacterial chromosomes, plasmids, or bacteriophages. The classifier was trained on a diverse collection of 23,211 chromosomal, plasmid, and bacteriophage sequences from hundreds of bacterial species. In order to adapt the classifier to incomplete sequences, each complete sequence was subsampled into 5,000 nucleotide fragments and further subdivided into k-mers. This three-class classifier succeeded in identifying chromosomes, plasmids, and bacteriophages using k-mer distributions of complete and partial genome sequences, including simulated metagenomic scaffolds with minimum performance of 0.939 area under the receiver operating characteristic curve (AUC). This classifier, implemented as SourceFinder, has been made available as an online web service to help the community with predicting the chromosomal, plasmid, and bacteriophage sources of assembled bacterial sequence data (https://cge.food.dtu.dk/ services/SourceFinder/). IMPORTANCE Extra-chromosomal genes encoding antimicrobial resistance, metal resistance, and virulence provide selective advantages for bacterial survival under stress conditions and pose serious threats to human and animal health. These accessory genes can impact the composition of microbiomes by providing selective advantages to their hosts. Accurately identifying extra-chromosomal elements in genome sequence data are critical for understanding gene dissemination trajectories and taking preventative measures. Therefore, in this study, we developed a random forest classifier for identifying the source of bacterial chromosomal, plasmid, and bacteriophage sequences.
Many animal species are susceptible to SARS-CoV-2 and could potentially act as reservoirs, yet transmission of the virus in non-human free-living animals has not been documented. White-tailed deer (Odocoileus virginianus), the predominant cervid in North America, are susceptible to SARS-CoV-2 infection, and experimentally infected fawns can transmit the virus. To test the hypothesis that SARS-CoV-2 may be circulating in deer, we tested 283 retropharyngeal lymph node (RPLN) samples collected from 151 free-living and 132 captive deer in Iowa from April 2020 through December of 2020 for the presence of SARS-CoV-2 RNA. Ninety-four of the 283 deer (33.2%; 95% CI: 28, 38.9) samples were positive for SARS-CoV-2 RNA as assessed by RT-PCR. Notably, between November 23, 2020 and January 10, 2021, 80 of 97 (82.5%; 95% CI 73.7, 88.8) RPLN samples had detectable SARS-CoV-2 RNA by RT-PCR. Whole genome sequencing of the 94 positive RPLN samples identified 12 SARS-CoV-2 lineages, with B.1.2 (n = 51; 54.5%), and B.1.311 (n = 19; 20%) accounting for ~75% of all samples. The geographic distribution and nesting of clusters of deer and human lineages strongly suggest multiple zooanthroponotic spillover events and deer-to-deer transmission. The discovery of sylvatic and enzootic SARS-CoV-2 transmission in deer has important implications for the ecology and long-term persistence, as well as the potential for spillover to other animals and spillback into humans. These findings highlight an urgent need for a robust and proactive “One Health” approach to obtaining a better understanding of the ecology and evolution of SARS-CoV-2. One-Sentence Summary SARS-CoV-2 was detected in one-third of sampled white-tailed deer in Iowa between September 2020 and January of 2021 that likely resulted from multiple human-to-deer spillover and deer-to-deer transmission events.
Since the beginning of the COVID-19 pandemic, SARS-CoV-2 has demonstrated its ability to rapidly and continuously evolve, leading to the emergence of thousands of different sequence variants, many with distinctive phenotypic properties. Fortunately, the broad application of next generation sequencing (NGS) across the globe has produced a wealth of SARS-CoV-2 genome sequences, offering a comprehensive picture of how this virus is evolving so that accurate diagnostics, reliable therapeutics, and prophylactic vaccines against COVID-19 can be developed and maintained. The millions of SARS-CoV-2 sequences deposited into genomic sequencing databases, including GenBank, BV-BRC, and GISAID, are annotated with the dates and geographic locations of sample collection, and can be aligned to and compared with the Wuhan-Hu-1 reference genome to extract their constellation of nucleotide and amino acid substitutions. By aggregating these data into concise datasets, the spread of variants through space and time can be assessed. Variant tracking efforts have initially focused on the Spike protein due to its critical role in viral tropism and antibody neutralization. To identify emerging variants of concern as early as possible, we developed a computational pipeline to process the genomic data and assign risk scores based on both epidemiological and functional parameters. Epidemiological dynamics are used to identify variants exhibiting substantial growth over time and spread across geographical regions. Experimental data that quantify Spike protein regions targeted by adaptive immunity and critical for other virus characteristics are used to predict variants with consequential immunogenic and pathogenic impacts. The growth assessment and functional impact scores are combined to produce a Composite Score for any set of Spike substitutions detected. With this systematic method to routinely score and rank emerging variants, we have established an approach to identify threatening variants early and prioritize them for experimental evaluation.
There is mounting evidence of SARS-CoV-2 spillover from humans into many domestic, companion, and wild animal species. Research indicates that humans have infected white-tailed deer, and that deer-to-deer transmission has occurred, indicating that deer could be a wildlife reservoir and a source of novel SARS-CoV-2 variants. We examined the hypothesis that the Omicron variant is actively and asymptomatically infecting the free-ranging deer of New York City. Between December 2021 and February 2022, 155 deer on Staten Island, New York, were anesthetized and examined for gross abnormalities and illnesses. Paired nasopharyngeal swabs and blood samples were collected and analyzed for the presence of SARS-CoV-2 RNA and antibodies. Of 135 serum samples, 19 (14.1%) indicated SARS-CoV-2 exposure, and 11 reacted most strongly to the wild-type B.1 lineage. Of the 71 swabs, 8 were positive for SARS-CoV-2 RNA (4 Omicron and 4 Delta). Two of the animals had active infections and robust neutralizing antibodies, revealing evidence of reinfection or early seroconversion in deer. Variants of concern continue to circulate among and may reinfect US deer populations, and establish enzootic transmission cycles in the wild: this warrants a coordinated One Health response, to proactively surveil, identify, and curtail variants of concern before they can spill back into humans.
AbstractGenetic variants of SARS-CoV-2 continue to dramatically alter the landscape of the COVID-19 pandemic. The recently described variant of concern designated Omicron (B.1.1.529) has rapidly spread worldwide and is now responsible for the majority of COVID-19 cases in many countries. Because Omicron was recognized very recently, many knowledge gaps exist about its epidemiology, clinical severity, and disease course. A genome sequencing study of SARS-CoV-2 in the Houston Methodist healthcare system identified 4,468 symptomatic patients with infections caused by Omicron from late November 2021 through January 5, 2022. Omicron very rapidly increased in only three weeks to cause 90% of all new COVID-19 cases, and at the end of the study period caused 98% of new cases. Compared to patients infected with either Alpha or Delta variants in our healthcare system, Omicron patients were significantly younger, had significantly increased vaccine breakthrough rates, and were significantly less likely to be hospitalized. Omicron patients required less intense respiratory support and had a shorter length of hospital stay, consistent with on average decreased disease severity. Two patients with Omicron “stealth” sublineage BA.2 also were identified. The data document the unusually rapid spread and increased occurrence of COVID-19 caused by the Omicron variant in metropolitan Houston, and address the lack of information about disease character among US patients.
Genetic variants of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) continue to dramatically alter the landscape of the coronavirus disease 2019 (COVID-19) pandemic. The recently described variant of concern designated Omicron (B.1.1.529) has rapidly spread worldwide and is now responsible for the majority of COVID-19 cases in many countries. Because Omicron was recognized recently, many knowledge gaps exist about its epidemiology, clinical severity, and disease course. A genome sequencing study of SARS-CoV-2 in the Houston Methodist health care system identified 4468 symptomatic patients with infections caused by Omicron from late November 2021 through January 5, 2022. Omicron rapidly increased in only 3 weeks to cause 90% of all new COVID-19 cases, and at the end of the study period caused 98% of new cases. Compared with patients infected with either Alpha or Delta variants in our health care system, Omicron patients were significantly younger, had significantly increased vaccine breakthrough rates, and were significantly less likely to be hospitalized. Omicron patients required less intense respiratory support and had a shorter length of hospital stay, consistent with on average decreased disease severity. Two patients with Omicron stealth sublineage BA.2 also were identified. The data document the unusually rapid spread and increased occurrence of COVID-19 caused by the Omicron variant in metropolitan Houston, Texas, and address the lack of information about disease character among US patients.
We seek to transform how new and emergent variants of pandemic-causing viruses, specifically SARS-CoV-2, are identified and classified. By adapting large language models (LLMs) for genomic data, we build genome-scale language models (GenSLMs) which can learn the evolutionary landscape of SARS-CoV-2 genomes. By pre-training on over 110 million prokaryotic gene sequences and fine-tuning a SARS-CoV-2-specific model on 1.5 million genomes, we show that GenSLMs can accurately and rapidly identify variants of concern. Thus, to our knowledge, GenSLMs represents one of the first whole genome scale foundation models which can generalize to other prediction tasks. We demonstrate scaling of GenSLMs on GPU-based supercomputers and AI-hardware accelerators utilizing 1.63 Zettaflops in training runs with a sustained performance of 121 PFLOPS in mixed precision and peak of 850 PFLOPS. We present initial scientific insights from examining GenSLMs in tracking evolutionary dynamics of SARS-CoV-2, paving the path to realizing this on large biological data.
The emergence of a novel pathogen in a susceptible population can cause rapid spread of infection. High prevalence of SARS-CoV-2 infection in white-tailed deer (Odocoileus virginianus) has been reported in multiple locations, likely resulting from several human-to-deer spillover events followed by deer-to-deer transmission. Knowledge of the risk and direction of SARS-CoV-2 transmission between humans and potential reservoir hosts is essential for effective disease control and prioritisation of interventions. Using genomic data, we reconstruct the transmission history of SARS-CoV-2 in humans and deer, estimate the case finding rate and attempt to infer relative rates of transmission between species. We found no evidence of direct or indirect transmission from deer to human. However, with an estimated case finding rate of only 4.2%, spillback to humans cannot be ruled out. The extensive transmission of SARS-CoV-2 within deer populations and the large number of unsampled cases highlights the need for active surveillance at the human-animal interface.
The COVID-19 pandemic has resulted in extensive surveillance of the genomic diversity of SARS-CoV-2. Sequencing data generated as part of these efforts can also capture the diversity of the SARS-CoV-2 virus populations replicating within infected individuals. To assess this within-host diversity of SARS-CoV-2 we quantified low frequency (minor) variants from deep sequence data of thousands of clinical samples collected by a large urban hospital system over the course of a year. Using a robust analytical pipeline to control for technical artefacts, we observe that at comparable viral loads, specimens from patients hospitalized due to COVID-19 had a greater number of minor variants than samples from outpatients. Since individuals with highly diverse viral populations could be disproportionate drivers of new viral lineages in the patient population, these results suggest that transmission control should pay special attention to patients with severe or protracted disease to prevent the spread of novel variants.
Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.
White-tailed deer (Odocoileus virginianus) are highly susceptible to infection by SARS-CoV-2, with multiple reports of widespread spillover of virus from humans to free-living deer. While the recently emerged SARS-CoV-2 B.1.1.529 Omicron variant of concern (VoC) has been shown to be notably more transmissible amongst humans, its ability to cause infection and spillover to non-human animals remains a challenge of concern. We found that 19 of the 131 (14.5%; 95% CI: 0.10–0.22) white-tailed deer opportunistically sampled on Staten Island, New York, between December 12, 2021, and January 31, 2022, were positive for SARS-CoV-2 specific serum antibodies using a surrogate virus neutralization assay, indicating prior exposure. The results also revealed strong evidence of age-dependence in antibody prevalence. A significantly (χ2, p < 0.001) greater proportion of yearling deer possessed neutralizing antibodies as compared with fawns (OR=12.7; 95% CI 4–37.5). Importantly, SARS-CoV-2 nucleic acid was detected in nasal swabs from seven of 68 (10.29%; 95% CI: 0.0–0.20) of the sampled deer, and whole-genome sequencing identified the SARS-CoV-2 Omicron VoC (B.1.1.529) is circulating amongst the white-tailed deer on Staten Island. Phylogenetic analyses revealed the deer Omicron sequences clustered closely with other, recently reported Omicron sequences recovered from infected humans in New York City and elsewhere, consistent with human to deer spillover. Interestingly, one individual deer was positive for viral RNA and had a high level of neutralizing antibodies, suggesting either rapid serological conversion during an ongoing infection or a “breakthrough” infection in a previously exposed animal. Together, our findings show that the SARS-CoV-2 B.1.1.529 Omicron VoC can infect white-tailed deer and highlights an urgent need for comprehensive surveillance of susceptible animal species to identify ecological transmission networks and better assess the potential risks of spillback to humans. Key Findings These studies provide strong evidence of infection of free-living white-tailed deer with the SARS-CoV-2 B.1.1.529 Omicron variant of concern on Staten Island, New York, and highlight an urgent need for investigations on human-to-animal-to-human spillovers/spillbacks as well as on better defining the expanding host-range of SARS-CoV-2 in non-human animals and the environment.
The continuing emergence of SARS-CoV-2 variants of concern (VOCs) presents a serious public health threat, exacerbating the effects of the COVID19 pandemic. Although millions of genomes have been deposited in public archives since the start of the pandemic, predicting SARS-CoV-2 clinical characteristics from the genome sequence remains challenging. In this study, we used a collection of over 29,000 high quality SARS-CoV-2 genomes to build machine learning models for predicting clinical detection cycle threshold (Ct) values, which correspond with viral load. After evaluating several machine learning methods and parameters, our best model was a random forest regressor that used 10-mer oligonucleotides as features and achieved an R2 score of 0.521 +/- 0.010 (95% confidence interval over 5 folds) and an RMSE of 5.7 +/- 0.034, demonstrating the ability of the models to detect the presence of a signal in the genomic data. In an attempt to predict Ct values for newly emerging variants, we predicted Ct values for Omicron variants using models trained on previous variants. We found that approximately 5% of the data in the model needed to be from the new variant in order to learn its Ct values. Finally, to understand how the model is working, we evaluated the top features and found that the model is using a multitude of k-mers from across the genome to make the predictions. However, when we looked at the top k-mers that occurred most frequently across the set of genomes, we observed a clustering of k-mers that span spike protein regions corresponding with key variations that are hallmarks of the VOCs including G339, K417, L452, N501, and P681, indicating that these sites are informative in the model and may impact the Ct values that are observed in clinical samples.