The presence of Vibrio parahaemolyticus (Vp) at various stages of seafood production has adversely affected public health and threatened the sustainability of the industry. Driven by the advancement of next-generation-sequencing technologies and public health data sharing initiative, an increasing volume of public Vp genomic data with metadata has become available, which serve as the foundation for building learning models to accurately differentiate isolation sources and further uncover gene-level determinants of pathogenic potential. The primary goal of this study was to develop and validate machine learning (ML) and deep learning (DL) algorithms to differentiate Vp strains from different isolation sources (clinical vs. environmental isolates) using pangenome assemblies and achieved robust and precise pathogenic potential prediction of Vp. The secondary goal of this study was to obtain critical biological insights revealing pathogenic potential-associated genes contributing to the isolation source difference from these established learning models. Based on the results, the developed learning models demonstrated strong performance, achieving an AUC greater than 0.95 in distinguishing clinical and environmental isolates using pangenome signals. Besides, the gene feature weight analysis from RF revealed the importance of specific accessory genes during Vp evolution including but not limited to functional unknown cloud genes, vspR, sctC5, and tdh1, which provides biological insights as potential future research directions. These findings essentially highlight critical importance of accessory and cloud genes in differentiating clinical and environmental isolates, and provide new insights into how recently acquired genes may contribute to pathogenic evolution of Vp. Additionally, the framework demonstrated in this study provides a cost-effective intelligent strategy by leveraging large public genomic datasets to support surveillance and risk assessment of seafood-associated pathogens.
The rapid rate of virus discovery renders manual curation by taxonomy experts increasingly impractical, creating a need for reliable software that can reproducibly assign viral contigs to taxa at all fifteen ranks of the virus taxonomy. We led an open community challenge for the computational taxonomic classification of viruses and assembled a dataset of virus sequences combining expert-curated and metagenomic sequences. Seventeen teams contributed a total of thirty-four automated, fully reproducible classification pipelines. Most tools correctly assigned viruses belonging to established species, genera, or families, but viruses that are unclassified at those lower ranks remain challenging. This study provides datasets, open-source software, novel approaches, and recommendations to benchmark computational taxonomic classification of viruses, and support organizing the many viruses discovered in big omics data.
Equine Juvenile Spinocerebellar Ataxia (EJSCA) is a novel autosomal recessive neurologic disease in Quarter Horses. Affected foals display a progressive proprioceptive ataxia by 1-5 weeks of age, leading to recumbency and necessitating euthanasia. Whole genome sequencing was performed on 7 EJSCA cases and unaffected horses that included 4 obligate carriers, 4 unaffected half or full-siblings, and 28 unrelated, unaffected control Quarter Horses. An 82 kb region of association was identified (EquCab3.0, chr11: 6963986-7045999), containing 9 candidate SNPs across four genes (FADS6, FDXR, GRIN2C and TMEM104). Decreased FDXR mRNA expression and a cryptic exon was identified in spinal cord tissue from EJSCA cases via RNA-sequencing. One of the 9 associated SNPs (FDXR-203 c.177 + 1778G > C) was the eighth base pair of this cryptic exon. Affected foals were all homozygous for the variant. Protein concentrations of FDXR were lower in EJSCA cases in spinal cord and liver compared to unaffected controls. The FDXR-203 c.177 + 1778G > C mutation represents the first non-coding neurological genetic variant in horses. Additionally, this is the first genetic cause of a degenerative axonopathy in the horse and a spontaneous disease model to study FDXR pathology in humans.
IntroductionNatural killer (NK) cells in mice and humans are key effectors of the innate immune system with complex immunoregulatory functions, and diverse subsets have been identified with distinct characteristics and roles. Companion dogs with spontaneous cancer have been validated as models of human disease, including cancer immunology and immunotherapy, and greater understanding of NK cell heterogeneity in dogs can inform NK biology across species and optimize NK immunotherapy for both dogs and people.MethodsHere, we assessed canine NK cell populations by single-cell RNA sequencing (scRNAseq) across blood, lung, liver, spleen, and placenta with comparison to human NK cells from blood and the same tissues to better characterize the differential gene expression of canine and human NK cells regarding ontogeny, heterogeneity, patterns of activation, inhibition, and tissue residence.ResultsOverall, we observed tissue-specific NK cell signatures consistent with immature NK cells in the placenta, mature and activated NK cells in the lung, and NK cells with a mixed activated and inhibited signature in the liver with significant cross-species homology.DiscussionTogether, our results point to heterogeneous canine NK populations highly comparable to human NK cells, and we provide a comprehensive atlas of canine NK cells across organs which will inform future cross-species NK studies and further substantiate the spontaneous canine model to optimize NK immunotherapy across species.
Our previous research observed that dietary supplementation of 500 mg/kg of Bacillus subtilis reduced the frequency of diarrhea and enhanced growth performance of weaned pigs experimentally infected with a pathogenic Escherichia coli (E. coli). This study aimed to further investigate the effects of dietary Bacillus subtilis or carbadox on the functional antimicrobial resistance gene (AMR) composition and abundance of the microbial community in feces collected from weaned pigs in this research project. The four experimental treatments were: 1) negative control, pigs fed with control diet without E. coli challenge, 2) positive control, pigs fed with control diet with E. coli challenge, 3) antibiotic group, pigs fed diet supplemented with 50 mg/kg of carbadox with E. coli challenge, and 4) probiotics group, pig fed diet supplemented with 500 mg/kg of Bacillus subtilis probiotics with E. coli challenge. A total of 32 fecal samples were directly collected from the rectum of pigs on day 21 after the first E. coli inoculation with 8 samples per treatment. Total microbial community DNA was extracted from fecal samples and then submitted to UC Davis Genome Center for Whole Genome- Shotgun Sequencing (NovaSeq S4 PE150) on the Illumina platform. Microbiota data analysis were conducted using Sourmash and R packages. AMR gene analysis was performed with the ATLAS pipeline and AMRFinderPlus. Supplementing Bacillus subtilis or antibiotics did not impact alpha diversity, including Chao1 and Shannon indices. The bacterial community composition in fecal samples collected from pigs in the antibiotics group were more separated (P < 0.05) from those in the other treatments. At the phylum level, supplementation of antibiotics reduced (P < 0.05) the relative abundance of Bacillota but increased (P < 0.10) the relative abundance of Bacteroidota and Eremiobacterota in feces of pigs compared with those in other groups. At the family level, supplementation of antibiotics had the lowest (P < 0.05) relative abundance of Lactobacillaceae, Streptococcaceae, Ruminococcaceae, CAG-239, and Burkholderiaceae among all treatments. However, pigs fed with antibiotics had the highest (P < 0.05) relative abundance of Erysipelotrichaceae, Turicibacteraceae, Clostridiaceae, Oscillospiraceae, Peptostreptococcaceae, and Succinivibrionaceae among all dietary treatments. A total of 77 AMR gene determinants were identified in fecal samples. 38 AMR genes that belong to 8 drug classes were shared by all treatments. The presence of aph(2’’)-lla were lower (P < 0.10) in Bacillus subtilis group compared with other treatments. The presence of tetB(P) were lower (P < 0.05) in the positive control among all treatments. The presence of cfx(A) were higher (P < 0.05) in negative and positive control than Bacillus subtilis treatment. In conclusion, dietary Bacillus subtilis and carbadox have different impacts on microbiota community and the presence of antimicrobial resistance gene in feces of weaned pigs.
Understanding genetic variation is necessary to unravel the complexities of evolution and diverse traits across species. Short genetic variants (<50 bp in length) represent key genomic variations that play a crucial role in shaping the genetic landscape of populations. Human short genetic variant databases are often considered the gold standard for variant repositories, and this review compares them with variant databases created for various animal species. This review examines the methodologies, data integration, and various applications that differ or align between human and animal databases. The goal is to identify challenges in leveraging genetic variation across species, recommend strategies for overcoming these challenges, and suggest future directions for research and database development.
With advancing genomic technologies, single-nucleotide polymorphism (SNP) arrays and whole genome sequencing (WGS) have become essential tools in equine genetic research. In this study, we assessed the concordance in SNP calls and trait-mapping efficacy by comparing data of 21 horses both genotyped on the Equine 670 K SNP array and sequenced at either ~12× or ~30× depth. Our analysis revealed that higher sequencing depths were significantly associated with fewer discordant calls between platforms. Additionally, we investigated the most frequent no-call and discordant positions and identified positions that were indels or multiallelic in the WGS. To assess the effectiveness of the 670 K SNP array vs. WGS in trait association studies, we mapped the chestnut coat color. Both technologies showed a clear peak at the expected locus, although neither association had loci reaching Bonferroni-corrected statistical significance, which was not statistically possible in this small group of horses. The findings of this study provide valuable insights for making informed decisions when selecting between SNP arrays and WGS at varying sequencing depths for equine genomic research applications.
Comparisons of long-read and short-read (meta)genome assemblies typically show that short-read sequence assemblies are less error-prone, but struggle to assemble complicated genome regions (e.g. repeats) compared to long-read sequence assemblies. Accurate metagenome assembly is especially challenging in diverse environments, such as soil, and long-read sequencing has been shown to improve assembly. Here, we use metagenomic data with paired long-read and short-read sequences to identify specific factors that impact genome assembly and assess their relative importance in a natural soil community. Our analysis suggests that low coverage and high sequence diversity are the two main factors leading to misassemblies in short-read data, and many of these "missed" regions tend to be variable parts of the genome, such as integrated viruses or defense system islands. Taken together, our results demonstrate that short-read metagenomes can possibly underestimate the diversity of these genome regions and that long-read sequencing can complement short-read metagenomes by improving assembly contiguity and the recovery of variable regions.
IntroductionNatural killer (NK) cells have great potential to extend the promise of cancer immunotherapy, but additional research is needed to improve their efficacy in solid cancers. Dogs develop spontaneous cancers with striking similarities to humans and can serve as a crucial link to bridge murine studies and human clinical trials to improve treatment outcomes across species and identify potential biomarkers of response.MethodsUsing single-cell RNA sequencing (scRNAseq), we integrated blood, tissue, and tumor samples from dog and human donors to compare NK cell gene expression and develop a canine sarcoma infiltrating NK signature. Canine tissue and tumor NK cell signatures were then used to contextualize NK cell changes in first-in-dog immunotherapy clinical trials.ResultsTumor infiltrating NK cells from both canine and human sarcomas exhibited enhanced migration with a simultaneously exhausted signature that most closely correlated transcriptionally with NK cells isolated from the liver. We also analyzed peripheral blood NK cells from dogs on first-in-dog clinical trials undergoing three distinct NK-targeting immunotherapy regimens, observing that dogs with favorable responses demonstrated increased NK proportions posttreatment. Genes upregulated in NK cells in the peripheral blood of good responders included genes associated with activated NK cells and revealed post-treatment gene expression changes in the blood as a predictor of response.DiscussionOverall, NK effector functions are well adapted to their tissue of residence but dysregulated in sarcoma infiltrating NK cells despite enhanced migration. We describe NK cell trends across canine clinical trials as a platform through which we can elucidate mechanisms of response and determine novel immunotherapy strategies to improve cancer outcomes in both humans and dogs.
The presence of Vibrio parahaemolyticus ( Vp ) at various stages of seafood production has adversely affected public health and threatened the sustainability of the industry. To address the critical public health threats posed by this prevalent seafood-borne pathogen, this research applied advanced machine learning (ML) and deep learning (DL) algorithms to predict the pathogenic potential of Vp using pangenome data. Utilizing comprehensive pangenomic assemblies and sophisticated ML/DL models, this study achieved robust and precise pathogenic potential prediction of Vp based on source attribution, which provides a novel reliable diagnostic tool facilitating conventional serotyping and virulence gene combination approaches. Based on results, non-core regions in Vp pangenome exhibited useful signals ML models can utilize in pathogenic potential determination process. Tree-based ensemble learning methods (Random Forest and Gradient Boosting Trees) have shown the distinguished performance with AUC score 0.97 based on selected pangenome matrix. Furthermore, Convolutional Neural Network successfully predicted the pathogenic potential of isolates with slightly better performance with AUC score 0.98 based on full pangenome. Critical biological insights revealing critical pathogenic potential-associated genes were retrieved from established ML/DL models: the gene feature weight analysis from Random Forest revealed the importance of accessory genes during Vp evolution (similarly highlighted by Gram-cam analysis of Convolutional Neural Network), which provided potential guidance for future research direction. ### Competing Interest Statement The authors have declared no competing interest.
ABSTRACT The evolution of oxygenic photosynthesis in the Cyanobacteria was one of the most transformative events in Earth history, eventually leading to the oxygenation of Earth’s atmosphere. However, it is difficult to understand how the earliest Cyanobacteria functioned or evolved on early Earth in part because we do not understand their ecology, including the environments in which they lived. Here, we use a cutting-edge bioinformatics tool to survey nearly 500,000 metagenomes for relatives of the taxa that likely bookended the evolution of oxygenic photosynthesis to identify the modern environments in which these organisms live. Ancestral state reconstruction suggests that the common ancestors of these organisms lived in terrestrial (soil and/or freshwater) environments. This restricted distribution may have increased the lag between the evolution of oxygenic photosynthesis and the oxygenation of Earth’s atmosphere.IMPORTANCECyanobacteria generate oxygen as part of their metabolism and are responsible for the rise of oxygen in Earth’s atmosphere over two billion years ago. However, we do not know how long this process may have taken. To help constrain how long this process would have taken, it is necessary to understand where the earliest Cyanobacteria may have lived. Here, we use a cutting-edge bioinformatics tool called branch water to examine the environments where modern Cyanobacteria and their relatives live to constrain those inhabited by the earliest Cyanobacteria. We find that these species likely lived in non-marine environments. This indicates that the rise of oxygen may have taken longer than previously believed.
BACKGROUND:Long-read sequencing (LRS) enables high-quality structural variant (SV) discovery. SV genotypers utilize these precise call sets to improve the recall and precision of genotyping in short-read sequencing (SRS) samples. With the extensive growth in publicly available SRS datasets, it is now possible to calculate accurate population allele frequencies of SVs. However, reprocessing hundreds of terabytes of raw SRS data to genotype new variants is impractical for population-scale studies, a computational challenge known as the N+1 problem (i.e., the challenge of re-genotyping an entire cohort for one additional variant). Overcoming this computational bottleneck is essential for analyzing new SVs from the growing number of pangenomes, public genomic databases, and pathogenic variant discovery studies. RESULTS:We propose the Great Genotyper, a population-scale genotyping workflow to address the N+1 problem. Applied to a human dataset, the workflow begins by preprocessing 4.2k short-read samples of a total of 183 TB raw data to create an 867-GB Counting Colored de Bruijn Graph (CCDG). The Great Genotyper uses this CCDG to genotype a list of phased or unphased variants, leveraging the CCDG population information to increase both precision and recall. The Great Genotyper offers the same accuracy as the state-of-the-art genotypers while achieving unprecedented performance. It took about 100 hours to genotype 4.5M variants across the 4.2k samples and calculate their population allele frequencies using 1 server with 32 cores and 145 GB of memory. The Great Genotyper opens the door to new ways to study SVs. For example, using the premade index, we demonstrate the Great Genotyper's application in finding pathogenic variants by calculating accurate allele frequency for novel SVs. Also, we used it to create a 4k reference panel by genotyping variants from the Human Pangenome Reference Consortium (HPRC). The new reference panel allows for SV imputation from genotyping microarrays. Moreover, we genotype the human GWAS Catalog and merge its variants with the 4k reference panel. We show 6,253 events of high linkage between the HPRC's SVs and nearby GWAS single-nucleotide polymorphisms, which can help in interpreting the effect of these SVs on gene functions. This analysis uncovers the detailed haplotype structure of the human fibrinogen locus and revives the pathogenic association of a 28-bp insertion in the FGA gene with thromboembolic disorders. CONCLUSION:The Great Genotyper solves the N+1 problem for population-scale genotyping of small and structural variants, offering both high accuracy and efficiency. Its ability to rapidly re-genotype large cohorts paves the road for several new studies of SVs.
The volume of biological data being generated by the scientific community is growing exponentially, reflecting technological advances and research activities. The National Institutes of Health's (NIH) Sequence Read Archive (SRA), which is maintained by the National Center for Biotechnology Information (NCBI) at the National Library of Medicine (NLM), is a rapidly growing public database that researchers use to drive scientific discovery across all domains of life. This increase in available data has great promise for pushing scientific discovery but also introduces new challenges that scientific communities need to address. As genomic datasets have grown in scale and diversity, a parade of new methods and associated software have been developed to address the challenges posed by this growth. These methodological advances are vital for maximally leveraging the power of next-generation sequencing (NGS) technologies. With the goal of laying a foundation for evaluation of methods for petabyte-scale sequence search, the Department of Energy (DOE) Office of Biological and Environmental Research (BER), the NIH Office of Data Science Strategy (ODSS), and NCBI held a virtual codeathon 'Petabyte Scale Sequence Search: Metagenomics Benchmarking Codeathon' on September 27 - Oct 1 2021, to evaluate emerging solutions in petabyte scale sequence search. The codeathon attracted experts from national laboratories, research institutions, and universities across the world to (a) develop benchmarking approaches to address challenges in conducting large-scale analyses of metagenomic data (which comprises approximately 20 that benefit from SRA-wide searches and the tools required to execute the search, and (c) produce community resources i.e. a public facing repository with information to rebuild and reproduce the problems addressed by each team challenge.
Traditional taxonomy provides a hierarchical organization of bacteria and archaea across taxonomic ranks from kingdom to subspecies. More recently, bacterial taxonomy has been more robustly quantified using comparisons of sequenced genomes, as in the Genome Taxonomy Database (GTDB), resolving down to genera and species. Such taxonomies have proven useful in many contexts, yet lack the flexibility and resolution of a more fine-grained approach. We apply our Life Identification Number (LIN) approach as a common, quantitative framework to tie existing (and future) bacterial taxonomies together, increase the resolution of genome-based discrimination of taxa, and extend taxonomic identification below the species level in a principled way. We utilize our existing concept of a LINgroup as an organizational concept for microorganisms that are closely related by overall genomic similarity, to help resolve some of the confusions and unforeseen negative effects of nomenclature changes of microbes due to genome-based reclassification. Our results obtained from experimentation demonstrate the value of LINs and LINgroups in mapping between taxonomies, translating between different nomenclatures, and integrating them into a single taxonomic framework.
Genetic selection has remarkably helped U.S. dairy farms to decrease their carbon footprint by more than doubling milk production per cow over time. Despite the environmental and economic benefits of improved feed and milk production efficiency, there is a critical need to explore phenotypical variance for feed utilization to advance the long-term sustainability of dairy farms. Feed is a major expense in dairy operations, and their enteric fermentation is a major source of greenhouse gases in agriculture. The challenges to expanding the phenotypic database, especially for feed efficiency predictions, and the lack of understanding of its drivers limit its utilization. Herein, we leveraged an artificial intelligence approach with feature engineering and ensemble methods to explore the predictive power of the rumen microbiome for feed and milk production efficiency traits, as rumen microbes play a central role in physiological responses in dairy cows. The novel ensemble method allowed to further identify key microbes linked to the efficiency measures. We used a population of 454 genotyped Holstein cows in the U.S. and Canada with individually measured feed and milk production efficiency phenotypes. The study underscored that the rumen microbiome is a major driver of residual feed intake ( RFI ), the most robust feed efficiency measure evaluated in the study, accounting for 36% of its variation. Further analyses showed that several alpha-diversity metrics were lower in more feed-efficient cows. For RFI, [Ruminococcus] gauvreauii group was the only genus positively associated with an improved feed efficiency status while seven other taxa were associated with inefficiency. The study also highlights that the rumen microbiome is pivotal for the unexplained variance in milk fat and protein production efficiency. Estimation of the carbon footprint of these cows shows that selection for better RFI could reduce up to 5 kg of diet consumed per cow daily, potentially reducing up to 37.5% of CH 4 . These findings shed light that the integration of artificial intelligence approaches, microbiology, and ruminant nutrition can be a path to further advance our understanding of the rumen microbiome on nutrient requirements and lactation performance of dairy cows to support the long-term sustainability of the dairy community.
Irber et al., (2024). sourmash v4: A multitool to quickly search, compare, and analyze genomic and metagenomic data sets. Journal of Open Source Software, 9(98), 6830, https://doi.org/10.21105/joss.06830
Background Natural killer (NK) cells are cytotoxic cells capable of recognizing heterogeneous cancer targets without prior sensitization, making them promising prospects for use in cellular immunotherapy. Companion dogs develop spontaneous cancers in the context of an intact immune system, representing a valid cancer immunotherapy model. Previously, CD5 depletion of peripheral blood mononuclear cells (PBMCs) was used in dogs to isolate a CD5dim-expressing NK subset prior to co-culture with an irradiated feeder line, but this can limit the yield of the final NK product. This study aimed to assess NK activation, expansion, and preliminary clinical activity in first-in-dog clinical trials using a novel system with unmanipulated PBMCs to generate our NK cell product.Methods Starting populations of CD5-depleted cells and PBMCs from healthy beagle donors were co-cultured for 14 days, phenotype, cytotoxicity, and cytokine secretion were measured, and samples were sequenced using the 3’-Tag-RNA-Seq protocol. Co-cultured human PBMCs and NK-isolated cells were also sequenced for comparative analysis. In addition, two first-in-dog clinical trials were performed in dogs with melanoma and osteosarcoma using autologous and allogeneic NK cells, respectively, to establish safety and proof-of-concept of this manufacturing approach.Results Calculated cell counts, viability, killing, and cytokine secretion were equivalent or higher in expanded NK cells from canine PBMCs versus CD5-depleted cells, and immune phenotyping confirmed a CD3-NKp46+ product from PBMC-expanded cells at day 14. Transcriptomic analysis of expanded cell populations confirmed upregulation of NK activation genes and related pathways, and human NK cells using well-characterized NK markers closely mirrored canine gene expression patterns. Autologous and allogeneic PBMC-derived NK cells were successfully expanded for use in first-in-dog clinical trials, resulting in no serious adverse events and preliminary efficacy data. RNA sequencing of PBMCs from dogs receiving allogeneic NK transfer showed patient-unique gene signatures with NK gene expression trends in response to treatment.Conclusions Overall, the use of unmanipulated PBMCs appears safe and potentially effective for canine NK immunotherapy with equivalent to superior results to CD5 depletion in NK expansion, activation, and cytotoxicity. Our preclinical and clinical data support further evaluation of this technique as a novel platform for optimizing NK immunotherapy in dogs.
1AbstractLong-read sequencing (LRS) enables variant calling of high-quality structural variants (SVs). Genotypers of SVs utilize these precise call sets to increase the recall and precision of genotyping in short-read sequencing (SRS) samples. With the extensive growth in availabilty of SRS datasets in recent years, we should be able to calculate accurate population allele frequencies of SV. However, reprocessing hundreds of terabytes of raw SRS data to genotype new variants is impractical for population-scale studies, a computational challenge known as the N+1 problem. Solving this computational bottleneck is necessary to analyze new SVs from the growing number of pangenomes in many species, public genomic databases, and pathogenic variant discovery studies.To address the N+1 problem, we propose The Great Genotyper, a population genotyping workflow. Applied to a human dataset, the workflow begins by preprocessing 4.2K short-read samples of a total of 183TB raw data to create an 867GB Counting Colored De Bruijn Graph (CCDG). The Great Genotyper uses this CCDG to genotype a list of phased or unphased variants, leveraging the CCDG population information to increase both precision and recall. The Great Genotyper offers the same accuracy as the state-of-the-art genotypers with the addition of unprecedented performance. It took 100 hours to genotype 4.5M variants in the 4.2K samples using one server with 32 cores and 145GB of memory. A similar task would take months or even years using single-sample genotypers.The Great Genotyper opens the door to new ways to study SVs. We demonstrate its application in finding pathogenic variants by calculating accurate allele frequency for novel SVs. Also, a premade index is used to create a 4K reference panel by genotyping variants from the Human Pangenome Reference Consortium (HPRC). The new reference panel allows for SV imputation from genotyping microarrays. Moreover, we genotype the GWAS catalog and merge its variants with the 4K reference panel. We show 6.2K events of high linkage between the HPRC’s SVs and nearby GWAS SNPs, which can help in interpreting the effect of these SVs on gene functions. This analysis uncovers the detailed haplotype structure of the human fibrinogen locus and revives the pathogenic association of a 28 bp insertion in the FGA gene with thromboembolic disorders.