Motivation:Phage therapy offers a viable alternative for bacterial infections amid rising antimicrobial resistance. Its success relies on selecting safe and effective phage candidates that require comprehensive genomic screening to identify potential risks. However, this process is often labor intensive and time-consuming, hindering rapid clinical deployment. Results:We developed Sphae, an automated bioinformatics pipeline designed to streamline the therapeutic potential of a phage in under 10 minutes. Using Snakemake workflow manager, Sphae integrates tools for quality control, assembly, genome assessment, and annotation tailored specifically for phage biology. Sphae automates the detection of key genomic markers, including virulence factors, antimicrobial resistance genes, and lysogeny indicators such as integrase, recombinase, and transposase, which could preclude therapeutic use. Among the 65 phage sequences analyzed, 28 showed therapeutic potential, 8 failed due to low sequencing depth, 22 contained prophage or virulent markers, and 23 had multiple phage genomes. This workflow produces a report to assess phage safety and therapy suitability quickly. Sphae is scalable and portable, facilitating efficient deployment across most high-performance computing and cloud platforms, accelerating the genomic evaluation process. Availability and implementation:Sphae source code is freely available at https://github.com/linsalrob/sphae, with installation supported on Conda, PyPi, Docker containers.
Prokaryotes dominate the biosphere and form diverse communities disrupted by invasion. Invaders and remaining community members experience resource surfeit, competition, and selective pressures. Little is known about invasion in natural microbial communities. We examined invasion by chemotaxis in a meso-tube system at taxonomic, functional, and genomic levels as communities sank, rose, and formed a chemotactic band that migrated for metres. The band velocity increased as the community migrated despite non-motile bacterial hitchhikers and up to 10⁶ viruses/ml. Migrating communities left complex residual communities in their wake, showing dynamic taxonomic composition and adaptation through increased migration-associated genes. Approximately 500 species migrated together, competing for dominance. This system offers a superior method for studying band and residual community dynamics, bacterial hitchhiking, viral transport, gene evolution, and survival strategies, revealing cohesive communities that persist over extended distances. Our methods and results provide an experimental foundation for investigating microbial invasion in multiple ecological settings.
Microorganisms found in natural environments are fundamental components of ecosystems and play vital roles in various ecological processes.Studying their genomes can provide valuable insights into the diversity, functionality, and evolution of microbial life, as well as their impacts on human health.Once the genetic material is extracted from environmental samples, it undergoes sequencing using technologies like whole genome sequencing (WGS).The raw sequence data is then analysed, and computational methods are applied to assemble the fragmented sequences and reconstruct the complete microbial genomes (Wick et al.,
Improvements in the accuracy and availability of long-read sequencing mean that complete bacterial genomes are now routinely reconstructed using hybrid (i.e. short- and long-reads) assembly approaches. Complete genomes allow a deeper understanding of bacterial evolution and genomic variation beyond single nucleotide variants. They are also crucial for identifying plasmids, which often carry medically significant antimicrobial resistance genes. However, small plasmids are often missed or misassembled by long-read assembly algorithms. Here, we present Hybracter which allows for the fast, automatic and scalable recovery of near-perfect complete bacterial genomes using a long-read first assembly approach. Hybracter can be run either as a hybrid assembler or as a long-read only assembler. We compared Hybracter to existing automated hybrid and long-read only assembly tools using a diverse panel of samples of varying levels of long-read accuracy with manually curated ground truth reference genomes. We demonstrate that Hybracter as a hybrid assembler is more accurate and faster than the existing gold standard automated hybrid assembler Unicycler. We also show that Hybracter with long-reads only is the most accurate long-read only assembler and is comparable to hybrid methods in accurately recovering small plasmids.
BACKGROUND:Modern sequencing technologies offer extraordinary opportunities for virus discovery and virome analysis. Annotation of viral sequences from metagenomic data requires a complex series of steps to ensure accurate annotation of individual reads and assembled contigs. In addition, varying study designs will require project-specific statistical analyses. FINDINGS:Here we introduce Hecatomb, a bioinformatic platform coordinating commonly used tasks required for virome analysis. Hecatomb means "a great sacrifice." In this setting, Hecatomb is "sacrificing" false-positive viral annotations using extensive quality control and tiered-database searches. Hecatomb processes metagenomic data obtained from both short- and long-read sequencing technologies, providing annotations to individual sequences and assembled contigs. Results are provided in commonly used data formats useful for downstream analysis. Here we demonstrate the functionality of Hecatomb through the reanalysis of a primate enteric and a novel coral reef virome. CONCLUSION:Hecatomb provides an integrated platform to manage many commonly used steps for virome characterization, including rigorous quality control, host removal, and both read- and contig-based analysis. Each step is managed using the Snakemake workflow manager with dependency management using Conda. Hecatomb outputs several tables properly formatted for immediate use within popular data analysis and visualization tools, enabling effective data interpretation for a variety of study designs. Hecatomb is hosted on GitHub (github.com/shandley/hecatomb) and is available for installation from Bioconda and PyPI.
Formalin-fixed paraffin-embedded (FFPE) samples are valuable but underutilized in single-cell omics research due to their low RNA quality. In this study, leveraging a recent advance in single-cell genomic technology, we introduce snPATHO-seq, a versatile method to derive high-quality single-nucleus transcriptomic data from FFPE samples. We benchmarked the performance of the snPATHO-seq workflow against existing 10x 3' and Flex assays designed for frozen or fresh samples and highlighted the consistency in snRNA-seq data produced by all workflows. The snPATHO-seq workflow also demonstrated high robustness when tested across a wide range of healthy and diseased FFPE tissue samples. When combined with FFPE spatial transcriptomic technologies such as FFPE Visium, the snPATHO-seq provides a multi-modal sampling approach for FFPE samples, allowing more comprehensive transcriptomic characterization. A combination of an FFPE nuclei preparation protocol and a probe-based transcriptomic profiling technique enables snRNA-seq characterization of archival human FFPE tissues, holding promise for retrospective studies involving aged clinical cohorts.
Despite mounting evidence of their importance in human health and ecosystem functioning, the definition and measurement of 'healthy microbiomes' remain unclear. More advanced knowledge exists on health associations for compounds used or produced by microbes. Environmental microbiome exposures (especially via soils) also help shape, and may supplement, the functional capacity of human microbiomes. Given the synchronous interaction between microbes, their feedstocks, and micro-environments, with functional genes facilitating chemical transformations, our objective was to examine microbiomes in terms of their capacity to process compounds relevant to human health. Here we integrate functional genomics and biochemistry frameworks to derive new quantitative measures of in silico potential for human gut and environmental soil metagenomes to process a panel of major compound classes (e.g., lipids, carbohydrates) and selected biomolecules (e.g., vitamins, short-chain fatty acids) linked to human health. Metagenome functional potential profile data were translated into a universal compound mapping 'landscape' based on bioenergetic van Krevelen mapping of function-level meta-compounds and corresponding functional relative abundances, reflecting imprinted genetic capacity of microbiomes to metabolize an array of different compounds. We show that measures of 'compound processing potential' associated with human health and disease (examining atherosclerotic cardiovascular disease, colorectal cancer, type 2 diabetes and anxious-depressive behavior case studies), and displayed seemingly predictable shifts along gradients of ecological disturbance in plant-soil ecosystems (three case studies). Ecosystem quality explained 60-92 % of variation in soil metagenome compound processing potential measures in a post-mining restoration case study dataset. With growing knowledge of the varying proficiency of environmental microbiota to process human health associated compounds, we might design environmental interventions or nature prescriptions to modulate our exposures, thereby advancing microbiota-oriented approaches to human health. Compound processing potential offers a simplified, integrative approach for applying metagenomics in ongoing efforts to understand and quantify the role of microbiota in environmental- and human-health.
Phages integrated into a bacterial genome – called prophages – continuously monitor the vigour of the host bacteria to determine when to escape the genome and to protect their host from other phage infections, and they may provide genes that promote bacterial growth. Prophages are essential to almost all microbiomes, including the human microbiome. However, most human microbiome studies have focused on bacteria, ignoring free and integrated phages, so we know little about how these prophages affect the human microbiome. To address this gap in our knowledge, we compared the prophages identified in 14 987 bacterial genomes isolated from human body sites to characterize prophage DNA in the human microbiome. Here, we show that prophage DNA is ubiquitous, comprising on average 1–5 % of each bacterial genome. The prophage content per genome varies with the isolation site on the human body, the health of the human and whether the disease was symptomatic. The presence of prophages promotes bacterial growth and sculpts the microbiome. However, the disparities caused by prophages vary throughout the body.
Motivation:Phage therapy is a viable alternative for treating bacterial infections amidst the escalating threat of antimicrobial resistance. However, the therapeutic success of phage therapy depends on selecting safe and effective phage candidates. While experimental methods focus on isolating phages and determining their lifecycle and host range, comprehensive genomic screening is critical to identify markers that indicate potential risks, such as toxins, antimicrobial resistance, or temperate lifecycle traits. These analyses are often labor-intensive and time-consuming, limiting the rapid deployment of phage in clinical settings. Results:We developed Sphae, an automated bioinformatics pipeline designed to streamline therapeutic potential of a phage in under ten minutes. Using Snakemake workflow manager, Sphae integrates tools for quality control, assembly, genome assessment, and annotation tailored specifically for phage biology. Sphae automates the detection of key genomic markers, including virulence factors, antimicrobial resistance genes, and lysogeny indicators like integrase, recombinase, and transposase, which could preclude therapeutic use. Benchmarked on 65 phage sequences, 28 phage samples showed therapeutic potential, 8 failed during assembly due to low sequencing depth, 22 samples included prophage or virulent markers, and the remaining 23 samples included multiple phage genomes per sample. This workflow outputs a comprehensive report, enabling rapid assessment of phage safety and suitability for phage therapy under these criteria. Sphae is scalable, portable, facilitating efficient deployment across most high-performance computing (HPC) and cloud platforms, expediting the genomic evaluation process. Availability:Sphae is source code and freely available at https://github.com/linsalrob/sphae, with installation supported on Conda, PyPi, Docker containers.
Phages integrated into a bacterial genome-called prophages-continuously monitor the health of the host bacteria to determine when to escape the genome, protect their host from other phage infections, and may provide genes that promote bacterial growth. Prophages are essential to almost all microbiomes, including the human microbiome. However, most human microbiome studies focus on bacteria, ignoring free and integrated phages, so we know little about how these prophages affect the human microbiome. We compared the prophages identified in 11,513 bacterial genomes isolated from human body sites to characterise prophage DNA in the human microbiome. Here, we show that prophage DNA comprised an average of 1-5% of each bacterial genome. The prophage content per genome varies with the isolation site on the human body, the health of the human, and whether the disease was symptomatic. The presence of prophages promotes bacterial growth and sculpts the microbiome. However, the disparities caused by prophages vary throughout the body.
Spatial transcriptomics is a rapidly evolving field, overwhelmed by a multitude of technologies. This study aims to offer a comparative analysis of datasets generated from leading in situ imaging platforms. We have generated spatial transcriptomics data from serial sections of prostate adenocarcinoma using the 10x Genomics Xenium and NanoString CosMx SMI platforms. Additionally, orthogonal single-nucleus RNA sequencing (snRNA-seq) was performed on the same FFPE tissue to establish a reference for the tumor’s transcriptional profiles. We assessed various technical aspects, such as reproducibility, sensitivity, dynamic range, cell segmentation, cell type annotation, and congruence with single-cell profiling. The practicality of assessing cellular organization and biomarker localization was evaluated. Although fewer genes are measured (CosMx: 960, Xenium: 377, with an overlap of 125), Xenium consistently demonstrates higher sensitivity, a broader dynamic range, and better alignment with single-cell reference profiles. Conversely, CosMx’s out-of-the-box segmentation outperformed Xenium’s, resulting in noticeable transcript misassignment in Xenium within certain tissue areas. However, the impact of this on the cells’ transcriptional profile was minimal. Together, this comprehensive comparison of two leading commercial platforms for spatial transcriptomics provides essential metrics for assessing their performance, offering invaluable insights for future research and technological advancements in this dynamic field.### Competing Interest StatementL.G.M., J.T.P., D.P.C., D.P.K., K.B.J., K.W., M.J.R., F.S-D., N.K.R., M.Z., I.S.V., S.R.V.K., L.M.B., J.L.W. and N.B. declare no professional or financial affiliations with 10X Genomics or NanoString Technologies. During the conduct of this study, L.G.M. served as a Scientific Advisor for Millenium Sciences (no longer in this role), Omniscope, and ArgenTag. N.B. is Chief Scientist at Deepcell. S.R.V.K. is the founder of and a consultant for Faeth Therapeutics and Transomic Technologies. None of the authors received any form of payment or compensation from these companies. Consumables used in this study from both companies were purchased at full price, acquired at a discounted rate, or provided free of charge, although not specifically for this study. Subscription to AtoMx was provided at no cost by NanoString Technologies.
The gut virome is an incredibly complex part of the gut ecosystem. Gut viruses play a role in many disease states, but it is unknown to what extent the gut virome impacts everyday human health. New experimental and bioinformatic approaches are required to address this knowledge gap. Gut virome colonization begins at birth and is considered unique and stable in adulthood. The stable virome is highly specific to each individual and is modulated by varying factors such as age, diet, disease state, and use of antibiotics. The gut virome primarily comprises bacteriophages, predominantly order Crassvirales, also referred to as crAss-like phages, in industrialized populations and other Caudoviricetes (formerly Caudovirales). The stability of the virome's regular constituents is disrupted by disease. Transferring the fecal microbiome, including its viruses, from a healthy individual can restore the functionality of the gut. It can alleviate symptoms of chronic illnesses such as colitis caused by Clostridiodes difficile. Investigation of the virome is a relatively novel field, with new genetic sequences being published at an increasing rate. A large percentage of unknown sequences, termed 'viral dark matter', is one of the significant challenges facing virologists and bioinformaticians. To address this challenge, strategies include mining publicly available viral datasets, untargeted metagenomic approaches, and utilizing cutting-edge bioinformatic tools to quantify and classify viral species. Here, we review the literature surrounding the gut virome, its establishment, its impact on human health, the methods used to investigate it, and the viral dark matter veiling our understanding of the gut virome.
Motivation: Microbial communities have a profound impact on both human health and various environments. Viruses infecting bacteria, known as bacteriophages or phages, play a key role in modulating bacterial communities within environments. High-quality phage genome sequences are essential for advancing our understanding of phage biology, enabling comparative genomics studies and developing phage-based diagnostic tools. Most available viral identification tools consider individual sequences to determine whether they are of viral origin. As a result of challenges in viral assembly, fragmentation of genomes can occur, and existing tools may recover incomplete genome fragments. Therefore, the identification and characterization of novel phage genomes remain a challenge, leading to the need of improved approaches for phage genome recovery. Results: We introduce Phables, a new computational method to resolve phage genomes from fragmented viral metagenome assemblies. Phables identifies phage-like components in the assembly graph, models each component as a flow network, and uses graph algorithms and flow decomposition techniques to identify genomic paths. Experimental results of viral metagenomic samples obtained from different environments show that Phables recovers on average over 49% more high-quality phage genomes compared to existing viral identification tools. Furthermore, Phables can resolve variant phage genomes with over 99% average nucleotide identity, a distinction that existing tools are unable to make. Availability and implementation: Phables is available on GitHub at https://github.com/Vini2/phables.
Phages dominate every ecosystem on the planet. While virulent phages sculpt the microbiome by killing their bacterial hosts, temperate phages provide unique growth advantages to their hosts through lysogenic conversion. Many prophages benefit their host, and prophages are responsible for genotypic and phenotypic differences that separate individual microbial strains. However, the microbes also endure a cost to maintain those phages: additional DNA to replicate and proteins to transcribe and translate. We have never quantified those benefits and costs. Here, we analysed over two and a half million prophages from over half a million bacterial genome assemblies. Analysis of the whole dataset and a representative subset of taxonomically diverse bacterial genomes demonstrated that the normalised prophage density was uniform across all bacterial genomes above 2 Mbp. We identified a constant carrying capacity of phage DNA per bacterial DNA. We estimated that each prophage provides cellular services equivalent to approximately 2.4 % of the cell's energy or 0.9 ATP per bp per hour. We demonstrate analytical, taxonomic, geographic, and temporal disparities in identifying prophages in bacterial genomes that provide novel targets for identifying new phages. We anticipate that the benefits bacteria accrue from the presence of prophages balance the energetics involved in supporting prophages. Furthermore, our data will provide a new framework for identifying phages in environmental datasets, diverse bacterial phyla, and from different locations.
Bacteroides, the prominent bacteria in the human gut, play a crucial role in degrading complex polysaccharides. Their abundance is influenced by phages belonging to the Crassvirales order. Despite identifying over 600 Crassvirales genomes computationally, only few have been successfully isolated. Continued efforts in isolation of more Crassvirales genomes can provide insights into phage-host-evolution and infection mechanisms. We focused on wastewater samples, as potential sources of phages infecting various Bacteroides hosts. Sequencing, assembly, and characterization of isolated phages revealed 14 complete genomes belonging to three novel Crassvirales species infecting Bacteroides cellulosilyticus WH2. These species, Kehishuvirus sp. 'tikkala' strain Bc01, Kolpuevirus sp. 'frurule' strain Bc03, and 'Rudgehvirus jaberico' strain Bc11, spanned two families, and three genera, displaying a broad range of virion productions. Upon testing all successfully cultured Crassvirales species and their respective bacterial hosts, we discovered that they do not exhibit co-evolutionary patterns with their bacterial hosts. Furthermore, we observed variations in gene similarity, with greater shared similarity observed within genera. However, despite belonging to different genera, the three novel species shared a unique structural gene that encodes the tail spike protein. When investigating the relationship between this gene and host interaction, we discovered evidence of purifying selection, indicating its functional importance. Moreover, our analysis demonstrated that this tail spike protein binds to the TonB-dependent receptors present on the bacterial host surface. Combining these observations, our findings provide insights into phage-host interactions and present three Crassvirales species as an ideal system for controlled infectivity experiments on one of the most dominant members of the human enteric virome.
Background Analysis of viral diversity using modern sequencing technologies offers extraordinary opportunities for discovery. However, these analyses present a number of bioinformatic challenges due to viral genetic diversity and virome complexity. Due to the lack of conserved marker sequences, metagenomic detection of viral sequences requires a non-targeted, random (shotgun) approach. Annotation and enumeration of viral sequences relies on rigorous quality control and effective search strategies against appropriate reference databases. Virome analysis also benefits from the analysis of both individual metagenomic sequences as well as assembled contigs. Combined, virome analysis results in large amounts of data requiring sophisticated visualization and statistical tools. Results Here we introduce Hecatomb, a bioinformatics platform enabling both read and contig based analysis. Hecatomb integrates query information from both amino acid and nucleotide reference sequence databases. Hecatomb integrates data collected throughout the workflow enabling analyst driven virome analysis and discovery. Hecatomb is available on GitHub at . Conclusions Hecatomb provides a single, modular software solution to the complex tasks required of many virome analysis. We demonstrate the value of the approach by applying Hecatomb to both a host-associated (enteric) and an environmental (marine) virome data set. Hecatomb provided data to determine true- or false-positive viral sequences in both data sets and revealed complex virome structure at distinct marine reef sites. ### Competing Interest Statement The authors have declared no competing interest. * AIDS : acquired immunodeficiency syndrome SIV : simian immunodeficiency virus HPC : high-performance computing NCBI : National Center for Biotechnology Information RPKM : reads per kilobase million FPKM : fragments per kilobase million SPM : sequences per million LCA : lowest common ancestor ICTV : International Committee on Taxonomy of Viruses PERMANOVA : permutational analysis of variance PCoA : principal coordinate analysis ANOVA : analysis of variance SIMPER : similarity percentag
The epidermal microbiome is a critical element of marine organismal immunity, but the epidermal virome of marine organisms remains largely unexplored. The epidermis of sharks represents a unique viromic ecosystem. Sharks secrete a thin layer of mucus which harbors a diverse microbiome, while their hydrodynamic dermal denticles simultaneously repel environmental microbes. Here, we sampled the virome from the epidermis of three shark species in the family Carcharhinidae: the genetically and morphologically similar Carcharhinus obscurus (n = 6) and Carcharhinus galapagensis (n = 10) and the outgroup Galeocerdo cuvier (n = 15). Virome taxonomy was characterized using shotgun metagenomics and compared with a suite of multivariate analyses. All three sharks retain species-specific but highly similar epidermal viromes dominated by uncharacterized bacteriophages which vary slightly in proportional abundance within and among shark species. Intraspecific variation was lower among C. galapagensis than among C. obscurus and G. cuvier. Using both the annotated and unannotated reads, we were able to determine that the Carcharhinus galapagensis viromes were more similar to that of G. cuvier than they were to that of C. obscurus, suggesting that behavioral niche may be a more prominent driver of virome than host phylogeny.
Acinetobacter baumannii is an opportunistic human pathogen responsible for numerous severe nosocomial infections. Genome analysis on the A. baumannii clinical isolate 04117201 revealed the presence of 13 two-component signal transduction systems (TCS). Of these, we examined the putative TCS named here as StkSR. The stkR response regulator was deleted via homologous recombination and its progeny, ΔstkR, was phenotypically characterized. Antibiogram analyses of ΔstkR cells revealed a two-fold increase in resistance to the clinically relevant polymyxins, colistin and polymyxin B, compared to wildtype. PAGE-separation of silver stained purified lipooligosaccharide isolated from ΔstkR and wildtype cells ruled out the complete loss of lipooligosaccharide as the mechanism of colistin resistance identified for ΔstkR. Hydrophobicity analysis identified a phenotypical change of the bacterial cells when exposed to colistin. Transcriptional profiling revealed a significant up-regulation of the pmrCAB operon in ΔstkR compared to the parent, associating these two TCS and colistin resistance. These results reveal that there are multiple levels of regulation affecting colistin resistance; the suggested ‘cross-talk’ between the StkSR and PmrAB two-component systems highlights the complexity of these systems.