High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.
Most microbes grow in spatially structured communities, and this profoundly shapes their ecology and evolution. At the microscale, short interaction ranges and steep nutrient gradients underlie cross-feeding, quorum sensing, and niche construction, generating spatial patterns that influence microbial behavior, community assembly, and stability. Here, we review theoretical and experimental evidence for how spatial organization drives eco-evolutionary processes, including founder effects during colonization, allele surfing during range expansion, emergent patterns that facilitate multilevel selection, and the exploration of rare epistatic genotypes. While the ecological and evolutionary consequences of spatial structure at the microscale are becoming clearer, linking these processes across scales to predict community- and ecosystem-level outcomes remains a major challenge. Addressing spatial interactions explicitly in microbiome research will be key. Recent advances in computational modeling, cultivation approaches, and omics now offer unprecedented opportunities to meet this challenge, providing fresh insights into how spatial structure governs the organization and dynamics of the microbial world across scales.
Understanding the adaptations of microorganisms to their environment is key to predicting the stability and dynamics of microbial communities. To uncover molecular mechanisms of environmental response, we extracted genomic features from 13,554 prokaryotic isolates, and trained machine learning models to identify which ones are most strongly associated with the microbial salinity, temperature, oxygen, and pH preferences. To extract these features in high throughput, including gene families, non-coding RNAs (ncRNAs), oligonucleotides, and amino acid usage, we built FxTractor, a scalable and adjustable pipeline available at: https://github.com/MGXlab/FxTractor. We validated the performance of our models with experimental data from a newly isolated deep-sea extremophile belonging to the genus Limnochorda that is not well-represented among the ML training sets, showing strong agreement between predictions and the conditions used to isolate this strain. Our analysis revealed specific gene and ncRNA families associated with each of the four environmental parameters, uncovering both established and potentially new molecular mechanisms. Examples include the bacterial large Signaling Recognition Particle in isolates that are able to grow at high temperatures (≥55°C), suggesting a role in translational pausing and structural stability under thermal stress. We also found the anti-hemB ncRNA to be associated with low-salinity (<0.7% NaCl), indicating a conserved antisense mechanism regulating the energetic costs of heme biosynthesis. Together, these findings provide new insights into microbe-environment interactions, and show how FxTractor enables high throughput discovery of genomic associations.
Abstract The microbiomes of wild animals, despite being integral to host survival, remain largely unexplored, particularly beyond the gut. Here, using > 450 museum-preserved dental calculus samples from 34 ecologically and phylogenetically diverse mammalian species, most of them previously unstudied, we determine the main drivers of oral microbiome evolution. We show that, similar to the gut, the oral microbiome is shaped by host ecology, particularly diet, and to a lesser extent host phylogenetic relationships. It may contribute to host adaptation, by synthesising essential micronutrients and degrading potentially harmful compounds. We also find that some oral microorganisms consistently associated with mammals throughout over evolutionary time scales and provide evidence for a likely oral origin of some rumen taxa. Together our findings demonstrate the importance of this mostly overlooked microbial community for mammalian biology.
ABSTRACT Background Urinary tract infections (UTIs) represent a major public health concern, increasingly complicated by rising antibiotic resistance, diminishing treatment efficacy, and increasing prevalence of recurrence. The urinary tract microbiome (urobiome) remains poorly characterized, despite its potential role in UTIs. Results To provide a comprehensive overview of the urobiome, we here integrated seven publicly available shotgun metagenomics studies, linking microbial composition to clinical infection status or diagnostics. Community-level analyses revealed distinct urobiome clusters, defined by one predominant bacterial taxon. Genome-resolved metagenomics allowed recovery of high-quality metagenome-assembled genomes (MAGs), enabling phylogenetic reconstruction, prediction of pathogenic potential, and profiling of antimicrobial resistance genes across multiple taxa. We then analyzed the pangenome of the clinically significant Escherichia coli species and found that its genomic variation is driven more by phylotype than isolation source or UTI status, supporting a model of opportunistic infection. Conclusions Taken together, our analyses represent a systematic, cross-study view of the urobiome that emphasizes the ecological complexity of the urobiome and the importance of integrating functional and phylogenetic information when studying UTIs.
Bacteriophages (phages) play essential roles in microbial systems, yet most phage proteins remain poorly characterised. Protein tertiary and quaternary structure information contributes valuable information about protein function. As many phage proteins function as homooligomers, complexes that consist of multiple identical subunits, there is great interest in computationally predicting their configurations. Here we present a computational framework, the Phage Homomer Level Estimate and Generation Method (PHLEGM) for inferring homooligomeric states directly from the protein sequence by combining AlphaFold-Multimer modelling with inter-subunit interface quality assessment. We proceeded to experimentally validate two out of nine predicted homooligomers using size exclusion chromatography and complementary hydrodynamic techniques. These efforts confirmed our predictions for a dimer and a trimer, highlighting the value of experimentally benchmarked computational predictions and showing the challenges of heterologous phage protein production. Applied to >22,000 phage protein sequences in the PHROGs database, our approach revealed extensive diversity in phage homooligomeric protein complexes. Benchmarking against protein language model-based predictors on a curated reference set of known phage homooligomers demonstrated superior accuracy of our structure-based method, achieving robust performance in classifying protein homooligomeric states, with the highest accuracy observed for trimers and higher-order complexes. These results highlight the value of computational predictions to decipher the complexities of the vast viral sequence space. All predicted complex structures and functional inferences are made publicly available to support structural and functional studies of phage proteins.
Bacteriophages play critical roles in microbial ecosystems, yet their dynamics in complex natural communities remain poorly understood compared to simplified laboratory systems. Here, we tracked viral dynamics in 20 compost-derived microbial communities over 1 year. Communities formed two alternative stable types, each dominated by distinct cellulose degraders and comprising hundreds of genera. In one community type, we observed massive, parallel outbreaks of Theomophage, a previously uncharacterized member of the Schitoviridae, reaching up to 74% of metagenomic reads-the largest bacteriophage outbreak documented to date. Despite extensive replication, Theomophage displayed notable genetic stability during outbreaks and over time. In contrast, the experimental migration of viral communities triggered rapid evolution driven by recombination and the accumulation of newly arising mutations, particularly after colonization of communities of the alternative type in which the phage was initially absent. These results reveal the spatial and temporal scales at which bacteriophage microdiversity evolves in complex ecosystems and show that viral mixing, likely common in nature, can rapidly accelerate phage evolution.
The rapid rate of virus discovery renders manual curation by taxonomy experts increasingly impractical, creating a need for reliable software that can reproducibly assign viral contigs to taxa at all fifteen ranks of the virus taxonomy. We led an open community challenge for the computational taxonomic classification of viruses and assembled a dataset of virus sequences combining expert-curated and metagenomic sequences. Seventeen teams contributed a total of thirty-four automated, fully reproducible classification pipelines. Most tools correctly assigned viruses belonging to established species, genera, or families, but viruses that are unclassified at those lower ranks remain challenging. This study provides datasets, open-source software, novel approaches, and recommendations to benchmark computational taxonomic classification of viruses, and support organizing the many viruses discovered in big omics data.
Abstract Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.
MOTIVATION:Oxford Nanopore sequencing enables long-read analysis for diverse applications, but artefacts introduced by Nanopore barcoding are poorly characterized and can compromise demultiplexing accuracy and downstream analyses. RESULTS:Using a rapid barcoding experiment on 66 diagnostic samples, we found that only 83% of reads followed the expected single-barcode configuration, while 17% showed complex barcode attachments. We observed similar patterns in public datasets, and also in native barcoding datasets where only 30%-70% of the reads had barcodes on both ends. Widely used demultiplexers, including Dorado, fail to resolve these cases, leaving ∼10% of our rapid barcoding reads partially trimmed and contaminated with adapter fragments. We developed Barbell, a pattern-aware demultiplexer that is designed to detect complex barcode configurations. Barbell reduced contaminated reads from >400 000 (Dorado/Flexiplex) to 166 (99.96% reduction), minimized barcode bleeding, and supports custom experimental designs such as dual-end barcodes and shorter barcodes (e.g. Illumina barcodes). We further show that such contamination is widespread in public databases, with Nanopore sequences detected in hundreds of NCBI entries, some of which are responsible for artificial taxonomic connections. AVAILABILITY AND IMPLEMENTATION:Barbell is open source and available at https://github.com/rickbeeloo/barbell.
Abstract Metagenomics has vastly expanded our knowledge of the human gut virome, yet the focus on dominant taxa has missed low-abundant but highly prevalent bacteriophages. Here, high-sensitivity taxonomic profiling of 53,976 samples from 42 countries, including non-industrialised populations and ancient DNA from five archaeological sites, reveals that only 0.3% of 64,238 phage genera are cosmopolitan, occurring in >20% of all subjects and on six continents. The most prevalent of these is Mushuvirus , with two distinct species detected in 60% of individuals, previously missed due to artefacts in reference genomes and their consistently low relative abundance (<0.8%). Population genomics shows that while one species is globally ubiquitous, the other predominates in individuals with West Eurasian ancestry. Together with 12 additional novel genera, we propose to classify these viruses into a new family Mushuviridae , found in ∼89% of studied humans. These transposable phages integrate into diverse hosts from two taxonomic orders of short-chain fatty acid–producing bacteria. Multi-omics experiments with Faecalibacterium demonstrate that hosts produce and secrete Mushuvirus -encoded receptor-binding proteins, which are actively diversified in lysogenic hosts. Together, these findings illuminate a globally distributed, active phage lineage that has persisted for millennia and continues to reshape the functional repertoire of bacteria central to human gut health.
Background Oxford Nanopore sequencing enables long-read sequencing across diverse applications, yet the experimental artifacts introduced by Nanopore barcoding are not well characterized. These artifacts can affect demultiplexing accuracy and downstream analyses. Results We performed a rapid barcoding experiment on 66 diagnostic samples and found that 83% of reads carried the expected single-barcode pattern, while 17% contained multiple barcodes or other artifacts. Current demultiplexers, including the widely used Dorado, fail to correctly handle these complex cases, leaving approximately 7% of reads partially trimmed and contaminated with adapter fragments. Additional issues include the presence of two barcodes at the same read end—either identical, originating from the same sample, or different, introduced after pooling. The latter can lead to barcode bleeding when the outer barcode is incorrectly selected. To address these challenges, we developed Barbell, a pattern-aware demultiplexer that detects all barcode configurations. Barbell reduces trimming errors by three orders of magnitude, minimizes barcode bleeding, and supports custom experimental setups such as shorter barcodes, dual-end barcodes, and custom flank sequences. Conclusions Our results highlight the impact of complex barcode attachments in Nanopore sequencing and demonstrate that Barbell drastically reduces their effects on downstream analyses. Barbell is open source and available at . ### Competing Interest Statement The authors have declared no competing interest. ZonMw, The Dutch Organisation for knowledge and innovation in health, healthcare and well-being, https://ror.org/01yaj9a77, 541003001 European Research Council, https://ror.org/0472cxd90, 865694 DiversiPHI, Deutsche Forschungsgemein schaft (DFG, German Research Foundation) under Germany’s Excellence Strategy - EXC 2051, 390713860 Alexander von Humboldt Foundation in the context of an Alexander von Humboldt-Professorship founded by German Federal Ministry of Education and Research.
Although virus ecogenomics has expanded access to and understanding of the virosphere, existing classification tools lack taxonomic resolution and are unable to scale to modern discovery-based datasets or classify previously unknown sequence space. Here we develop vConTACT3-a machine learning-based tool that improves scalability and accuracy of virus taxonomy. By optimizing gene-sharing thresholds and leveraging adaptive, realm-specific cut-offs, vConTACT3 expands classification to both eukaryote and prokaryote viruses for four of the six officially recognized realms, and establishes accurate hierarchical taxonomy from genus to order. Specifically, vConTACT3 achieves >95% agreement with official taxonomy for 35,545 and 13,524 public prokaryotic and eukaryotic virus genomes, respectively, to surpass vConTACT2 across most realms, while still uniquely classifying previously uncharacterized taxa, and doing so even faster. vConTACT3 application provides taxonomy assignments for tens of thousands of unclassified taxa rapidly, automatically and systematically; evaluates virus sequence space to reveal support for fewer taxonomic ranks than currently available and identifies taxonomically challenging areas across the virosphere.
The usage of synonymous codons varies along the genome, with strong biases in conserved and highly expressed genes that are optimized for translation. The extent of codon usage adaptation across genes and co-adaptation between genes, as well as the influence of gene function on these patterns, remain important open questions. Here, we show that codon usage is highly non-random in most bacterial genes, with at least ~20-46% of the gene families presenting synchronized codon usage evolution. We show that co-adapting genes are co-expressed, co-regulated, metabolically connected, and functionally associated. Codon usage adaptations of key marker genes highlight differences in their expression context between microbes with alternative ecological strategies. This underappreciated regulatory dimension has important implications for function discovery and engineering. ### Competing Interest Statement The authors have declared no competing interest. Dutch Research Council, https://ror.org/04jsz6e67, ALWGR.2017.002 European Research Council, Consolidator grant 865694 Deutsche Forschungsgemeinschaft, 390713860 Alexander von Humboldt Foundation, https://ror.org/012kf4317
Viruses are key players in diverse ecosystems, but studying their impacts is technically and taxonomically challenging. Taxonomic complexities derive from undersampling, diverse DNA and RNA genomes with multiple evolutionary origins, and lack of a universal barcode gene. While virus ecogenomics has expanded access to and understanding of the virosphere, available classification tools poorly scale to modern discovery-based datasets, lack taxonomic resolution, and/or are unable to classify novel sequence space. Here we develop, benchmark, and release vConTACT3, a machine learning-based tool that improves scalability and accuracy, adds extensive user-requested features, expands classification to both eukaryote and prokaryote viruses for 4/6 officially recognized realms, and establishes accurate hierarchical taxonomy from genus to order. Application to 48,069 public virus genomes provided new taxonomy assignments for thousands of taxa, revealed support for fewer taxonomic ranks than currently available, and systematically identified taxonomically problematic areas across the virosphere. ### Competing Interest Statement The authors have declared no competing interest. U.S. National Science Foundation, https://ror.org/021nxhr62, DBI-2149506, DBI-2022070 Ohio Supercomputer Center, https://ror.org/01apna436 Deutsche Forschungsgemeinschaft, https://ror.org/018mejw64, EXC 2051 – Project-ID 390713860 European Research Council, https://ror.org/0472cxd90, Consolidator grant 865694 Alexander von Humboldt Foundation Biotechnology and Biological Sciences Research Council, BBS/E/F/000PR13631, BBS/E/F/000PR13633, BB/X011011/1, BBS/E/F/000PR13634, BBS/E/F/000PR13635, BBS/E/F/000PR13636
The early growth phase is a critical period for the development of the chicken gut microbiome. In this study, the spatiotemporal diversity of the gastrointestinal microbiota, shifts in taxonomic composition, and relative abundances of the main bacterial taxa were characterized in Kadaknath, a high-value indigenous Indian chicken breed, using sequencing of the V3–V4 region 16S rRNA gene. To assess microbiome composition and bacterial abundance shifts, three chickens per growth phase (3, 28, and 35 days) were sampled, with microbiota analyzed from three gut regions (crop, small intestine, and ceca) per bird. The results revealed Firmicutes as the most abundant phylum and Lactobacillus as the dominant genus across all stages. Lactobacillus was particularly abundant in the crop at early stages (3 and 28 days), while the ceca exhibited a transition towards the dominance of genus Phocaeicola by day 35. Microbial richness and evenness increased with age, reflecting microbiome maturation, and the analyses of the microbial community composition revealed distinct spatiotemporal differences, with the ceca on day 35 showing the highest differentiation. Pathogen analysis highlighted a peak in poultry-associated taxa Campylobacter, Staphylococcus, and Clostridium paraputrificum in 3-day-old Kadaknath, particularly in the small intestine, underscoring the vulnerability of early growth stages. These findings provide critical insights into age-specific microbiome development and early life-stage susceptibility to pathogens, emphasizing the need for targeted interventions to optimize poultry health management and growth performance.
Bacteriophages, lytic or lysogenic, play critical roles in structuring different soil bacteriomes and driving their functionality. Lysogeny is favored in the plant rhizosphere and may play a major role in plant-rhizobacteria assembly and function. However, the ecological footprint and consequence of prophage activity in the rhizosphere are poorly understood. Here, we conducted a 35-day pot experiment to examine how prophage induction influences soybean rhizosphere viromes and bacterial communities, along with associated changes in nutrient cycling and plant development. The results showed that mitomycin C-induced prophage induction triggered immense viral production, altering virome structure-with more observed species richness in the rhizosphere. We observed a greater impact on the rhizosphere virome than on the bulk soil virome. The resulting lysis decreased the soil organic matter content but significantly increased dissolved organic carbon and nitrate contents in the soil, which improved soil nutrient conditions and stimulated soybean root development. Prophage induction markedly influenced the rhizobacterial community structure, resulting in reduced community diversity. The enrichment of fast-growing bacterial populations was stimulated, suggesting that viral lysis increased microbial activities and accelerated nutrient turnover. The bacterial interaction network was drastically shifted, with complexity being decreased in the bulk soil and increased in the rhizosphere, potentially stimulating the differentiation of the bacterial communities. Together, our results demonstrated that induction of prophages can cause extensive nutrient turnover and variations in plant-rhizobacteria interactions, driving the rhizobacterial community assembly process. This study provides novel insights into the mechanisms of phages controlling microbial function in primary production and soil carbon storage by modulating microbial traits (e.g., carbon use efficiency, growth rate, death, and community assembly) and via processes like the viral shunt.
Viromics produces millions of viral genomes and fragments annually, overwhelming traditional sequence comparison methods. Here we introduce Vclust, an approach that determines average nucleotide identity by Lempel-Ziv parsing and clusters viral genomes with thresholds endorsed by authoritative viral genomics and taxonomy consortia. Vclust demonstrates superior accuracy and efficiency compared to existing tools, clustering millions of genomes in a few hours on a mid-range workstation.
Untargeted metabolomics can comprehensively map the chemical space of a biome, but is limited by low annotation rates (< 10 Chemical characteristics vectors allow sample comparison with interpretable chemical information. By leveraging molecular fingerprints or compound classes, CCVs utilize “chemical dark matter” that would otherwise be excluded. This approach enhances the interpretability of untargeted metabolomic data, revealing key chemical patterns across biomes.