Consortia of microbial isolates, also known as synthetic communities (SynComs), are increasingly used to study and harness microbe-microbe and microbe-host interactions. Since “synthetic” potentially evokes negative connotations, we propose adopting the term “Defined Microbial Community” for practical applications.
In October 2024, Parties to the United Nations Convention on Biological Diversity agreed to a new multilateral mechanism to fund biodiversity conservation through the sharing of benefits from open biodiversity data. Biological databases hosting genetic and other biological data, known as digital sequence information (DSI), are central to the implementation of the mechanism. This paper assesses the new international agreement and its implications for DSI databases. We walk through the database provisions in COP16 Decision 16/2, which include notifying users and submitters about the mechanism, improving metadata on geographical location of sample collection, and consistency with open access, as well as consideration of the FAIR, CARE, and TRUST principles. Drawing on surveys, interviews, and a workshop with biological database managers, we identify practical and scalable measures including updating terms of use, revising submission procedures, and strengthening user communication. We also propose approaches to capture and report non-monetary benefits such as capacity building, publications, interoperability, and training. These actions illustrate how DSI databases can remain open, sustainable, and globally connected while supporting benefit-sharing from the use of DSI on genetic resources.
Addressing the growing threat of antimicrobial resistance (AMR) requires the development of large-scale resources that link bacterial genomic data with phenotypic AMR profiles. Such datasets are essential for advancing genotype-based predictions of resistance to uncover novel resistance mechanisms, as well as identifying and tracking global trends. Here, we describe the development of the "Comprehensive Assessment of Bacterial-Based AMR prediction from GEnotypes" (CABBAGE) database, linking bacterial genomes to associated antibiotic susceptibility data and relevant metadata across WHO Bacterial Priority Pathogens, sourced from both publications and existing databases, and curated into a format that is compatible with, and extends, both NCBI and ENA formats. The resulting CABBAGE database, comprising over 170 000 unique sequenced isolates and approximately 1.7 million genome-phenotype pairs linked to extensive metadata, represents the largest database of its kind, consolidating existing AMR phenotype-genotype data into a single unified format. CABBAGE encompasses a broad range of antimicrobials, facilitating the analysis of global resistance trends as well as benchmarks of genotype-to-phenotype predictive methods, and empowering further research uses. The database is freely accessible via the Antimicrobial Resistance Portal at EMBL-EBI and is currently being integrated with the BioSample database, enabling easy access for the AMR research community.
The HMMER web server, available at https://www.ebi.ac.uk/Tools/hmmer, provides online access to tools from the HMMER software suite (http://hmmer.org/) for protein analysis using profile hidden Markov models. Users can perform sequence similarity searches against a range of regularly updated protein sequence databases or annotate protein sequences with domains and families using profile HMM libraries from protein family databases. Since the 2018 update, the continued exponential growth of sequence databases has necessitated substantial infrastructural improvements to maintain search performance speed and service reliability. To achieve this, the web interface has been completely reengineered using modern web technologies (JavaScript and React), providing users with an enhanced experience, including session-based search history and streamlined results visualization. The web application programming interface has been rewritten to better support programmatic access with updated endpoints and JSON-based responses. The infrastructure has been redesigned to efficiently handle searches against much larger databases through horizontal scaling and asynchronous job processing. Target database offerings have been updated to reflect current usage patterns and data availability. The HMMER web server is free and open to all users, and there is no login requirement.
The wide geographic distribution of microorganisms, combined with their vast taxonomic and functional diversity, make them indispensable reservoirs of genetic variation that sustain ecosystem resilience and fuel biotechnological innovation. However, to use this diversity, microbiologists must navigate a complex legal and regulatory landscape governed by multiple United Nations treaties and their respective access and benefit-sharing frameworks as well as regulatory frameworks specific to particular ecosystems, biosecurity, pathogens, and intellectual property. This complex regulatory web is also actively growing and changing, which makes it immensely challenging for a "regular" microbiologist to navigate. For policymakers and negotiators, it is also difficult to appreciate the full complexity that practitioners experience. This policy briefing provides a concise regulatory guide for practitioners and policymakers alike, summarized in a graphical overview, to provide more clarity and understanding for those at the edge of decision-making and practice.
Lichens are symbiotic associations between filamentous fungi and photosynthetic micro-organisms, such as green algae and/or cyanobacteria, that result in a single anatomically complex structure that can thrive in environments inhospitable to most organisms, including arctic tundra, high mountains, and deserts. Recent evidence suggests that lichens may be even more complex than previously appreciated, containing multiple microbial constituents, but how genomes of the principal fungal symbiont (which provides the majority of biomass in lichen tissue) have been shaped during evolution is largely unexplored. Recently, giant transposable elements called Starships have been found in many genomes of filamentous fungi, but to which extent they occur in lichen-forming fungi is not known. In this report, we describe a Starship element from the lichen fungus Xanthoria parietina . This element, named Tangerine , contains several genes that have signatures of horizontal gene transfer from nonlichen-forming fungi, most likely from black yeasts of the Chaetothyriales, that are often lichen-associated. Repetitive sequences carried by Tangerine , and found in other sites in Xanthoria genomes, are affected by repeat-induced point mutation, a mechanism of genome defense against transposable elements, consistent with fungal sexual reproduction which always precedes new lichen formation by X. parietina . Tangerine ’s “captain” belongs to a newly defined family of tyrosine recombinases specific to lichen-forming Lecanoromycetes. Several other captain clades have signatures of horizontal gene transfer between distantly related lichen-forming fungi and nonmycobiont lichen-associated fungi. We speculate that Starships may play a significant, yet hitherto unrecognized role, in lichen genome evolution and provide a roadmap for further investigation.
Indigenous groups across the world have been underrepresented in the ongoing efforts to map human microbiome diversity. This study investigates the faecal microbiome and resistome diversities of Indigenous Orang Asli communities in Malaysia with different lifestyles and level of urbanisation. Healthy adults (>18 years) from three Indigenous communities: Temuan (urban), Temiar (semi-urban), and Jahai (rural hunter-gatherer), as well as urban Malays, were recruited, and their stools samples were subjected to shotgun metagenomic sequencing. Microbial alpha diversity (Shannon) decreased but antibiotic resistance genes (ARGs) diversity increased as the degree of urbanisation in the groups increased (P < 0.05). The groups contributed to 13% of the variation observed in the microbial composition (PERMANOVA, P = 0.001), and 14.5% in resistome composition (PERMANOVA, P = 0.001). Romboutsia timonensis was significantly depleted in Jahai compared to Malay (FDR = 0.04). Shared ARGs conferring resistance to beta-lactams (cfxA), tetracyclines (tet), and macrolides (erm) were observed across all groups, irrespective of geographical location, ethnicity and lifestyle. These results establish a baseline framework for urbanisation-linked microbial shifts in the under-studied Malaysian gut microbiome and provide foundational evidence of antimicrobial resistance patterns and underscores the need for broader inclusion of underrepresented populations in national surveillance and stewardship efforts.Indigenous groups across the world have been underrepresented in the ongoing efforts to map human microbiome diversity. This study investigates the faecal microbiome and resistome diversities of Indigenous Orang Asli communities in Malaysia with different lifestyles and level of urbanisation. Healthy adults (>18 years) from three Indigenous communities: Temuan (urban), Temiar (semi-urban), and Jahai (rural hunter-gatherer), as well as urban Malays, were recruited, and their stools samples were subjected to shotgun metagenomic sequencing. Microbial alpha diversity (Shannon) decreased but antibiotic resistance genes (ARGs) diversity increased as the degree of urbanisation in the groups increased (P < 0.05). The groups contributed to 13% of the variation observed in the microbial composition (PERMANOVA, P = 0.001), and 14.5% in resistome composition (PERMANOVA, P = 0.001). Romboutsia timonensis was significantly depleted in Jahai compared to Malay (FDR = 0.04). Shared ARGs conferring resistance to beta-lactams (cfxA), tetracyclines (tet), and macrolides (erm) were observed across all groups, irrespective of geographical location, ethnicity and lifestyle. These results establish a baseline framework for urbanisation-linked microbial shifts in the under-studied Malaysian gut microbiome and provide foundational evidence of antimicrobial resistance patterns and underscores the need for broader inclusion of underrepresented populations in national surveillance and stewardship efforts.
Domestication represents one of the largest biological shifts of life on Earth, and for many animal species, behavioral selection is thought to facilitate early stages of the process. The gut microbiome of animals can respond to environmental changes and have diverse and powerful effects on host behavior. As such, we hypothesize that selection for tame behavior during early domestication, may have indirectly selected on certain gut microbiota that contribute to the behavioral plasticity necessary to adapt to the new social environment. Here, we explore the gut microbiome of foxes from the tame and aggressive strains of the "Russian-Farm-Fox-Experiment". Microbiota profiles reveal a significant depletion of bacteria in the tame fox population that have been associated with aggressive and fear-related behaviors in other mammals. Our metagenomic survey allows for the reconstruction of microbial pathways enriched in the gut of tame foxes, such as glutamate degradation, which converge with host genetic and physiological signals, revealing a potential role of functional host-microbiota interactions that could influence behaviors associated with domestication. Overall, by characterizing how compositional and functional potential of the gut microbiota and host behaviors co-vary during early animal domestication, we provide further insight into our mechanistic understanding of this adaptive, eco-evolutionary process.
The early-life development of the gut microbiome in broiler chickens is a dynamic ecological process with significant implications for host physiology and productivity. Using 388 genome-resolved metagenomic and 61 metatranscriptomic samples across two replicated trials, we analysed the compositional and functional succession of the caecal microbiome in chickens from hatching to slaughter age. We reconstructed 822 bacterial genomes and distilled gene annotations into comprehensive metabolic traits that captured the functional capacities of each genome. We observed that the increase in microbial diversity with chicken age was accompanied by a decline in community-level average metabolic capacity, driven by a shift from metabolically versatile generalists (Lachnospiraceae) to hitherto uncultured, genome-reduced specialists (RF39, RF32, and UBA1242). However, the specific identity of the dominant genome-reduced specialists varied among individuals, resulting in contrasting associations with host body weight. At slaughter age, only 10 UBA660 (RF39) bacteria were positively associated with body weight, while other genome-reduced lineages, such as UBA1242 (Christensenellales), were among 190 negatively associated bacteria. Gene expression analyses revealed that despite their reduced functional repertoire, UBA660 exhibited greater metabolic activity than UBA1242, particularly in the production of two key metabolites for host nutrition and intestinal homeostasis: the essential amino acid lysine and the signaling molecule indole-3-acetate. These findings provide new insights into the functional ecology of the chicken gut microbiome and highlight the relevance of cultivation approaches to retrieve underexplored and uncultured bacterial taxa, which could open new avenues for microbiome-based strategies aimed at improving poultry growth and health in intensive production systems.
The identification of amplicon sequence variants from DNA metabarcoding data is a common method for revealing the taxonomic makeup of environmental samples, and for allowing comparative studies between similar datasets. A significant hurdle to the large-scale calling of amplicon sequence variants from publicly available nucleotide datasets is the heterogeneous presence of primer sequences in reads, the removal of which is a necessary pre-processing step for this form of analysis. Furthermore, as the details of the experimental primers are rarely captured in the metadata associated with the sequence records, there is a need for a method that can automatically infer the presence and identity of primers in sequencing data. In this work, we introduce the PrIMER infereNce TOolkit (PIMENTO), a Python package that uses a dual-strategy approach for identifying primers that are present in sequencing reads to enable their removal, and therefore facilitate amplicon sequence variant calling at scale.
Abstract Metagenomic assembly can be a computationally intensive step in microbiome analysis, with memory requirements that vary widely depending on input data characteristics. In workflow systems like Galaxy and large-scale platforms like MGnify, which run thousands of heterogeneous jobs, inaccurate memory allocation drives job failures and costly retries when underestimated, and reduces throughput when overestimated. Current approaches rely primarily on heuristic rules based on input file size or sample metadata, which often fail to generalize across diverse datasets. In this study, we present a machine learning-based framework for predicting memory requirements of metagenomic assembly using metaSPAdes. We analyzed 300 assembly jobs from diverse biomes and evaluated 18 predictive models using combinations of input file size, biome classification, and sequence-derived k-mer features. K-mer profiles were computed from raw sequencing data and summarized into statistical descriptors capturing sequence complexity and diversity. Model performance was assessed using both conventional regression metrics and a production-oriented cost function that accounts for retry policies and resource waste in high-performance computing environments. Our results show that machine learning models can outperform commonly used heuristics. In particular, models incorporating biome information achieved the best performance and can be tuned to favor conservative predictions that reduce job failure rates. Simpler models based solely on input file size also performed competitively, offering a practical alternative for systems with limited feature availability. When evaluated under realistic workload distributions, predictive approaches reduced total memory waste by several million gigabyte-hours per 1,000 jobs compared to static allocation strategies. These findings demonstrate that data-driven resource prediction can substantially improve efficiency in metagenomic workflows. The proposed framework is adaptable to different computational environments and provides a foundation for integrating predictive resource allocation into large-scale bioinformatics platforms beyond Galaxy.
Abstract Protein language model embeddings are increasingly used to organise biological sequences, yet how biological meaning is encoded within embedding neighbourhoods remains poorly understood. Using two independent hierarchical enzyme systems, carbohydrate-active enzymes and peptidases, we investigated how biological interpretation changes across embedding organisations aligned to different levels of biological hierarchy. Different embedding organisations give rise to distinct neighbourhood semantics. When aligned to membership-boundary resolution, embeddings robustly separated artefacts and unrelated proteins from members of the target category. However, embeddings aligned to functional-grouping resolution maintained compositional neighbourhood structure for multi-domain proteins spanning more than one functional or catalytic group. Finally, embeddings aligned to local-family resolution recovered compact family-like neighbourhoods, including families withheld from training, while weakening broader membership-boundary and functional-grouping relationships. Moreover, embeddings optimised toward the same level of biological organisation retain different biological relationships depending on optimisation trajectory employed. Together, our results show that proximity in protein embedding space has no fixed biological interpretation. Instead, biological meaning emerges across embedding resolutions through selective preservation of different forms of biological organisation.
GENCODE produces comprehensive reference gene annotation for human and mouse. Entering its twentieth year, the project remains highly active as new technologies and methodologies allow us to catalog the genome at ever-increasing granularity. In particular, long-read transcriptome sequencing enables us to identify large numbers of missing transcripts and to substantially improve existing models, and our long non-coding RNA catalogs have undergone a dramatic expansion and reconfiguration as a result. Meanwhile, we are incorporating data from state-of-the-art proteomics and Ribo-seq experiments to fine-tune our annotation of translated sequences, while further insights into function can be gained from multi-genome alignments that grow richer as more species' genomes are sequenced. Such methodologies are combined into a fully integrated annotation workflow. However, the increasing complexity of our resources can present usability challenges, and we are resolving these with the creation of filtered genesets such as MANE Select and GENCODE Primary. The next challenge is to propagate annotations throughout multiple human and mouse genomes, as we enter the pangenome era. Our resources are freely available at our web portal www.gencodegenes.org, and via the Ensembl and UCSC genome browsers.
Microbiome research has grown substantially over the past decade in terms of the range of biomes sampled, identified taxa, and the volume of data derived from the samples. In particular, experimental approaches such as metagenomics, metabarcoding, metatranscriptomics and metaproteomics have provided profound insights into the vast, hitherto unknown, microbial biodiversity. The ELIXIR Marine Metagenomics Community, initiated amongst researchers focusing on marine microbiomes, has concentrated on promoting standards around microbiome-derived sequence analysis, as well as understanding the gaps in methods and reference databases, and identifying solutions to the computational overheads of performing such analyses. Nevertheless, the methods used and the challenges faced are not confined to marine microbiome studies, but are broadly applicable to other biomes. Thus, expanding this Marine Metagenomics Community to a more inclusive ELIXIR Microbiome Community will enable it to encompass a broader range of biomes and link expertise across ‘omics technologies. Furthermore, engaging with a large number of researchers will improve the efficiency and sustainability of bioinformatics infrastructure and resources for microbiome research (standards, data, tools, workflows, training), which will enable a deeper understanding of the function and taxonomic composition of the different microbial communities.
Microbiome research has grown substantially over the past decade in terms of the range of biomes sampled, identified taxa, and the volume of data derived from the samples. In particular, experimental approaches such as metagenomics, metabarcoding, metatranscriptomics and metaproteomics have provided profound insights into the vast, hitherto unknown, microbial biodiversity. The ELIXIR Marine Metagenomics Community, initiated amongst researchers focusing on marine microbiomes, has concentrated on promoting standards around microbiome-derived sequence analysis, as well as understanding the gaps in methods and reference databases, and identifying solutions to the computational overheads of performing such analyses. Nevertheless, the methods used and the challenges faced are not confined to marine microbiome studies, but are broadly applicable to other biomes. Thus, expanding this Marine Metagenomics Community to a more inclusive ELIXIR Microbiome Community will enable it to encompass a broader range of biomes and link expertise across 'omics technologies. Furthermore, engaging with a large number of researchers will improve the efficiency and sustainability of bioinformatics infrastructure and resources for microbiome research (standards, data, tools, workflows, training), which will enable a deeper understanding of the function and taxonomic composition of the different microbial communities.
The calling of amplicon sequence variants from DNA metabarcoding data is a common method of revealing the taxonomic makeup of environmental samples. A significant hurdle to the large-scale calling of amplicon sequence variants from publicly available nucleotide datasets is the presence of primer sequences in reads, the removal of which is a necessary pre-processing step for this form of analysis. Further, as the details of which primers were used is rarely associated with the sequence records, there is a need for a method that can automatically infer the presence and identity of primers in sequencing data. In this work, we introduce PIMENTO, a Python package which uses a dual-strategy approach for identifying primers that are present in sequencing reads to enable their removal, and therefore facilitate amplicon sequence variant calling at scale. ### Competing Interest Statement The authors have declared no competing interest. European Union, https://ror.org/019w4f821, 101094227, 101112823 European Bioinformatics Institute, https://ror.org/02catss52
SUMMARY:In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION:The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.
The use of 16S rRNA metabarcoding for functional prediction is limited by several biases. Shallow shotgun sequencing is a cost-effective and taxonomically high-resolution alternative to 16S rRNA metabarcoding, but the low sequencing depth limits functional inference. Our BioSIFTR tool maps shallow shotgun sequencing reads against single-biome databases and extrapolates into their precalculated functional profiles. We used three datasets from red junglefowl, mice, and human gut, containing matched deep shotgun metagenomic and 16S rRNA metabarcoding data for taxonomic and functional benchmarking. An additional human gut deep shotgun sequencing data set was subsampled to 1 M reads and analysed with BioSIFTR to replicate previously obtained results. BioSIFTR taxonomic and functional profiles closely agree with the results of the full deep sequencing data in all biomes. We also replicated differences in the human gut microbiome between high and low trimethylamine N-oxide producing participants, using only < 2 % of the original deep sequencing data. The BioSIFTR tool is a powerful approach which approximates the functional information of a deep-sequenced metagenome while using only a fraction of the data. Shallow shotgun sequencing combined with BioSIFTR could be a stand-in replacement for 16S rRNA metabarcoding with an increased taxonomic and functional resolution, and lower bias. ![Figure][1] ### Competing Interest Statement The authors have declared no competing interest. European Union’s Horizon 2020 research and innovation programme, 952914 The Research Council of Finland, 338818 Finnish Cultural Foundation, https://ror.org/027xav248, 210944 Medical Research Council, https://ror.org/03x94j517, MC\_PC\_21045 [1]: pending:yes
The European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) is one of the world's leading sources of public biomolecular data. Based at the Wellcome Genome Campus in Hinxton, UK, EMBL-EBI is one of six sites of the European Molecular Biology Laboratory, Europe's only intergovernmental life sciences organization. This overview summarizes the latest developments in services that EMBL-EBI data resources provide to scientific communities globally (https://www.ebi.ac.uk/services).
Phages and bacteria are locked in a molecular arms race, with phage anti-defence proteins (ADPs) enabling them to evade bacterial immune systems. To streamline access to information on ADPs, we developed the Encyclopaedia of Viral Anti-DefencE Systems (EVADES), an online resource containing sequences, structures, protein family annotations, and mechanisms of action (MoA) for 257 ADPs. Through computational structural analysis we established putative MoAs for 21 uncharacterised ADPs. We demonstrate the utility of EVADES by exploring three hypotheses and demonstrate that: (i) ADPs acting as DNA mimics exhibit broad-spectrum activity against defences; (ii) protein sharing across defence systems enables multi-defence inhibitory activity of ADPs; and (iii) eukaryotic dsDNA viruses encode ADP homologs, suggesting conserved immune evasion strategies across domains. These results highlight the broader relevance of phage ADPs in understanding interactions between prokaryotic or eukaryotic viruses and their hosts. EVADES is freely accessible at . Highlights ### Competing Interest Statement The authors have declared no competing interest. European Union’s Horizon 2020 research and innovation programme, 945405
Paul D. Thomas合作论文数Artificial Intelligence Center6