Sewage metagenomics is a powerful tool for proactive pathogen surveillance and understanding microbial community dynamics. To support such efforts, we present a highly curated and accessible longitudinal dataset of 239 sewage samples collected from five European cities. The dataset, processed through metagenomic sequencing, includes rich analytical outputs such as taxonomic profiles, identified antimicrobial resistance genes, assembled contigs with annotated origins, metagenome-assembled genomes with functional gene annotations, and metadata. Given the computational intensity and time required to reproduce such analyses, we share this dataset to promote reuse and advance research. In addition to the metagenomic data, qPCR was used to identify specific pathogens, and Hi-C sequencing was performed on a subset of the samples to strengthen genomic linkage analysis. Central to this resource is a publicly available PostgreSQL database, designed to facilitate efficient exploration and reuse of the data. This comprehensive database allows users to perform targeted queries, subset data, and streamline access to this extensive resource.
Abstract Antimicrobial exposure can alter gut resistance reservoirs, but bulk metagenomics alone often cannot distinguish whether observed changes reflect expansion of bacterial hosts, altered abundance of plasmid-derived sequences, or redistribution of mobile elements across host backgrounds. Here, we combined longitudinal bulk short-read metagenomics with selected bulk long-read and single-cell shotgun metagenomic sequencing to analyse faecal samples from six Danish pigs over 11 weeks, including an unplanned tiamulin exposure affecting the three pigs housed on the right side of the stable. We constructed a catalogue of 885 plasmid-derived sequences collapsed into 195 bins. Twenty-eight bins and 212 contigs carried resistance annotations, including ribosomal-target markers relevant to pleuromutilin exposure. Single-cell evidence linked subsets of plasmid-derived bins and contigs to bacterial host taxa, enabling host-resolved inspection of resistance-associated plasmid-derived features in longitudinal bulk metagenomes. The microbiome-wide plasmid-derived-sequence prevalence screen identified two bins with post-event associations, whereas resistance-gene abundance and host-attributed plasmid-derived-sequence abundance screens identified no significant host-resolved associations. Because exposure was unplanned and confounded with pen side and disease signs, treatment-response results are exploratory. The main contribution is a single-cell-informed microbial ecology workflow for linking plasmid-derived resistance features to host backgrounds and longitudinal abundance patterns in complex gut
Objectives Outline the basis for integrated One Health AMR surveillance in the Nordic region by mapping existing surveillance systems and research assets, identifying key challenges to cross-border alignment, and proposing practical steps toward coordinated cross-sector, cross-country surveillance. Study design Mapping and review of Nordic AMR infrastructure. Methods We mapped AMR data sources and surveillance infrastructure across the Nordic countries and compiled the findings into an online resource (www.nomoreamr.org). We also assessed requirements for linking national systems, including ethical, legal, and data-sharing considerations relevant to establishing an integrated regional framework. Results The Nordic countries have well-established AMR surveillance systems supported by digital infrastructure, longstanding public health collaboration, and similar socio-economic organization. However, surveillance is currently conducted independently, with limited cross-border data sharing and coordination. Synchronizing national systems would strengthen regional preparedness by enabling earlier detection of emerging threats and supporting more consistent, evidence-based policies and coordinated antimicrobial stewardship. A shared Nordic surveillance network could also serve as an adaptable model for other regions. Conclusions Establishing an integrated Nordic One Health AMR surveillance is feasible but requires structured linkage of existing national systems and clear ethical and legal frameworks for data access and sharing. The compiled mapping resource can support the technical and governance steps needed to advance regional integration.
Abstract Metagenomics is a widely used approach in microbiome research. However, a major limitation of metagenomic datasets is their compositional nature, which prevents direct quantification of absolute abundances and complicates cross-sample comparisons. Existing strategies for absolute quantification typically require additional experiments or spike-in controls. Here, we introduce the MetaGenome Calibrator (MGCalibrator), a new tool that enables spike-in free, absolute abundance estimation based on routine DNA concentration measurements. We validated the accuracy of absolute abundances obtained with MGCalibrator against qPCR for 5 targets. Our results show a strong correlation with qPCR data, indicating that MGCalibrator enables qPCR-like trend analyses. For Bacteroides dorei , the estimated abundances were highly similar between the two methods (r2 = 0.98, y = 1.00x). For other targets like crAssphage or the bacterial 16S rRNA gene, qPCR values were underrepresented by a factor of 7 or overrepresented by a factor of 4. Benchmarking with synthetic microbiome data demonstrated that our method accurately determines copy numbers in sequencing datasets, and application to whole-cell mock community samples produced expected values based on known extraction biases. In an extraction-bias-free experiment, MGCalibrator accurately quantified genome copy numbers within a twofold range in 98% of cases and determined 16S rRNA gene copies within 1.6-fold or less. Finally, we applied MGCalibrator to track temporal trends in antibiotic resistance genes (ARGs) in wastewater treatment plants in two Dutch provincial capitals. We observed an overall increase in ARGs—such as sul2 in Utrecht and qnrS5 in Houtrust—likely driven by rising bacterial loads. Our findings demonstrate that MGCalibrator provides robust calibration of metagenomic data, paving the way for metagenomics to play a central role in future surveillance by enabling trend analysis across thousands of genetic targets, similar to the capabilities of qPCR for individual genes. The source code and documentation for MGCalibrator are available at github.com/NimroddeWit/MGCalibrator.
Abstract Antimicrobial resistance databases are central to genomic surveillance, but resistance determinants remain distributed across resources with different scopes, structures, and annotations. We developed PanRes, a curated resistance database of 11,717 genes integrating acquired and latent determinants of antibiotic, biocide, and metal resistance within a unified ontology. We predicted representative protein structures and clustered them by structural similarity, grouping proteins into 598 structurally conserved clusters coherent despite sequence divergence. Their structure-guided alignments were used to build Hidden Markov Models (HMMs) for remote homology search. In wastewater metagenomes from seven European cities, PanRes 3D-based HMMs expanded detection beyond high-confidence BLAST, with 35.2% of retained hits identified only by the HMMs and generally showing greater divergence from known proteins. For beta-lactamases, several proteins retained beta-lactamase-like folds and catalytic geometry despite weak sequence similarity. PanRes is available through an interactive web platform ( https://panres.rambio.dk/ ), a structure-informed resource for exploring the whole resistome.
Antimicrobial resistance (AMR) poses a major global public health threat, and ongoing surveillance of antimicrobial resistance genes (ARGs) is critical to mitigate current and future risks. Sewage-based ARG surveillance is gaining traction, but insight into how it compares to surveillance by clinical bacterial isolates is limited, especially when it comes to ARG mutational variants. We compared ARGs identified in clinical bacterial isolates (n = 2,989) with those detected in sewage metagenomes (n = 468) across 33 countries. ARG variant detection data from clinical isolates and sewage metagenomes shared some regional patterns in detection, but many ARG variants were detected exclusively in either sewage metagenomes or clinical isolates. We found that across all samples, only 69% of ARG clusters detected in clinical isolates were also detected via read mapping in sewage. Some ARGs highly prevalent in clinical isolates were not detected in sewage. Among clinically widespread ARGs, prevalence varied across bacterial species and clinical isolate types depending on whether the ARGs were also detected in sewage. This could indicate that sewage surveillance is better suited for detection of clinically relevant ARGs prevalent in certain bacterial species and infection sites than others. Spearman correlation between ARG abundance in sewage and the proportion of clinical isolates from the same country with detection was 0.28 overall, with stronger correlations for certain ARGs. The results demonstrate that sewage ARG profiles correlate, to some extent, to the clinical AMR landscape, but do not capture the full spectrum of clinically relevant ARGs at currently realistic sequencing depths.IMPORTANCEAntimicrobial resistance (AMR) is a major public health threat. Surveillance of AMR is important and can be conducted via the detection of antimicrobial resistance genes (ARGs). Sewage can be used as a medium for surveillance as an alternative to analyzing individual bacterial isolates from health clinics. We compared detection in large global data collections of sewage metagenomes and clinical isolates. We found that while there were significant positive correlations between findings in sewage and clinical isolates, some widespread clinical ARGs were not detectable in sewage. This should be considered if establishing sewage surveillance systems.
Hepatitis A and E viruses (HAV, HEV) are the main causes of enterically transmitted hepatitis. Many infections remain undiagnosed due to their mild clinical course or asymptomatic presentation, and limited testing. Yet, knowledge of circulating genotype diversity is needed to understand their epidemiology and guide interventions. Wastewater surveillance may complement clinical monitoring by capturing infections missed through routine diagnostics. We applied capture-based metagenomic sequencing to characterize HAV and HEV genetic diversity in wastewater from 62 cities across 38 countries (2017-2019), complemented with longitudinal sampling from five European cities (2020-2021). Rocahepevirus ratti (rat HEV) genotype C1 was detected in 70% of cities, extending its known geographic range by 12 countries. Rat HEV sequences clustered by city and country, though some lineages spanned multiple continents. Paslahepevirus balayani (human HEV), predominantly genotype 3, was most prevalent in European cities. HAV genotype distribution largely reflected regional endemicity, although some subgenotypes were detected in regions where they are rarely reported clinically. Faecal source analysis suggested that rat HEV detected in wastewater originates primarily from rodent contamination rather than human infection. These findings reveal the global distribution of HAV and HEV genotypes and the widespread occurrence of rat HEV, demonstrating the value of wastewater metagenomics for population-level monitoring of both human and zoonotic hepatitis viruses.
Antimicrobial resistance genes (ARGs) have rapidly emerged and spread globally, but the pathways driving their spread remain poorly understood. We analyzed 1240 sewage samples from 351 cities across 111 countries, comparing ARGs known to be mobilized with those identified through functional metagenomics (FG). FG ARGs showed stronger associations with bacterial taxa than the acquired ARGs. Network analyses further confirmed this and showed potential for source attribution of both known and novel ARGs. The FG resistome was more evenly dispersed globally, whereas the acquired resistome followed distinct geographical patterns. City-wise distance-decay analyses revealed that the FG ARGs showed significant decay within countries but not across regions or globally. In contrast, acquired ARGs showed decay at both national and regional scales. At the variant level, both ARG groups had significant national and regional distance-decay effects, but only FG ARGs at a global scale. Additionally, we observed stronger distance effects in Sub-Saharan Africa and East Asia compared to North America. Our findings suggest that differential selection and niche competition, rather than dispersal, shape the global resistome patterns. A limited number of bacterial taxa may act as reservoirs of latent FG ARGs, highlighting the need of targeted surveillance to mitigate future resistance threats.
Understanding global viral dynamics is critical for public health. Traditional surveillance focuses on individual pathogens and symptomatic cases, which may miss asymptomatic infections or newly emerging viruses, delaying detection and response. Wastewater-based epidemiology has been used to track pathogens through targeted molecular assays, but its reliance on predefined targets limits detection of the full viral spectrum. Here, we analyse longitudinal wastewater samples from 62 cities across six continents (2017-2019) using metagenomics and capture-based sequencing with probes targeting viruses associated with gastrointestinal disease. We detect over 2500 viral species spanning 122 families, many with human, animal, or plant health relevance. The bacteriophage family Microviridae and plant virus family Virgaviridae dominate the metagenomic dataset, while Astroviridae and Picornaviridae prevail in the capture-based sequence dataset. Virus distributions are broadly similar across continents at the family and genus levels, yet distinct city-level fingerprints reveal geographical and temporal variation, enabling spatiotemporal surveillance of viruses such as astroviruses and enteroviruses. Global wastewater-based epidemiology enables early detection of emerging viruses, including Echovirus 30 in Europe and Tomato brown rugose fruit virus. These findings highlight the potential of wastewater sequencing for the early detection of emerging viruses and population-wide virome monitoring across diverse hosts.
Relational databases offer an efficient solution for storing and retrieving complex data sets, yet the requirement for SQL programming expertise presents a significant challenge for many life science users. We explore whether a cutting-edge large language model can effectively translate plain English queries into SQL scripts (Text-to-SQL), thereby simplifying database interaction and eliminating the typical usage barriers. A complex database comprising 19 interconnected tables of metagenomic analyses from 239 sewage samples across five European cities was available. A large language model was provided with details of the database’s structure and background information on its contents. We evaluated the functionalities of this “SewageGPT” tool and assessed the accuracy of its responses to complex questions and visualisation of results. Providing a detailed description of the database enabled SewageGPT to accurately respond to complex inquiries, accelerating the database querying process. Knowledge of the database content proved beneficial, as it minimized the risk of ambiguities in queries; however, ambiguities can lead to incorrect responses. Therefore, human oversight remains crucial, particularly for questions that lack detail or involve ambiguities. The integration of state-of-the-art large language models with direct database connectivity substantially enhances the efficiency of query generation, statistical analysis and visualization of the results.
Single-cell sequencing may serve as a powerful complementary technique to shotgun metagenomics to study microbiomes. This emerging technology allows the separation of complex microbial communities into individual bacterial cells, enabling high-throughput sequencing of genetic material from thousands of singular bacterial cells in parallel. Here, we validated the use of microfluidics and semi-permeable capsules (SPCs) technology (Atrandi) to isolate individual bacterial cells from sewage and pig fecal samples. Our method involves extracting and amplifying single bacterial DNA within individual SPCs, followed by combinatorial split-and-pool single-amplified genome (SAG) barcoding and short-read sequencing. We tested two different sequencing approaches with different numbers of SPCs from the same sample for each sequencing run. Using a deep sequencing approach, we detected 1,796 and 1,220 SAGs, of which 576 and 599 were used for further analysis from one sewage and one fecal sample, respectively. In shallow sequencing data, we aimed for 10-times more cells and detected 12,731 and 17,909 SAGs, of which we used 2,456 and 1,599 for further analysis for sewage and fecal samples, respectively. Additionally, we identified the top 10 antimicrobial resistance genes (ARGs) in both sewage and feces samples and linked them to their individual host bacterial species.
We report the discovery of a persistent presence of Vibrio cholerae at very low abundance in the inlet of a single wastewater treatment plant in Copenhagen, Denmark at least since 2015. Remarkably, no environmental or locally transmitted clinical case of V. cholerae has been reported in Denmark for more than 100 years. We, however, have recovered a near-complete genome out of 115 metagenomic sewage samples taken over the past 8 years, despite the extremely low relative abundance of one V. cholerae read out of 500,000 sequenced reads. Due to the very low relative abundance, routine screening of the individual samples did not reveal V. cholerae. The recovered genome lacks the gene responsible for cholerae toxin production, but although this strain may not pose an immediate public health risk, our finding illustrates the importance, challenges, and effectiveness of wastewater-based pathogen surveillance.
Motivation Analyzing metagenomic data can be highly valuable for understanding the function and distribution of antimicrobial resistance genes (ARGs). However, there is a need for standardized and reproducible workflows to ensure the comparability of studies, as the current options involve various tools and reference databases, each designed with a specific purpose in mind.Results In this work, we have created the workflow ARGprofiler to process large amounts of raw sequencing reads for studying the composition, distribution, and function of ARGs. ARGprofiler tackles the challenge of deciding which reference database to use by providing the PanRes database of 14 078 unique ARGs that combines several existing collections into one. Our pipeline is designed to not only produce abundance tables of genes and microbes but also to reconstruct the flanking regions of ARGs with ARGextender. ARGextender is a bioinformatic approach combining KMA and SPAdes to recruit reads for a targeted de novo assembly. While our aim is on ARGs, the pipeline also creates Mash sketches for fast searching and comparisons of sequencing runs.Availability and implementation The ARGprofiler pipeline is a Snakemake workflow that supports the reuse of metagenomic sequencing data and is easily installable and maintained at https://github.com/genomicepidemiology/ARGprofiler.
Additional file 1. General summary of simulated community 1a. MAGICIAN output summarizing general information on one of the simulation runs for community 1a. BBTools’ stats output is summarized in the “BBstats” sheet, CheckM output is summarized in the “CheckM” sheet, and dRep compare output is summarized in the “dRep” sheet. Explanations of terminology are given in the “explanations” sheet.
Sewage metagenomics has risen to prominence in urban population surveillance of pathogens and antimicrobial resistance (AMR). Unknown species with similarity to known genomes cause database bias in reference-based metagenomics. To improve surveillance, we designed this study to recover sewage genomes and develop a quantification and correlation workflow for these genomes and AMR over time. We used longitudinal sewage sampling in seven treatment plants from five major European cities to explore the utility of catch-all sequencing of these population-level samples. Using metagenomic assembly methods, we recovered 2,332 metagenome-assembled genomes (MAGs) from prokaryotic species, 1,334 of which were previously undescribed. These genomes account for ∼69% of sequenced DNA and provide insight into sewage microbial dynamics. Rotterdam (Netherlands) and Copenhagen (Denmark) showed strong seasonal microbial community shifts, while Bologna, Rome, (Italy) and Budapest (Hungary) had occasional blooms of Pseudomonas -dominated communities, accounting for up to ∼95% of sample DNA. Seasonal shifts and blooms present challenges for effective sewage surveillance. We find that bacteria of known shared origin, like human gut microbiota, form communities, suggesting the potential for source-attributing novel species and their ARGs through network community analysis. This could significantly improve AMR tracking in urban environments. ### Competing Interest Statement The authors have declared no competing interest.
Our 24-month study used metagenomics to investigate antimicrobial resistance (AMR) abundance in raw sewage from wastewater treatment works (WWTWs) in two municipalities in Gauteng Province, South Africa. At the AMR class level, data showed similar trends at all WWTWs, showing that aminoglycoside, beta-lactam, sulfonamide and tetracycline resistance was most abundant. AMR abundance differences were shown between municipalities, where Tshwane Metropolitan Municipality (TMM) WWTWs showed overall higher abundance of AMR compared to Ekurhuleni Metropolitan Municipality (EMM) WWTWs. Also, within each municipality, there were differing trends in AMR abundance. Notably, within TMM, certain AMR classes (macrolides and macrolides_streptogramin B) were in higher abundance at a WWTW serving an urban high-income area, while other AMR classes (aminoglycosides) were in higher abundance at a WWTW serving a semi-urban low income area. At the AMR gene level, all WWTWs samples showed the most abundance for the sul1 gene (encoding sulfonamide resistance). Following this, the next 14 most abundant genes encoded resistance to sulfonamides, aminoglycosides, macrolides, tetracyclines and beta-lactams. Notably, within TMM, some macrolide-encoding resistance genes (mefC, msrE, mphG and mphE) were in highest abundance at a WWTW serving an urban high-income area; while sul1, sul2 and tetC genes were in highest abundance at a WWTW serving a semi-urban low income area. Differential abundance analysis of AMR genes at WWTWs, following stratification of data by season, showed some notable variance in six AMR genes, of which blaKPC-2 and blaKPC-34 genes showed the highest prevalence of seasonal abundance differences when comparing data within a WWTW. The general trend was to see higher abundances of AMR genes in colder seasons, when comparing seasonal data within a WWTW. Our study investigated wastewater samples in only one province of South Africa, from WWTWs located within close proximity to one another. We would require a more widespread investigation at WWTWs distributed across all regions/provinces of South Africa, in order to describe a more comprehensive profile of AMR abundance across the country.
Background The possibility of recovering metagenome-assembled genomes (MAGs) from sequence reads allows for further insights into microbial communities and their members, possibly even analyzing such sequences with tools designed for single-isolate genomes. As result quality depends on sequence quality, performance of tools for single-isolate genomes on MAGs should be tested beforehand. Bioinformatics can be leveraged to quickly create varied synthetic test sets with known composition for this purpose. Results We present MAGICIAN, a flexible, user-friendly pipeline for the simulation of MAGs. MAGICIAN combines a synthetic metagenome simulator with a metagenomic assembly and binning pipeline to simulate MAGs based on user-supplied input genomes, allowing users to test performance of tools on MAGs while having a ground truth to compare results to. Using MAGICIAN, we found that even very slight (1%) changes in depth of coverage can drastically affect whether a genome can be recovered. We also demonstrate the use of simulated MAGs by evaluating the suitability of such genomes obtained with MAGICIAN’s current default pipeline for analysis with the antimicrobial resistance gene identification tool ResFinder. Conclusions Using MAGICIAN, it is possible to simulate MAGs which, while generally high in quality, reflect issues encountered with real-world data, thus providing realistic best-case data. Evaluating the results of ResFinder analysis of these genomes revealed a risk for plausible-looking false positives, which underlines the need for pipeline validation so that researchers are aware of the potential issues when interpreting real-world data. Furthermore, the effects of fluctuations in depth of coverage on genome recovery in our simulated “random sequencing” warrant further investigation and indicate random subsampling of reads may affect discovery of more genomes.
The rapid spread of antimicrobial resistance (AMR) is a threat to global health, and the nature of co-occurring antimicrobial resistance genes (ARGs) may cause collateral AMR effects once antimicrobial agents are used. Therefore, it is essential to identify which pairs of ARGs co-occur. Given the wealth of next-generation sequencing data available in public repositories, we have investigated the correlation between ARG abundances in a collection of 214,095 metagenomic data sets. Using more than 6.76∙108 read fragments aligned to acquired ARGs to infer pairwise correlation coefficients, we found that more ARGs correlated with each other in human and animal sampling origins than in soil and water environments. Furthermore, we argued that the correlations could serve as risk profiles of resistance co-occurring to critically important antimicrobials (CIAs). Using these profiles, we found evidence of several ARGs conferring resistance for CIAs being co-abundant, such as tetracycline ARGs correlating with most other forms of resistance. In conclusion, this study highlights the important ARG players indirectly involved in shaping the resistomes of various environments that can serve as monitoring targets in AMR surveillance programs. IMPORTANCE:Understanding the collateral effects happening in a resistome can reveal previously unknown links between antimicrobial resistance genes (ARGs). Through the analysis of pairwise ARG abundances in 214K metagenomic samples, we observed that the co-abundance is highly dependent on the environmental context and argue that these correlations can be used to show the risk of co-selection occurring in different settings.