Microbiomes are information-rich biological systems, yet most computational analyses still reduce communities to cohort-specific abundance tables. Here we introduce MGM2, a multimodal foundation model pretrained on 1,821,291 MicrobeAtlas samples and 225,067 OTUs clustered at 99% sequence similarity. MGM2 couples NTv3-derived microbial sequence embeddings with abundance conditioning and community-semantic alignment to learn transferable sample- and token-level representations. Frozen MGM2 representations outperformed DeepPhylo by 0.06 to 0.21 macro-AUROC across five temporally held-out MGnify hierarchy levels, with the largest gains for rare and fine-grained labels. In fecal microbiota transplantation, MGM2-XLarge achieved a response ROC AUC of 0.79 and reduced post-transplant Bray-Curtis distance by 15% relative to the recipient baseline. The same representation supported ASV-level trend forecasting across 24 wastewater treatment plants. Sparse autoencoder analysis resolved MGM2-XLarge token states into a 4,096-feature dictionary spanning taxonomic identity, abundance state, ecological context and technical variation. MGM2 therefore provides a sequence-aware and interpretable representation layer for microbiome classification, paired-community prediction, forecasting and feature discovery.
ABSTRACT Microbiome sequencing has advanced faster than microbiome understanding. Although large‐scale 16S, metagenomic, metatranscriptomic, and proteomic datasets have accumulated rapidly, most analyses remain cohort‐specific and association‐driven, limiting mechanistic insight, cross‐study transferability, and robustness to technical confounding. Foundation models offer a new computational framework by learning reusable biological representations from large unlabeled datasets. In this Review, we present microbiome foundation models as a hierarchy spanning biological scales. Sequence‐centric models capture the syntax and semantics of DNA and proteins for taxonomic inference, functional annotation, and generative design. Community‐centric models learn ecological structure from abundance profiles, while addressing compositionality, sparsity, and the unordered nature of microbial communities. Emerging multimodal frameworks integrate sequence‐derived functional potential with community‐level ecological dynamics under host and environmental context. We discuss key design choices, including tokenization, representation granularity, self‐supervised objectives, and evaluation strategies, and highlight challenges in interpretability, domain shift, causal reasoning, and biological validation. Finally, we propose a transition from static representation learning toward intervention‐aware microbiome world models capable of simulation, digital twinning, and generative microbiome engineering.
Gut microbiota and bile acids have been reported to affect sepsis progression, but the underlying mechanisms remain largely unknown. Here we investigated gut microbiota-bile acid interplay in two paediatric sepsis cohorts. Integration of bile acid-targeted metabolomics with gut metagenome data from paediatric sepsis patients identified deoxycholic acid 3-sulfate (DCA-3S) as significantly associated with paediatric sepsis progression. In vitro and in vivo experiments identified Enterococcus raffinosus as the primary producer of DCA-3S, contributing at least 80% of its total production, challenging the conventional notion of hepato-centric bile acid sulfation pathways. Intervention experiments in mouse and intestinal organoid models revealed that DCA-3S administration effectively alleviated sepsis by improving intestinal barrier function and attenuating inflammatory response. Collectively, our findings highlight a previously unrecognized microbial contribution to bile acid sulfation and position DCA-3S as a promising diagnostic and therapeutic biomarker for paediatric sepsis.
Background Micro-ecological islands provide unique habitats for microbes and play a crucial role in the functioning of aquatic ecosystems. Microbes settle on these micro-ecological islands, forming distinct microbial communities. Previous studies have provided some understanding of the colonization processes and regulatory mechanisms of protozoa in microbial communities. However, these islands are also subject to colonization by a variety of microbes beyond protozoa, and comprehensive cross-kingdom studies and their potential mechanisms remain largely unexplored. Results Using polyurethane foam units (PFU) to simulate micro-ecological islands, we studied the colonization dynamics of microbes in two distinct aquatic ecosystems, the Yangtze River and East Lake. Over 10-day colonization survey was conducted, we applied eDNA-PFU technology combined with metagenomic sequencing to comprehensively identify species present in the microbial communities, including bacteria, fungi, flagellates, protozoa, and metazoa. We found that microeukaryotes, rather than prokaryotes, were the primary colonizers in these two aquatic ecosystems. Our study reveals a colonization process of microeukaryotes in PFUs, profoundly influenced by their motility modes. Additionally, we propose a hypothetical food web framework within micro-ecological islands that maintains community stability, representing the most fundamental biological interactions. Conclusions Overall, this study enriches our understanding of micro-ecological islands and provides deeper insights into the colonization processes and regulatory mechanisms of microbial communities. It highlights the practical significance of micro-ecological islands in biological resource management, environmental protection, and biodiversity conservation.
Secondary metabolites in hemp enhance its pharmaceutical and nutraceutical value, yet the epigenetic regulatory network underlying secondary metabolite biosynthesis remains poorly understood in hemp. Here, we profiled the inflorescences of two cultivars with different trichome density by integrating metabolomics, transcriptomics and ATAC-seq. Multi-omics data revealed pronounced differences in metabolites (491 differentially accumulated metabolites (DAMs)), transcripts (8343 differentially expressed genes (DEGs)), and chromatin accessibility (11376 different accessibility genes, (DAGs)) between two cultivars. Integrated analyses reveal that increased chromatin accessibility at the promoters of several flavonoid-biosynthetic genes up-regulated their expression, resulting in the accumulation of flavonoids. Although chromatin accessibility of cannabinoid biosynthetic gene promoters modulates content, differential chromatin accessibility of the promoter of fatty acid biosynthetic and trichome density (trichome initiation, MeJA signaling, and identity of floral organ) related genes constitutes the key driver underlying cannabinoid divergence between two cultivars. Our study advances the understanding of epigenetic regulation of plant secondary metabolites and offers a novel strategy for enhancing cannabinoid and flavonoid content in Cannabis, providing efficient and precise molecular markers for the selection and breeding of new cannabis varieties.
Microbial communities are integral to human health, biotechnology, and environmental systems, yet their analysis is hindered by data heterogeneity and batch effects across studies. Traditional supervised methods often fail to capture universal patterns, limiting their utility in diverse contexts. Here, we present the Microbial General Model (MGM), the first large-scale foundation model for microbiome analysis, pretrained on 260,000 samples using transformer-based language modeling. MGM employs self-attention mechanisms and autoregressive pre-training to learn contextualized representations of microbial compositions, enabling robust transfer learning for downstream tasks. Benchmark evaluations demonstrate MGM's superior performance over conventional methods (average ROC-AUC = 0.99 vs. 0.68-0.97) in microbial community classification, with enhanced generalization across geographic regions. MGM also captures spatial and temporal microbial dynamics, as evidenced by its application to a longitudinal infant cohort, where it delineated delivery mode-specific microbiome trajectories and identified keystone genera such as Bacteroides and Bifidobacterium in vaginal deliveries and Haemophilus in cesarean deliveries. Furthermore, through prompt-guided generation, MGM produced realistic microbial profiles conditioned on disease labels. By integrating self-supervised learning with domain-specific fine-tuning, MGM advances the scalability and precision of microbiome analyses, offering a unified framework for diagnostics, ecological studies, and therapeutic discovery.
Radiation therapy is widely used for the treatment of brain tumors and metastases from extracranial malignancies; however, it may also cause damage to normal brain tissue, potentially resulting in radiation-induced brain injury (RBI). Emerging evidence highlights the microbiota–gut–brain axis (MGBA) as a critical mediator of bidirectional communication between the gut microbiota and the brain, playing an important role in central nervous system (CNS) homeostasis and pathology. This review aims to summarize current evidence regarding the potential involvement of the MGBA in the pathogenic mechanisms of RBI, with particular emphasis on bidirectional interactions along this axis. We focus on underlying mechanisms, including neuroimmune and inflammatory responses, signal transduction, DNA damage, and oxidative stress. By integrating these perspectives, this review seeks to provide a novel conceptual framework for understanding RBI and to identify potential directions for future MGBA-targeted interventions.
Abstract End-to-end omics analysis requires more than selecting tools: a usable agent must bind data correctly, execute long workflows without blocking, preserve provenance and recover the biological conclusions that motivate an analysis. Existing LLM-driven bioinformatics agents automate parts of this process, but their operational dependencies and conclusion-level validity are often unclear. Here we present MetaClaw, an auditable agent that maps a user request to a registered workflow, executes standardized upstream processing on FlowHub, and runs study-specific downstream analyses in network-isolated OpenClaw containers. A YAML registry and an explicit plan–submit–poll–finalise lifecycle record file bindings, parameters, scripts, environments and outputs in per-job bundles. Across the full cohorts of four published studies (769 metagenomic profiles), MetaClaw recovered 4/4 sorghum marker groups, 3/3 RRMS features, 4/5 canonical CRC markers among the top 20 classifier features and 5/5 permafrost marker groups. In 45 model-by-prompt runs, upstream completion was consistent whereas downstream validity depended on the backend and instruction detail; three decoy-tested endpoints showed no significant differences. In 48 ablation sessions, removing the registry, planning loop or manifest caused distinct losses, with registry removal increasing time, tool calls and token cost. MetaClaw therefore connects standardized upstream execution, local analytical flexibility and conclusion-level validation in a rerunnable framework for metagenomic and microbiome multi-omics analysis.
The extreme climatic conditions of the Tibetan Plateau foster unique microbial communities, especially in the sediment ecosystem. A thorough understanding of these communities could facilitate revealing their microbial diversity, biological resources, and response to climate change. Here, we have constructed the Tibetan Plateau Microbial Catalog of Sediment (TPMC-S) based on 248 metagenomic sediment samples from the Tibetan Plateau. We identified 511,056,752 nonredundant genes and recovered 13,696 metagenome-assembled genomes with enormous phylogenetic novelty (over 90% novel species), far exceeding other contemporary Tibetan microbial catalogs and expanding the microbial functional diversity. We also revealed that similarities of sediment microbial communities followed the distance-decay relationship. Furthermore, sediments contained a high proportion of evolutionarily "possible ancient species (PAS)" compared with paired aquatic samples, especially ancient archaeal lineages, suggesting a microbial "sedimentary archive" in sediment. Finally and most importantly, Asgardarchaeota, including 2 potentially novel genera, were identified from the sediments, and their latest divergence predated the uplift of the Tibetan Plateau, while they still gained functions to adapt to extreme environments. Our findings positioned the Tibetan Plateau as both a genomic repository of microbial antiquity, especially Asgardarchaeota, and an active arena for modern extremophile innovation, providing insights for deciphering microbial resilience strategies in climate-sensitive ecosystems and informing novel bioprospecting efforts.
Forensic microbiology leverages postmortem microbiome succession as a promising biomarker for estimating the postmortem interval (PMI). However, current methods are constrained by sparse sampling (typically 3–5 time points) and poor cross-anatomical generalizability, leading to imprecise PMI estimates with errors often exceeding ± 3 days, particularly in cases of dismembered remains. To overcome these, we developed mHolmes, a transformer-based digital twin framework powered by transfer learning. Trained on high-resolution data from 34 cadavers over 21 days, mHolmes achieves accurate daily predictions of microbial dynamics, reducing errors and demonstrating high accuracy (MAE < ± 2 days) in cross-anatomical forecasting (e.g., hip to face). Shapley Additive exPlanations (SHAP) analysis ensures interpretability by identifying seven key bacterial classes as conserved biomarkers. This study highlights mHolmes as a robust, high-precision tool that addresses critical bottlenecks, enabling reliable PMI estimation from incomplete data with significant applications in forensic investigations, such as body part matching and daily-resolution timeline reconstruction.
BACKGROUND:Patients undergoing cardiac surgery with cardiopulmonary bypass (CSCPB) are at substantial postoperative risk, which may be influenced by alterations in gut microbiota and metabolites. The roles of these biological changes in postoperative outcomes remain inadequately explored. METHODS:We collected 54 preoperative samples and 33 postoperative samples from 60 CSCPB patients. Metagenomic and metabolomic sequencing were performed to identify the gut microbiota and serum and fecal metabolites. We examined the dynamic pattern of these microbiota and metabolites, as well as their associations with the postoperative risks. Additionally, we developed a predictive model for postoperative risk based on preoperative microbiome and metabolome data. RESULTS:We revealed significant alterations of gut microbiota ( P = 0.012), serum metabolites ( P = 3.50 e-10 ), and fecal metabolites ( P = 0.0081) in patients following CSCPB, among which lysophosphatidylcholines (LPCs) exhibited notable changes. Particularly, we identified a potential regulatory function of the microbiota on LPC metabolism, which further influences the postoperative risk. The predictive model for intensive care unit stay duration achieved a mean absolute error of 1.27 days and an R² of 0.63, suggesting its utility in assessing postoperative risk. Also, our study provides a valuable resource (catalogue GM3C) for further investigation into potential medical targets in CSCPB patients, comprising more than 2,000 metagenome-assembled genomes and 3 million unigenes. CONCLUSIONS:Our study reveals that the gut microbiome and LPC-centered metabolism form a functional network influencing postoperative risk in CSCPB patients. These findings underscore the role of gut-derived signals in modulating noninfectious inflammatory responses and host imbalance, offering a multiomics framework for decoding systemic complications beyond classical sepsis paradigms. TRIAL REGISTRATION:ClinicalTrials.gov (NCT04032938). Registered 25 July 2019, https://clinicaltrials.gov/study/NCT04032938#study-record-dates .
The global dissemination of antibiotic resistance genes (ARGs) represents a critical challenge to One Health. Existing ARG risk assessment tools (e.g. MetaCompare, ARRI) are constrained by short-read sequencing data, limiting their utility for long-read platforms. To address this gap, we developed the Long-read based Antibiotic Resistome Risk Assessment Pipeline (L-ARRAP), which calculates the Long-read based Antibiotic Resistome Risk Index (L-ARRI) to quantify antibiotic resistome risks. Building upon our previous ARRI framework, L-ARRAP leverages long-read sequencing advantages to concurrently identify ARGs, mobile genetic elements, and human bacterial pathogens, integrating their interactions for risk scoring. Our results showed that L-ARRAP was not only able to accurately identify ARGs and evaluate the antibiotic resistance risk scores in samples of hospital wastewater (HWW), Chaohu lake, and human fecal samples, but also significantly distinguish the ARG risk in HWW samples between before and after disinfection groups, demonstrating the performance of L-ARRAP. Furthermore, L-ARRAP scores exhibited strong concordance with those generated by our laboratory-adapted MetaCompare variant (L-MetaCompare), corroborating its methodological reliability. Overall, to our knowledge, L-ARRAP is the first assessment pipeline of antibiotic resistome for long sequencing reads and has a great potential for monitoring the risk of ARGs in various environmental niches.
Molecular Dynamics (MD) simulations with first-principles accuracy are widely applied in various fields, including materials science and molecular pharmacology. Current research focus on reducing the solution time of ab initio molecular dynamics (AIMD) from both algorithmic and software perspectives. However, these optimizations are still far from meeting the ever-increasing performance demands for AIMD. In this paper, a fine-grained pipeline architecture is proposed to further optimize the solution time of MD. We design a high-utilization systolic line that eliminates the injection and evacuation time during each matrix operation. Besides, in conjunction with the MD dataflow, a computation migration strategy is introduced to reduce the storage overhead. Furthermore, we leverage dataflow rearrangement and preloading to eliminate the matrix transpose costs. The MD-pipe architecture has been implemented and verified using AMD VPK180 FPGA and ASIC synthesis tools. Evaluation results show that our work on FPGA and ASIC outperforms the NVIDIA A100 GPU by 2.97x and 23.77x respectively. Moreover, the ASIC implementation achieves a simulation speed of 67.6 mu s/ day, outperforming state-of-the-art work on Fugaku supercomputer by 454 times under extremely strong scaling deployment (one-atom-per-core).
AbstractMicrobial communities significantly impact medicine, biotechnology, and agriculture. Advanced sequencing technologies have generated extensive microbiome data, enabling the discovery of substantial evolutionary and ecological patterns. However, traditional supervised learning methods struggle to capture universal patterns in microbial community data, largely due to the large data heterogeneity and profound batch effects among samples, rendering it difficult to classify samples as well as detect biomarkers from millions of samples, not to say the intricate but important dynamic patterns from a variety of contextualized sceneries. In this study, we introduce the Microbial General Model (MGM), the first microbiome community foundation model pre-trained on a dataset of 263,302 microbiome samples using language modeling techniques. MGM demonstrated significant improvements in microbial community classification compared to traditional machine learning methods. Additionally, MGM has enabled contextualized classification, effectively overcomes cross-regional limitations, showing enhanced performance on intercontinental datasets through transfer learning. Furthermore, fine-tuning MGM on a longitudinal infant dataset revealed distinct keystone genera during development, withBacteroidesandBifidobacteriumexhibiting higher attention weights in vaginal deliveries, andHaemophilusin cesarean deliveries. Finally, through in silico modeling, the model also uncovered novel microbial dynamic patterns in a Crohn’s disease cohort following antibiotic treatment. In conclusion, by leveraging self-attention and autoregressive pre-training, MGM serves as a versatile model for various downstream microbiome tasks and holds significant potential for achieving contextualized aims.Key pointsThe Microbial General Model (MGM) is a foundation model with millions of parameters pre-trained on sub-million microbial community data.MGM outperforms traditional methods in various microbiome classification and prediction tasks, such as microbial community classification.MGM effectively captures the spatial and temporal dynamics of microbial communities.MGM could detect the effects of perturbation on microbial community through in silico experiments.
The ability to accurately predict the dynamic evolution of microbial communities is critical for advancing personalized medicine, precision intervention, and ecological system management. However, the irregular sampling, high missingness, and complex temporal behaviors that characterize longitudinal microbiome datasets present substantial challenges to reliable forecasting. Here we propose MicroProphet, a personalized digital twin framework capable of accurately forecasting microbial abundance trajectories from incomplete longitudinal observations without the need for data interpolation. By leveraging a time-aware Transformer architecture, MicroProphet reconstructs individualized microbial trajectories using as little as the initial 30% of time points, capturing critical transitional states through its attention mechanism. We demonstrate its robust cross-ecosystem generalizability across synthetic communities, human gut microbiomes, infant gut development, and corpse decomposition. In clinical contexts, MicroProphet enables early identification of disease-related microbial shifts and supports intervention timing optimization, exemplified in inflammatory bowel disease and antibiotic perturbation responses. By transforming incomplete and sparse data into actionable forecasts, MicroProphet establishes a foundation for real-time microbial monitoring, therapeutic decision support, and precision ecological management, paving the way for broader applications of digital twin systems in biology and personalized healthcare. Highlights ### Competing Interest Statement The authors have declared no competing interest. National Key R&D Program of China, 2023YFA1800900, 2018YFC0910502 the National Natural Science Foundation of China, 32071465, 31871334, 81827901
Biosynthetic gene clusters (BGCs), key in synthesizing microbial secondary metabolites, are mostly hidden in microbial genomes and metagenomes. To unearth this vast potential, we present BGC-Prophet, a transformer-based language model for BGC prediction and classification. Leveraging the transformer encoder, BGC-Prophet captures location-dependent relationships between genes. As one of the pioneering ultrahigh-throughput tools, BGC-Prophet significantly surpasses existing methods in efficiency and fidelity, enabling comprehensive pan-phylogenetic and whole-metagenome BGC screening. Through the analysis of 85 203 genomes and 9428 metagenomes, BGC-Prophet has profiled an extensive array of sub-million BGCs. It highlights notable enrichment in phyla like Actinomycetota and the widespread distribution of polyketide, NRP, and RiPP BGCs across diverse lineages. It reveals enrichment patterns of BGCs following important geological events, suggesting environmental influences on BGC evolution. BGC-Prophet’s capabilities in detection of BGCs and evolutionary patterns offer contributions to deeper understanding of microbial secondary metabolites and application in synthetic biology.
Increasing evidence has established the gut microbiota as a central driver of BA modification. Gut microbiota is involved in extensive BA modifications through various processes, including deconjugation, 7α-dehydroxylation, oxidation, epimerization, as well as emerging pathways like re-conjugation and succinylation. This review examined these microbial transformations, delineating the specific microbial species and metabolic enzymes involved. Focusing on the association between microbial modified BAs and human physiology, we investigated the existing therapeutic strategies targeting the microbiota-BA axis. These strategies can be categorized into two main domains: regulation of microbial composition (e.g., probiotics, FMT) and BA modifications (e.g., enzyme inhibitors). BA receptors, including canonical receptors (FXR, TGR5) and non-canonical (PXR, CAR, VDR, S1PR2, RORγt) receptors, further mediated the effects of BAs on various physiological functions, such as BA homeostasis, intestinal barrier integrity, and immune responses. We also summarized emerging tools, such as reverse metabolomics, source tracking algorithms, and organoid models, which facilitated knowledge discovery of novel BA modifications and their biological roles. Collectively, this review offers a comprehensive perspective on microbial BA modifications and underscores the significant potential of the gut microbiota-BA axis for disease management.
Fermented food remains poorly understood, largely due to the lack of knowledge about microbes in food fermentation. Here, this study constructed Moutai Fermented Grain Catalog (MTFGC), a representative liquor fermented by one of the most complex fermentations. MTFGC comprised 8,379,551 non-redundant genes and 5,159 metagenome-assembled genomes, with 20% species and 20% genes being novel. Additionally, 25,625 biosynthetic gene clusters (BGCs) and 28 BGC-enriched species were identified. Moreover, the microbial community assembly was deterministic, with significant species and gene changes in early fermentation stages, while stabilizing in later stages. Further BGC-knockout experiments verified Bacillus licheniformis, a BGC-enriched species, employed its BGCs for synthesizing the aroma-related lipopeptide lichenysin. This study has established the largest genetic resource for fermented food, uncovering its uniqueness and high metabolic potential. These findings facilitate the transition potential from traditional fermentation to precision-driven synthetic biology in food systems.