QSProteome ( https://QSProteome.org ) is a community-scale platform for modeling, evaluating, and refining quaternary protein structure. The resource hosts 35,528 unique modeled assemblies spanning over 42,000 genes, covering nearly all curated complexes in BioCyc and ComplexPortal databases. Each model is displayed on an interactive page with 3D visualization, chain-level confidence metrics, structural alignments, functional annotations, and automated stoichiometry checks against database expectations. A cloud-based server supports continuous user uploads and automated processing pipelines-enabling submission and validation of >54,000 models within 14 weeks. To promote iterative refinement, QSProteome includes a gamified re-curation workflow that transitions users from training modules into live curation, enabling the community-led assessment of 1,547 ABC transporter complexes. Together, these components form a dynamic, scalable infrastructure for proteome-scale structural biology. By unifying modeling, validation, and annotation in a reusable, searchable, and community-extensible framework, QSProteome enables proteome-scale structure accessibility and reuse-powering discovery, annotation, and collaborative refinement across the structural biology community.
Affordable sequencing has flooded public databases with bacterial genomes; yet, species-scale maps that connect gene content variation to metabolic functions essential to biotechnology/system biology remain scarce. We address this gap by building a pangenome-wide gene-protein-reaction association and applying it to 2377 Escherichia coli genomes to reconstruct a pangenome-scale metabolic model (panGEM). We validate panGEM against Biolog carbon source utilization assays, achieving ≈0.99 precision in growth/no-growth predictions. Using panGEM, we identify >11,000 rare metabolic genes, yet only 35 metabolic reactions are rare. To explain the mismatch, we examined rare genes and found that most are pseudogenes or diverged orthologs acquired by horizontal gene transfer (HGT). Results indicate a recurrent loss-reacquisition cycle in which a core allele is lost/pseudogenized and its function is restored by HGT, preserving function without expanding the reactome, generating genetic heterogeneity in a small subset (~3.6%) of reactions, marking selection pressure hotspots of metabolism. Thus, pangenome annotation reveals the evolutionary dynamics that shape the genetic basis of metabolism.
Energy homeostasis facilitated by the interplay of substrate-level and oxidative phosphorylation is crucial for bacterial adaptation to diverse substrates and environments. To investigate how bioenergetic systems optimize under metabolically restrictive conditions, we evolved Electron Transport System (ETS) variants with distinct proton-pumping efficiencies on succinate and glycerol. These substrates impose unique metabolic constraints: succinate requires complete gluconeogenesis, while glycerol supports mixed glycolytic and gluconeogenic fluxes. Multi-scale computational analysis of the strains revealed (a) Growth optimization across carbon substrates for multiple ETS variants, (b) A conserved aerobicity stimulon comprising seven independently regulated gene groups that are co-regulated with increasing aerobic capacities, (c) Proteome reallocation linked to aerobicity, validated using genome-scale metabolism and expression modeling, and (d) Carbon source-specific compensatory mutations. These findings define the aerobicity stimulon and establish a unifying framework for understanding bacterial respiratory flexibility, demonstrating how transcriptional networks and metabolic systems integrate to achieve energy homeostasis and bioenergetic resilience.
Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida. Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.IMPORTANCEPseudomonas putida has become an organism of interest for biotechnological applications, but a species-level understanding of its metabolic diversity remains incomplete. In this study, we analyzed 164 P. putida strains using a combination of genome sequencing, phenotypic profiling, and metabolic modeling. Our results indicate that while many metabolic pathways are conserved, notable differences exist across strains, particularly in aromatic compound degradation. These observations may inform future strain selection and engineering strategies tailored to specific industrial or environmental goals. In addition, the genome-scale models and phenotypic data generated here can serve as a foundation for broader studies of metabolism and functional variation within this species.
Abstract Streptococcus pyogenes (group A Streptococcus , GAS) causes over 700 million infections annually and has resurged globally as a cause of invasive disease. However, the genomic basis by which distinct GAS lineages produce diverse clinical phenotypes remains incompletely understood. Here, we analyzed 399 quality-controlled complete genomes spanning 103 emm types and 27 countries and applied non-negative matrix factorization (NMF) to decompose the GAS pangenome into co-inherited gene modules capturing both clonal lineages and mobile elements. This framework revealed that the deepest division in the accessory genome is not defined by emm type, but by a 27-gene Sda-1-encoding prophage linked to the hypervirulent M1T1 pandemic clone. This NET-degrading module unites emm1 , emm12 , and emm77 lineages, including fixation in emm77/ST63, extending its distribution beyond classically invasive lineages. Across lineages, virulence determinants segregate into distinct combinations, indicating that invasive disease arises through convergent but mechanistically distinct programs. Consistent with this, highly invasive ST52- emm28 lacks Sda1 and deploys an alternative repertoire. The decomposition also resolves a gene module corresponding to the emm-pattern D regulon associated with skin tropism, linking accessory genome structure to host niche adaptation. We further identify three sequence-divergent speC paralogs on independent prophages. These findings define modular GAS virulence architectures and establish NMF-based pangenome analysis as a framework for genomic surveillance beyond single-locus typing. Importance Group A Streptococcus causes diseases ranging from strep throat to life-threatening flesh-eating infections, with severe cases increasing in recent years and no vaccine currently available. The species is conventionally classified by its M protein gene, but this single-locus system does not capture the extensive genomic diversity driven by prophages that transfer toxin and immune-evasion genes between bacterial lineages. Here, we applied unsupervised machine learning to decompose the pangenome of nearly 400 complete genomes into modules of co-inherited genes. We find that the deepest division among lineages is determined by a prophage carrying a gene that enables evasion of neutrophil extracellular traps, and identify this module at fixation in a lineage not previously known to carry it. We also show that a key toxin gene exists as three independently evolved variants on distinct prophages, and that invasive lineages deploy distinct combinations of virulence factors. This framework enables genomic surveillance based on virulence architecture rather than single-gene typing.
Understanding the functions of DNAs, RNAs, and proteins is fundamental to advancing life science research and enabling translational applications such as drug discovery and precision medicine. While deep learning methods have shown promise in biomolecular function prediction, they typically constrain outputs to predefined categories and require training separate models for each task. Existing multi-task learning methods operate on a fixed set of predefined tasks and require model retraining when new tasks arise. Furthermore, current approaches produce one-shot, static outputs, lacking the capacity for iterative refinement or deeper exploration of predictions. This position paper argues that multi-modal large language models (LLMs) are essential for enabling free-form and interactive prediction of biomolecular functions, and zero-shot generalization to new tasks without model retraining. These models can generate coherent and context-aware text outputs that reflect the complexity and nuance of diverse functional roles. Importantly, they can generalize to novel biomolecules whose functions are unknown or poorly characterized, and they enable generalization to new tasks through prompt-driven adaptation, eliminating the need for task-specific retraining. Additionally, multi-modal LLMs enable interactive, multi-turn dialogue, allowing users to iteratively refine queries, clarify contexts, and explore hypotheses in a dynamic and responsive manner. By leveraging these capabilities, multi-modal LLMs provide a scalable, adaptable, and generalizable framework for advancing biomolecular function prediction and accelerating biological discovery.
ABSTRACT Enterococci are Gram-positive opportunistic pathogens responsible for a wide range of nosocomial infections. One enterococcocal species, Enterococcus faecium , is steadily increasing in prevalence and has been listed among major multidrug-resistant ESKAPE pathogens. To gain systems-level insights into its metabolism and support discovery of potential therapeutic targets, we constructed iDR479, a comprehensive manually curated genome-scale metabolic model (GEM) to serve as a digital twin for E. faecium TX0016 (strain DO). The reconstruction was curated through extensive homology searches and literature evidence, and further refined and gap-filled through experimental validation. Phenotypic profiling using Biolog microarrays enabled assessment of carbon source utilization, while amino acid leave-out growth assays allowed the evaluation of auxotrophies. The final refined model is 100% accurate in predicting amino acid auxotrophy and 85% accurate in predicting growth on sole carbon sources. Discrepancies between model predictions and experimental phenotypes identified specific knowledge gaps across metabolic pathways, including unresolved carbon source utilization phenotypes, e.g., psicose, sorbitol, and palatinose utilization. Those gaps will guide future experimental characterization. Additionally, gene essentiality analysis was conducted to evaluate the predictive capacity of iDR479 model. Since no experimental gene essentiality data are currently available for E. faecium , model predictions were compared against Tn-seq experimental results from E. faecalis MMH594. Under simulated rich medium conditions, iDR479 achieved 86.7% concordance with the experimental essentiality results of E. faecalis MMH594. iDR479 thus provides a framework for studying E. faecium , offers insights into its metabolic network, and serves as a source for guiding future research and identification of therapeutic targets.
Escherichia coli strains are widely used across numerous industrial and biotechnological applications. Yet their performance varies substantially in ways that can not be anticipated from genome annotation. Because transcriptional regulatory networks (TRNs) govern cellular functions such as motility, stress responses, metabolic flexibility, and production efficiency, differences in TRN organization and use may underlie many observed phenotypic differences. To investigate TRN differences between strains, we generated a compendium of 433 matched RNA-Seq profiles for six commonly used industrial E. coli strains (BL21, C, Crooks, MG1655, W, and W3110) and applied iModulon analysis to compare the state of their TRNs under similar growth conditions. This analysis revealed that core regulatory programs with similar functions are wired differently across the strains, and that the strains engage these programs in distinct ways when exposed to the same environmental challenges. Together, these findings highlight transcriptional regulation diversity underlying phenotypic expression among industrial E. coli strains. By providing an integrated view of TRN differences across widely used hosts, this work offers a fundamental basis for interpreting strain-specific behaviors and supports more informed approaches to strain selection and optimization.
The rapid growth of bacterial gene expression databases has enabled computational inference of transcriptional regulatory networks (TRNs), yet it remains unclear why mathematically simple models often capture their apparent complexity. Using a 1035-sample E. coli expression database, we identify two transcriptome principles that support successful TRN inference. First, regulons defined from measured binding sites show limited overlap in gene membership, consistent with statistical independence exhibited by many successful inference methods. Second, 21% of genes, or 877 genes, exhibit regulator "dominance," in which expression strongly correlates with a single regulator activity and receives minimal contributions from other regulators under most conditions. We formalize these properties with quantitative metrics and provide a reference catalog of dominantly regulated E. coli genes. Regulator dominance explains differences between expression-inferred and binding site-defined regulons, and removing dominated genes sharply reduces inference performance, suggesting that simply regulated promoter subsets are central to effective TRN inference.
As space agencies prepare for complex missions such as human exploration of Mars and life detection, Planetary Protection (PP) risk assessments would benefit from extending beyond traditional microbial monitoring methods like the Standard Spore Assay (SSA). Metagenomics offers a powerful complementary approach, enabling detection of non-culturable organisms and providing functional insights essential for risk-informed decision-making. However, operational best practices would include rigorous validation, standardization, and integration into PP frameworks. In November 2024, NASA Ames Research Center hosted a workshop with experts from academia, government, and industry to assess the readiness of metagenomic methods for PP and outline a roadmap for implementation. Discussions addressed six key areas: (1) safetycritical decision making, (2) microbial dark matter, (3) ultra-low biomass metagenomics, (4) bioinformatics and databases, (5) technology development and automation, and (6) human health and the built environment. Cross-cutting recommendations included adopting tiered testing frameworks, developing shared reference standards, standardizing protocols and metadata, building probabilistic risk models, and investing in automation and contamination-aware technologies. These actions aim to enable metagenomics as a robust tool for Planetary Protection in future space missions.
Auto-brewery syndrome (ABS) is a rarely diagnosed disorder of alcohol intoxication due to gut microbial ethanol production. Despite case reports and a small cohort study, the microbiological profiles of patients remain poorly understood. Here we conducted an observational study of 22 patients with ABS and 21 unaffected household partners. Faecal samples from individuals with ABS during a flare produced more ethanol in vitro, which could be reduced by antibiotic treatment. Gut microbiome analysis using metagenomics revealed an enrichment of Proteobacteria, including Escherichia coli and Klebsiella pneumoniae. Genes in metabolic pathways associated with ethanol production were enriched, including the mixed-acid fermentation pathway, heterolactic fermentation pathway and ethanolamine utilization pathway. Faecal metabolomics revealed increased acetate levels associated with ABS, which correlated with blood alcohol concentrations. Finally, one patient was treated with faecal microbiota transplantation, with positive correlations between gut microbiota composition and function, and symptoms. These findings can inform future clinical interventions for ABS. Gut microorganisms capable of producing ethanol cause alcohol intoxication, which can be corrected via faecal microbiota transplantation.
The earliest responses of pathogenic bacteria to antibiotics can affect the outcome of an infection. While long-term adaptations have been extensively studied, the immediate transcriptional changes that unfold immediately following antibiotic exposure remain poorly understood. Here, we applied iModulon analysis to time-resolved transcriptomic data from Escherichia coli exposed to subinhibitory concentrations of two antibiotics (ampicillin and ciprofloxacin), capturing transcriptional regulatory changes occurring within the first 30 min of exposure. This analysis proposes an integrated, three-phase response model: an immediate and sustained primary response that broadly activates stress programs, a transient secondary response that restores redox balance, and a tertiary response that supports long-term survival through metabolic remodeling and antibiotic-specific defenses. These results highlight a coordinated and dynamic regulatory strategy describing how metabolic, redox, and stress responses are integrated to manage the physiological challenges of antibiotic stress. By disentangling these overlapping transcriptional regulatory programs, this work offers a genome-scale understanding of how early regulatory programs are engaged immediately after antibiotic exposure. Together, these findings provide a structured framework for characterizing complex transcriptomic responses and generating testable hypotheses about the regulatory logic that shapes the understudied early phase of antibiotic exposure.IMPORTANCEInitial bacterial responses to antibiotics are important for survival and can influence the development of tolerance and resistance. However, this period remains poorly understood, in part, because the transcriptional responses that unfold within minutes of antibiotic exposure are complex and difficult to interpret. In this study, we applied novel data generation and data analytics approaches to resolve the regulatory structure of the initial response of Escherichia coli to two antibiotics. We identify a three-phase process that explains how E. coli coordinates stress responses, maintains redox homeostasis, and initiates downstream protective programs. The novel transcriptomic analytics elucidate independently regulated sets of genes that constitute cellular processes. By identifying the regulatory modules that change over this initial timescale, we can deconvolute the response based on first principles of cellular physiology.
Diatoms are vital primary producers in marine ecosystems and play a key role in blue carbon sequestration. Although their circadian rhythms have been studied in isolation, how these rhythms are modulated by co-occurring bacteria remains unknown. Using a defined co-culture of the diatom Phaeodactylum tricornutum with the marine bacterium Aliivibrio fischeri, we observed time-dependent physiological and transcriptomic changes in P. tricornutum, including a 16.3% reduction in rhythmic genes with prolonged culture time. Genome-scale metabolic modeling suggested a biphasic response, with predicted biomass flux increasing by 81% at the early co-culture stage but decreasing by 87% during prolonged co-culture. Deconvolution of the transcriptome via AI-driven independent component analysis identified gene modules associated with silica transport and senescence-related responses under co-culture conditions. Together, these findings establish a systems-level framework that links interspecies interactions between diatoms and bacteria, providing mechanistic insights into how microbial associations influence phytoplankton chronobiology and rhythmic regulation.
Red blood cells (RBCs) are transcriptionally silent yet dynamically remodel metabolism in response to oxygen tension. Using ultra-pure human RBCs, we generated the deepest contamination-free proteome to date (3,775 proteins) and mapped the oxygen-dependent interactome. These datasets reveal an oxygen-responsive metabolon centered on the Band 3 (SLC4A1) N-terminus. We identify biliverdin reductase B (BLVRB) as a previously unrecognized Band 3 interactor that dissociates under hypoxia, coincident with increased Band 3-deoxyhemoglobin contacts. This reversible assembly functions as an oxygen-sensitive switch coordinating redox and glycolytic remodeling. Humanized mice lacking Band 3 N-terminal segments exhibit impaired oxygen-dependent regulation of BLVRB binding to band 3, impaired hypoxic activation of glycolysis, reduced 2,3-bisphosphoglycerate synthesis, and diminished exercise tolerance, demonstrating physiological relevance. Population-scale cis-pQTLs for SLC4A1 and BLVRB suggest functions beyond canonical heme catabolism. Mechanistically, biochemical analyses in vitro suggest that hemoglobin β (HBB), Band 3, and BLVRB can undergo S-nitrosation and may participate in trans-nitrosation reactions with the glycolytic enzyme GAPDH, whose modification at C152 inhibits enzymatic activity in vitro. Collectively, these findings define a Band 3-BLVRB axis that integrates oxygen-dependent protein interactions with thiol-based redox chemistry, providing a framework for understanding how an anucleate cell achieves metabolic adaptability through reversible protein-protein interactions and post-translational modification. These findings suggest that perturbation of the Band 3-BLVRB axis may influence oxygen delivery and metabolic flexibility during hypoxic stress, with potential relevance to high-altitude adaptation, exercise physiology, and cardiopulmonary disease.
Streptococcus pyogenes can cause a wide variety of acute infections throughout the body of its human host [...]
Streptomyces albidoflavus is a widely used strain for natural product discovery and production through heterologous biosynthetic gene clusters (BGCs). However, the transcriptional regulatory network (TRN) and its impact on secondary metabolism remain poorly understood. Here, we characterize the TRN using independent component analysis on 218 RNA sequencing (RNA-seq) transcriptomes across 88 unique growth conditions. We identify 78 independently modulated sets of genes (iModulons) that quantitatively describe the TRN across diverse conditions. Our analyses reveal (1) TRN adaptation to different growth conditions, (2) conserved and unique characteristics of the TRN across diverse lineages, (3) transcriptional activation of several endogenous BGCs, including surugamide, minimycin, and paulomycin, and (4) inferred functions of 40% of uncharacterized genes in the S. albidoflavus genome. These findings provide a comprehensive and quantitative understanding of the S. albidoflavus TRN, offering a knowledge base for further exploration and experimental validation.
Creatine is an important energy storage molecule produced exclusively in vertebrates and is crucial for muscle development. It is particularly valuable as a food supplement, especially for plant-based diets. Here, we present an alternative to chemical synthesis by developing a biosynthetic process using an Escherichia coli cell factory expressing a heterologous pathway. We employed a model-driven growth-coupled selection approach combined with adaptive laboratory evolution to overcome metabolic bottlenecks in the heterologous synthesis of creatine. We developed a novel growth-coupling strategy to optimize an important glycine amidinotransferase step guided by genome-scale modeling. We also improved creatine tolerance of E. coli by adaptive evolution. Several design-build-test-learn cycles of evolution and selection resulted in a 58 % increase in titer over the baseline strain from glycine and arginine. This study highlights the advantage of combining production with growth for efficient cell factory generation driven by evolutionary engineering and computational biology.
Generating longitudinal and multi-layered big biological data is crucial for effectively implementing artificial intelligence (AI) and systems biology approaches in characterising whole-body biological functions in health and complex disease states. Big biological data consists of multi-omics, clinical, wearable device, and imaging data, and information on diet, drugs, toxins, and other environmental factors. Given the significant advancements in omics technologies, human metabologenomics, and computational capabilities, several multi-omics studies are underway. Here, we first review the recent application of AI and systems biology in integrating and interpreting multi-omics data, highlighting their contributions to the creation of digital twins and the discovery of novel biomarkers and drug targets. Next, we review the multi-omics datasets generated worldwide to reveal interactions across multiple biological layers of information over time, which enhance precision health and medicine. Finally, we address the need to incorporate big biological data into clinical practice, supporting the development of a clinical decision support system essential for AI-driven hospitals and creating the foundation for an AI and systems biology-based healthcare model.
Sequenced genomes for thousands of strains of a bacterial species allow for a comprehensive analysis of its pangenome. We present a pangenome study of Escherichia coli metabolism by formulating gene to protein to reaction associations (GPRs) for about 2,700 metabolic reactions in 2,377 fully sequenced strains. On one hand, these GPRs reconstruct strain specific networks that allow computational predictions (and experimental validation) of metabolic phenotypes, while on the other hand, they give the genetic basis for a given metabolic reaction in every strain. A pangenome wide analysis of GPRs shows that: 1) We can reveal the genetic basis for a specific metabolic property at the species level; 2) The genetic basis for many metabolic reactions is diverse; 3) Many rare genes show variation in the genes genomic neighborhood which often contain genes from transposable elements, 4) Many rare genes show large scale fragmentation and horizontal gene transfer (>11,000 rare genes in 2,377 strains); and 5) The aromatic amino acids and branched chain amino acids pathways are enriched with rare genes, with Acetolactate synthase having 29 distinct genes. Thus, analysis of GPRs across the pangenome reveals a complex dynamic evolutionary history of metabolism, revealing the role of conserved, fragmented, and horizontally transferred metabolic genes. ### Competing Interest Statement The authors have declared no competing interest.