Abstract Microbial diversity is often assumed to be limited by the number of available resources, yet many communities persist well beyond that expectation. Understanding the mechanisms that enable such coexistence remains a central question in microbial ecology. Here, using a four-species bacterial consortium, we asked whether coexistence can emerge from interactions between species rather than from the external environment alone. Across 31 simple nutrient conditions, including 16 single-resource environments, all four species persisted and repeatedly reached stable coexistence. We then chose 27 additional conditions to further probe the boundaries of coexistence by varying resource concentrations, temporal dynamics, nutrient complexity and relief of auxotrophy-associated dependencies, and only observed the extinction of one species in one of these conditions. Although the community composition in each environment was largely shaped by species’ fitness on the supplied resources, experimental assays and consumer-resource modeling showed that the coexistence was not explained by resource supply, but rather by cross-feeding and niche partitioning of metabolic byproducts. These metabolic interactions were strong enough to sustain coexistence even for species unable to use the supplied resources directly. Furthermore, robust coexistence across environments appears to be an emergent property of microbial communities, ingrained in members’ metabolic byproduct profiles and niche differences. Our findings demonstrate how microbes can increase the chemical complexity of their environment sufficiently to maintain coexistence well beyond what is expected from external resource supply. Significance Understanding the drivers of microbial diversity is essential for managing natural ecosystems and designing synthetic microbiomes. This study challenges the conventional application of the competitive exclusion principle, demonstrating that a four-species consortium can coexist across 31 chemically and metabolically diverse one- and two-carbon source environments. By systematically testing and ruling out alternative stabilizing mechanisms, we show that co-existence is an emergent property of the consortium, sustained by metabolic cross-feeding and niche partitioning. Guided by computational models, we identify hallmarks of robust co-existence in simple environments, including high variance in resource affinities and growth on partner-derived metabolites. Our work demonstrates how microbes modify their environment to sustain high diversity and provides principles for designing synthetic microbiomes that persist across environments.
Microbial biotechnology has the potential to address several societal issues through the sustainable production of industrially relevant compounds. Despite decades of successful cases, rational engineering of microbial metabolism is still a complex process due to the fine balance between nutrient supply, allocation of cellular resources, energy demand and redox balancing. In this work, we implemented a text-mining workflow for metabolic engineering and compiled a database of experimentally validated strain design strategies from over 15,000 research articles, which includes information on host strain, target compounds, and gene modifications. This large dataset reveals trends on the selection of suitable hosts for different kinds of products and their respective gene targets. Despite the wide variety of microbes and products, we observe a conserved set of target metabolic genes associated with central carbon metabolism, especially in upper glycolysis, pentose-phosphate pathway, citric acid cycle, and fermentative pathways. The most distinguishing feature among strain design strategies seems not to be which genes are targeted, but rather the direction in which they are modified (increased or decreased expression). Controlling flux at key branching points and redox balancing reactions is thus a critical engineering step to steer metabolism. Our collection of 25 years of literature can provide a stepping stone for starting new strain design projects without reinventing the wheel.
The development of sustainable biotechnological processes requires a transition from the traditional fermentation of refined substrates toward the valorization of waste materials such as lignocellulosic biomass. Although these so-called recalcitrant substrates cannot be degraded by model industrial organisms, they can be degraded by microbial consortia through a process of anaerobic digestion, where different community members are able to break down polysaccharides of varied complexity. Among these microbes, Ruminiclostridium cellulolyticum stands out as a promising candidate for fermentation of lignocellulose due to its ability to degrade both cellulose and hemicellulose. In this work, we present an updated genome-scale metabolic model for R. cellulolyticum strain H10. The model was manually curated with experimental data, and the pathways for degradation of cellulose and hemicellulose (arabinoxylan and xyloglucan) were reconstructed and annotated with full detail. The model enables the simulation of the fermentation profile of lignocellulosic materials of various compositions, facilitating the use of this organism as a potential workhorse for sustainable biotechnology, and it provides a valuable template for the reconstruction and optimization of lignocellulose degradation pathways in related organisms. IMPORTANCE:In this work, we present a manually curated genome-scale metabolic model for Ruminiclostridium cellulolyticum, one of the few species known to fully degrade cellulose and hemicellulose. The model was extensively curated with experimental data obtained from the literature, covering approximately 25 years of research on this organism. We use this model to simulate the fermentation of mixed lignocellulosic polysaccharides and observe a good agreement with experimental data. This organism is therefore a promising microbial cell factory for sustainable transformation of lignocellulosic residues into valuable industrial products.
Deregulated signal transduction and energy metabolism are hallmarks of cancer and both play a fundamental role in tumorigenesis. While it is increasingly recognised that signalling and metabolism are highly interconnected, the underpinning mechanisms of their co-regulation are still largely unknown. Here we designed and acquired proteomics, phosphoproteomics, and metabolomics experiments in fumarate hydratase (FH) deficient cells and developed a computational modelling approach to identify putative regulatory phosphorylation-sites of metabolic enzymes. We identified previously reported functionally relevant phosphosites and potentially novel regulatory residues in enzymes of the central carbon metabolism. In particular, we showed that pyruvate dehydrogenase (PDHA1) enzymatic activity is inhibited by increased phosphorylation in FH-deficient cells, restricting carbon entry from glucose to the tricarboxylic acid cycle. Moreover, we confirmed PDHA1 phosphorylation in human FH-deficient tumours. Our work provides a novel approach to investigate how post-translational modifications of enzymes regulate metabolism and could have important implications for understanding the metabolic transformation of FH-deficient cancers with potential clinical applications.
Genome-scale metabolic modeling is a powerful framework for predicting metabolic phenotypes of any organism with an annotated genome. For two decades, this framework has been used for the rational design of microbial cell factories. In the last decade, the range of applications has exploded, and new frontiers have emerged, including the study of the gut microbiome and its health implications and the role of microbial communities in global ecosystems. However, all the critical steps in this framework, from model construction to simulation, require the use of powerful linear optimization solvers, with the choice often relying on commercial solvers for their well-known computational efficiency. In this work, I benchmark a total of six solvers (two commercial and four open source) and measure their performance to solve linear and mixed-integer linear problems of increasing complexity. Although commercial solvers are still the fastest, at least two open-source solvers show comparable performance. These results show that genome-scale metabolic modeling does not need to be hindered by commercial licensing schemes and can become a truly open science framework for solving urgent societal challenges.IMPORTANCEModeling the metabolism of organisms and communities allows for computational exploration of their metabolic capabilities and testing their response to genetic and environmental perturbations. This holds the potential to address multiple societal issues related to human health and the environment. One of the current limitations is the use of commercial optimization solvers with restrictive licenses for academic and non-academic use. This work compares the performance of several commercial and open-source solvers to solve some of the most complex problems in the field. Benchmarking results show that, although commercial solvers are indeed faster, some of the open-source options can also efficiently tackle the hardest problems, showing great promise for the development of open science applications.
Abstract Cheese fermentation and flavour formation are governed by complex biochemical reactions driven by polymicrobial activity. While the compositional dynamics of cheese microbiomes is relatively well mapped, the mechanistic role of microbial interactions in flavour formation is yet unknown. We microbially and metabolically characterised a year-long Cheddar cheese process using a commonly used starter culture containing Streptococcus thermophilus and Lactococcus strains. By using an experimental strategy whereby certain strains were left out from the starting mixture, we identified the critical role of S. thermophilus in boosting Lactococcus growth and in shaping flavour compound profile. Controlled milk fermentations with systematic exclusion of single Lactococcus strains, combined with genomics, genome-scale metabolic modelling, and metatranscriptomics, indicated that proteolytic activity of S. thermophilus relieves nitrogen limitation for Lactococcus and boosts de novo nucleotide biosynthesis. While S. thermophilus had large contribution to the flavour profile, L. cremoris also played a role by limiting diacetyl and acetoin formation which leads to off-flavour when in excess. This off-flavour control could be attributed to different metabolic re-routing of citrate between L. cremoris and other L. lactis strains. Further, closely related L. lactis strains exhibited different interaction patterns with S. thermophilus highlighting the importance of strain-specificity in cheese-making. Overall, our results bring forward the critical role of competitive and cooperative microbial interactions shaping cheese flavour profile.
Cheese fermentation and flavour formation are the result of complex biochemical reactions driven by the activity of multiple microorganisms. Here, we studied the roles of microbial interactions in flavour formation in a year-long Cheddar cheese making process, using a commercial starter culture containing Streptococcus thermophilus and Lactococcus strains. By using an experimental strategy whereby certain strains were left out from the starter culture, we show that S. thermophilus has a crucial role in boosting Lactococcus growth and shaping flavour compound profile. Controlled milk fermentations with systematic exclusion of single Lactococcus strains, combined with genomics, genome-scale metabolic modelling, and metatranscriptomics, indicated that S. thermophilus proteolytic activity relieves nitrogen limitation for Lactococcus and boosts de novo nucleotide biosynthesis. While S. thermophilus had large contribution to the flavour profile, Lactococcus cremoris also played a role by limiting diacetyl and acetoin formation, which otherwise results in an off-flavour when in excess. This off-flavour control could be attributed to the metabolic re-routing of citrate by L. cremoris from diacetyl and acetoin towards α -ketoglutarate. Further, closely related Lactococcus lactis strains exhibited different interaction patterns with S. thermophilus , highlighting the significance of strain specificity in cheese making. Our results highlight the crucial roles of competitive and cooperative microbial interactions in shaping cheese flavour profile.
The increasing world population and living standards urgently necessitate the transition towards a sustainable food system. One solution is microbial protein, i.e. using microbial biomass as alternative protein source for human nutrition, particularly based on renewable electron and carbon sources that do not require arable land. Upcoming green electrification and carbon capture initiatives enable this, yielding new routes to H2, CO2 and CO2-derived compounds like methane, methanol, formic- and acetic acid. Aerobic hydrogenotrophs, methylotrophs, acetotrophs and microalgae are the usual suspects for nutritious and palatable biomass production on these compounds. Interestingly, these compounds are largely un(der)explored for purple non-sulfur bacteria, even though these microbes may be suitable for growing aerobically and phototrophically on these substrates. Currently, selecting the best strains, metabolisms and cultivation conditions for nutritious and palatable microbial food mainly starts from empirical growth experiments, and mostly does not stretch beyond bulk protein. We propose a more target-driven and efficient approach starting from the genome-embedded potential to tuning towards, for instance, essential amino- and fatty acids, vitamins, taste,... Genome-scale metabolic models combined with flux balance analysis will facilitate this, narrowing down experimental variations and enabling to get the most out of the 'best' combinations of strain and electron and carbon sources.
The repeated, rapid and often pronounced patterns of evolutionary divergence observed in insular plants, or the ‘plant island syndrome’, include changes in leaf phenotypes, growth, as well as the acquisition of a perennial lifestyle. Here, we sequence and describe the genome of the critically endangered, Galápagos-endemic species Scalesia atractyloides Arnot., obtaining a chromosome-resolved, 3.2-Gbp assembly containing 43,093 candidate gene models. Using a combination of fossil transposable elements, k -mer spectra analyses and orthologue assignment, we identify the two ancestral genomes, and date their divergence and the polyploidization event, concluding that the ancestor of all extant Scalesia species was an allotetraploid. There are a comparable number of genes and transposable elements across the two subgenomes, and while their synteny has been mostly conserved, we find multiple inversions that may have facilitated adaptation. We identify clear signatures of selection across genes associated with vascular development, growth, adaptation to salinity and flowering time, thus finding compelling evidence for a genomic basis of the island syndrome in one of Darwin’s giant daisies.
Recent studies have brought forward the critical role of emergent properties in shaping microbial communities and the ecosystems of which they are a part. Emergent properties—patterns or functions that cannot be deduced linearly from the properties of the constituent parts—underlie important ecological characteristics such as resilience, niche expansion and spatial self-organization. While it is clear that emergent properties are a consequence of interactions within the community, their non-linear nature makes mathematical modelling imperative for establishing the quantitative link between community structure and function. As the need for conservation and rational modulation of microbial ecosystems is increasingly apparent, so is the consideration of the benefits and limitations of the approaches to model emergent properties. Here we review ecosystem modelling approaches from the viewpoint of emergent properties. We consider the scope, advantages and limitations of Lotka–Volterra, consumer–resource, trait-based, individual-based and genome-scale metabolic models. Future efforts in this research area would benefit from capitalizing on the complementarity between these approaches towards enabling rational modulation of complex microbial ecosystems.
Tumor relapse from treatment-resistant cells (minimal residual disease, MRD) underlies most breast cancer-related deaths. Yet, the molecular characteristics defining their malignancy have largely remained elusive. Here, we integrated multi-omics data from a tractable organoid system with a metabolic modeling approach to uncover the metabolic and regulatory idiosyncrasies of the MRD. We find that the resistant cells, despite their non-proliferative phenotype and the absence of oncogenic signaling, feature increased glycolysis and activity of certain urea cycle enzyme reminiscent of the tumor. This metabolic distinctiveness was also evident in a mouse model and in transcriptomic data from patients following neo-adjuvant therapy. We further identified a marked similarity in DNA methylation profiles between tumor and residual cells. Taken together, our data reveal a metabolic and epigenetic memory of the treatment-resistant cells. We further demonstrate that the memorized elevated glycolysis in MRD is crucial for their survival and can be targeted using a small-molecule inhibitor without impacting normal cells. The metabolic aberrances of MRD thus offer new therapeutic opportunities for post-treatment care to prevent breast tumor recurrence.
Resource competition and metabolic cross-feeding are among the main drivers of microbial community assembly. Yet the degree to which these two conflicting forces are reflected in the composition of natural communities has not been systematically investigated. Here, we use genome-scale metabolic modelling to assess the potential for resource competition and metabolic cooperation in large co-occurring groups (up to 40 members) across thousands of habitats. Our analysis reveals two distinct community types, which are clustered at opposite ends of a spectrum in a trade-off between competition and cooperation. At one end are highly cooperative communities, characterized by smaller genomes and multiple auxotrophies. At the other end are highly competitive communities, which feature larger genomes and overlapping nutritional requirements, and harbour more genes related to antimicrobial activity. The latter are mainly present in soils, whereas the former are found in both free-living and host-associated habitats. Community-scale flux simulations show that, whereas competitive communities can better resist species invasion but not nutrient shift, cooperative communities are susceptible to species invasion but resilient to nutrient change. We also show, by analysing an additional data set, that colonization by probiotic species is positively associated with the presence of cooperative species in the recipient microbiome. Together, our results highlight the bifurcation between competitive and cooperative metabolism in the assembly of natural communities and its implications for community modulation.
Microbial communities often undergo intricate compositional changes yet also maintain stable coexistence of diverse species. The mechanisms underlying long-term coexistence remain unclear as system-wide studies have been largely limited to engineered communities, ex situ adapted cultures or synthetic assemblies. Here, we show how kefir, a natural milk-fermenting community of prokaryotes (predominantly lactic and acetic acid bacteria) and yeasts (family Saccharomycetaceae), realizes stable coexistence through spatiotemporal orchestration of species and metabolite dynamics. During milk fermentation, kefir grains (a polysaccharide matrix synthesized by kefir microorganisms) grow in mass but remain unchanged in composition. In contrast, the milk is colonized in a sequential manner in which early members open the niche for the followers by making available metabolites such as amino acids and lactate. Through metabolomics, transcriptomics and large-scale mapping of inter-species interactions, we show how microorganisms poorly suited for milk survive in—and even dominate—the community, through metabolic cooperation and uneven partitioning between grain and milk. Overall, our findings reveal how inter-species interactions partitioned in space and time lead to stable coexistence. Using kefir as a natural model microbial ecosystem, the authors apply metabolomics, transcriptomics and large-scale mapping of inter-species interactions to study the drivers of stable coexistence of species in space and time.
ABSTRACT Tumor relapse is responsible for most breast cancer related deaths 1,2 . The disease recurrence stems from treatment refractory cancer cells that persist as minimal residual disease (MRD) for years following initial therapy 3 . Yet, the molecular characteristics defining the malignancy of MRD remain elusive due to difficulties in observing these rare cells in patients or in model organisms. Here, we use a tractable organoid system and multi-omics analysis to show that the dormant MRD cells retain metabolic peculiarities reminiscent of the tumor state. While the MRD cells were distinct from both normal and tumor cells at a global transcriptomic level, their metabolomic and lipidomic profile markedly resembled that of the tumor state. The MRD cells particularly exhibited a de-regulated urea cycle and elevated glycolysis. We find the latter being crucial for their survival and could be selectively targeted using a small molecule inhibitor of glycolytic activity. We validated these metabolic peculiarities of the MRD cells in corresponding tissues obtained from the mouse model as well as in transcriptomic data from patients following neo-adjuvant therapy. Together, our results show that the treatment surviving MRD cells retain features of the tumor state over an extended period suggestive of an oncogenic memory. In accord, we found striking similarity in DNA methylation profiles between the tumor and the MRD cells. The distinction of MRD from normal breast cells comes as a surprise, considering their phenotypic similarity with regards to proliferation and polarized epithelial organization. The metabolic aberrances of the MRD cells offer a therapeutic opportunity towards tackling emergence of breast tumor recurrence in post-treatment care.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Article Figures and data Abstract eLife digest Introduction Results Discussion Materials and methods Appendix 1 Appendix 2 Data availability References Decision letter Author response Article and author information Metrics Abstract To capture the functional diversity of microbiota, one must identify metabolic functions and species of interest within hundreds or thousands of microorganisms. We present Metage2Metabo (M2M) a resource that meets the need for de novo functional screening of genome-scale metabolic networks (GSMNs) at the scale of a metagenome, and the identification of critical species with respect to metabolic cooperation. M2M comprises a flexible pipeline for the characterisation of individual metabolisms and collective metabolic complementarity. In addition, M2M identifies key species, that are meaningful members of the community for functions of interest. We demonstrate that M2M is applicable to collections of genomes as well as metagenome-assembled genomes, permits an efficient GSMN reconstruction with Pathway Tools, and assesses the cooperation potential between species. M2M identifies key organisms by reducing the complexity of a large-scale microbiota into minimal communities with equivalent properties, suitable for further analyses. eLife digest All the microbes that live in a specific environment, for example an organ, are collectively called the microbiota. In humans, the microbiota of the gut has been extensively studied by sequencing the DNA of the different microbes to identify them and determine the roles they play in health and disease. The DNA sequences of all the members of the microbiota is called the metagenome. The chemical reactions that the gut microbiota perform to produce energy and make the biomolecules they need to survive are collectively referred to as the metabolism of these microbes. Studying the metabolism of the gut microbiota can help researchers understand the roles of the different microbes. However, the large variety of species in the gut microbiota and gaps in the information about them render these studies difficult, despite technology improving quickly. To tackle this issue, Belcour, Frioux et al developed a new piece of software called Metage2Metabo (M2M) that simulates the metabolism of the gut microbiota and describes the metabolic relationships between the different microbes. Metage2Metabo analyses the roles of the metabolic genes of a large number of microbe species to establish how they complement each other metabolically. Then, it can calculate the minimum number of species needed to perform a metabolic role of interest within that microbiota, and which key species are associated with that role. To test the new software, Belcour, Frioux et al. used Metage2Metabo to analyse genomes from the human gut microbiota and from the cow rumen (one of the cow's stomachs). They showed that even if the metagenome was incomplete, the software was able to make stable predictions of key species involved in metabolic complementarity. Additionally, they also illustrated how the method can be used to study the gut microbiota of individuals. This work presents a new method for determining the metabolic relationships between species within a microbiota. The software is highly flexible and could be adapted to identify key members within different communities. In the context of the gut microbiota, the predictions of Metage2Metabo could shed lights on the interactions between the host and the microbes and contribute to a better understanding of microbe environments. Introduction Understanding the interactions between organisms within microbiomes is crucial for ecological (Tara Oceans coordinators et al., 2015) and health (Integrative HMP (iHMP) Research Network Consortium, 2014) applications. Improvements in metagenomics, and in particular the development of methods to assemble individual genomes from metagenomes, have given rise to unprecedented amounts of data which can be used to elucidate the functioning of microbiomes. Hundreds or thousands of genomes can now be reconstructed from various environments (Pasolli et al., 2019; Forster et al., 2019; Zou et al., 2019; Stewart et al., 2018; Almeida et al., 2020), either with the help of reference genomes or through metagenome-assembled genomes (MAGs), paving the way for numerous downstream analyses. Some major interactions between species occur at the metabolic level. This is the case for negative interactions such as exploitative competition (e.g. for nutrient resources), or for positive interactions such as cross-feeding or syntrophy (Coyte and Rakoff-Nahoum, 2019) that we will refer to with the generic term of cooperation. In order to unravel such interactions between species, it is necessary to go beyond functional annotation of individual genomes and connect metagenomic data to metabolic modelling. The main challenges impeding mathematical and computational analysis and simulation of metabolism in microbiomes are the scale of metagenomic datasets and the incompleteness of their data. Genome-scale metabolic networks (GSMNs) integrate all the expected metabolic reactions of an organism. Thiele and Palsson, 2010 defined a precise protocol for their reconstruction, associating the use of automatic methods and thorough curation based on expertise, literature, and mathematical analyses. There now exists a variety of GSMN reconstruction implementations: all-in-one platforms such as Pathway Tools (Karp et al., 2016), CarveMe (Machado et al., 2018) or KBase that provides narratives from metagenomic datasets analysis up to GSMN reconstruction with ModelSEED (Henry et al., 2010; Seaver et al., 2020). In addition, a variety of toolboxes (Aite et al., 2018; Wang et al., 2018; Schellenberger et al., 2011), or individual tools perform targeted refinements and analyses on GSMNs (Prigent et al., 2017; Thiele et al., 2014; Vitkin and Shlomi, 2012). Reconstructed GSMNs are a resource to analyse the metabolic complementarity between species, which can be seen as a representation of the putative cooperation within communities (Opatovsky et al., 2018). SMETANA (Zelezniak et al., 2015) estimates the cooperation potential and simulates flux exchanges within communities. MiSCoTo (Frioux et al., 2018) computes the metabolic potential of interacting species and performs community reduction. NetCooperate (Levy et al., 2015) predicts the metabolic complementarity between species. In addition, a variety of toolboxes have been proposed to study communities of organisms using GSMNs (Kumar et al., 2019; Sen and Orešič, 2019), most of them relying on constraint-based modelling (Chan et al., 2017; Zomorrodi and Maranas, 2012; Khandelwal et al., 2013). However, these tools can only be applied to communities with few members, as the computational cost scales exponentially with the number of included members (Kumar et al., 2019). Only recently has the computational bottleneck started to be addressed (Diener et al., 2020). In addition, current methods require GSMNs of high quality in order to produce accurate mathematical predictions and quantitative simulations. Reaching this level of quality entails manual modifications to the models using human expertise, which is not feasible at a large scale in metagenomics. Automatic reconstruction of GSMNs scales to metagenomic datasets, but it comes with the cost of possible missing reactions and inaccurate stoichiometry that impede the use of constraint-based modelling (Bernstein et al., 2019). Therefore, development of tools tailored to the analysis of large communities is needed. Here, we describe Metage2Metabo (M2M), a software system for the characterisation of metabolic complementarity starting from annotated individual genomes. M2M capitalises on the parallel reconstruction of GSMNs and a relevant metabolic modelling formalism to scale to large microbiotas. It comprises a pipeline for the individual and collective analysis of GSMNs and the identification of communities and key species ensuring the producibility of metabolic compounds of interest. M2M automates the systematic reconstruction of GSMNs using Pathway Tools or relies on GSMNs provided by the user. The software system uses the algorithm of network expansion (Ebenhöh et al., 2004) to capture the set of producible metabolites in a GSMN. This choice answers the needs for stoichiometry inaccuracy handling, and the robustness of the algorithm was demonstrated by the stability of the set of reachable metabolites despite missing reactions (Handorf et al., 2005; Kruse and Ebenhöh, 2008). Consequently, M2M scales metabolic modelling to metagenomics and large collections of (metagenome-assembled) genomes. We applied M2M on a collection of 1520 draft bacterial reference genomes from the gut microbiota (Zou et al., 2019) in order to illustrate the range of analyses the tool can produce. This demonstrates that M2M efficiently reconstructs metabolic networks for all genomes, identifies potential metabolites produced by cooperating bacteria, and suggests minimal communities and key species associated to their production. We then compared metabolic network reconstruction applied to the gut reference genomes to the results obtained with a collection of 913 cow rumen MAGs (Stewart et al., 2018). In addition, we tested the robustness of metabolic prediction with respect to genome incompleteness by degrading the rumen MAGs. The comparison of outputs from the pipeline indicates stability of the results with moderately degraded genomes, and the overall suitability of M2M to MAGs. Finally, we demonstrated the applicability of M2M in practice to metagenomic data of individuals. To that purpose, we reconstructed communities for 170 samples of healthy and diabetic individuals (Forslund et al., 2015; Diener et al., 2020). We show how M2M can help connect sequence analyses to metabolic screening in metagenomic datasets. Results M2M pipeline and key species M2M is a flexible software solution that performs automatic GSMN reconstruction and systematic screening of metabolic capabilities for up to thousands of species for which an annotated genome is available. The tool computes both the individual and collective metabolic capabilities to estimate the complementarity between the metabolisms of the species. Then based on a determined metabolic objective which can be ensuring the producibility of metabolites that need cooperation, that we call cooperation potential, M2M performs a community reduction step that aims at identifying a minimal community fulfilling the metabolic objective, as well as the set of associated key species. M2M's main pipeline (Figure 1a) consists in five main steps that can be performed sequentially or independently: (i) reconstruction of metabolic networks for all annotated genomes, (ii) computation of individual and (iii) collective metabolic capabilities, (iv) calculation of the cooperation potential, and (v) identification of minimal communities and key species for a targeted set of compounds. Figure 1 Download asset Open asset Overview of the Metage2Metabo (M2M) pipeline. (a) Main steps of the M2M pipeline and associated tools. The software's main pipeline (m2m workflow) takes as inputs a collection of annotated genomes that can be reference genomes or metagenomic-assembled genomes. The first step of M2M consists in reconstructing metabolic networks with Pathway Tools (step 0). This first step can be bypassed and genome-scale metabolic networks (GSMNs) can be directly loaded in M2M. The resulting metabolic networks are analysed to identify individual (step 1) and collective (step 2) metabolic capabilities. The added-value of cooperation is calculated (step 3) and used as a metabolic objective to compute a minimal community and key species (step 4). Optionally, one can customise the metabolic targets for community reduction. The pipeline without GSMN reconstruction can be called with m2m metacom, and each step can also be called independently (m2m iscope, m2m cscope, m2m addedvalue, m2m mincom). (b) Description of key species. Community reduction performed at step 4 can lead to multiple equivalent communities. M2M provides one minimal community and efficiently computes the full set of species that occur in all minimal communities, without the need for a full enumeration, thanks to projection modes. It is possible to distinguish the species occurring in every minimal community (essential symbionts), from those occurring in some (alternative symbionts). Altogether, these two groups form the key species. Sets of producible metabolites for individual or communities of species are computed using the network expansion algorithm (Ebenhöh et al., 2004) that is implemented in Answer Set Programming in dependencies of M2M. Network expansion enables the calculation of the scope of one or several metabolic networks in given nutritional conditions, described as seed compounds. The scope therefore represents the metabolic potential or reachable metabolites in these conditions (see Materials and methods). M2M calculates individual scopes for all metabolic networks, and the community scope comprising all reachable metabolites for the interacting species. Network expansion is also used in the community reduction optimisation implemented in MiSCoTo (Frioux et al., 2018), the dependency of M2M, as reduced communities are expected to produce the metabolites of interest. The inputs to the whole workflow are a set of annotated genomes, a list of nutrients representing a growth medium, and optionally a list of targeted compounds to be produced by selected communities that will bypass the default objective of ensuring the producibility of the cooperation potential. Users can use the annotation pipeline of their choice prior running M2M. The whole pipeline is called with the command m2m workflow but each step can also be run individually as described in Table 1. Table 1 List and description of Metage2Metabo (M2M) commands. CommandActionm2m workflowRuns the whole m2m workflowm2m metacomRuns the workflow with already-reconstructed metabolic networksm2m reconReconstructs metabolic networks using Pathway Toolsm2m iscopeComputes scopes for individual metabolic networksm2m cscopeComputes the community scopem2m addedvalueComputes the cooperation potentialm2m mincomSelects a minimal community and computes key speciesm2m seedsCreates a SBML file for nutrientsm2m testRuns m2m workflow on a sample datasetm2m_analysisRuns additional analyses on community selection A main characteristic of M2M is to provide at the end of the pipeline a set of key species associated to a metabolic function together with one minimal community predicted to satisfy this function. We define as key species organisms whose GSMNs are selected in at least one of the minimal communities predicted to fulfill the metabolic objective. Among key species, we distinguish those that occur in every minimal community, suggesting that they possess key functions associated to the objective, from those that occur only in some communities. We call the former essential symbionts, and the latter alternative symbionts. These terms were inspired by the terminology used in flux variability analysis (Orth et al., 2010) for the description of reactions in all optimal flux distributions. If interested, one can compute the enumeration of all minimal communities with m2m_analysis, which will provide the total number of minimal communities as well as the composition of each. Figure 1b illustrates these concepts with an initial community formed of eight species. There are four minimal communities satisfying the metabolic objective. Each includes three species, and in particular, the yellow one is systematically a member. Therefore, the yellow species is an essential symbiont whereas the four other species involved in minimal communities constitute the set of alternative symbiont. As key species represent the diversity associated to all minimal communities, it is likely that their number is greater than the size of a minimal community, as this is the case in Figure 1b. M2M connects metagenomics to metabolism with GSMN reconstruction, metabolic complementarity screening and community reduction In order to illustrate its applicability to real data, M2M was applied to a collection of 1520 bacterial high-quality draft reference genomes from the gut microbiota presented in Zou et al., 2019. The genomes were derived from cultured bacteria, isolated from faecal samples covering typical gut phyla (Costea et al., 2018): 796 Firmicutes, 447 Bacteroidetes, 235 Actinobacteria, 36 Proteobacteria, and 6 Fusobacteria. The dereplicated genomes represent 338 species. The genomes were already annotated and could therefore directly enter M2M pipeline. The full workflow (from GSMN reconstruction to key species computation) took 155 min on a cluster with 72 CPUs and 144 Gb of memory. We illustrate in the next paragraphs the scalability of M2M and the range of analyses it proposes by applying the pipeline to this collection of genomes. GSMN reconstruction GSMNs were automatically reconstructed for the 1520 isolate-based genomes using their published annotation. A total of 3932 unique reactions and 4001 metabolites were included in the reconstructed GSMNs (Table 2). The reconstructed gut metabolic networks contained on average 1144 (±255) reactions and 1366 (±262) metabolites per genome. Of the reactions, 74.6% were associated to genes, the remaining being spontaneous reactions or reactions added by the PathoLogic algorithm (they can be removed in M2M using the –noorphan option). Table 2 Results of the genome-scale metabolic network (GSMN) reconstruction step and metabolic potential analysis for the three datasets presented in the article (Avg = Average, '±' precedes standard deviation). Gut datasetRumen datasetDiabetes datasetInitial dataDraft reference genomesMAGsMAGsNumber of genomes1520913778GSMN reconstructionAll reactions393244185554All metabolites400144665386Avg reactions per GSMN1144 (±255)1155 (±199)1640 (±368)Avg metabolites per GSMN1366 (±262)1422 (±212)1925 (±361)Avg genes per mn596 (±150)543 (±107)1658 (±469)% reactions associated to genes74.6 (±2.17)73.8 (±2.61)79.57 (±1.60)Avg pathways per mn163 (±49)146 (±32)220 (±58)Metabolic potentialNumber of seeds9326175Avg scope per mn286 (±70)101 (±44)508 (±83)Union of individual scopes8283681326 The metabolic potential, or scope, was computed for each individual GSMN (Table 2). Nutrients in this experiment were components of a classical diet (see Materials and methods). The union of all individual scopes is of size 828 (21% of all compounds included in the GSMNs), indicating a small part of the metabolism reachable in the chosen nutritional environment (Supplementary file 1 - Table 1). Appendix 1—figure 1h,i displays the distributions of the scopes. Across all GSMNs, individual scopes are overall stable in size. The core set of producible metabolites is small and a variety of metabolites are only reachable by a small number of organisms (Appendix 1—figure 1i). The overall small size of metabolic potentials can be explained by the restricted amount of seeds used for computation. Among metabolites that are reachable by all or almost all metabolic networks, the primary metabolism is highly represented, as expected, with metabolites derived from common sugars (glucose, fructose), pyruvate, 2-oxoglutaric acid, amino acids… On the other hand but not surprisingly, metabolites that predicted to be reached by a limited number of individual producers include compounds from secondary metabolism: (fatty) acids (e.g. oxalate, maleate, allantoate, hydroxybutanoate, methylthiopropionate), derivatives of amino acids, amines (spermidine derivatives). Cooperation potential Metabolic cooperation enables the activation of more reactions in GSMNs than what can be expected when networks are considered in isolation. By taking into account the complementarity between GSMNs in each dataset, it is possible to capture the putative benefit of metabolic cooperation on the diversity of producible metabolites. Running m2m cscope predicted 156 new metabolites as producible by the gut collection of GSMNs if cooperation is allowed. We analysed the composition of the 156 newly producible metabolites for the gut dataset using the ontology provided for metabolic compounds in the MetaCyc database. 80.1% of them could be grouped into six categories: amino acids and derivatives (5 metabolites), aromatic compounds (11), carboxy acids (14), coenzyme A (CoA) derivatives (10), lipids (28), sugar derivatives (58). The groups were used in the subsequent analyses. The remaining 30 compounds were highly heterogeneous, we therefore restrained our subsequent analyses to subcategories of biochemically homogeneous targets. We paid a particular attention to the predicted producibility of short-chain fatty acids (SCFAs) among genomes of the gut collection. We analysed formate, acetate, propionate, and butyrate in individual and collective metabolic potentials (Supplementary file 1 - Table 25). A total of 543 metabolic networks are predicted to be able to produce all four molecules in a cooperative context, 74% of them belonging the Firmicutes, as expected. Surprisingly, predicted individual producers of the four SCFAs (n = 128) are mostly Bacteroidetes (70%) suggesting the dependency of Firmicutes to interactions in order to permit the producibility of SCFAs in this experimental setting. The same observations are made when focusing on butyrate alone, that has the particularity of belonging to the seeds. As Bacteroidetes are not the main butyrate producers in the gut, the predictions of such producibility is likely an artefact relying on alternative pathways, and further emphasises the fact that owning the genetic material for a function does not entail its expression. Key species associated to groups of metabolites M2M proposes by default one community composition for an objective defined by enabling the producibility of metabolic end-products. Given the functional redundancy of gut bacteria (Moya and Ferrer, 2016), there could be thousands of bacterial composition combinations, and it is computationally costly to enumerate them. To circumvent this restriction, M2M identifies key species without the need for all possible combinations of species to be enumerated, consequentially reducing computational time. Key species include all species occurring in at least one minimal community for the production of chosen end-products. They can be distinguished in two categories: essential symbionts occurring in all minimal communities, and alternative symbionts occurring in some minimal communities. To explore the spectrum of possible key species, we ran M2M community reduction step (m2m mincom command) with the above six metabolic target groups. This allowed us to compute predicted key species for each of them (Table 3). The contents of key species for each of the six groups of targets as well as for the complete set of targets is displayed in Supplementary file 1 (Tables 6 to 12). To our surprise, the size of the minimal community is relatively small for each group of metabolites (between 4 and 11), compared to the initial community of 1520 GSMNs. The number of identified key species varies between 59 and 227, which might be closer to the total taxonomic diversity found in the human gut microbiome. This strong reduction compared to the initial number of 1520 GSMNs used for the analysis illustrates the existence of groups of bacteria with specific metabolic capabilities. In particular, essential symbionts are likely of high importance for the functions as they are found in each solution. More generally, compositions vary across the target categories: a high proportion of key species for the production of lipids targets are Bacteroidetes, whereas Firmicutes were more often key species for aminoacids and derivatives production. The propensity of Bacteroidetes to metabolise lipids has been proposed previously, it has for example been observed in the Bacteroides enterotype for functions related to lipolysis (Vieira-Silva et al., 2016). Table 3 Community reduction analysis of the target categories in the gut. All minimal communities were enumerated, starting from the set of 1520 genome-scale metabolic networks (GSMNs). KS: key species, ES: essential symbionts, AS: alternative symbionts, Firm.: Firmicutes, Bact.: Bacteroidetes, Acti.: Actinobacteria, Prot.: Proteobacteria, Fuso.: Fusobacteria. Firm.Bact.Acti.Prot.Fuso.TotalAminoacids and derivatives (5 targets)4 bact. per community120,329 communitiesKS142520276227ES000000AS142520276227Aromatic compounds (11 targets)5 bact. per community950 communitiesKS520020072ES200103AS500019069Carboxyacids (14 targets)9 bact. per community48,412 communitiesKS1613028259ES200204AS1413026255CoA derivatives (10 targets)5 bact. per community95,256 communitiesKS106050171174ES000011AS106050170173Lipids (28 targets)7 bact. per community58,520 communitiesKS314022200185ES300104AS014022190181Sugar derivatives (58 targets)11 bact. per community7,860,528 communitiesKS113078230142ES500005AS63078230137 Analysis of minimal communities identifies groups of organisms with equivalent roles To go further, we enumerated all minimal communities for each individual group of targets using m2m_analysis. The number of optimal solutions is large, reaching more than 7 million equivalent minimal communities for the sugar-derived metabolites (Table 3). Our analysis of key species indicates that the large number of optimal communities is due to combinatorial choices among a rather small number of bacteria (Table 3). In order to visualise the association of GSMNs in minimal communities, we created for each target set a graph whose nodes are the key species (Supplementary file 1 - Tables 6 to 12), and whose edges represent the association between two species if they co-occur in at least one of the enumerated communities. Graphs were very dense: 185 nodes, 6888 edges for the lipids, 142 nodes, and 6602 edges for the sugar derivatives. This density is expected given the large number of optimal communities and the comparatively small number of key species. The graphs were compressed into power graphs to capture the combinatorics of association within minimal communities. Power graphs enable a lossless compression of re-occurring motifs within a graph: cliques, bicliques, and star patterns (Royer et al., 2008). The increased readability of power graphs permits pinpointing metabolic equivalency between members of the key species with respect to the target compound families. Figure 2 presents the compressed graphs for each set of targets. Graph nodes are the key species, coloured by their phylum. Nodes are included into power nodes that are connected by power edges, illustrating the redundant metabolic function(s) that species provide to the community when considering specific end-products. GSMNs belonging to a power node play the same role in the construction of the minimal communities. In this visualisation, essential symbionts are easily identifiable, either into power nodes with loops (Figure 2a,e) or as individual nodes connected to power nodes (Figure 2a,c,d,f). Figure 2 with 6 supplements see all Download asset Open asset Power graph analysis of predicted microbial associations within communities for the human gut dataset. Each category of metabolites predicted as newly producible in the gut was defined as a target set for community selection among the 1520 genome-scale metabolic networks (GSMNs) from the gut microbiota reference genomes dataset. For each metabolic group, key species and the full enumeration of all minimal communities were computed. Association graphs were built to associate members that are found together in at least one minimal community among the enumeration. These graphs were compressed as power graphs to identify patterns of associations and groups of equivalence within key species. Power graphs a., b., c., d., e., f. were generated for the sets of lipids, aminoacids and derivatives, carboxy-acids, sugar derivatives, aromatic compounds, and coenzyme A derivative compounds, respectively. Node colour describes the phylum associated to the GSMN. Figure (a) has an additional description to ease readability. Edges symbolise conjunctions ('AND') and the co-occurrences of nodes in regular power nodes (as in power node 1, 2, 4) symbolise disjunctions ('OR') related to alternative symbionts. Power nodes with a loop (e.g. power node 5) indicate conjunctions. Therefore, each enumerated minimal community for lipid production is composed of the two Firmicutes and the Proteobacteria from power node 5, the Firmicutes node 3 (the four of them being the essential symbionts), and one Proteobacteria from power node 4, one Actinobacteria from power node 2 and 1 Bacteroidetes from power node 1. Members from an inner power node are interchangeable with respect to the metabolic objective. A version of the figures with species identification is available in Figure 2—figure supplement 1, Figure 2—figure supplement 2, Figure 2—figure supplement 3, Figure 2—figure supplement 4, Figure 2—figure supplement 5, Figure 2—figure supplement 6 (see Supplementary file 1 - Table 4 for a mapping between identifiers and taxonomy). Power graphs can be generated with m2m_analysis. The figures display one visual representation for each power graph although such representations are not unique. The number of power edges is minimal, which leads to nesting of (power) nodes. We observe that power nodes often contain GSMNs from the same phylum, indicating that phylogenetic groups encode redundant functions. Figure 3a has additional comm
Predicting phenotype from genotype is the holy grail of quantitative systems biology. Kinetic models of metabolism are among the most mechanistically detailed tools for phenotype prediction. Kinetic models describe changes in metabolite concentrations as a function of enzyme concentration, reaction rates, and concentrations of metabolic effectors uniquely enabling integration of multiple omics data types in a unifying mechanistic framework. While development of such models for Escherichia coli has been going on for almost twenty years, multiple separate models have been established and systematic independent benchmarking studies have not been performed on the full set of models available. In this study we compared systematically all recently published kinetic models of the central carbon metabolism of Escherichia coli . We assess the ease of use of the models, their ability to include omics data as input, and the accuracy of prediction of central carbon metabolic flux phenotypes. We conclude that there is no clear winner among the models when considering the resulting tradeoffs in performance and applicability to various scenarios. This study can help to guide further development of kinetic models, and to demonstrate how to apply such models in real-world setting, ultimately enabling the design of efficient cell factories.Author summary Kinetic modeling is a promising method to predict cell metabolism. Such models provide mechanistic description of how concentrations of metabolites change in the cell as a function of time, cellular environment and the genotype of the cell. In the past years there have been several kinetic models published for various organisms. We want to assess how reliably models of Escherichia coli metabolism could predict cellular metabolic state upon genetic or environmental perturbations. We test selected models in the ways that represent common metabolic engineering practices including deletion and overexpression of genes. Our results suggest that all published models have tradeoffs and the model to use should be chosen depending on the specific application. We show in which cases users could expect the best performance from published models. Our benchmarking study should help users to make a better informed choice and also provides systematic training and testing dataset for model developers.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.