Genome scale metabolic models (GEMs) and other constraint-based models (CBMs) play a pivotal role in understanding biological phenotypes and advancing research in areas like metabolic engineering, human disease modelling, drug discovery, and personalized medicine. Despite their growing application, a significant challenge remains in ensuring the reproducibility of GEMs, primarily due to inconsistent reporting and inadequate model documentation of model results. Addressing this gap, we introduce FROG analysis, a community driven initiative aimed at standardizing reproducibility assessments of CBMs and GEMs. The FROG framework encompasses four key analyses including Flux variability, Reaction deletion, Objective function, and Gene deletion to produce standardized, numerically reproducible FROG reports. These reports serve as reference datasets, enabling model evaluators, curators, and independent researchers to verify the reproducibility of GEMs systematically. BioModels, a leading repository of systems biology models, has integrated FROG analysis into its curation workflow, enhancing the reproducibility and reusability of submitted GEMs. In our study evaluating 65 GEM submissions from the community, approximately 40% reproduced without intervention, 28% requiring minor adjustments, and 32% needing input from authors. The standardization introduced by FROG analysis facilitated the detection and resolution of issues, ultimately leading to the successful reproduction of all models. By establishing a standardized and comprehensive approach to evaluating GEM reproducibility, FROG analysis significantly contributes to making CBMs and GEMs more transparent, reusable, and reliable for the broader scientific community. ### Competing Interest Statement The authors have declared no competing interest.
1 Unseen Bio ApS, Copenhagen, Denmark 2 Microbiome Sciences Group, Department of Archaeogenetics, Max Planck Institute for Evolutionary Anthropology, Leipzig, Germany 3 Associated Research Group of Archaeogenetics, Leibniz Institute for Natural Product Research and Infection Biology Hans Knöll Institute, Jena, Germany 4 Department of Microbiology, Tumor and Cell Biology, Karolinska Institute, Solna, Sweden 5 Department of Paleobiotechnology, Leibniz Institute for Natural Product Research and Infection Biology Hans Knöll Institute, Jena, Germany ¶ Corresponding author DOI: 10.21105/joss.05627
1 Abstract Metagenomic classification tackles the problem of characterising the taxonomic source of all DNA sequencing reads in a sample. A common approach to address the differences and biases between the many different taxonomic classification tools is to run metagenomic data through multiple classification tools and databases. This, however, is a very time-consuming task when performed manually - particularly when combined with the appropriate preprocessing of sequencing reads before the classification. Here we present nf-core/taxprofiler, a highly parallelised read-processing and taxonomic classification pipeline. It is designed for the automated and simultaneous classification and/or profiling of both short- and long-read metagenomic sequencing libraries against a 11 taxonomic classifiers and profilers as well as databases within a single pipeline run. Implemented in Nextflow and as part of the nf-core initiative, the pipeline benefits from high levels of scalability and portability, accommodating from small to extremely large projects on a wide range of computing infrastructure. It has been developed following best-practise software development practises and community support to ensure longevity and adaptability of the pipeline, to help keep it up to date with the field of metagenomics.
eQuilibrator (equilibrator.weizmann.ac.il) is a database of biochemical equilibrium constants and Gibbs free energies, originally designed as a web-based interface. While the website now counts around 1,000 distinct monthly users, its design could not accommodate larger compound databases and it lacked a scalable Application Programming Interface (API) for integration into other tools developed by the systems biology community. Here, we report on the recent updates to the database as well as the addition of a new Python-based interface to eQuilibrator that adds many new features such as a 100-fold larger compound database, the ability to add novel compounds, improvements in speed and memory use, and correction for Mg2+ ion concentrations. Moreover, the new interface can compute the covariance matrix of the uncertainty between estimates, for which we show the advantages and describe the application in metabolic modelling. We foresee that these improvements will make thermodynamic modelling more accessible and facilitate the integration of eQuilibrator into other software platforms.
Computational models have great potential to accelerate bioscience, bioengineering, and medicine. However, it remains challenging to reproduce and reuse simulations, in part, because the numerous formats and methods for simulating various subsystems and scales remain siloed by different software tools. For example, each tool must be executed through a distinct interface. To help investigators find and use simulation tools, we developed BioSimulators (https://biosimulators.org), a central registry of the capabilities of simulation tools and consistent Python, command-line and containerized interfaces to each version of each tool. The foundation of BioSimulators is standards, such as CellML, SBML, SED-ML and the COMBINE archive format, and validation tools for simulation projects and simulation tools that ensure these standards are used consistently. To help modelers find tools for particular projects, we have also used the registry to develop recommendation services. We anticipate that BioSimulators will help modelers exchange, reproduce, and combine simulations.
Background The long-term management of irritable bowel syndrome (IBS) poses many challenges. In short-term studies, eHealth interventions have been demonstrated to be safe and practical for at-home monitoring of the effects of probiotic treatments and a diet low in fermentable oligosaccharides, disaccharides, monosaccharides, and polyols (FODMAPs). IBS has been linked to alterations in the microbiota. Objective The aim of this study was to determine whether a web-based low-FODMAP diet (LFD) intervention and probiotic treatment were equally good at reducing IBS symptoms, and whether the response to treatments could be explained by patients’ microbiota. Methods Adult IBS patients were enrolled in an open-label, randomized crossover trial (for nonresponders) with 1 year of follow-up using the web application IBS Constant Care (IBS CC). Patients were recruited from the outpatient clinic at the Department of Gastroenterology, North Zealand University Hospital, Denmark. Patients received either VSL#3 for 4 weeks (2 × 450 billion colony-forming units per day) or were placed on an LFD for 4 weeks. Patients responding to the LFD were reintroduced to foods high in FODMAPs, and probiotic responders received treatments whenever they experienced a flare-up of symptoms. Treatment response and symptom flare-ups were defined as a reduction or increase, respectively, of at least 50 points on the IBS Severity Scoring System (IBS-SSS). Web-based ward rounds were performed daily by the study investigator. Fecal microbiota were analyzed by shotgun metagenomic sequencing (at least 10 million 2 × 100 bp paired-end sequencing reads per sample). Results A total of 34 IBS patients without comorbidities and 6 healthy controls were enrolled in the study. Taken from participating subjects, 180 fecal samples were analyzed for their microbiota composition. Out of 21 IBS patients, 12 (57%) responded to the LFD and 8 (38%) completed the reintroduction of FODMAPs. Out of 21 patients, 13 (62%) responded to their first treatment of VSL#3 and 7 (33%) responded to multiple VSL#3 treatments. A median of 3 (IQR 2.25-3.75) probiotic treatments were needed for sustained symptom control. LFD responders were reintroduced to a median of 14.50 (IQR 7.25-21.75) high-FODMAP items. No significant difference in the median reduction of IBS-SSS for LFD versus probiotic responders was observed, where for LFD it was –126.50 (IQR –196.75 to –76.75) and for VSL#3 it was –130.00 (IQR –211.00 to –70.50; P>.99). Responses to either of the two treatments were not able to be predicted using patients’ microbiota. Conclusions The web-based LFD intervention and probiotic treatment were equally efficacious in managing IBS symptoms. The response to treatments could not be explained by the composition of the microbiota. The IBS CC web application was shown to be practical, safe, and useful for clinical decision making in the long-term management of IBS. Although this study was underpowered, findings from this study warrant further research in a larger sample of patients with IBS to confirm these long-term outcomes. Trial Registration ClinicalTrials.gov NCT03586622; https://clinicaltrials.gov/ct2/show/NCT03586622
For over 10 years, ModelSEED has been a primary resource for the construction of draft genome-scale metabolic models based on annotated microbial or plant genomes. Now being released, the biochemistry database serves as the foundation of biochemical data underlying ModelSEED and KBase. The biochemistry database embodies several properties that, taken together, distinguish it from other published biochemistry resources by: (i) including compartmentalization, transport reactions, charged molecules and proton balancing on reactions; (ii) being extensible by the user community, with all data stored in GitHub; and (iii) design as a biochemical ‘Rosetta Stone’ to facilitate comparison and integration of annotations from many different tools and databases. The database was constructed by combining chemical data from many resources, applying standard transformations, identifying redundancies and computing thermodynamic properties. The ModelSEED biochemistry is continually tested using flux balance analysis to ensure the biochemical network is modeling-ready and capable of simulating diverse phenotypes. Ontologies can be designed to aid in comparing and reconciling metabolic reconstructions that differ in how they represent various metabolic pathways. ModelSEED now includes 33,978 compounds and 36,645 reactions, available as a set of extensible files on GitHub, and available to search at https://modelseed.org/biochem and KBase.
Summary We achieve a significant improvement in thermodynamic-based flux analysis (TFA) by introducing multivariate treatment of thermodynamic variables and leveraging component contribution, the state-of-the-art implementation of the group contribution methodology. Overall, the method greatly reduces the uncertainty of thermodynamic variables. Results We present multiTFA, a Python implementation of our framework. We evaluated our application using the core E. coli model and achieved a median reduction of 6.8 kJ/mol in reaction Gibbs free energy ranges, while three out of 12 reactions in glycolysis changed from reversible to irreversible. Availability and implementation Our framework along with documentation is available on https://github.com/biosustain/multitfa .
ABSTRACT Standardization of data and models facilitates effective communication, especially in computational systems biology. However, both the development and consistent use of standards and resources remains challenging. As a result, the amount, quality, and format of the information contained within systems biology models are not consistent and therefore present challenges for widespread use and communication. Here, we focused on these standards, resources, and challenges in the field of metabolic modeling by conducting a community-wide survey. We used this feedback to (1) outline the major challenges that our field faces and to propose solutions and (2) identify a set of features that defines what a “gold standard” metabolic network reconstruction looks like concerning content, annotation, and simulation capabilities. We anticipate that this community-driven outline will help the long-term development of community-inspired resources as well as produce high-quality, accessible models. More broadly, we hope that these efforts can serve as blueprints for other computational modeling communities to ensure continued development of both practical, usable standards and reproducible, knowledge-rich models.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Computational systems biology methods enable rational design of cell factories on a genome-scale and thus accelerate the engineering of cells for the production of valuable chemicals and proteins. Unfortunately, the majority of these methods' implementations are either not published, rely on proprietary software, or do not provide documented interfaces, which has precluded their mainstream adoption in the field. In this work we present cameo, a platform-independent software that enables in silico design of cell factories and targets both experienced modelers as well as users new to the field. It is written in Python and implements state-of-the-art methods for enumerating and prioritizing knockout, knock-in, overexpression, and down-regulation strategies and combinations thereof. Cameo is an open source software project and is freely available under the Apache License 2.0. A dedicated Web site including documentation, examples, and installation instructions can be found at http://cameo.bio . Users can also give cameo a try at http://try.cameo.bio .
Several studies have shown that neither the formal representation nor the functional requirements of genome-scale metabolic models (GEMs) are precisely defined. Without a consistent standard, comparability, reproducibility, and interoperability of models across groups and software tools cannot be guaranteed. Here, we present memote (https://github.com/opencobra/memote) an open-source software containing a community-maintained, standardized set of metabolic model tests. The tests cover a range of aspects from annotations to conceptual integrity and can be extended to include experimental datasets for automatic model validation. In addition to testing a model once, memote can be configured to do so automatically, i.e., while building a GEM. A comprehensive report displays the model’s performance parameters, which supports informed model development and facilitates error detection. Memote provides a measure for model quality that is consistent across reconstruction platforms and analysis software and simplifies collaboration within the community by establishing workflows for publicly hosted and version controlled models.
Motivation: Extensive drug treatment gene expression data have been generated in order to identify biomarkers that are predictive for toxicity or to classify compounds. However, such patterns are often highly variable across compounds and lack robustness. We and others have previously shown that supervised expression patterns based on pathway concepts rather than unsupervised patterns are more robust and can be used to assess toxicity for entire classes of drugs more reliably. Results: We have developed a database, ToxDB, for the analysis of the functional consequences of drug treatment at the pathway level. We have collected 2694 pathway concepts and computed numerical response scores of these pathways for 437 drugs and chemicals and 7464 different experimental conditions. ToxDB provides functionalities for exploring these pathway responses by offering tools for visualization and differential analysis allowing for comparisons of different treatment parameters and for linking this data with toxicity annotation and chemical information. Database URL: http://toxdb.molgen.mpg.de