The U.S. Department of Energy’s Systems Biology Knowledgebase (KBase; www.kbase.us) is an open, collaborative platform that integrates data, models, and analysis tools to accelerate discovery in microbiology, plant biology, and environmental systems. Recently, KBase expanded as a comprehensive, multi-omics ecosystem. KBase enables representation of scientific samples, long-read sequence analysis, protein structure integration, and scalable modeling of microbial communities across diverse environments. KBase also generates digital notebooks as citable, executable research objects that link data, methods, and interpretation. KBase also supports a global education community focused on training the next generation of scientists to use high-performance computational tools. Together, these advances position KBase as a central hub for open, reproducible systems biology. In turn, this enables us to integrate many of the emerging advances in data federation, semantic interoperability, and agent-assisted analysis, paving the way for KBase to support the next generation of AI-driven discovery tools.
We introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command-line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelization on high-performance cluster systems, which enables high-throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application.
High-performance computing environments face increasing challenges from diverse scientific workflows, imposing conflicting demands for stability, customization, and reproducibility that traditional monolithic software stacks cannot accommodate. We present a comprehensive approach to seamless end-to-end containerized HPC environments which decomposes the technical challenge into five manageable areas: specification and construction of environments, session provisioning, scheduler integration, system integration, and security. We develop and evaluate prototypes across these five technical areas, demonstrating practical feasibility through Spack-based environment construction with CI/CD pipelines, transparent session access via PAM and Kubernetes, and flexible job execution using Slurm's native container support. Through this work, we demonstrate that comprehensive containerization of HPC environments can be achieved using open standards, providing enhanced reproducibility and flexibility without sacrificing user experience.
Leveraging the use of multiplex multi-omic networks, key insights into genetic and epigenetic mechanisms supporting biofuel production have been uncovered. Here, we introduce RWRtoolkit, a multiplex generation, exploration, and statistical package built for R and command line users. RWRtoolkit enables the efficient exploration of large and highly complex biological networks generated from custom experimental data and/or from publicly available datasets, and is species agnostic. A range of functions can be used to find topological distances between biological entities, determine relationships within sets of interest, search for topological context around sets of interest, and statistically evaluate the strength of relationships within and between sets. The command-line interface is designed for parallelisation on high performance cluster systems, which enables high throughput analysis such as permutation testing. Several tools in the package have also been made available for use in reproducible workflows via the KBase web application. ### Competing Interest Statement The authors have declared no competing interest.
Accessible and easy-to-use standardized bioinformatics workflows are necessary to advance microbiome research from observational studies to large-scale, data-driven approaches. Standardized multi-omics data enables comparative studies, data reuse, and applications of machine learning to model biological processes. To advance broad accessibility of standardized multi-omics bioinformatics workflows, the National Microbiome Data Collaborative (NMDC) has developed the Empowering the Development of Genomics Expertise (NMDC EDGE) resource, a user-friendly, open-source web application (https://nmdc-edge.org). Here, we describe the design and main functionality of the NMDC EDGE resource for processing metagenome, metatranscriptome, natural organic matter, and metaproteome data. The architecture relies on three main layers (web application, orchestration, and execution) to ensure flexibility and expansion to future workflows. The orchestration and execution layers leverage best practices in software containers and accommodate high-performance computing and cloud computing services. Further, we have adopted a robust user research process to collect feedback for continuous improvement of the resource. NMDC EDGE provides an accessible interface for researchers to process multi-omics microbiome data using production-quality workflows to facilitate improved data standardization and interoperability.
Bridging molecular information to ecosystem-level processes would provide the capacity to understand system vulnerability and, potentially, a means for assessing ecosystem health. Here, we present an integrated dataset containing environmental and metagenomic information from plant-associated microbial communities, plant transcriptomics, plant and soil metabolomics, and soil chemistry and activity characterization measurements derived from the model tree species Populus trichocarpa. Soil, rhizosphere, root endosphere, and leaf samples were collected from 27 different P. trichocarpa genotypes grown in two different environments leading to an integrated dataset of 318 metagenomes, 98 plant transcriptomes, and 314 metabolomic profiles that are supported by diverse soil measurements. This expansive dataset will provide insights into causal linkages that relate genomic features and molecular level events to system-level properties and their environmental influences.
Microbial communities have evolved to colonize all ecosystems of the planet, from the deep sea to the human gut. Microbes survive by sensing, responding, and adapting to immediate environmental cues. This process is driven by signal transduction proteins such as histidine kinases, which use their sensing domains to bind or otherwise detect environmental cues and "transduce" signals to adjust internal processes. We hypothesized that an ecosystem's unique stimuli leave a sensor "fingerprint," able to identify and shed insight on ecosystem conditions. To test this, we collected 20,712 publicly available metagenomes from Host-associated, Environmental, and Engineered ecosystems across the globe. We extracted and clustered the collection's nearly 18M unique sensory domains into 113,712 similar groupings with MMseqs2. We built gradient-boosted decision tree machine learning models and found we could classify the ecosystem type (accuracy: 87%) and predict the levels of different physical parameters (R2 score: 83%) using the sensor cluster abundance as features. Feature importance enables identification of the most predictive sensors to differentiate between ecosystems which can lead to mechanistic interpretations if the sensor domains are well annotated. To demonstrate this, a machine learning model was trained to predict patient's disease state and used to identify domains related to oxygen sensing present in a healthy gut but missing in patients with abnormal conditions. Moreover, since 98.7% of identified sensor domains are uncharacterized, importance ranking can be used to prioritize sensors to determine what ecosystem function they may be sensing. Furthermore, these new predictive sensors can function as targets for novel sensor engineering with applications in biotechnology, ecosystem maintenance, and medicine.IMPORTANCEMicrobes infect, colonize, and proliferate due to their ability to sense and respond quickly to their surroundings. In this research, we extract the sensory proteins from a diverse range of environmental, engineered, and host-associated metagenomes. We trained machine learning classifiers using sensors as features such that it is possible to predict the ecosystem for a metagenome from its sensor profile. We use the optimized model's feature importance to identify the most impactful and predictive sensors in different environments. We next use the sensor profile from human gut metagenomes to classify their disease states and explore which sensors can explain differences between diseases. The sensors most predictive of environmental labels here, most of which correspond to uncharacterized proteins, are a useful starting point for the discovery of important environment signals and the development of possible diagnostic interventions.
The National Microbiome Data Collaborative (NMDC) Data Portal (https://data.microbiomedata.org) supports microbiome multi-omics data exploration and access through an integrated, distributed data framework aligned with the FAIR (Findable, Accessible, Interoperable and Reusable) data principles (1). The NMDC Data Portal currently hosts 10.2 terabytes of multi-omics microbiome data, spanning five data types (metagenomes, metatranscriptomes, metaproteomes, metabolomes, and natural organic matter characterizations), generated at two Department of Energy User Facilities, the Joint Genome Institute (JGI) at Lawrence Berkeley National Laboratory (LBNL) and the Environmental Molecular Systems Laboratory (EMSL) at Pacific Northwest National Laboratory (PNNL). A flexible data schema (https://github.com/microbiomedata/nmdc-schema) leveraging community-driven standards underpins how data is managed and integrated. Annotated multi-omic data products are produced by the NMDC workflows and linked through common biosamples to enable search capabilities based on environmental context, instrumentation, and functional attributes. As a pilot system, the NMDC Data Portal offers download capabilities and several search components, including interactive geographic visualization of samples; environmental classification distribution visualized through an interactive Sankey diagram; time-series slider to select longitudinal samples of interest; and an upset plot displaying the number of multi-omics data generated from the same biosample within a study.
The nascent field of microbiome science is transitioning from a descriptive approach of cataloging taxa and functions present in an environment to applying multi-omics methods to investigate microbiome dynamics and function. A large number of new tools and algorithms have been designed and used for very specific purposes on samples collected by individual investigators or groups. While these developments have been quite instructive, the ability to compare microbiome data generated by many groups of researchers is impeded by the lack of standardized application of bioinformatics methods. Additionally, there are few examples of broad bioinformatics workflows that can process metagenome, metatranscriptome, metaproteome and metabolomic data at scale, and no central hub that allows processing, or provides varied omics data that are findable, accessible, interoperable and reusable (FAIR). Here, we review some of the challenges that exist in analyzing omics data within the microbiome research sphere, and provide context on how the National Microbiome Data Collaborative has adopted a standardized and open access approach to address such challenges.
Containers have provided a popular new paradigm for managing software and services. However, in HPC the use of containers has historically been more difficult due to multi-tenancy, security, and performance requirements; consequently several custom HPC container runtimes have emerged from the community. The resulting fractured ecosystem presents challenges both for HPC container framework maintainers and for users. In this paper, we describe work at NERSC to adapt Podman, a popular OCI-compliant container framework developed by Red Hat, Inc., for use in HPC. Podman has several key features which make it appealing for use in an HPC environment: its rootless container mode addresses many security concerns, it has a standardized command interface which will be familiar to users of established popular container runtimes, it is daemonless, and it is open-source and community supported. Additional innovations at NERSC have enabled Podman to achieve the good scaling behavior required by HPC applications.
Thermodynamics plays a crucial role in regulating the metabolic processes in all living organisms. Accurate determination of biochemical and biophysical properties is important to understand, analyze, and synthetically design such metabolic processes for engineered systems. In this work, we extensively performed first-principles quantum mechanical calculations to assess its accuracy in estimating free energy of biochemical reactions and developed automated quantum-chemistry (QC) pipeline (https://appdev.kbase.us/narrative/45710) for the prediction of thermodynamics parameters of biochemical reactions. We benchmark the QC methods based on density functional theory (DFT) against different basis sets, solvation models, pH, and exchange-correlation functionals using the known thermodynamic properties from the NIST database. Our results show that QC calculations when combined with simple calibration yield a mean absolute error in the range of 1.60–2.27 kcal/mol for different exchange-correlation functionals, which is comparable to the error in the experimental measurements. This accuracy over a diverse set of metabolic reactions is unprecedented and near the benchmark chemical accuracy of 1 kcal/mol that is usually desired from DFT calculations.
HPC centers face increasing demand for software flexibility, and there is growing consensus that Linux containers are a promising solution. However, existing container build solutions require root privileges and cannot be used directly on HPC resources. This limitation is compounded as supercomputer diversity expands and HPC architectures become more dissimilar from commodity computing resources. Our analysis suggests this problem can best be solved with low-privilege containers. We detail relevant Linux kernel features, propose a new taxonomy of container privilege, and compare two open-source implementations: mostly-unprivileged rootless Podman and fully-unprivileged Charliecloud. We demonstrate that low-privilege container build on HPC resources works now and will continue to improve, giving normal users a better workflow to securely and correctly build containers. Minimizing privilege in this way can improve HPC user and developer productivity as well as reduce support workload for exascale applications.
For over 10 years, ModelSEED has been a primary resource for the construction of draft genome-scale metabolic models based on annotated microbial or plant genomes. Now being released, the biochemistry database serves as the foundation of biochemical data underlying ModelSEED and KBase. The biochemistry database embodies several properties that, taken together, distinguish it from other published biochemistry resources by: (i) including compartmentalization, transport reactions, charged molecules and proton balancing on reactions; (ii) being extensible by the user community, with all data stored in GitHub; and (iii) design as a biochemical ‘Rosetta Stone’ to facilitate comparison and integration of annotations from many different tools and databases. The database was constructed by combining chemical data from many resources, applying standard transformations, identifying redundancies and computing thermodynamic properties. The ModelSEED biochemistry is continually tested using flux balance analysis to ensure the biochemical network is modeling-ready and capable of simulating diverse phenotypes. Ontologies can be designed to aid in comparing and reconciling metabolic reconstructions that differ in how they represent various metabolic pathways. ModelSEED now includes 33,978 compounds and 36,645 reactions, available as a set of extensible files on GitHub, and available to search at https://modelseed.org/biochem and KBase.
Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.
Microbiome samples are inherently defined by the environment in which they are found. Therefore, data that provide context and enable interpretation of measurements produced from biological samples, often referred to as metadata, are critical. Important contributions have been made in the development of community-driven metadata standards; however, these standards have not been uniformly embraced by the microbiome research community. To understand how these standards are being adopted, or the barriers to adoption, across research domains, institutions, and funding agencies, the National Microbiome Data Collaborative (NMDC) hosted a workshop in October 2019. This report provides a summary of discussions that took place throughout the workshop, as well as outcomes of the working groups initiated at the workshop.
A better understanding of the genetic and metabolic mechanisms that confer stress resistance and tolerance in plants is key to engineering new crops through advanced breeding technologies. This requires a systems biology approach that builds on a genome-wide understanding of the regulation of gene expression, plant metabolism, physiology and growth. In this study, we examine the response to drought stress in Sorghum, as we leverage the tools for transcriptomics and plant metabolic modeling we have implemented at the U.S. Department of Energy Systems Biology Knowledgebase (KBase). KBase enables researchers worldwide to collaborate and advance research by uploading private or public data into the KBase Narrative Interface, analyzing it using a rich, extensible array of computational and data-analytics tools, and securely sharing scientific workflows and conclusions. We demonstrate how to use the current RNA-seq tools in KBase, applicable to both plants and microbes, to assemble and quantify long transcripts and identify differentially expressed genes effectively. More specifically, we demonstrate the utility of the platform by identifying key genes differentially expressed during drought-stress in Sorghum bicolor, an important sustainable production crop plant. We then show how we can use KBase tools to predict the membership of genes in metabolic pathways and examine expression data in the context of metabolic subsystems. We demonstrate the power of the platform by making the data, analysis and interpretation available to the biologists in the reproducible, re-usable, point-and-click format of a KBase Narrative thus promoting FAIR (Findable, Accessible, Interoperable and Reusable) guiding principles for scientific data management and stewardship.
The National Microbiome Data Collaborative (NMDC) is a multi-organizational effort to integrate microbiome data across diverse areas in environmental science. Data provided by the NMDC can then undergo advanced analysis and provide new insights into metagenomics, metatranscriptomics, metaproteomics, and metabolomics. To address these challenges, we have developed our schema using the Linked data Modeling Language (LinkML). This allows us to easily map data to existing standards and ontologies.
To harness the potential of microbiome science across the broad range of relevant disciplines, new approaches to data infrastructure and transdisciplinary collaboration are necessary. The National Microbiome Data Collaborative (NMDC) is a new initiative to support microbiome data exploration and discovery through a collaborative, integrative data science ecosystem. To harness the potential of microbiome science across the broad range of relevant disciplines, new approaches to data infrastructure and transdisciplinary collaboration are necessary. The National Microbiome Data Collaborative is a new initiative to support microbiome data exploration and discovery through a collaborative, integrative data science ecosystem.
At the National Energy Research Scientific Computing (NERSC) Center, interactive access to high-performance computing and data through Jupyter is a priority. We will discuss the nuts and bolts of how Jupyter is deployed at NERSC, and how we've adapted to engage the Jupyter ecosystem and open-source community to deliver this key capability to our users. Jupyter is a major component in our Superfacility initiative, which aims to connect experimental and observational big data facilities (telescopes, microscopes, genome sequencers, light sources, etc.) with next-generation supercomputing and data capabilities at NERSC.