EcoCyc is a bioinformatics database (DB) available at EcoCyc.org that describes the genome and the biochemical machinery of Escherichia coli K-12 MG1655. The long-term goal of the project was to describe the complete molecular catalog of the E. coli cell, as well as the functions of each of its molecular parts, to facilitate a system-level understanding of E. coli . EcoCyc is an electronic reference source for E. coli biologists and for biologists who work with related microorganisms. The database includes information pages on each E. coli gene product, metabolite, reaction, operon, and metabolic pathway. The database also includes information on the regulation of gene expression, E. coli gene essentiality, and nutrient conditions that do or do not support the growth of E. coli . The website and downloadable software contain tools for the analysis of high-throughput data sets. In addition, a steady-state metabolic flux model is generated from each new version of EcoCyc and can be executed via EcoCyc.org. The model can predict metabolic flux rates, nutrient uptake rates, and growth rates for different gene knockouts and nutrient conditions. Data generated from a whole-cell model that is parameterized from the latest data on EcoCyc is also available. This review outlines the data content of EcoCyc and the procedures by which this content is generated.
We present a tool for multi-omics data analysis that enables simultaneous visualization of up to four types of omics data on organism-scale metabolic network diagrams. The tool's interactive web-based metabolic charts depict the metabolic reactions, pathways, and metabolites of a single organism as described in a metabolicpathway database for that organism; the charts are constructed using automated graphical layout algorithms. The multi-omics visualization facility paints each individual omics dataset onto a different "visual channel" of the metabolic-network diagram. For example, a transcriptomics dataset might be displayed by coloring the reaction arrows within the metabolic chart, while a companion proteomics dataset is displayed as reaction arrow thicknesses, and a complementary metabolomics dataset is displayed as metabolite node colors. Once the network diagrams are painted with omics data, semantic zooming provides more details within the diagram as the user zooms in. Datasets containing multiple time points can be displayed in an animated fashion. The tool will also graph data values for individual reactions or metabolites designated by the user. The user can interactively adjust the mapping from data value ranges to the displayed colors and thicknesses to provide more informative diagrams.
CyanoCyc is a web portal that integrates an exceptionally rich database collection of information about cyanobacterial genomes with an extensive suite of bioinformatics tools. It was developed to address the needs of the cyanobacterial research and biotechnology communities. The 277 annotated cyanobacterial genomes currently in CyanoCyc are supplemented with computational inferences including predicted metabolic pathways, operons, protein complexes, and orthologs; and with data imported from external databases, such as protein features and Gene Ontology (GO) terms imported from UniProt. Five of the genome databases have undergone manual curation with input from more than a dozen cyanobacteria experts to correct errors and integrate information from more than 1,765 published articles. CyanoCyc has bioinformatics tools that encompass genome, metabolic pathway and regulatory informatics; omics data analysis; and comparative analyses, including visualizations of multiple genomes aligned at orthologous genes, and comparisons of metabolic networks for multiple organisms. CyanoCyc is a high-quality, reliable knowledgebase that accelerates scientists’ work by enabling users to quickly find accurate information using its powerful set of search tools, to understand gene function through expert mini-reviews with citations, to acquire information quickly using its interactive visualization tools, and to inform better decision-making for fundamental and applied research.
ChatGPT and Bard (now called Gemini), two conversational AI models developed by OpenAI and Google AI, respectively, have garnered considerable attention for their ability to engage in natural language conversations and perform various language-related tasks. While the versatility of these chatbots in generating text and simulating human-like conversations is undeniable, we wanted to evaluate their effectiveness in retrieving biological knowledge for curation and research purposes. To do so we asked each chatbot a series of questions and scored their answers based on their quality. Out of a maximal score of 24, ChatGPT scored 5 and Bard scored 13. The encountered issues included missing information, incorrect answers, and instances where responses combine accurate and inaccurate details. Notably, both tools tend to fabricate references to scientific papers, undermining their usability. In light of these findings, we recommend that biologists continue to rely on traditional sources while periodically assessing the reliability of ChatGPT and Bard. As ChatGPT aptly suggested, for specific and up-to-date scientific information, established scientific journals, databases, and subject-matter experts remain the preferred avenues for trustworthy data.
The model organism Escherichia coli K-12 has one of the most extensively annotated genomes in terms of functional characterization, yet a significant number of genes, ∼35%, are still considered poorly characterized. Initially genes without known functional understanding were given 'y' gene names. However, due to inconsistency in changing 'y' names to non-'y' names over the years, gene name alone does not provide sufficient information as to the characterization level of genes. Attempts to characterize y-ome genes, i.e. those that lack experimental evidence for function, are ongoing, and recent categorization based on the level of experimental evidence has helped clarify those genes that are well characterized versus uncharacterized. EcoCyc, the most comprehensive, curated genome database for E. coli K-12 substr. MG1655, has updated this approach by expanding the categories to include Partially characterized genes using a set of computational rules that includes keywords, experimental evidence codes and Gene Ontology terms. Approximately half of the previously categorized y-ome genes are now categorized as Partially characterized, leaving 15.5% (738) as Uncharacterized genes in EcoCyc. This new categorization scheme is searchable in the EcoCyc database, will be updated as new experimental evidence is curated and provides important information for research decisions.
The Comparative Genome Dashboard is a web-based software tool for interactive exploration of the similarities and differences in gene functions between organisms. It provides a high-level graphical survey of cellular functions, and enables the user to drill down to examine subsystems of interest in greater detail. At its highest level the Comparative Dashboard contains panels for cellular systems such as biosynthesis, energy metabolism, transport, and response to stimulus. Each panel contains a set of bar graphs that plot the numbers of compounds or gene products for each organism across a set of subsystems of that panel. Users can interactively drill down to focus on subsystems of interest and see grids of compounds produced or consumed by each organism, specific GO term assignments, pathway diagrams, and links to more detailed comparison pages. For example, the dashboard enables users to compare the cofactors that a set of organisms can synthesize, the metal ions that they are able to transport, their DNA damage repair capabilities, their biofilm-formation genes, and their viral response proteins. The dashboard enables users to quickly perform comprehensive comparisons at varying levels of detail.
Porphyromonas gingivalis is an oral human pathogen. The bacterium destroys dental tissue and is a serious health problem worldwide. Experimental data and bioinformatic analysis revealed that the pathogen produces three types of lipopolysaccharides (LPS): normal (O-type), anionic (A-type), and capsular (K-type). The enzymes involved in the production of all three types of lipopolysaccharide have been largely identified for the first two and partially for the third type. In the current work, we use bioinformatics tools to predict biosynthetic pathways for the production of the normal (O-type) lipopolysaccharide in the W50 strain Porphyromonas gingivalis and compare the pathway with other putative pathways in fully sequenced and completed genomes of other pathogenic strains. Selected enzymes from the pathway have been modeled and putative structures are presented. The pathway for the A-type antigen could not be predicted at this time due to two mutually exclusive structures proposed in the literature. The pathway for K-type antigen biosynthesis could not be predicted either due to the lack of structural data for the antigen. However, pathways for the synthesis of lipid A, its core components, and the O-type antigen ligase reaction have been proposed based on a combination of experimental data and bioinformatic analyses. The predicted pathways are compared with known pathways in other systems and discussed. It is the first report in the literature showing, in detail, predicted pathways for the synthesis of selected LPS components for the model W50 strain of P. gingivalis.
Background Enrichment or over-representation analysis is a common method used in bioinformatics studies of transcriptomics, metabolomics, and microbiome datasets. The key idea behind enrichment analysis is: given a set of significantly expressed genes (or metabolites), use that set to infer a smaller set of perturbed biological pathways or processes, in which those genes (or metabolites) play a role. Enrichment computations rely on collections of defined biological pathways and/or processes, which are usually drawn from pathway databases. Although practitioners of enrichment analysis take great care to employ statistical corrections (e.g., for multiple testing), they appear unaware that enrichment results are quite sensitive to the pathway definitions that the calculation uses. Results We show that alternative pathway definitions can alter enrichment p -values by up to nine orders of magnitude, whereas statistical corrections typically alter enrichment p -values by only two orders of magnitude. We present multiple examples where the smaller pathway definitions used in the EcoCyc database produces stronger enrichment p -values than the much larger pathway definitions used in the KEGG database; we demonstrate that to attain a given enrichment p -value, KEGG-based enrichment analyses require 1.3–2.0 times as many significantly expressed genes as does EcoCyc-based enrichment analyses. The large pathways in KEGG are problematic for another reason: they blur together multiple (as many as 21) biological processes. When such a KEGG pathway receives a high enrichment p -value, which of its component processes is perturbed is unclear, and thus the biological conclusions drawn from enrichment of large pathways are also in question. Conclusions The choice of pathway database used in enrichment analyses can have a much stronger effect on the enrichment results than the statistical corrections used in these analyses.
We present new developments in the bioinformatics algorithms that reconstruct the metabolic network of an organism from its sequenced genome, and that convert that metabolic network into a quantitative metabolic model. These algorithms are implemented in the Pathway Tools software. The Pathway Tools metabolic reconstruction process begins with an annotated genome; its first step infers the subset of metabolic reactions in the MetaCyc database that are catalyzed by enzymes within the genome. Reactome inference is based on the enzyme name and/or EC numbers within the genome annotation. The second step uses the inferred reactome to infer the subset of MetaCyc metabolic pathways present in the organism using a pathway scoring algorithm. Recent improvements to these algorithms are based on a large increase in the number of EC numbers in recent years, growth in the reaction (+30%) and pathway (+22%) content of the MetaCyc database, and development of a new pathway scoring algorithm that makes use of key reactions that must be either present or absent, and pathway taxonomic range information. The resulting metabolic reconstruction is available in the form of a Pathway/Genome Database (PGDB) that can be queried and visualized via an extensive set of online tools. The BioCyc.org website contains such PGDBs for 14,700 organisms.The MetaFlux component of Pathway Tools enables creation and execution of quantitative steady‐state metabolic flux models. Several steps are used to refine a PGDB to a quantitative steady‐state metabolic flux model. First, we must identify a set of nutrients that support growth of the organism, and we must identify a set of biomass metabolites that form the molecular building blocks of the organism. Next we ensure that the reactome within the PGDB can synthesize each of the biomass metabolites from the provided nutrients. Often the reactome contains gaps due to incompleteness in the genome annotation. We present a new taxonomic gap‐filling algorithm with significantly improved accuracy: when evaluated on randomly degraded variants of the Escherichia coli metabolic model, it shows an average F1 score of 99.0, compared to 91.0 for the previous Pathway Tools gap filler, and compared to 80.3 for a basic gap filler. Once gaps have been filled in the reaction network, MetaFlux can execute the model, using flux‐balance analysis to compute flux rates for all reactions in the model. Those fluxes can be interpreted by visualizing them on a whole‐organism metabolic network diagram. MetaFlux can also simulate the effects of single and double‐gene knock‐out mutants on the model. And MetaFlux can execute groups of models that simulate the growth of organism communities.Support or Funding InformationThis work was supported by the National Institutes of Health under grants GM080746 and GM075742. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
MOTIVATION:Biological systems function through dynamic interactions among genes and their products, regulatory circuits and metabolic networks. Our development of the Pathway Tools software was motivated by the need to construct biological knowledge resources that combine these many types of data, and that enable users to find and comprehend data of interest as quickly as possible through query and visualization tools. Further, we sought to support the development of metabolic flux models from pathway databases, and to use pathway information to leverage the interpretation of high-throughput data sets.RESULTS:In the past 4 years we have enhanced the already extensive Pathway Tools software in several respects. It can now support metabolic-model execution through the Web, it provides a more accurate gap filler for metabolic models; it supports development of models for organism communities distributed across a spatial grid; and model results may be visualized graphically. Pathway Tools supports several new omics-data analysis tools including the Omics Dashboard, multi-pathway diagrams called pathway collages, a pathway-covering algorithm for metabolomics data analysis and an algorithm for generating mechanistic explanations of multi-omics data. We have also improved the core pathway/genome databases management capabilities of the software, providing new multi-organism search tools for organism communities, improved graphics rendering, faster performance and re-designed gene and metabolite pages.AVAILABILITY:The software is free for academic use; a fee is required for commercial use. See http://pathwaytools.com.CONTACT:pkarp@ai.sri.com.SUPPLEMENTARY INFORMATION:Supplementary data are available at Briefings in Bioinformatics online.
MetaCyc (https://MetaCyc.org) is a comprehensive reference database of metabolic pathways and enzymes from all domains of life. It contains more than 2570 pathways derived from >54 000 publications, making it the largest curated collection of metabolic pathways. The data in MetaCyc is strictly evidence-based and richly curated, resulting in an encyclopedic reference tool for metabolism. MetaCyc is also used as a knowledge base for generating thousands of organism-specific Pathway/Genome Databases (PGDBs), which are available in the BioCyc (https://BioCyc.org) and other PGDB collections. This article provides an update on the developments in MetaCyc during the past two years, including the expansion of data and addition of new features.
BioCyc.org is a web portal that contains Pathway/Genome Databases for over 14,500 organisms, including 2400 human microbiome members. The BioCyc databases are unique in integrating a diverse range of data and providing a high level of curation for important microbes as well as model eukaryotic organisms and Homo sapiens. Each Pathway/Genome Database describes the genome and proteome of an organism, as well as its biochemical pathways and (for a small number of organisms) its regulatory network. BioCyc bioinformatics tools combine unparalleled breadth and user friendliness and include a unique set of visualization tools to speed comprehension of its extensive and complex data. BioCyc also includes MetaCyc, the largest collection of metabolic pathways and enzymes currently available, with over 2600 pathways. It contains the equivalent of about 9,400 textbook-pages of mini-review summaries for enzymes and pathways. The data in MetaCyc has been manually curated from over 58,000 published papers. Analysis tools in BioCyc include comparisons of pathways and genomes among multiple organisms, alignment of orthologous genes in a multi-genome browser, and searching for minimal-cost metabolic network routes that connect two substrates. Tools are also provided for visualizing omics data – such as metabolomics and gene expression data – on individual pathway diagrams, on personalized groups of multiple pathway diagrams called pathway collages, and on diagrams depicting the full metabolic network of an organism. A new visualization tool in BioCyc is the Omics Dashboard, an interactive tool for exploration and analysis of gene-expression and metabolomics datasets through a hierarchy of cellular systems and subsystems. This highly visual tool enables the user to survey their data at a very high level of abstraction, and to drill down to examine details at the pathway, subsystem, and gene levels. Patterns in the data can often be perceived quickly and easily. Another new tool is the MultiOmics Explainer, which helps researchers understand and interpret the results of their omics experiments in the context of what is known about an organism's metabolic and regulatory network. At this time this tool is not available online and requires installation of the software on a local computer. Another recent addition to BioCyc involves manual curation of databases for important pathogens. So far, the project includes Mycobacterium tuberculosis (H37Rv), Staphylococcus aureus (NCTC 8325) and Salmonella enterica (14028S) The BioCyc databases can be accessed through the BioCyc.org web site and can be downloaded. In addition, the Pathway Tools software can be downloaded and installed locally to create BioCyc-like databases for more genomes, curate existing databases, or perform additional analyses that are not available online. Access to MetaCyc and EcoCyc is unrestricted. Unlimited access to other databases requires a subscription. Support or Funding Information National Institute of General Medical Sciences of the National Institutes of Health (NIH) [GM080746, GM077678, GM75742] A portion of the initial display of the Omics Dashboard, showing the average gene expression levels of genes, at different time points, in pathways from different subsystems. Another example output of the Omics Dashboard, showing gene-expression profiles at the regulon level. A typical pathway diagram from a BioCyc database. The Multi-Genome Browser enables comparison of orthologous regions from different organisms. Genes in the Genome Browser are color-coded to indicate transcription units. Reaction diagrams provide color-coded atom mapping between reactants and products. This abstract is from the Experimental Biology 2019 Meeting. There is no full text article associated with this abstract published in The FASEB Journal.
BioCyc.org is a microbial genome Web portal that combines thousands of genomes with additional information inferred by computer programs, imported from other databases and curated from the biomedical literature by biologist curators. BioCyc also provides an extensive range of query tools, visualization services and analysis software. Recent advances in BioCyc include an expansion in the content of BioCyc in terms of both the number of genomes and the types of information available for each genome; an expansion in the amount of curated content within BioCyc; and new developments in the BioCyc software tools including redesigned gene/protein pages and metabolite pages; new search tools; a new sequence-alignment tool; a new tool for visualizing groups of related metabolic pathways; and a facility called SmartTables, which enables biologists to perform analyses that previously would have required a programmer's assistance.
The EcoCyc model-organism database collects and summarizes experimental data for Escherichia coli K-12. EcoCyc is regularly updated by the manual curation of individual database entries, such as genes, proteins, and metabolic pathways, and by the programmatic addition of results from select high-throughput analyses. Updates to the Pathway Tools software that supports EcoCyc and to the web interface that enables user access have continuously improved its usability and expanded its functionality. This article highlights recent improvements to the curated data in the areas of metabolism, transport, DNA repair, and regulation of gene expression. New and revised data analysis and visualization tools include an interactive metabolic network explorer, a circular genome viewer, and various improvements to the speed and usability of existing tools.
EcoCyc (EcoCyc.org) is a freely accessible, comprehensive database that collects and summarizes experimental data for Escherichia coli K-12, the best-studied bacterial model organism. New experimental discoveries about gene products, their function and regulation, new metabolic pathways, enzymes and cofactors are regularly added to EcoCyc. New SmartTable tools allow users to browse collections of related EcoCyc content. SmartTables can also serve as repositories for user- or curator-generated lists. EcoCyc now supports running and modifying E. coli metabolic models directly on the EcoCyc website.
BioCyc.org is a genomic resource that contains more than 7600 Pathway/Genome Databases (PGDBs) for organisms whose genomes have been completely sequenced. Although most PGDBs are bacterial, PGDBs exist for key experimental organisms, including humans (HumanCyc), yeast (YeastCyc), and Arabidopsis (AraCyc).Each BioCyc database integrates the genome sequence of one organism with its predicted metabolic network, including metabolic pathways, reactions, and chemical compounds, all presented via a user‐friendly graphical interface. The BioCyc databases contain additional computed information such as predicted operons and candidate enzymes for catalyzing predicted enzymatic steps that have no assigned enzymes in the genome annotation.The BioCyc.org website enables users to perform many types of searches and data analyses online. Analysis tools include comparisons of pathways and genomes among multiple organisms, alignment of orthologous genes in a multi‐genome browser, and searching for minimal‐cost metabolic network routes that connect two substrates. Tools are also provided for visualizing omics data – such as metabolomics and gene expression data – on individual pathway diagrams, on personalized groups of multiple pathway diagrams called pathway collages, and on diagrams depicting the full metabolic network of an organism.A powerful tool in BioCyc is Smart Tables, which are user‐defined groups of metabolites or genes that can be transformed easily into other types of objects, analyzed, saved in a BioCyc account, and shared with other users.All BioCyc databases save one were created computationally using MetaCyc as a reference. MetaCyc (MetaCyc.org) is a comprehensive manually‐curated metabolic database whose contents are derived from experimentally studies. It currently contains more than 2,400 experimentally determined metabolic pathways from more than 2,700 organisms, curated from more than 47,800 scientific publications. It also contains extensive information on metabolic enzymes, reactions, and substrates.The BioCyc databases can be accessed through the BioCyc.org web site and can be downloaded. In addition, the Pathway Tools software can be downloaded and installed locally to create BioCyc‐like databases for more genomes, curate existing databases, or perform additional analyses that are not available online.Support or Funding InformationNational Institute of General Medical Sciences of the National Institutes of Health (NIH) [GM080746, GM077678, GM75742].
EcoCyc (EcoCyc.org) is a freely accessible, comprehensive database that collects and summarizes experimental data for Escherichia coli K-12, the best-studied bacterial model organism. New experimental discoveries about gene products, their function and regulation, new metabolic pathways, enzymes and cofactors are regularly added to EcoCyc. New SmartTable tools allow users to browse collections of related EcoCyc content. SmartTables can also serve as repositories for user- or curator-generated lists. EcoCyc now supports running and modifying E. coli metabolic models directly on the EcoCyc website.
Pathway Tools is a bioinformatics software environment with a broad set of capabilities. The software provides genome-informatics tools such as a genome browser, sequence alignments, a genome-variant analyzer and comparative-genomics operations. It offers metabolic-informatics tools, such as metabolic reconstruction, quantitative metabolic modeling, prediction of reaction atom mappings and metabolic route search. Pathway Tools also provides regulatory-informatics tools, such as the ability to represent and visualize a wide range of regulatory interactions. This article outlines the advances in Pathway Tools in the past 5 years. Major additions include components for metabolic modeling, metabolic route search, computation of atom mappings and estimation of compound Gibbs free energies of formation; addition of editors for signaling pathways, for genome sequences and for cellular architecture; storage of gene essentiality data and phenotype data; display of multiple alignments, and of signaling and electron-transport pathways; and development of Python and web-services application programming interfaces. Scientists around the world have created more than 9800 Pathway/Genome Databases by using Pathway Tools, many of which are curated databases for important model organisms.
Pathway Tools is a bioinformatics software environment with a broad set of capabilities. The software provides genome-informatics tools such as a genome browser, sequence alignments, a genome-variant analyzer, and comparative-genomics operations. It offers metabolic-informatics tools, such as metabolic reconstruction, quantitative metabolic modeling, prediction of reaction atom mappings, and metabolic route search. Pathway Tools also provides regulatory-informatics tools, such as the ability to represent and visualize a wide range of regulatory interactions. The software creates and manages a type of organism-specific database called a Pathway/Genome Database (PGDB), which the software enables database curators to interactively edit. It supports web publishing of PGDBs and provides a large number of query, visualization, and omics-data analysis tools. Scientists around the world have created more than 45,000 PGDBs by using Pathway Tools, many of which are curated databases for important model organisms. Those PGDBs can be exchanged using a peer-to-peer database-sharing system called the PGDB Registry.
Background The MetaCyc and KEGG projects have developed large metabolic pathway databases that are used for a variety of applications including genome analysis and metabolic engineering. We present a comparison of the compound, reaction, and pathway content of MetaCyc version 16.0 and a KEGG version downloaded on Feb-27-2012 to increase understanding of their relative sizes, their degree of overlap, and their scope. To assess their overlap, we must know the correspondences between compounds, reactions, and pathways in MetaCyc, and those in KEGG. We devoted significant effort to computational and manual matching of these entities, and we evaluated the accuracy of the correspondences. Results KEGG contains 179 module pathways versus 1,846 base pathways in MetaCyc; KEGG contains 237 map pathways versus 296 super pathways in MetaCyc. KEGG pathways contain 3.3 times as many reactions on average as do MetaCyc pathways, and the databases employ different conceptualizations of metabolic pathways. KEGG contains 8,692 reactions versus 10,262 for MetaCyc. 6,174 KEGG reactions are components of KEGG pathways versus 6,348 for MetaCyc. KEGG contains 16,586 compounds versus 11,991 for MetaCyc. 6,912 KEGG compounds act as substrates in KEGG reactions versus 8,891 for MetaCyc. MetaCyc contains a broader set of database attributes than does KEGG, such as relationships from a compound to enzymes that it regulates, identification of spontaneous reactions, and the expected taxonomic range of metabolic pathways. MetaCyc contains many pathways not found in KEGG, from plants, fungi, metazoa, and actinobacteria; KEGG contains pathways not found in MetaCyc, for xenobiotic degradation, glycan metabolism, and metabolism of terpenoids and polyketides. MetaCyc contains fewer unbalanced reactions, which facilitates metabolic modeling such as using flux-balance analysis. MetaCyc includes generic reactions that may be instantiated computationally. Conclusions KEGG contains significantly more compounds than does MetaCyc, whereas MetaCyc contains significantly more reactions and pathways than does KEGG, in particular KEGG modules are quite incomplete. The number of reactions occurring in pathways in the two DBs are quite similar.
Peter D Karp合作论文数Artificial Intelligence Center, SRI International36