As lipidomics approaches its 25th anniversary, we explore how lipid research has matured over the years while highlighting emerging innovations that are expanding our ability to study these diverse, life-critical biomolecules. In particular, we showcase the community-driven, open-access databases, software, and educational resources made freely available through the ELIXIR Core Data Resource LIPID MAPS for the benefit of both established and new researchers.
Searching and learning from aggregated public metabolomics data spanning thousands of studies remained largely inaccessible. Here we present StructureMASST, a web-based application enabling scalable, structure-centric searches across public metabolomics repositories using molecule names or chemical representations. It queries a precomputed knowledgebase of 2.19 billion spectral matches and 420 million metadata links, supports modification-tolerant and mass-shift searches, and maps chemical structures across taxonomy, biological context and environmental conditions to accelerate discovery.
Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.
Metabolites are referenced in spectral, structural and pathway databases with a diverse array of schemas, including various internal database identifiers and large tables of common name synonyms. Cross-linking metabolite identifiers is a required step for meta-analysis of metabolomic results across studies but made difficult due to the lack of a consensus identifier system. We have implemented metLinkR, an R package that leverages RefMet and RaMP-DB to automate and simplify cross-linking metabolite identifiers across studies and generating common names. MetLinkR accepts as input metabolite common names and identifiers from five different databases (HMDB, KEGG, ChEBI, LIPIDMAPS and PubChem) to exhaustively search for possible overlap in supplied metabolites from input data sets. In an example of 13 metabolomic data sets totaling 10,400 metabolites, metLinkR identified and provided common names for 1377 metabolites in common between at least 2 data sets in less than 18 min and produced standardized names for 74.4% of the input metabolites. In another example comprising five data sets with 3512 metabolites, metLinkR identified 715 metabolites in common between at least two data sets in under 12 min and produced standardized names for 82.3% of the input metabolites. Outputs of MetLInR include output tables and metrics allowing users to readily double check the mappings and to get an overview of chemical classes represented. Overall, MetLinkR provides a streamlined solution for a common task in metabolomic epidemiology and other fields that meta-analyze metabolomic data. The R package, vignette and source code are freely downloadable at https://github.com/ncats/metLinkR.
Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.
The Playbook Workflow Builder (PWB) is a web-based platform to dynamically construct and execute bioinformatics workflows by utilizing a growing network of input datasets, semantically annotated API endpoints, and data visualization tools contributed by an ecosystem of collaborators. Via a user-friendly user interface, workflows can be constructed from contributed building-blocks without technical expertise. The output of each step of the workflow is added into reports containing textual descriptions, figures, tables, and references. To construct workflows, users can click on cards that represent each step in a workflow, or construct workflows via a chat interface that is assisted by a large language model (LLM). Completed workflows are compatible with Common Workflow Language (CWL) and can be published as research publications, slideshows, and posters. To demonstrate how the PWB generates meaningful hypotheses that draw knowledge from across multiple resources, we present several use cases. For example, one of these use cases prioritizes drug targets for individual cancer patients using data from the NIH Common Fund programs GTEx, LINCS, Metabolomics, GlyGen, and ExRNA. The workflows created with PWB can be repurposed to tackle similar use cases using different inputs. The PWB platform is available from: https://playbook-workflow-builder.cloud/.
The Data Distillery Knowledge Graph (DDKG) is a framework for semantic integration and querying of biomedical data across domains. Built for the NIH Common Fund Data Ecosystem, it supports translational research by linking clinical and experimental datasets in a unified graph model. Clinical standards such as ICD-10, SNOMED, and DrugBank are integrated through UMLS, while genomics and basic science data are structured using ontologies and standards such as HPO, GENCODE, Ensembl, STRING, and ClinVar. The DDKG uses a property graph architecture based on the UBKG infrastructure and supports ontology-based ingestion, identifier normalization, and graph-native querying. The system is modular and can be extended with new datasets or schema modules. We demonstrate its utility for informatics queries across eight use cases, including regulatory variant analysis, tissue-specific expression, biomarker discovery, and cross-species variant prioritization. The DDKG is accessible via a public interface, a programmatic API, and downloadable builds for local use.
Many biomedical research projects produce large-scale datasets that may serve as resources for the research community for hypothesis generation, facilitating diverse use cases. Towards the goal of developing infrastructure to support the findability, accessibility, interoperability, and reusability (FAIR) of biomedical digital objects and maximally extracting knowledge from data, complex queries that span across data and tools from multiple resources are currently not easily possible. By utilizing existing FAIR application programming interfaces (APIs) that serve knowledge from many repositories and bioinformatics tools, different types of complex queries and workflows can be created by using these APIs together. The Playbook Workflow Builder (PWB) is a web-based platform that facilitates interactive construction of workflows by enabling users to utilize an ever-growing network of input datasets, semantically annotated API endpoints, and data visualization tools contributed by an ecosystem. Via a user-friendly web-based user interface (UI), workflows can be constructed from contributed building-blocks without technical expertise. The output of each step of the workflows are provided in reports containing textual descriptions, as well as interactive and downloadable figures and tables. To demonstrate the ability of the PWB to generate meaningful hypotheses that draw knowledge from across multiple resources, we present several use cases. For example, one of these use cases sieves novel targets for individual cancer patients using data from the GTEx, LINCS, Metabolomics, GlyGen, and the ExRNA Communication Consortium (ERCC) Common Fund (CF) Data Coordination Centers (DCCs). The workflows created with the PWB can be published and repurposed to tackle similar use cases using different inputs. The PWB platform is available from: . ### Competing Interest Statement The authors have declared no competing interest.
Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, we've developed pan-repository universal identifiers and harmonized cross-repository metadata. This novel ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplifies data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.
LIPID MAPS (LIPID Metabolites and Pathways Strategy), www.lipidmaps.org, provides a systematic and standardized approach to organizing lipid structural and biochemical data. Founded 20 years ago, the LIPID MAPS nomenclature and classification has become the accepted community standard. LIPID MAPS provides databases for cataloging and identifying lipids at varying levels of characterization in addition to numerous software tools and educational resources, and became an ELIXIR-UK data resource in 2020. This paper describes the expansion of existing databases in LIPID MAPS, including richer metadata with literature provenance, taxonomic data and improved interoperability to facilitate FAIR compliance. A joint project funded by ELIXIR-UK, in collaboration with WikiPathways, curates and hosts pathway data, and annotates lipids in the context of their biochemical pathways. Updated features of the search infrastructure are described along with implementation of programmatic access via API and SPARQL. New lipid-specific databases have been developed and provision of lipidomics tools to the community has been updated. Training and engagement have been expanded with webinars, podcasts and an online training school.
Background Biomedical research often involves contextual integration of multimodal and multiomic data in search of mechanisms for improved diagnosis, treatment, and monitoring. Researchers need to access information from diverse sources, comprising data in various and sometimes incongruent formats. The downstream processing of the data to decipher mechanisms by reconstructing networks and developing quantitative models warrants considerable effort.Results MetGENE is a knowledge-based, gene-centric data aggregator that hierarchically retrieves information about the gene(s), their related pathway(s), reaction(s), metabolite(s), and metabolomic studies from standard data repositories under one dashboard to enable ease of access through centralization of relevant information. We note that MetGENE focuses only on those genes that encode for proteins directly associated with metabolites. All other gene-metabolite associations are beyond the current scope of MetGENE. Further, the information can be contextualized by filtering by species, anatomy (tissue), and condition (disease or phenotype).Conclusions MetGENE is an open-source tool that aggregates metabolite information for a given gene(s) and presents them in different computable formats (e.g., JSON) for further integration with other omics studies. MetGENE is available at https://bdcw.org/MetGENE/index.php.
Progress in mass spectrometry lipidomics has led to a rapid proliferation of studies across biology and biomedicine. These generate extremely large raw datasets requiring sophisticated solutions to support automated data processing. To address this, numerous software tools have been developed and tailored for specific tasks. However, for researchers, deciding which approach best suits their application relies on ad hoc testing, which is inefficient and time consuming. Here we first review the data processing pipeline, summarizing the scope of available tools. Next, to support researchers, LIPID MAPS provides an interactive online portal listing open-access tools with a graphical user interface. This guides users towards appropriate solutions within major areas in data processing, including (1) lipid-oriented databases, (2) mass spectrometry data repositories, (3) analysis of targeted lipidomics datasets, (4) lipid identification and (5) quantification from untargeted lipidomics datasets, (6) statistical analysis and visualization, and (7) data integration solutions. Detailed descriptions of functions and requirements are provided to guide customized data analysis workflows.
We report the development of a spectral knowledgebase named ADAP-KDB for tracking and prioritizing unknown gas chromatography-mass spectrometry (GC-MS) spectra in the NIH's Metabolomics Data Repository-a national and international repository for metabolomics data. ADAP-KDB consists of two parts. One part is a computational workflow that preprocesses raw mass spectrometry data and derives consensus mass spectra. The other part is a web portal for users to browse the consensus spectra and match query spectra against them. For each consensus spectrum, the Gini-Simpson diversity index and the p-value from χ2 goodness-of-fit test are calculated to measure its statistical significance, which enables prioritization of unknown mass spectra for subsequent costly compound identification.
Access to web-based platforms has enabled scientists to perform research remotely. A critical aspect of mass spectrometry data analysis is the inspection, analysis, and visualization of the raw data to validate data quality and confirm statistical observations. We developed the GNPS Dashboard, a web-based data visualization tool, to facilitate synchronous collaborative inspection, visualization, and analysis of private and public mass spectrometry data remotely.
We present LipidFinder 2.0, incorporating four new modules that apply artefact filters, remove lipid and contaminant stacks, in-source fragments and salt clusters, and a new isotope deletion method which is significantly more sensitive than available open-access alternatives. We also incorporate a novel false discovery rate (FDR) method, utilizing a target-decoy strategy, which allows users to assess data quality. A renewed lipid profiling method is introduced which searches three different databases from LIPID MAPS and returns bulk lipid structures only, and a lipid category scatter plot with color blind friendly pallet. An API interface with XCMS Online is made available on LipidFinder’s online version. We show using real data that LipidFinder 2.0 provides a significant improvement over non-lipid metabolite filtering and lipid profiling, compared to available tools. Availability LipidFinder 2.0 is freely available at https://github.com/ODonnell-Lipidomics/LipidFinder and http://lipidmaps.org/resources/tools/lipidfinder . Contact lipidfinder@cardiff.ac.uk Supplementary information Supplementary data are available at Bioinformatics online.
AbstractArchived metabolomics data represent a broad resource for the scientific community. However, the absence of tools for the meta‐analysis of heterogeneous data types makes it challenging to perform direct comparisons in a single and cohesive workflow. Here, we present a framework for the meta‐analysis of metabolic pathways and interpretation with proteomic and transcriptomic data. This framework facilitates the comparison of heterogeneous types of metabolomics data from online repositories (eg, XCMS Online, Metabolomics Workbench, GNPS, and MetaboLights) representing tens of thousands of studies, as well as locally acquired data. As a proof of concept, we apply the workflow for the meta‐analysis of (a) independent colon cancer studies, further interpreted with proteomics and transcriptomics data, (b) multimodal data from Alzheimer's disease and mild cognitive impairment studies, demonstrating its high‐throughput capability for the systems level interpretation of metabolic pathways. Moreover, the platform has been modified for improved knowledge dissemination through a collaboration with Metabolomics Workbench and LIPID MAPS. We envision that this meta‐analysis tool combined with our in‐source fragmentation/annotation (ISA) technology will help overcome the primary bottleneck in analyzing diverse datasets and facilitate the full exploitation of archival metabolomics data for addressing a broad array of questions in metabolism research and systems biology.
A comprehensive and standardized system to report lipid structures analyzed by MS is essential for the communication and storage of lipidomics data. Herein, an update on both the LIPID MAPS classification system and shorthand notation of lipid structures is presented for lipid categories Fatty Acyls (FA), Glycerolipids (GL), Glycerophospholipids (GP), Sphingolipids (SP), and Sterols (ST). With its major changes, i.e., annotation of ring double bond equivalents and number of oxygens, the updated shorthand notation facilitates reporting of newly delineated oxygenated lipid species as well. For standardized reporting in lipidomics, the hierarchical architecture of shorthand notation reflects the diverse structural resolution powers provided by mass spectrometric assays. Moreover, shorthand notation is expanded beyond mammalian phyla to lipids from plant and yeast phyla. Finally, annotation of atoms is included for the use of stable isotope-labeled compounds in metabolic labeling experiments or as internal standards. This update on lipid classification, nomenclature, and shorthand annotation for lipid mass spectra is considered a standard for lipid data presentation.
With the advent of high throughput mass spectrometric methods, metabolomics has emerged as an essential area of research in biomedicine with the potential to provide deep biological insights into normal and diseased functions in physiology. However, to achieve the potential offered by metabolomics measures, there is a need for biologist-friendly integrative analysis tools that can transform data into mechanisms that relate to phenotypes. Here, we describe MetENP, an R package, and a user-friendly web application deployed at the Metabolomics Workbench site extending the metabolomics enrichment analysis to include species-specific pathway analysis, pathway enrichment scores, gene-enzyme information, and enzymatic activities of the significantly altered metabolites. MetENP provides a highly customizable workflow through various user-specified options and includes support for all metabolite species with available KEGG pathways. MetENPweb is a web application for calculating metabolite and pathway enrichment analysis. Availability and Implementation The MetENP package is freely available from Metabolomics Workbench GitHub: (<https://github.com/metabolomicsworkbench/MetENP>), the web application, is freely available at (<https://www.metabolomicsworkbench.org/data/analyze.php>) ### Competing Interest Statement The authors have declared no competing interest.