Finding relevant chemicals in the vast (known) chemical space is a major challenge for environmental and exposomics studies leveraging nontarget high resolution mass spectrometry (NT-HRMS) methods. Chemical databases now contain hundreds of millions of chemicals, yet many are not relevant. This article details an extensive collaborative, open science effort to provide a dynamic collection of chemicals for environmental, metabolomics, and exposomics research, along with supporting information about their relevance to assist researchers in the interpretation of candidate hits. The PubChemLite for Exposomics collection is compiled from ten annotation categories within PubChem, enhanced with patent, literature and annotation counts, predicted partition coefficient (logP) values, as well as predicted collision cross section (CCS) values using CCSbase. Monthly versions are archived on Zenodo under a CC-BY license, supporting reproducible research, and a new interface has been developed, including historical trends of patent and literature data, for researchers to browse the collection. This article details how PubChemLite can support researchers in environmental and exposomics studies, describes efforts to increase the availability of experimental CCS values, and explores known limitations and potential for future developments. The data and code behind these efforts are openly available. PubChemLite can be browsed at https://pubchemlite.lcsb.uni.lu.
Environmental sciences, including environmental chemistry and toxicology, are highly interdisciplinary fields that integrate researchers with various backgrounds and expertise. This interdisciplinary aspect is critical to addressing issues of chemical pollution, environmental sustainability, and health. However, a standardized method for reporting chemical data is needed to address these issues effectively. This becomes increasingly important as both the number of chemical structures and our reliance on and use of computational analysis and cheminformatics tools grow. This paper provides background, examples, and recommendations on how to report chemical data in a findable, accessible, interoperable, and reproducible (FAIR) manner within environmental science disciplines. Ultimately, the goal is to broaden the scope and applicability of environmental research to help the entire community tackle the issues of chemical pollution and sustainability in a comprehensive manner.
As the occurrence of human diseases and conditions increase, questions continue to arise about their linkages to chemical exposure, especially for per-and polyfluoroalkyl substances (PFAS). Currently, many chemicals of concern have limited experimental information available for their use in analytical assessments. Here, we aim to increase this knowledge by providing the scientific community with multidimensional characteristics for 175 PFAS and their resulting 281 ion types. Using a platform coupling reversed-phase liquid chromatography (RPLC), electrospray ionization (ESI) or atmospheric pressure chemical ionization (APCI), drift tube ion mobility spectrometry (IMS), and mass spectrometry (MS), the retention times, collision cross section (CCS) values, and m/z ratios were determined for all analytes and assembled into an openly available multidimensional dataset. This information will provide the scientific community with essential characteristics to expand analytical assessments of PFAS and augment machine learning training sets for discovering new PFAS.
Integration of glycan-related databases between different research fields is essential in glycoscience. It requires knowledge across the breadth of science because most glycans exist as glycoconjugates. On the other hand, especially between chemistry and biology, glycan data has not been easy to integrate due to the huge variety of glycan structure representations. We have developed WURCS (Web 3.0 Unique Representation of Carbohydrate Structures) as a notation for representing all glycan structures uniquely for the purpose of integrating data across scientific data resources. While the integration of glycan data in the field of biology has been greatly advanced, in the field of chemistry, progress has been hampered due to the lack of appropriate rules to extract sugars from chemical structures. Thus, we developed a unique algorithm to determine the range of structures allowed to be considered as sugars from the structural formulae of compounds, and we developed software to extract sugars in WURCS format according to this algorithm. In this manuscript, we show that our algorithm can extract sugars from glycoconjugate molecules represented at the molecular level and can distinguish them from other biomolecules, such as amino acids, nucleic acids, and lipids. Available as software, MolWURCS is freely available and downloadable ( https://gitlab.com/glycoinfo/molwurcs ).
Studying glycans and their functions in the body aids in the understanding of disease mechanisms and developing new treatments. This necessitates resources that provide comprehensive glycan data integrated with relevant information from other scientific fields such as genomics, genetics, proteomics, metabolomics, and chemistry. The present paper describes two resources at the U.S. National Center for Biotechnology Information (NCBI), the NCBI Glycans and PubChem, which provide glycan-related information useful for the glycoscience research community. The NCBI Glycans ( https://www.ncbi.nlm.nih.gov/glycans/ ) is a dedicated website for glycobiology data content at NCBI and provides quick access to glycan-related information scattered across multiple NCBI databases as well as other information resources external to NCBI. Importantly, the NCBI Glycans hosts the official web page for the symbol nomenclature for glycans (SNFG), which is the standard graphical representation of glycan structures recommended for scientific publication. On the other hand, PubChem ( https://pubchem.ncbi.nlm.nih.gov ) is a research-focused, large-scale public chemical database, containing a substantial number of glycan-containing records and is integrated with important glycoscience resources like GlyTouCan, GlyCosmos, and GlyGen. PubChem organizes glycan-related information within multiple data collections (i.e., Substance, Compound, Protein, Gene, Pathway, and Taxonomy) and provides various tools and services that allow users to access them both interactively through a web browser and programmatically through a REST-ful interface, including PUG-View. The NCBI Glycans and PubChem highlight glycan-related data and improve their accessibility, helping scientists exploit these data in their research.
Alzheimer’s Disease (AD) is a complex and multifactorial neurodegenerative disease, which is currently diagnosed via clinical symptoms and non-specific biomarkers (such as Aβ1-42, t-Tau, and p-Tau) measured in cerebrospinal fluid (CSF), which alone do not provide sufficient insights into disease progression. In this pilot study, these biomarkers were complemented with small molecule analysis using non-target high resolution mass spectrometry (NT-HRMS) coupled to liquid chromatography (LC) on the CSF of three groups; AD, Mild Cognitive Impairment (MCI) due to AD, and a non-demented control group (ND). An open source cheminformatics pipeline based on MS-DIAL and patRoon was enhanced using CSF- and AD-specific suspect lists to assist in data interpretation. ChemRICH analysis revealed a significant increase of hydroxybutyrates in AD, including 3-hydroxy butanoic acid (BHBA), which was found at higher levels in AD compared to MCI and ND. Furthermore, a highly sensitive target LC-MS method was used to quantify 35 bile acids (BAs) in the CSF, revealing several statistically significant differences including higher dehydrolithocholic acid levels and decreased conjugated BAs levels in AD. This work provides several promising small molecule hypotheses that could be used to help track the progression of AD in CSF samples.
PubChem ( https://pubchem.ncbi.nlm.nih.gov ) is a public chemical information resource containing more than 100 million unique chemical structures. One of the most requested tasks in PubChem and other chemical databases is to search chemicals by name (also commonly called a "chemical synonym"). PubChem performs this task by looking up chemical synonym-structure associations provided by individual depositors to PubChem. In addition, these synonyms are used for many purposes, including creating links between chemicals and PubMed articles (using Medical Subject Headings (MeSH) terms). However, these depositor-provided name-structure associations are subject to substantial discrepancies within and between depositors, making it difficult to unambiguously map a chemical name to a specific chemical structure. The present paper describes PubChem's crowdsourcing-based synonym filtering strategy, which resolves inter- and intra-depositor discrepancies in synonym-structure associations as well as in the chemical-MeSH associations. The PubChem synonym filtering process was developed based on the analysis of four crowd-voting strategies, which differ in the consistency threshold value employed (60% vs 70%) and how to resolve intra-depositor discrepancies (a single vote vs. multiple votes per depositor) prior to inter-depositor crowd-voting. The agreement of voting was determined at six levels of chemical equivalency, which considers varying isotopic composition, stereochemistry, and connectivity of chemical structures and their primary components. While all four strategies showed comparable results, Strategy I (one vote per depositor with a 60% consistency threshold) resulted in the most synonyms assigned to a single chemical structure as well as the most synonym-structure associations disambiguated at the six chemical equivalency contexts. Based on the results of this study, Strategy I was implemented in PubChem's filtering process that cleans up synonym-structure associations as well as chemical-MeSH associations. This consistency-based filtering process is designed to look for a consensus in name-structure associations but cannot attest to their correctness. As a result, it can fail to recognize correct name-structure associations (or incorrect ones), for example, when a synonym is provided by only one depositor or when many contributors are incorrect. However, this filtering process is an important starting point for quality control in name-structure associations in large chemical databases like PubChem.
Dynamic changes in protein glycosylation impact human health and disease progression. However, current resources that capture disease and phenotype information focus primarily on the macromolecules within the central dogma of molecular biology (DNA, RNA, proteins). To gain a better understanding of organisms, there is a need to capture the functional impact of glycans and glycosylation on biological processes. A workshop titled “Functional impact of glycans and their curation” was held in conjunction with the 16th Annual International Biocuration Conference to discuss ongoing worldwide activities related to glycan function curation. This workshop brought together subject matter experts, tool developers, and biocurators from over 20 projects and bioinformatics resources. Participants discussed four key topics for each of their resources: (i) how they curate glycan function-related data from publications and other sources, (ii) what type of data they would like to acquire, (iii) what data they currently have, and (iv) what standards they use. Their answers contributed input that provided a comprehensive overview of state-of-the-art glycan function curation and annotations. This report summarizes the outcome of discussions, including potential solutions and areas where curators, data wranglers, and text mining experts can collaborate to address current gaps in glycan and glycosylation annotations, leveraging each other’s work to improve their respective resources and encourage impactful data sharing among resources. Database URL: https://wiki.glygen.org/Glycan_Function_Workshop_2023
PubChem (https://pubchem.ncbi.nlm.nih.gov) is a large and highly-integrated public chemical database resource at NIH. In the past two years, significant updates were made to PubChem. With additions from over 130 new sources, PubChem contains >1000 data sources, 119 million compounds, 322 million substances and 295 million bioactivities. New interfaces, such as the consolidated literature panel and the patent knowledge panel, were developed. The consolidated literature panel combines all references about a compound into a single list, allowing users to easily find, sort, and export all relevant articles for a chemical in one place. The patent knowledge panels for a given query chemical or gene display chemicals, genes, and diseases co-mentioned with the query in patent documents, helping users to explore relationships between co-occurring entities within patent documents. PubChemRDF was expanded to include the co-occurrence data underlying the literature knowledge panel, enabling users to exploit semantic web technologies to explore entity relationships based on the co-occurrences in the scientific literature. The usability and accessibility of information on chemicals with non-discrete structures (e.g. biologics, minerals, polymers, UVCBs and glycans) were greatly improved with dedicated web pages that provide a comprehensive view of all available information in PubChem for these chemicals.
The GlyCosmos Glycoscience Portal (https://glycosmos.org) and PubChem (https://pubchem.ncbi.nlm.nih.gov/) are major portals for glycoscience and chemistry, respectively. GlyCosmos is a portal for glycan-related repositories, including GlyTouCan, GlycoPOST, and UniCarb-DR, as well as for glycan-related data resources that have been integrated from a variety of 'omics databases. Glycogenes, glycoproteins, lectins, pathways, and disease information related to glycans are accessible from GlyCosmos. PubChem, on the other hand, is a chemistry-based portal at the National Center for Biotechnology Information. PubChem provides information not only on chemicals, but also genes, proteins, pathways, as well as patents, bioassays, and more, from hundreds of data resources from around the world. In this work, these 2 portals have made substantial efforts to integrate their complementary data to allow users to cross between these 2 domains. In addition to glycan structures, key information, such as glycan-related genes, relevant diseases, glycoproteins, and pathways, was integrated and cross-linked with one another. The interfaces were designed to enable users to easily find, access, download, and reuse data of interest across these resources. Use cases are described illustrating and highlighting the type of content that can be investigated. In total, these integrations provide life science researchers improved awareness and enhanced access to glycan-related information.
Transformation product (TP) information is essential to accurately evaluate the hazards compounds pose to human health and the environment. However, information about TPs is often limited, and existing data is often not fully Findable, Accessible, Interoperable and Reusable (FAIR). FAIRifying existing TP knowledge is a relatively easy path towards improving access to data for identification workflows and for machine learning-based algorithms. ShinyTPs was developed to curate existing transformation information derived from text-mined data within the PubChem database. The application (available as an R package) visualizes the text-mined chemical names to facilitate user validation of the automatically extracted reactions. ShinyTPs was applied to a case study using 436 tentatively identified compounds to prioritize TP retrieval. This resulted in the extraction of 645 reactions (associated with 496 compounds), of which 319 reactions were not previously available in PubChem. The curated reactions were added to the PubChem Transformations library, which was used as a TP suspect list for identification of TPs using the open-source workflow patRoon. In total 72 compounds from the library were tentatively identified, 18% of which were curated using ShinyTPs, showing that the app can help support TP identification in non-target analysis workflows.
This is a MetFrag database file constructed from the "Molecule contains PFAS parts larger than CF2/CF3" subnode of the OECD PFAS Definition node in the PFAS and Fluorinated Organic Compounds in PubChem Tree on the Classification Browser in PubChem. This file was constructed by downloading the node contents, selecting the columns of interest, changing the headers to MetFrag-compatible headers and adding exact mass, PubMed ID (PMID) counts and patent counts to the file (via this package). Entries containing Xe and Pr were removed; charges were also removed from formulas to avoid issues with MetFragCL. The construction of the tree is documented here.
PubChem (https://pubchem.ncbi.nlm.nih.gov) is a popular chemical information resource that serves a wide range of use cases. In the past two years, a number of changes were made to PubChem. Data from more than 120 data sources was added to PubChem. Some major highlights include: the integration of Google Patents data into PubChem, which greatly expanded the coverage of the PubChem Patent data collection; the creation of the Cell Line and Taxonomy data collections, which provide quick and easy access to chemical information for a given cell line and taxon, respectively; and the update of the bioassay data model. In addition, new functionalities were added to the PubChem programmatic access protocols, PUG-REST and PUG-View, including support for target-centric data download for a given protein, gene, pathway, cell line, and taxon and the addition of the 'standardize' option to PUG-REST, which returns the standardized form of an input chemical structure. A significant update was also made to Pub-ChemRDF. The present paper provides an overview of these changes.
This is the dataset used in the publication of "Resource Description Framework (RDF) Modeling of Named Entity Co-occurrences in Biomedical Literature and Its Integration with PubChemRDF"
Nonulosonic acids or non-2-ulosonic acids (NulOs) are an ancient family of 2-ketoaldonic acids (α-ketoaldonic acids) with a 9-carbon backbone. In nature, these monosaccharides occur either in a 3-deoxy form (referred to as “sialic acids”) or in a 3,9-dideoxy “sialic-acid-like” form. The former sialic acids are most common in the deuterostome lineage, including vertebrates, and mimicked by some of their pathogens. The latter sialic-acid-like molecules are found in bacteria and archaea. NulOs are often prominently positioned at the outermost tips of cell surface glycans, and have many key roles in evolution, biology and disease. The diversity of stereochemistry and structural modifications among the NulOs contributes to more than 90 sialic acid forms and 50 sialic-acid-like variants described thus far in nature. This paper reports the curation of these diverse naturally occurring NulOs at the NCBI sialic acid page (https://www.ncbi.nlm.nih.gov/glycans/sialic.html) as part of the NCBI-Glycans initiative. This includes external links to relevant Carbohydrate Structure Databases. As the amino and hydroxyl groups of these monosaccharides are extensively derivatized by various substituents in nature, the Symbol Nomenclature For Glycans (SNFG) rules have been expanded to represent this natural diversity. These developments help illustrate the natural diversity of sialic acids and related NulOs, and enable their systematic representation in publications and online resources.
This is the repository for regular updates of the PubChemLite for Exposomics data collection. PubChemLite for Exposomics is a subset of PubChem selected from major categories of the Table of Contents page at the PubChem Classification Browser, described in DOI:10.1186/s13321-021-00489-0. PubChemLite for Exposomics is compiled from 10 categories: AgroChemInfo, BioPathway, DrugMedicInfo, FoodRelated, PharmacoInfo, SafetyInfo, ToxicityInfo, KnownUse, DisorderDisease, Identification. PubChemCIDs have been collapsed by InChIKey first block, reporting the structure from the most annotated CID, plus related CIDs. Entries that will be ignored by MetFrag (salts, disconnected substances) or cause errors (e.g. transition metals) have been removed. The Patent and PubMed ID counts are extracted from files on the PubChem FTP site. The `AnnoTypeCount' term counts how many of the categories are represented, the subsequent column (named per category) counts the number of annotation categories available in the next sub-category of the TOC entry. These files can be used `as is' as localCSV for MetFrag Command Line.
Alzheimer’s Disease (AD) is a complex and multifactorial neurodegenerative disease. The current diagnosis relies on non-specific biomarkers (Aβ1-42, t-Tau, and p-Tau) measured in cerebrospinal fluid (CSF), which do not provide sufficient insights into disease progression. Studying the exposome could reveal new disease-specific biomarkers for more accurate diagnosis. In this pilot study, exposomics was performed on the CSF of three groups; AD, Mild Cognitive Impairment (MCI) due to AD, and a non-demented control group (ND), using non-target high resolution mass spectrometry (NT-HRMS) coupled with liquid chromatography (LC). An open-source cheminformatics pipeline was developed using MS-DIAL and patRoon with PubChemLite for Exposomics, plus CSF- and AD-specific suspect lists. Fifteen statistically significant chemicals (nine Level l, six Level 2a) from diverse classes (amino acids, gut metabolites, sugars, environmental chemicals) were identified. Most of the relevant chemicals (thirteen out of fifteen) were detected using the Hydrophilic Interaction LC (HILIC) method. Environmental and lifestyle factors may explain some chemical differences found across groups, such as the higher levels of indole-3-acetic acid found in the AD and MCI compared to the ND group. This work provides a strong methodological basis and several promising hypotheses to upscale these efforts on larger AD cohort numbers in future studies.
The term "exposome" is defined as a comprehensive study of life-course environmental exposures and the associated biological responses. Humans are exposed to many different chemicals, which can pose a major threat to the well-being of humanity. Targeted or non-targeted mass spectrometry techniques are widely used to identify and characterize various environmental stressors when linking exposures to human health. However, identification remains challenging due to the huge chemical space applicable to exposomics, combined with the lack of sufficient relevant entries in spectral libraries. Addressing these challenges requires cheminformatics tools and database resources to share curated open spectral data on chemicals to improve the identification of chemicals in exposomics studies. This article describes efforts to contribute spectra relevant for exposomics to the open mass spectral library MassBank (https://www.massbank.eu) using various open source software efforts, including the R packages RMassBank and Shinyscreen. The experimental spectra were obtained from ten mixtures containing toxicologically relevant chemicals from the US Environmental Protection Agency (EPA) Non-Targeted Analysis Collaborative Trial (ENTACT). Following processing and curation, 5582 spectra from 783 of the 1268 ENTACT compounds were added to MassBank, and through this to other open spectral libraries (e.g., MoNA, GNPS) for community benefit. Additionally, an automated deposition and annotation workflow was developed with PubChem to enable the display of all MassBank mass spectra in PubChem, which is rerun with each MassBank release. The new spectral records have already been used in several studies to increase the confidence in identification in non-target small molecule identification workflows applied to environmental and exposomics research.