Supplementary Data from Reproducibility of the Blood and Urine Exposome: A Systematic Literature Review and Meta-Analysis
The aim of this literature review was to identify and provide a summary update on the validity and applicability of the most promising dietary biomarkers reflecting the intake of important foods in the Western diet for application in epidemiological studies. Many dietary biomarker candidates, reflecting intake of common foods and their specific constituents, have been discovered from intervention and observational studies in humans, but few have been validated. The literature search was targeted for biomarker candidates previously reported to reflect intakes of specific food groups or components that are of major importance in health and disease. Their validity was evaluated according to 8 predefined validation criteria and adapted to epidemiological studies; we summarized the findings and listed the most promising food intake biomarkers based on the evaluation. Biomarker candidates for alcohol, cereals, coffee, dairy, fats and oils, fruits, legumes, meat, seafood, sugar, tea, and vegetables were identified. Top candidates for all categories are specific to certain foods, have defined parent compounds, and their concentrations are unaffected by nonfood determinants. The correlations of candidate dietary biomarkers with habitual food intake were moderate to strong and their reproducibility over time ranged from low to high. For many biomarker candidates, critical information regarding dose response, correlation with habitual food intake, and reproducibility over time is yet unknown. The nutritional epidemiology field will benefit from the development of novel methods to combine single biomarkers to generate biomarker panels in combination with self-reported data. The most promising dietary biomarker candidates that reflect commonly consumed foods and food components for application in epidemiological studies were identified, and research required for their full validation was summarized.
Metabolites produced by the gut microbiota play an important role in the cross-talk with the human host. Many microbial metabolites are biologically active and can pass the gut barrier and make it into the systemic circulation, where they form the gut microbial exposome, i.e. the totality of gut microbial metabolites in body fluids or tissues of the host. A major difficulty faced when studying the microbial exposome and its role in health and diseases is to differentiate metabolites solely or partially derived from microbial metabolism from those produced by the host or coming from the diet. Our objective was to collect data from the scientific literature and build a database on gut microbial metabolites and on evidence of their microbial origin. Three types of evidence on the microbial origin of the gut microbial exposome were defined: (1) metabolites are produced in vitro by human faecal bacteria; (2) metabolites show reduced concentrations in humans or experimental animals upon treatment with antibiotics; (3) metabolites show reduced concentrations in germ-free animals when compared with conventional animals. Data was manually collected from peer-reviewed publications and inserted in the Exposome-Explorer database. Furthermore, to explore the chemical space of the microbial exposome and predict metabolites uniquely formed by the microbiota, genome-scale metabolic models (GSMMs) of gut bacterial strains and humans were compared. A total of 1848 records on one or more types of evidence on the gut microbial origin of 457 metabolites was collected in Exposome-Explorer. Data on their known precursors and concentrations in human blood, urine and faeces was also collected. About 66% of the predicted gut microbial metabolites (n = 1543) were found to be unique microbial metabolites not found in the human GSMM, neither in the list of 457 metabolites curated in Exposome-Explorer, and can be targets for new experimental studies. This new data on the gut microbial exposome, freely available in Exposome-Explorer ( http://exposome-explorer.iarc.fr/ ), will help researchers to identify poorly studied microbial metabolites to be considered in future studies on the gut microbiota, and study their functionalities and role in health and diseases.
Background: The NORMAN Association (https://www.norman-.network.com/) initiated the NORMAN Suspect List Exchange (NORMAN-SLE; https://www.norman-.network.com/nds/SLE/) in 2015, following the NORMAN collaborative trial on non-target screening of environmental water samples by mass spectrometry. Since then, this exchange of information on chemicals that are expected to occur in the environment, along with the accompanying expert knowledge and references, has become a valuable knowledge base for "suspect screening" lists. The NORMAN-SLE now serves as a FAIR (Findable, Accessible, Interoperable, Reusable) chemical information resource worldwide. Results: The NORMAN-SLE contains 99 separate suspect list collections (as of May 2022) from over 70 contributors around the world, totalling over 100,000 unique substances. The substance classes include per- and polyfluoroalkyl substances (PFAS), pharmaceuticals, pesticides, natural toxins, high production volume substances covered under the European REACH regulation (EC: 1272/2008), priority contaminants of emerging concern (CECs) and regulatory lists from NORMAN partners. Several lists focus on transformation products (TPs) and complex features detected in the environment with various levels of provenance and structural information. Each list is available for separate download. The merged, curated collection is also available as the NORMAN Substance Database (NORMAN SusDat). Both the NORMAN-SLE and NORMAN SusDat are integrated within the NORMAN Database System (NDS). The individual NORMAN-SLE lists receive digital object identifiers (DOIs) and traceable versioning via a Zenodo community (https:// zenodo.org/communities/norman-.sle), with a total of > 40,000 unique views, > 50,000 unique downloads and 40 citations (May 2022). NORMAN-SLE content is progressively integrated into large open chemical databases such as PubChem (https://pubchem.ncbi.nlm.nih.gov/) and the US EPA's CompTox Chemicals Dashboard (https://comptox. epa.gov/dashboard/), enabling further access to these lists, along with the additional functionality and calculated properties these resources offer. PubChem has also integrated significant annotation content from the NORMAN-SLE, including a classification browser (https://pubchem.ncbi.nlm.nih.gov/classification/#hid=101). Conclusions: The NORMAN-SLE offers a specialized service for hosting suspect screening lists of relevance for the environmental community in an open, FAIR manner that allows integration with other major chemical resources. These efforts foster the exchange of information between scientists and regulators, supporting the paradigm shift to the "one substance, one assessment" approach. New submissions are welcome via the contacts provided on the NORMAN-SLE website (https://www.norman-.network.com/nds/SLE/).
Endogenous and exogenous metabolite concentrations may be susceptible to variation over time. This variability can lead to misclassification of exposure levels and in turn to biased results. To assess the reproducibility of metabolites, the intraclass correlation coefficient (ICC) is computed. A literature search in three databases from 2000 to May 2021 was conducted to identify studies reporting ICCs for blood and urine metabolites. This review includes 192 studies, of which 31 studies are included in the meta-analyses. The ICCs of 359 single metabolites are reported, and the ICCs of 10 metabolites were meta-analyzed. The reproducibility of the single metabolites ranges from poor to excellent and is highly compound-dependent. The reproducibility of bisphenol A (BPA), mono-ethyl phthalate (MEP), mono-n-butyl phthalate (MnBP), mono-2-ethylhexyl phthalate (MEHP), mono(2-ethyl-5-hydroxyhexyl) phthalate (MEHHP), mono-benzyl phthalate (MBzP), mono-(2-ethyl-5-oxohexyl) phthalate (MEOHP), methylparaben, and propylparaben is poor to moderate (ICC median: 0.32; range: 0.15-0.49), and for 25-hydroxyvitamin D [25(OH)D], it is excellent (ICC: 0.95; 95% CI, 0.90-0.99). Pharmacokinetics, mainly the half-life of elimination and exposure patterns, can explain reproducibility. This review describes the reproducibility of the blood and urine exposome, provides a vast dataset of ICC estimates, and hence constitutes a valuable resource for future reproducibility and clinical epidemiologic studies.
Objective: In 2016, the International Agency for Research on Cancer, part of the World Health Organization, released the Exposome-Explorer, the first database dedicated to biomarkers of exposure for environmental risk factors for diseases. The database contents resulted from a manual literature search that yielded over 8,500 citations, but only a small fraction of these publications were used in the final database. Manually curating a database is time-consuming and requires domain expertise to gather relevant data scattered throughout millions of articles. This work proposes a supervised machine learning pipeline to assist the manual literature retrieval process.Methods: The manually retrieved corpus of scientific publications used in the Exposome-Explorer was used as training and testing sets for the machine learning models (classifiers). Several parameters and algorithms were evaluated to predict an article’s relevance based on different datasets made of titles, abstracts and metadata.Results: The top performance classifier was built with the Logistic Regression algorithm using the title and abstract set, achieving an F2-score of 70.1%. Furthermore, we extracted 1,143 entities from these articles with a classifier trained for biomarker entity recognition. Of these, we manually validated 45 new candidate entries to the database.Conclusion: Our methodology reduced the number of articles to be manually screened by the database curators by nearly 90%, while only misclassifying 22.1% of the relevant articles. We expect that this methodology can also be applied to similar biomarkers datasets or be adapted to assist the manual curation process of similar chemical or disease databases.
Exposome-Explorer (http://exposome-explorer.iarc.fr) is a database of dietary and pollutant biomarkers measured in population studies. In its first release, Exposome-Explorer contained comprehensive information on 692 biomarkers of dietary and pollution exposures extracted from the analysis of 480 peer-reviewed publications. Today, Exposome-Explorer has been further expanded and contains a total of 908 biomarkers. Two additional types of information have been collected. First, 185 candidate dietary biomarkers having 403 associations with food intake (as measured by metabolomic studies) have been identified and added. Second, 1356 associations between dietary biomarkers and cancer risk in epidemiological studies, which were collected from 313 publications, have also been added to the database. Classifications for both foods and compounds have been revised, and new classifications for biospecimens, analytical methods and cancers have been implemented. Finally, the web interface has been redesigned to significantly improve the user experience.
The Human Metabolome Database or HMDB (www.hmdb.ca) is a web-enabled metabolomic database containing comprehensive information about human metabolites along with their biological roles, physiological concentrations, disease associations, chemical reactions, metabolic pathways, and reference spectra. First described in 2007, the HMDB is now considered the standard metabolomic resource for human metabolic studies. Over the past decade the HMDB has continued to grow and evolve in response to emerging needs for metabolomics researchers and continuing changes in web standards. This year's update, HMDB 4.0, represents the most significant upgrade to the database in its history. For instance, the number of fully annotated metabolites has increased by nearly threefold, the number of experimental spectra has grown by almost fourfold and the number of illustrated metabolic pathways has grown by a factor of almost 60. Significant improvements have also been made to the HMDB's chemical taxonomy, chemical ontology, spectral viewing, and spectral/text searching tools. A great deal of brand new data has also been added to HMDB 4.0. This includes large quantities of predicted MS/MS and GC-MS reference spectral data as well as predicted (physiologically feasible) metabolite structures to facilitate novel metabolite identification. Additional information on metabolite-SNP interactions and the influence of drugs on metabolite levels (pharmacometabolomics) has also been added. Many other important improvements in the content, the interface, and the performance of the HMDB website have been made and these should greatly enhance its ease of use and its potential applications in nutrition, biochemistry, clinical chemistry, clinical genetics, medicine, and metabolomics science.
Exposome-Explorer (http://exposome-explorer.iarc. fr) is the first database dedicated to biomarkers of exposure to environmental risk factors. It contains detailed information on the nature of biomarkers, their concentrations in various human biospecimens, the study population where measured and the analytical techniques used for measurement. It also contains correlations with external exposure measurements and data on biological reproducibility over time. The data in Exposome-Explorer was manually collected from peer-reviewed publications and organized to make it easily accessible through a web interface for in-depth analyses. The database and the web interface were developed using the Ruby on Rails framework. A total of 480 publications were analyzed and 10 510 concentration values in blood, urine and other biospecimens for 692 dietary and pollutant biomarkers were collected. Over 8000 correlation values between dietary biomarker levels and food intake as well as 536 values of biological reproducibility over time were also compiled. Exposome-Explorer makes it easy to compare the performance between biomarkers and their fields of application. It should be particularly useful for epidemiologists and clinicians wishing to select panels of biomarkers that can be used in biomonitoring studies or in exposome-wide association studies, thereby allowing them to better understand the etiology of chronic diseases.
Scope The Phenol‐Explorer web database details 383 polyphenol metabolites identified in human and animal biofluids from 221 publications. Here, we exploit these data to characterize and visualize the polyphenol metabolome, the set of all metabolites derived from phenolic food components. Methods and results Qualitative and quantitative data on 383 polyphenol metabolites as described in 424 human and animal intervention studies were systematically analyzed. Of these metabolites, 301 were identified without prior enzymatic hydrolysis of biofluids, and included glucuronide and sulfate esters, glycosides, aglycones, and O ‐methyl ethers. Around one‐third of these compounds are also known as food constituents and corresponded to polyphenols absorbed without further metabolism. Many ring‐cleavage metabolites formed by gut microbiota were noted, mostly derived from hydroxycinnamates, flavanols, and flavonols. Median maximum plasma concentrations ( C max ) of all human metabolites were 0.09 and 0.32 μM when consumed from foods or dietary supplements, respectively. Median time to reach maximum plasma concentration in humans ( T max ) was 2.18 h. Conclusion These data show the complexity of the polyphenol metabolome and the need to take into account biotransformations to understand in vivo bioactivities and the role of dietary polyphenols in health and disease.
SCOPE:The Phenol-Explorer web database (http://www.phenol-explorer.eu) was recently updated with new data on polyphenol retention due to food processing. Here, we analyze these data to investigate the effect of different variables on polyphenol content and make recommendations aimed at refining estimation of intake in epidemiological studies. METHODS AND RESULTS:Data on the effects of processing upon 161 polyphenols compiled for the Phenol-Explorer database were analyzed to investigate the effects of polyphenol structure, food, and process upon polyphenol loss. These were expressed as retention factors (RFs), fold changes in polyphenol content due to processing. Domestic cooking of common plant foods caused considerable losses (median RF = 0.45-0.70), although variability was high. Food storage caused fewer losses, regardless of food or polyphenol (median RF = 0.88, 0.95, 0.92 for ambient, refrigerated, and frozen storage, respectively). The food under study was often a more important determinant of retention than the process applied. CONCLUSION:Phenol-Explorer data enable polyphenol losses due to processing from many different foods to be rapidly compared. Where experimentally determined polyphenol contents of a processed food are not available, only published RFs matching at least the food and polyphenol of interest should be used when building food composition tables for epidemiological studies.
DrugBank (http://www.drugbank.ca) is a comprehensive online database containing extensive biochemical and pharmacological information about drugs, their mechanisms and their targets. Since it was first described in 2006, DrugBank has rapidly evolved, both in response to user requests and in response to changing trends in drug research and development. Previous versions of DrugBank have been widely used to facilitate drug and in silico drug target discovery. The latest update, DrugBank 4.0, has been further expanded to contain data on drug metabolism, absorption, distribution, metabolism, excretion and toxicity (ADMET) and other kinds of quantitative structure activity relationships (QSAR) information. These enhancements are intended to facilitate research in xenobiotic metabolism (both prediction and characterization), pharmacokinetics, pharmacodynamics and drug design/discovery. For this release, >1200 drug metabolites (including their structures, names, activity, abundance and other detailed data) have been added along with >1300 drug metabolism reactions (including metabolizing enzymes and reaction types) and dozens of drug metabolism pathways. Another 30 predicted or measured ADMET parameters have been added to each DrugCard, bringing the average number of quantitative ADMET values for Food and Drug Administration-approved drugs close to 40. Referential nuclear magnetic resonance and MS spectra have been added for almost 400 drugs as well as spectral and mass matching tools to facilitate compound identification. This expanded collection of drug information is complemented by a number of new or improved search tools, including one that provides a simple analyses of drug–target, –enzyme and –transporter associations to provide insight on drug–drug interactions.
Phenol-Explorer is an open-access web database on polyphenols, a major group of phytochemicals abundant in plant foods. Version 2.0 of the database was released in late 2011 and includes comprehensive qualitative and quantitative data on the ‘polyphenol metabolome’ (i.e. all metabolites derived from the over 500 polyphenols known in foods) in humans and experimental animals. Such databases are necessary for the screening of metabolomic profiles and the identification of potential biomarkers of food consumption. The aim of this study was to analyse these new data to characterize and visualize the polyphenol metabolome. The update was implemented by the compilation of data on 383 polyphenol metabolites from 221 original intervention studies. Research articles were first screened for suitability using pre-defined criteria and then entered into a relational database via Microsoft Access. The polyphenol metabolome was then analyzed via a series of database queries and open-source visualization software. Data were mainly obtained in human and rat models, and profiles of metabolites were similar between these species. The highest Cmax values (maximum plasma concentration) were found in rats, as higher doses of pure polyphenols could be administered, although in both species, administration of pure polyphenols or polyphenol supplements led to much higher plasma concentrations than administration of foods. Conversely, Tmax (time to reach Cmax) was species-dependent and always shorter in the rat. Additionally, the ensemble of all studies administering pure compounds to humans and animals allowed an insight into precursor-metabolite specificity. 5-O-Caffeoylquinic acid, catechin and epicatechin gave rise to the broadest range of metabolites, while hippuric, ferulic, 4- hydroxybenzoic, dihydrocaffeic and vanillic acids were the metabolites derived from the largest number of precursors. Knowledge of polyphenol metabolism is crucial to understanding their in vivo bioactivities and the polyphenol metabolome is an important component of the information-rich food metabolome, which encompasses all metabolites derived from exposure to the diet. We gratefully acknowledge Danone Research, the French National Institute of Cancer and the University of Barcelona for financing the project.
Polyphenols are a major class of bioactive phytochemicals whose consumption may play a role in the prevention of a number of chronic diseases such as cardiovascular diseases, type II diabetes and cancers. Phenol-Explorer, launched in 2009, is the only freely available web-based database on the content of polyphenols in food and their in vivo metabolism and pharmacokinetics. Here we report the third release of the database (Phenol-Explorer 3.0), which adds data on the effects of food processing on polyphenol contents in foods. Data on >100 foods, covering 161 polyphenols or groups of polyphenols before and after processing, were collected from 129 peer-reviewed publications and entered into new tables linked to the existing relational design. The effect of processing on polyphenol content is expressed in the form of retention factor coefficients, or the proportion of a given polyphenol retained after processing, adjusted for change in water content. The result is the first database on the effects of food processing on polyphenol content and, following the model initially defined for Phenol-Explorer, all data may be traced back to original sources. The new update will allow polyphenol scientists to more accurately estimate polyphenol exposure from dietary surveys.
The Human Metabolome Database (HMDB) (www.hmdb.ca) is a resource dedicated to providing scientists with the most current and comprehensive coverage of the human metabolome. Since its first release in 2007, the HMDB has been used to facilitate research for nearly 1000 published studies in metabolomics, clinical biochemistry and systems biology. The most recent release of HMDB (version 3.0) has been significantly expanded and enhanced over the 2009 release (version 2.0). In particular, the number of annotated metabolite entries has grown from 6500 to more than 40,000 (a 600% increase). This enormous expansion is a result of the inclusion of both 'detected' metabolites (those with measured concentrations or experimental confirmation of their existence) and 'expected' metabolites (those for which biochemical pathways are known or human intake/exposure is frequent but the compound has yet to be detected in the body). The latest release also has greatly increased the number of metabolites with biofluid or tissue concentration data, the number of compounds with reference spectra and the number of data fields per entry. In addition to this expansion in data quantity, new database visualization tools and new data content have been added or enhanced. These include better spectral viewing tools, more powerful chemical substructure searches, an improved chemical taxonomy and better, more interactive pathway maps. This article describes these enhancements to the HMDB, which was previously featured in the 2009 NAR Database Issue. (Note to referees, HMDB 3.0 will go live on 18 September 2012.).
Phenol-Explorer, launched in 2009, is the only comprehensive web-based database on the content in foods of polyphenols, a major class of food bioactives that receive considerable attention due to their role in the prevention of diseases. Polyphenols are rarely absorbed and excreted in their ingested forms, but extensively metabolized in the body, and until now, no database has allowed the recall of identities and concentrations of polyphenol metabolites in biofluids after the consumption of polyphenol-rich sources. Knowledge of these metabolites is essential in the planning of experiments whose aim is to elucidate the effects of polyphenols on health. Release 2.0 is the first major update of the database, allowing the rapid retrieval of data on the biotransformations and pharmacokinetics of dietary polyphenols. Data on 375 polyphenol metabolites identified in urine and plasma were collected from 236 peer-reviewed publications on polyphenol metabolism in humans and experimental animals and added to the database by means of an extended relational design. Pharmacokinetic parameters have been collected and can be retrieved in both tabular and graphical form. The web interface has been enhanced and now allows the filtering of information according to various criteria. Phenol-Explorer 2.0, which will be periodically updated, should prove to be an even more useful and capable resource for polyphenol scientists because bioactivities and health effects of polyphenols are dependent on the nature and concentrations of metabolites reaching the target tissues. The Phenol-Explorer database is publicly available and can be found online at http://www.phenol-explorer.eu. Database URL: http://www.phenol-explorer.eu.
The Human Metabolome Database (HMDB) (www. hmdb.ca) is a resource dedicated to providing scientists with the most current and comprehensive coverage of the human metabolome. Since its first release in 2007, the HMDB has been used to facilitate research for nearly 1000 published studies in metabolomics, clinical biochemistry and systems biology. The most recent release of HMDB (version 3.0) has been significantly expanded and enhanced over the 2009 release (version 2.0). In particular, the number of annotated metabolite entries has grown from 6500 to more than 40 000 (a 600% increase). This enormous expansion is a result of the inclusion of both ‘detected’ metabolites (those with measured concentrations or experimental confirmation of their existence) and ‘expected’ metabolites (those for which biochemical pathways are known or human intake/exposure is frequent but the compound has yet to be detected in the body). The latest release also has greatly increased the number of metabolites with biofluid or tissue concentration data, the number of compounds with reference spectra and the number of data fields per entry. In addition to this expansion in data quantity, new database visualization tools and new data content have been added or enhanced. These include better spectral viewing tools, more powerful chemical substructure searches, an improved chemical taxonomy and better, more interactive pathway maps. This article describes these enhancements to the HMDB, which was previously featured in the 2009 NAR Database Issue. (Note to referees, HMDB 3.0 will go live on 18
The Yeast Metabolome Database (YMDB, http://www.ymdb.ca) is a richly annotated ‘metabolomic’ database containing detailed information about the metabolome of Saccharomyces cerevisiae. Modeled closely after the Human Metabolome Database, the YMDB contains >2000 metabolites with links to 995 different genes/proteins, including enzymes and transporters. The information in YMDB has been gathered from hundreds of books, journal articles and electronic databases. In addition to its comprehensive literature-derived data, the YMDB also contains an extensive collection of experimental intracellular and extracellular metabolite concentration data compiled from detailed Mass Spectrometry (MS) and Nuclear Magnetic Resonance (NMR) metabolomic analyses performed in our lab. This is further supplemented with thousands of NMR and MS spectra collected on pure, reference yeast metabolites. Each metabolite entry in the YMDB contains an average of 80 separate data fields including comprehensive compound description, names and synonyms, structural information, physico-chemical data, reference NMR and MS spectra, intracellular/extracellular concentrations, growth conditions and substrates, pathway information, enzyme data, gene/protein sequence data, as well as numerous hyperlinks to images, references and other public databases. Extensive searching, relational querying and data browsing tools are also provided that support text, chemical structure, spectral, molecular weight and gene/protein sequence queries. Because of S. cervesiae's importance as a model organism for biologists and as a biofactory for industry, we believe this kind of database could have considerable appeal not only to metabolomics researchers, but also to yeast biologists, systems biologists, the industrial fermentation industry, as well as the beer, wine and spirit industry.