Microgravity is a condition that may affect gastrointestinal function and metabolism during spaceflight. Despite attempts to keep normal dietary habits at the ISS, the food selection and meal timing may also be perturbed during the confinement. To understand the impact of spaceflight on human metabolism, we characterized the plasma metabolome of astronauts (n = 52) before, during, and after missions on board the International Space Station (ISS) using liquid chromatography coupled with mass spectrometry. Here we show that spaceflight affected around 40 circulating compounds. In this longitudinal assessment, the metabolic changes were already observed shortly after launch and subsided within a few days after landing. The flight-induced changes in metabolites reflected increased protein fermentation by the gut microbiota, possibly reflecting prolonged intestinal transit time caused by microgravity. Minor diet-related changes related to intakes of caffeine, fish, and some fats were also observed but affected less than one third of all metabolites responding to spaceflight. We did not observe major sex-specific differences in metabolism. Increasing the consumption of slow-fermented carbohydrates at ISS to reduce protein fermentation might improve gastrointestinal health in astronauts in future human long-term spaceflights.
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and streamlined workflows. Since its introduction in 2005, the xcms R package has become one of the most widely used tools for LC-MS data preprocessing. Developed through an open-source, community-driven approach, xcms has maintained long-term stability while continuously expanding its capabilities and accessibility. We present recent advancements that position xcms as a central component of a modular and interoperable software ecosystem for metabolomics data analysis. Key improvements include enhanced scalability, enabling the processing of large-scale experiments with thousands of samples on standard computing hardware. These developments empower users to build comprehensive, customizable, and reproducible workflows tailored to diverse experimental designs and analytical needs. An expanding collection of tutorials, documentation, and teaching materials further supports both new and experienced users in leveraging the broader R and Bioconductor ecosystems. These resources facilitate the integration of statistical modeling, visualization tools, and domain-specific packages, extending the reach and impact of xcms workflows. Together, these enhancements solidify xcms as a cornerstone of modern metabolomics research.
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing data set scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and streamlined workflows. Since its introduction in 2005, the xcms R package has become one of the most widely used tools for LC-MS data preprocessing. Developed through an open-source, community-driven approach, xcms maintains long-term stability while continuously expanding its capabilities and accessibility. We present recent advancements that position xcms as a central component of a modular and interoperable software ecosystem for metabolomics data analysis. Key improvements include enhanced scalability, enabling the processing of large-scale experiments with thousands of samples on standard computing hardware. These developments empower users to build comprehensive, customizable, and reproducible workflows tailored to diverse experimental designs and analytical needs. An expanding collection of tutorials, documentation, and teaching materials further supports both new and experienced users in leveraging broader R and Bioconductor ecosystems. These resources facilitate the integration of statistical modeling, visualization tools, and domain-specific packages, extending the reach and impact of xcms workflows. Together, these enhancements solidify xcms as a cornerstone of modern metabolomics research.
The ability to answer complex biological questions in metabolomics relies on the acquisition of high-quality data. However, due to the complex nature of liquid chromatography-mass spectrometry acquisition, data quality checks are often not done comprehensively and only at the postprocessing step. This can be too late to mitigate analytical problems such as loss of m/z calibration, retention time drift and severe ion suppression. It is often not practically or economically feasible to reanalyze samples, and interpretation of the acquired compromised data, if at all possible, is limited, despite the considerable expenses incurred to obtain them. We therefore introduce QC4Metabolomics, a real-time quality control monitoring software for untargeted metabolomics data. QC4Metabolomics monitors files as they are acquired or retrospectively by tracking any user-defined compound(s) and extracting diagnostic information such as observed m/z, retention time, intensity and peak shape, and presents the results on a web dashboard. QC4Metabolomics also monitors the levels of common or user-defined contaminants. We report herein real-world examples where QC4Metabolomics easily reveals analytical problems retrospectively that could have been immediately addressed with real-time monitoring, so that the samples would have been analyzed without any quality control issues. The Shiny app is available as open-source code at https://github.com/stanstrup/QC4Metabolomics. Docker images and a docker-compose setup file are also provided for easy deployment, along with demo data. The documentation can be found at https://stanstrup.github.io/QC4Metabolomics.
Feature-based molecular networking (FBMN) is a popular analysis approach for liquid chromatography–tandem mass spectrometry-based non-targeted metabolomics data. While processing liquid chromatography–tandem mass spectrometry data through FBMN is fairly streamlined, downstream data handling and statistical interrogation are often a key bottleneck. Especially users new to statistical analysis struggle to effectively handle and analyze complex data matrices. Here we provide a comprehensive guide for the statistical analysis of FBMN results, focusing on the downstream analysis of the FBMN output table. We explain the data structure and principles of data cleanup and normalization, as well as uni- and multivariate statistical analysis of FBMN results. We provide explanations and code in two scripting languages (R and Python) as well as the QIIME2 framework for all protocol steps, from data clean-up to statistical analysis. All code is shared in the form of Jupyter Notebooks ( https://github.com/Functional-Metabolomics-Lab/FBMN-STATS ). Additionally, the protocol is accompanied by a web application with a graphical user interface ( https://fbmn-statsguide.gnps2.org/ ) to lower the barrier of entry for new users and for educational purposes. Finally, we also show users how to integrate their statistical results into the molecular network using the Cytoscape visualization tool. Throughout the protocol, we use a previously published environmental metabolomics dataset for demonstration purposes. Together, the protocol, code and web application provide a complete guide and toolbox for FBMN data integration, cleanup and advanced statistical analysis, enabling new users to uncover molecular insights from their non-targeted metabolomics data. Our protocol is tailored for the seamless analysis of FBMN results from Global Natural Products Social Molecular Networking and can be easily adapted to other mass spectrometry feature detection, annotation and networking tools. Feature-based molecular networking is used to analyze non-targeted liquid chromatography–tandem mass spectrometry metabolomics data. This protocol includes instructions, ready-made code and a web app ( https://fbmn-statsguide.gnps2.org/ ) for statistical analysis of feature-based molecular networking results.
Precision nutrition requires precise tools to monitor dietary habits. Yet current dietary assessment instruments are subjective, limiting our understanding of the causal relationships between diet and health. Biomarkers of food intake (BFIs) hold promise to increase the objectivity and accuracy of dietary assessment, enabling adjustment for compliance and misreporting. Here, we update current concepts and provide a comprehensive overview of BFIs measured in urine and blood. We rank BFIs based on a four-level utility scale to guide selection and identify combinations of BFIs that specifically reflect complex food intakes, making them applicable as dietary instruments. We discuss the main challenges in biomarker development and illustrate key solutions for the application of BFIs in human studies, highlighting different strategies for selecting and combining BFIs to support specific study designs. Finally, we present a roadmap for BFI development and implementation to leverage current knowledge and enable precision in nutrition research.
BACKGROUND:Studies suggest that dairy-derived calcium supplements have additional beneficial properties compared with other calcium supplements in relation to bone health.OBJECTIVES:We investigated the postprandial calcium absorption from a milk-derived calcium permeate (CP) compared with calcium carbonate (CC).METHODS:In this randomized double-blinded cross-over study, 10 healthy postmenopausal females (age 50-65 y) received maltodextrin (placebo), 800 mg calcium from CP or from CC provided in 6 capsules on separate days. A fasting blood sample was collected at baseline, 60, 120, 240, and 360 min after ingestion. At baseline and 360 min, spot-urine samples were collected. Serum-ionized calcium, intact parathyroid hormone, phosphorus, and magnesium were analyzed, as were urinary calcium, phosphorus, and magnesium. A linear mixed model was applied.RESULTS:Serum-ionized calcium concentration after the CC supplement was higher at 240 min compared with the CP supplement [between-group difference; 95% confidence interval (CI): 0.039 mmol/L; 95% CI: 0.017-0.061; P = 0.00078]. Serum-ionized calcium concentration after the CC supplement was significantly higher than placebo at all postprandial time points except at 60 min. Urinary calcium concentration in 360 min spot urine was higher after intake of CC compared with CP [between-group difference; 95% CI: 2.47 mmol/L; 95% CI: 1.90-3.03; P = 0.0042].CONCLUSIONS:Postprandial calcium absorption from CP was lower than that of CC, and concurrently, urinary concentration reflected increased serum appearance by CC compared with CP, highlighting different metabolic responses. The long-term and clinical implications should be studied further.
The exposure of human DNA to genotoxic compounds induces the formation of covalent DNA adducts, which may contribute to the initiation of carcinogenesis. Liquid chromatography (LC) coupled with high-resolution mass spectrometry (HRMS) is a powerful tool for DNA adductomics, a new research field aiming at screening known and unknown DNA adducts in biological samples. The lack of databases and bioinformatics tool in this field limits the applicability of DNA adductomics. Establishing a comprehensive database will make the identification process faster and more efficient and will provide new insight into the occurrence of DNA modification from a wide range of genotoxicants. In this paper, we present a four-step approach used to compile and curate a database for the annotation of DNA adducts in biological samples. The first step included a literature search, selecting only DNA adducts that were unequivocally identified by either comparison with reference standards or with nuclear magnetic resonance (NMR), and tentatively identified by tandem HRMS/MS. The second step consisted in harmonizing structures, molecular formulas, and names, for building a systematic database of 279 DNA adducts. The source, the study design and the technique used for DNA adduct identification were reported. The third step consisted in implementing the database with 303 new potential DNA adducts coming from different combinations of genotoxicants with nucleobases, and reporting monoisotopic masses, chemical formulas, .cdxml files, .mol files, SMILES, InChI, InChIKey and IUPAC nomenclature. In the fourth step, a preliminary spectral library was built by acquiring experimental MS/MS spectra of 15 reference standards, generating in silico MS/MS fragments for all the adducts, and reporting both experimental and predicted fragments into interactive web datatables. The database, including 582 entries, is publicly available (https://gitlab.com/nexs-metabolomics/projects/dna_adductomics_database). This database is a powerful tool for the annotation of DNA adducts measured in (HR)MS. The inclusion of metadata indicating the source of DNA adducts, the study design and technique used, allows for prioritization of the DNA adducts of interests and/or to enhance the annotation confidence. DNA adducts identification can be further improved by integrating the present database with the generation of authentic MS/MS spectra, and with user-friendly bioinformatics tools.
Liquid chromatography-mass spectrometry (LC-MS)-based untargeted metabolomics experiments have become increasingly popular because of the wide range of metabolites that can be analyzed and the possibility to measure novel compounds. LC-MS instrumentation and analysis conditions can differ substantially among laboratories and experiments, thus resulting in non-standardized datasets demanding customized annotation workflows. We present an ecosystem of R packages, centered around the MetaboCoreUtils, MetaboAnnotation and CompoundDb packages that together provide a modular infrastructure for the annotation of untargeted metabolomics data. Initial annotation can be performed based on MS1 properties such as m/z and retention times, followed by an MS2-based annotation in which experimental fragment spectra are compared against a reference library. Such reference databases can be created and managed with the CompoundDb package. The ecosystem supports data from a variety of formats, including, but not limited to, MSP, MGF, mzML, mzXML, netCDF as well as MassBank text files and SQL databases. Through its highly customizable functionality, the presented infrastructure allows to build reproducible annotation workflows tailored for and adapted to most untargeted LC-MS-based datasets. All core functionality, which supports base R data types, is exported, also facilitating its re-use in other R packages. Finally, all packages are thoroughly unit-tested and documented and are available on GitHub and through Bioconductor.
Free radical mechanisms may be involved in the teratogenesis of diabetes. The contribution of oxidative stress in diabetic complications was investigated from the standpoint of oxidative damage to DNA, lipids, and proteins in the livers and embryos of pregnant diabetic rats. Diabetes was induced prior to pregnancy by the administration of streptozotocin (45 mg/kg). Two groups of diabetic rats were studied, one without any supplementation (D) and another treated during pregnancy with vitamin E (150 mg/d by gavage) (D + E). A control group was also included (C). The percentage of malformations in D rats were 44%, higher than the values observed in C (7%) and D + E (12%) animals. D Group rats showed a higher concentration of thiobarbituric acid reactive substances in the mother’s liver, however, treatment with vitamin E decreased this by 58%. The levels of protein carbonyls in the liver of C, D, and D + E groups were similar. The “total levels” of the DNA adducts measured, both in liver and embryos C groups were similar to the D groups. Treatment of D groups with vitamin E reduced the levels by 17% in the liver and by 25% in the embryos. In terms of the “total levels” of DNA adducts, the embryos in diabetic pregnancy appear to be under less oxidative stress when compared with the livers of their mothers. Graziewicz et al. (Free Radical Biology & Medicine, 28:75–83, 1999) suggested “that Fapyadenine is a toxic lesion that moderately arrests DNA synthesis depending on the neighboring nucleotide sequence and interactions with the active site of DNA polymerase.” Thus the increased levels of Fapyadenine in the diabetic livers and embryos may similarly arrest DNA polymerase, and in the case of this occurring in the embryos, contribute to the congenital malformations. It is now critical to probe the molecular mechanisms of the oxidative stress-associated development of diabetic congenital malformations.
The fatty acid (FA) composition of milk from six European areas, as well as the alteration in the FA profile during cheese production, was studied using both a targeted GC‐FID and an untargeted GC‐MS approach. By applying principal component, partial least square discriminant and chemical similarity enrichment analysis, a discrimination of the geographical areas could be achieved highlighting important FA classes such as odd‐ and branched‐chain FAs for the differentiation. The FA profile remained constant during cheese production, and aroma compounds have been identified as biomarkers for the ripening methods used, namely foil and smear ripening.
Prediction of retention times (RTs) is increasingly considered in untargeted metabolomics to complement MS/MS matching for annotation of unidentified peaks. We tested the performance of PredRet (http://predret.org/) topredict RTs for plant food bioactive metabolites in a data sharing initiative containing entry sets of 29-103 compounds (totalling 467 compounds, >30 families) across 24 chromatographic systems (CSs). Between 27 and 667 predictions were obtained with a median prediction error of 0.03-0.76 min and interval width of 0.33-8.78 min. An external validation test of eight CSs showed high prediction accuracy. RT prediction was dependent on shape and type of LC gradient, and number of commonly measured compounds. Our study highlights PredRet's accuracy and ability to transpose RT data acquired from one CS to another CS. We recommend extensive RT data sharing in PredRet by the community interested in plant food bioactive metabolites to achieve a powerful community-driven open-access tool for metabolomics annotation.
Hairy root (HR) cultures are quickly evolving as a fundamental research tool and as a bio-based production system for secondary metabolites. In this study, an efficient protocol for establishment and elicitation of anthocyanin-producing HR cultures from black carrot was established. Taproot and hypocotyl explants of four carrot cultivars were transformed using wild-type Rhizobium rhizogenes. HR growth performance on plates was monitored to identify three fast-growing HR lines, two originating from root explants (lines NB-R and 43-R) and one from a hypocotyl explant (line 43-H). The HR biomass accumulated 25- to 30-fold in liquid media over a 4 week period. Nine anthocyanins and 24 hydroxycinnamic acid derivatives were identified and monitored using UPLC-PDA-TOF during HR growth. Adding ethephon, an ethylene-releasing compound, to the HR culture substantially increased the anthocyanin content by up to 82% in line 43-R and hydroxycinnamic acid concentrations by >20% in line NB-R. Moreover, the activities of superoxide dismutase and glutathione S-transferase increased in the HRs in response to ethephon, which could be related to the functionality and compartmentalization of anthocyanins. These findings present black carrot HR cultures as a platform for the in vitro production of anthocyanins and antioxidants, and provide new insight into the regulation of secondary metabolism in black carrot.
AbstractApples are a rich source of polyphenols and fiber. Proanthocyanidins (PAs), the largest polyphenolic class in apples, can reach the colon almost intact where they interact with the gut microbiota producing simple phenolic acids. These metabolites have the potential to modulate gut microbiota composition and activity and impact on host physiology. A randomized, controlled, crossover, dietary intervention study was performed to determine the broad effects of whole apple intake on fecal gut microbiota composition and activity. Forty heathy mildly hypercholesterolemic volunteers (23 women, 17 men), with a mean BMI (± SD) 25.3 ± 3.7 kg/m2 and age 51 ± 11 years, consumed 2 apples/day (Renetta Canada, rich in PAs), or a sugar matched control apple beverage, for 8 weeks separated by a 4-week washout period in a random order. Fecal and 24-h urine samples were collected before and after each treatment. The broad effects of apple intake on fecal gut microbiota composition were explored by the high throughput sequencing (HTS) of 16S rRNA gene lllumina MiSeq sequencing (V3-V4 region). Sequencing data analysis was performed using the Quantitative Insight Into Microbial Ecology (QIIME) open-source pipeline version 1.9.1. Specific bacterial groups were also enumerated using the quantitative Fluorescence In Situ Hybridization (FISH). Furthermore, the potential formation of microbial polyphenol metabolites, after apple intake, was explored in urine using Liquid Chromatography (LC) High-Resolution Mass Spectrometry (HRMS) metabolomics. Preliminary analysis showed no changes in gut microbiota abundances measured by Illumina MiSeq, after correction for multiple testing. Apple intake significantly decreased Enterobacteriaceae population (P = 0.04) compared to the control beverage, as determined with FISH. Twenty-four polyphenol microbial metabolites were identified in higher concentrations in the apple group (P < 0.05) compared to the control, including valerolactones, valeric and phenolic acids. In conclusion, preliminary data suggest that the daily intake of 2 Renetta Canada apples significantly decreased Enterobacteriaceae population, a family known for its pathogenic members, in healthy mildly hypercholesterolemic subjects. Moreover, several polyphenol microbial metabolites were identified, suggesting that microbial activity is crucial and a prerequisite for the absorption of apple polyphenols, producing active metabolites with potential health benefits.
The influence of grape maturity on wine volatome was investigated using HS-SPME-GC x GC-TOFMS. Shiraz wines were made from grapes harvested from four different vineyards from two berry maturity levels. A total of 1276 putative compounds were detected in at least one of the wine samples and 175 showed significant trends related to grape maturity. The first two dimensions of the Principal component analysis accounted for 57% of the variation and separated the samples according to the harvest date. Wines from the first harvest date were characterised by an abundance of lipoxygenase derived compounds, norisoprenoids and sulfur-containing compounds whereas a significant increase in some acetate esters was observed in wines produced from the more mature grapes. This study demonstrated a common evolution of grape volatiles for Shiraz inside the same mesoclimate. During the late ripening stage of the grape, a direct nexus between sugar concentration and wine volatile evolution was not observed.
Metabolomics aims to measure and characterise the complex composition of metabolites in a biological system. Metabolomics studies involve sophisticated analytical techniques such as mass spectrometry and nuclear magnetic resonance spectroscopy, and generate large amounts of high-dimensional and complex experimental data. Open source processing and analysis tools are of major interest in light of innovative, open and reproducible science. The scientific community has developed a wide range of open source software, providing freely available advanced processing and analysis approaches. The programming and statistics environment R has emerged as one of the most popular environments to process and analyse Metabolomics datasets. A major benefit of such an environment is the possibility of connecting different tools into more complex workflows. Combining reusable data processing R scripts with the experimental data thus allows for open, reproducible research. This review provides an extensive overview of existing packages in R for different steps in a typical computational metabolomics workflow, including data processing, biostatistics, metabolite annotation and identification, and biochemical network and pathway analysis. Multifunctional workflows, possible user interfaces and integration into workflow management systems are also reviewed. In total, this review summarises more than two hundred metabolomics specific packages primarily available on CRAN, Bioconductor and GitHub.
Apples are one of the most commonly consumed fruits and their high polyphenol content is considered one of the most important determinants of their health-promoting activities. Here we studied the nutrikinetics of apple polyphenols by UHPLC-HRMS metabolite fingerprinting, comparing bioavailability when consumed in a natural or a polyphenol-enriched cloudy apple juice. Twelve men and women participated in an acute single blind controlled crossover study in which they consumed 250 mL of cloudy apple juice (CM), Crispy Pink apple variety, or 250 mL of the same juice enriched with 750 mg of an apple polyphenol extract (PM). Plasma and whole blood were collected at time 0, 1, 2, 3 and 5 h. Urine was collected at time 0 and 0-2, 2-5, 5-8, and 8-24 h after juice consumption. Faecal samples were collected from each individual during the study for 16S rRNA gene profiling. As many as 110 metabolites were significantly elevated following intake of polyphenol enriched cloudy apple juice, with large inter-individual variations. The comparison of the average area under the curve of circulating metabolites in plasma and in urine of volunteers consuming either the CM or the PAJ demonstrated a stable metabotype, suggesting that an increase in polyphenol concentration in fruit does not limit their bioavailability upon ingestion. Faecal bacteria were correlated with specific microbial catabolites derived from apple polyphenols. Human metabolism of apple polyphenols is a co-metabolic process between human encoded activities and those of our resident microbiota. Here we have identified specific blood and urine metabolic biomarkers of apple polyphenol intake and identified putative associations with specific genera of faecal bacteria, associations which now need confirmation in specifically designed mechanistic studies.
Compound identification is the main hurdle in LC-HRMS-based metabolomics, given the high number of 'unknown' metabolites. In recent years, numerous in silico fragmentation simulators have been developed to simplify and improve mass spectral interpretation and compound annotation. Nevertheless, expert mass spectrometry users and chemists are still needed to select the correct entry from the numerous candidates proposed by automatic tools, especially in the plant kingdom due to the huge structural diversity of natural compounds occurring in plants. In this work, we propose the use of a supervised machine learning approach to predict molecular substructures from isotopic patterns, training the model on a large database of grape metabolites. This approach, called 'Compounds Characteristics Comparison' (CCC) emulates the experience of a plant chemist who 'gains experience' from a (proof-of-principle) dataset of grape compounds. The results show that the CCC approach is able to predict with good accuracy most of the sub-structures proposed. In addition, after querying MS/MS spectra in Metfrag 2.2 and applying CCC predictions as scoring terms with real data, the CCC approach helped to give a better ranking to the correct candidates, improving users' confidence in candidate selection. Our results demonstrated that the proposed approach can complement current identification strategies based on fragmentation simulators and formula calculators, assisting compound identification. The CCC algorithm is freely available as R package (https://github.com/lucanard/CCC) which includes a seamless integration with Metfrag. The CCC package also permits uploading additional training data, which can be used to extend the proposed approach to other systems biological matrices. List of abbreviations: Acidic: acidic moiety; aliph: aliphatic chain; AUC: area under the ROC curve; bs: best glycosidic structure; CCC: Compounds' Characteristics Comparison; Cees: Carbons estimation errors; CO: Carbon to Oxygen ratio; Het: Heterocyclic moiety; IMD: Isotopic Mass Defect (and Pattern); LC-HRMS: Liquid Chromatography - High Resolution Mass Spectrometry; md: mass defect; MM: Monoisotopic Mass; MS: Mass Spectrometry; MSE: Mean Squared Error; nC: number of Carbons; NN: Nitrogen; pC: percentage of Carbon mass on the total mass; Pho: Phosphate; PLSr: Partial Least Square regression; ppm: parts per million; QSRR: Quantitative structure-retention relationship; RMD: Relative Mass Defect; ROC: Receiver Operating Characteristics; rRMD: residual Relative Mass Defect; RT: retention time; Sul: Sulphur; UPLC-ESI-Q-TOF-MS: Ultra Performance Liquid Chromatography - ElectroSpray Ionization -Quadropole - Time of Flight - Mass Spectrometry; VAT: Vitis arizonica Texas.