AbstractImaging mass spectrometry is a powerful technology enabling spatial metabolomics, yet metabolites can be assigned only to a fraction of the data generated. METASPACE-ML is a machine learning-based approach addressing this challenge which incorporates new scores and computationally-efficient False Discovery Rate estimation. For training and evaluation, we use a comprehensive set of 1710 datasets from 159 researchers from 47 labs encompassing both animal and plant-based datasets representing multiple spatial metabolomics contexts derived from the METASPACE knowledge base. Here we show that, METASPACE-ML outperforms its rule-based predecessor, exhibiting higher precision, increased throughput, and enhanced capability in identifying low-intensity and biologically-relevant metabolites.
Spatial metabolomics using imaging mass spectrometry (MS) enables untargeted and label-free metabolite mapping in biological samples. Despite the range of available imaging MS protocols and technologies, our understanding of metabolite detection under specific conditions is limited due to sparse empirical data and predictive theories. Consequently, challenges persist in designing new experiments, and accurately annotating and interpreting data. In this study, we systematically measured the detectability of 172 biologically-relevant metabolites across common imaging MS protocols using custom reference samples. We evaluated 24 MALDI-imaging MS protocols for untargeted metabolomics, and demonstrated the applicability of our findings to complex biological samples through comparison with animal tissue data. We showcased the potential for extending our results to further analytes by predicting metabolite detectability based on molecular properties. Additionally, our interlaboratory comparison of 10 imaging MS technologies, including MALDI, DESI, and IR-MALDESI, showed extensive metabolite coverage and comparable results, underscoring the broad applicability of our findings within the imaging MS community. We share our results and data through a new interactive web application integrated with METASPACE. This resource offers an extensive catalogue of detectable metabolite ions, facilitating protocol selection, supporting data annotation, and benefiting future untargeted spatial metabolomics studies.### Competing Interest StatementT.A. holds imaging mass spectrometry patents and leads a startup on single-cell metabolomics at BioInnovation Institute. M.A.M. is an employee and B.S. is a consultant of TransMIT GmbH. J.O. is employed at Bruker Daltonics GmbH & Co. KG.
While heterogeneity is a key feature of cancer, understanding metabolic heterogeneity at the single-cell level remains a challenge. Here we present 13C-SpaceM, a method for spatial single-cell isotope tracing that extends the previously published SpaceM method with detection of 13C6-glucose-derived carbons in esterified fatty acids. We validated 13C-SpaceM on spatially heterogeneous models using liver cancer cells subjected to either normoxia-hypoxia or ATP citrate lyase depletion. This revealed substantial single-cell heterogeneity in labelling of the lipogenic acetyl-CoA pool and in relative fatty acid uptake versus synthesis hidden in bulk analyses. Analysing tumour-bearing brain tissue from mice fed a 13C6-glucose-containing diet, we found higher glucose-dependent synthesis of saturated fatty acids and increased elongation of essential fatty acids in tumours compared with healthy brains. Furthermore, our analysis uncovered spatial heterogeneity in lipogenic acetyl-CoA pool labelling in tumours. Our method enhances spatial probing of metabolic activities in single cells and tissues, providing insights into fatty acid metabolism in homoeostasis and disease. Buglakova et al. present 13C-SpaceM, a method that combines stable isotope tracing with imaging mass spectrometry thus enabling spatial analysis of lipid dynamics with near-single-cell resolution in tissues.
On-tissue chemical derivatization is a valuable tool for expanding compound coverage in untargeted metabolomic studies with matrix-assisted laser desorption/ionization mass spectrometry imaging (MALDI-MSI). Applying multiple derivatization agents in parallel increases metabolite coverage even further but results in large and more complex datasets that can be challenging to analyze. In this work, we present a pipeline to provide rigorous annotations for on-tissue derivatized MSI data using Metaspace. To test and validate the pipeline, maize roots were used as a model system to obtain MSI datasets after chemical derivatization with four different reagents, Girard's T and P for carbonyl groups, coniferyl aldehyde for primary amines, and 2-picolylamine for carboxylic acids. Using this pipeline helped us annotate 631 unique metabolites from the CornCyc/BraChem database compared to 256 in the underivatized dataset, yet, at the same time, shortening the processing time compared to manual processing and providing robust and systematic scoring and annotation. We have also developed a method to remove false derivatized annotations, which can clean 5-25% of false derivatized annotations from the derivatized data, depending on the reagent. Taken together, our pipeline facilitates the use of broadly targeted spatial metabolomics using multiple derivatization reagents.
Metabolite annotation from imaging mass spectrometry (imaging MS) data is a difficult undertaking that is extremely resource intensive. Here, we adapted METASPACE, cloud software for imaging MS metabolite annotation and data interpretation, to quickly annotate microbial specialized metabolites from high-resolution and high-mass accuracy imaging MS data. Compared with manual ion image and MS1 annotation, METASPACE is faster and, with the appropriate database, more accurate. We applied it to data from microbial colonies grown on agar containing 10 diverse bacterial species and showed that METASPACE was able to annotate 53 ions corresponding to 32 different microbial metabolites. This demonstrates METASPACE to be a useful tool to annotate the chemistry and metabolic exchange factors found in microbial interactions, thereby elucidating the functions of these molecules.
Mass spectrometry imaging (MSI) is a powerfuland convenient method for revealing the spatial chemical composition ofdifferent biological samples. Molecular annotation of the detected signals isonly possible if a high mass accuracy is maintained over the entire image andthe m/z range. However, the heterogeneous molecular composition of biologicalsamples could lead to small fluctuations in the detected m/z-values, calledmass shift. The use of internal calibration is known to offer the best solutionto avoid, or at least to reduce, mass shifts. Their “a priori” selection for aglobal MSI acquisition is prone to false positive detection and therefore topoor recalibration. To fill this gap, this work describes an algorithm thatrecalibrates each spectrum individually by estimating its mass shift with thehelp of a list of pixel specific internal calibrating ions, automaticallygenerated in a data-adaptive manner(https://github.com/LaRoccaRaphael/MSI_recalibration). Through a practicalexample, we applied the methodology to a zebrafish whole body section acquiredat high mass resolution to demonstrate the impact of mass shift on dataanalysis and the capability of our algorithm to recalibrate MSI data. Inaddition, we illustrate the broad applicability of the method by recalibrating31 different public MSI datasets from METASPACE from various samples and typesof MSI and show that our recalibration significantly increases the numbers ofMETASPACE annotations (gaining from 20 up to 400 additional annotations),particularly the high-confidence annotations with a low false discovery rate.
Serverless computing has recently gained much attention as a feasible alternative to always-on IaaS for data processing. However, existing severless frameworks are not (yet) usable enough to reach out to a large number of users. To wit, they still require developers to specify the number of serverless functions for a simple sort job. We report our experience in designing Primula, a serverless sort operator that abstracts away users from the complexities of resource provisioning, skewed data and stragglers, yielding the most accessible sort primitive to date. Our evaluation on the IBM Cloud platform demonstrates the usability of Primula without abandoning performance (e.g., 3x faster than a serverless Spark backend and 62% slower than a hybrid serverless/IaaS solution).
Motivation: Imaging mass spectrometry (imaging MS) is a prominent technique for capturing distributions of molecules in tissue sections. Various computational methods for imaging MS rely on quantifying spatial correlations between ion images, referred to as co-localization. However, no comprehensive evaluation of co-localization measures has ever been performed; this leads to arbitrary choices and hinders method development. Results: We present ColocML, a machine learning approach addressing this gap. With the help of 42 imaging MS experts from nine laboratories, we created a gold standard of 2210 pairs of ion images ranked by their co-localization. We evaluated existing co-localization measures and developed novel measures using term frequency-inverse document frequency and deep neural networks. The semi-supervised deep learning Pi model and the cosine score applied after median thresholding performed the best (Spearman 0.797 and 0.794 with expert rankings, respectively). We illustrate these measures by inferring co-localization properties of 10 273 molecules from 3685 public METASPACE datasets.
BACKGROUND:Imaging mass spectrometry (imaging MS) is an enabling technology for spatial metabolomics of tissue sections with rapidly growing areas of applications in biology and medicine. However, imaging MS data is polluted with off-sample ions caused by sample preparation, particularly by the MALDI (matrix-assisted laser desorption/ionization) matrix application. Off-sample ion images confound and hinder statistical analysis, metabolite identification and downstream analysis with no automated solutions available. RESULTS:We developed an artificial intelligence approach to recognize off-sample ion images. First, we created a high-quality gold standard of 23,238 expert-tagged ion images from 87 public datasets from the METASPACE knowledge base. Next, we developed several machine and deep learning methods for recognizing off-sample ion images. The following methods were able to reproduce expert judgements with a high agreement: residual deep learning (F1-score 0.97), semi-automated spatio-molecular biclustering (F1-score 0.96), and molecular co-localization (F1-score 0.90). In a test-case study, we investigated off-sample images corresponding to the most common MALDI matrix (2,5-dihydroxybenzoic acid, DHB) and characterized properties of matrix clusters. CONCLUSIONS:Overall, our work illustrates how artificial intelligence approaches enabled by open-access data, web technologies, and machine and deep learning open novel avenues to address long-standing challenges in imaging MS.
Mass spectrometry imaging (MSI) is a powerful and convenient method to reveal the spatial chemical composition of different biological samples. The molecular annotation of the detected signals is only possible when high mass accuracy is maintained across the entire image and the m/z range. However, the heterogeneous molecular composition of biological samples could result in fluctuations in the detected m/z -values, called mass shift. Mass shifts impact the interpretability of the detected signals by decreasing the number of annotations and by affecting the spatial consistency and accuracy of ion images. The use of internal calibration is known to offer the best solution to avoid, or at least to reduce, mass shifts. The selection of internal calibrating signals for a global MSI acquisition is not trivial, prone to false positive detection of calibrating signals and therefore to poor recalibration. To fill this gap, this work describes an algorithm that recalibrates each spectrum individually by estimating its mass shift with the help of a list of internal calibrating ions generated automatically in a data-adaptive manner. The method exploits RANSAC ( Random Sample Consensus ) algorithm, to select, in a robust manner, the experimental signal corresponding to internal calibrating signals by filtering out calibration points with infrequent mass errors and by using the remaining points to estimate a linear model of the mass shifts. We applied the method to a zebrafish whole body section acquired at high mass resolution to demonstrate the impact of mass shift on data analysis and the capacity of our algorithm to recalibrate MSI data. We illustrate the broad applicability of the method by recalibrating 31 different public MSI datasets from METASPACE from various samples and types of MSI and show that our recalibration significantly increases the numbers of METASPACE annotations, especially the high-confident annotations at a low false discovery rate.
Metabolites, lipids, and other small molecules are key constituents of tissues supporting cellular programs in health and disease. Here, we present METASPACE, a community-populated knowledge base of spatial metabolomes from imaging mass spectrometry data. METASPACE is enabled by a high-performance engine for metabolite annotation in a confidence-controlled way that makes results comparable between experiments and laboratories. By sharing their results publicly, engine users continuously populate a knowledge base of annotated spatial metabolomes in tissues currently including over 3000 datasets from human cancer cohorts, whole-body sections of animal models, and various organs. The spatial metabolomes can be visualized, explored and shared using a web app as well as accessed programmatically for large-scale analysis. By using novel computational methods inspired by natural language processing, we illustrate that METASPACE provides molecular coverage beyond the capacity of any individual laboratory and opens avenues towards comprehensive metabolite atlases on the levels of tissues and organs.
Motivation Imaging mass spectrometry (imaging MS) is a prominent technique for capturing distributions of molecules in tissue sections. Various computational methods for imaging MS rely on quantifying spatial correlations between ion images, referred to as co-localization. However, no comprehensive evaluation of co-localization measures has ever been performed; this leads to arbitrary choices and hinders method development. Results We present ColocAI, an artificial intelligence approach addressing this gap. With the help of 42 imaging MS experts from 9 labs, we created a gold standard of 2210 pairs of ion images ranked by their co-localization. We evaluated existing co-localization measures and developed novel measures using tf-idf and deep neural networks. The semi-supervised deep learning Pi model and the cosine score applied after median thresholding performed the best (Spearman 0.797 and 0.794 with expert rankings respectively). We illustrate these measures by inferring co-localization properties of 10273 molecules from 3685 public METASPACE datasets. Availability and Implementation Contact theodore.alexandrov{at}embl.de
Motivation Imaging mass spectrometry (imaging MS) is a powerful technology for revealing localizations of hundreds of molecules in tissue sections. However, imaging MS data is polluted with off-sample ions caused by caused by sample preparation, particularly by the MALDI matrix application. The presence of the off-sample ion images confounds and hinders metabolite identification and downstream analysis. Results We created a high-quality gold standard of 23238 manually tagged ion images from 87 public datasets from the METASPACE knowledge base. We developed several machine and deep learning methods for recognizing off-sample ion images. Deep residual learning performed the best with the F1 score of 0.97. Spatio-molecular biclustering method achieved the F1 scores of 0.96 and 0.93 in semi- and fully-automated scenarios, respectively. Molecular co-localization method achieved the F1 score of 0.90. We investigated the clusters of the DHB matrix, the most common MALDI matrix, and characterized parameters of a clusters combinatorial model. This work addresses an important issue in imaging MS and illustrates how public data, modern web technologies, and machine and deep learning open novel avenues in imaging MS. Availability and Implementation Data and source code are available at: <https://github.com/metaspace2020/offsample>. Contact theodore.alexandrov{at}embl.de