A recurrent neural network (RNN) system based on a Hopfield neural network (HNN) was developed to extract essential spectral patterns from the time-of-flight secondary ion mass spectrometry (ToF-SIMS) spectra of peptide samples. Because ToF-SIMS produces various fragment ions from organic molecules, the interpretation of ToF-SIMS spectra is generally complicated. ToF-SIMS is useful for peptide analysis because it detects specific amino acid fragment ions from peptides that indicate peptide information. However, the ToF-SIMS spectra also contain fragment ions that do not preserve the main structures of the original molecules, which makes them difficult to interpret. Therefore, it is crucial to extract essential spectral patterns from the ToF-SIMS spectra of organic materials. Peptides were selected as the target organic materials for this study due to their systematic chemical structures. A modified HNN was trained on the ToF-SIMS spectra of each peptide, and the trained HNNs were then used to recall patterns for various peptide ToF-SIMS spectra. The results show that the modified HNN recall essential spectra containing specific ions, including the protonated molecular ions and amino acid fragment ions of target peptides. Furthermore, the HNN results revealed differences and similarities between peptides with similar and different amino acid sequences. Thus, this study demonstrates the effectiveness of the HNN in interpreting complex spectra and its potential for preprocessing data for further analysis.
Evaluating the relationship between crystal structures of V and hydrogen diffusion to extract key features is important for understanding the relationship between hydrogen diffusion behavior and crystal structure in metals. We applied an image fusion method to surface image datasets of the vanadium alloy (a single-phase bcc V-10 mol% Fe alloy) sample, obtained through different measurement methods, such as optical microscopy, scanning electron microscopy/energy-dispersive X-rays, and electron backscatter diffraction, to construct a multimodal dataset comprising hydrogen distribution and crystallographic orientation image data. The analysis of fused multimodal data by two unsupervised learning methods, such as principal component analysis and multivariate curve resolution, revealed that the hydrogen diffusion behavior differed depending on the crystallographic orientation. For example, grains oriented along the [111] and [101] directions exhibit greater hydrogen diffusion behavior than those oriented along the [001] direction. This trend was consistently observed across various analysis methods, but not when analyzed individually. In addition, it was also suggested that the shift from a pure orientation or the presence of other orientations changes the hydrogen diffusion behavior. Through multimodal data analysis of vanadium alloys, key characteristics of crystal orientation that affect hydrogen diffusion rate were extracted.
Time-of-flight secondary ion mass spectrometry (ToF-SIMS) is a powerful tool for imaging molecules in biological tissues owing to its high spatial resolution and sensitivity. Effective detection of key molecules in a sample is crucial for detailed evaluation of complex samples such as tissues. In this study, a target is a biomolecule, allantoin (C4H6N4O3), and the allantoin [M + H]+ and [M-H]- are adequately detected using ToF-SIMS with a Bi cluster ion beam from an allantoin control sample. However, the detection of ions related to allantoin permeated in human skin tissue is not straightforward because there are interfering mass peaks in the ToF-SISM spectra of control skin samples that make allantoin detection challenging, and the allantoin fragment ions have the same chemical structure as the fragment ion from other biomolecules in the tissue. In order to sufficiently detect the allantoin-related ions, we focused on ion beams with a higher ionization yield, such as gas cluster ion beams (GCIBs). As a result, a ToF-SIMS with water GCIB was significantly more effective in detecting the allantoin [M + H]+ and [M-H]- in the skin samples in both positive and negative ion spectra. The results revealed that the water GCIB approach is better suited for studying biological samples, as it effectively distinguishes the mass peaks of allantoin related ions using ToF-SIMS.
The interpretation of time-of-flight secondary ion mass spectrometry (ToF-SIMS) data is often complicated because ToF-SIMS has a high sensitivity for detecting extremely low amounts of molecules and generally produces numerous types of fragment ions from each molecule. Although machine learning techniques have been applied to such complex ToF-SIMS data interpretation to classify the components in a sample, identifying unknown molecules is often difficult, even after classification or segmentation of complex datasets. We developed a new secondary ion mass spectrometry (SIMS) identification system based on full ToF-SIMS spectra by applying a supervised machine learning method, random forest (RF), with effective teaching information to express common organic molecules. We automatically extracted chemical structures for unknown material identification from string-converted molecules using a simplified molecular-input line-entry system. The ToF-SIMS spectra of 32 organic molecules, including peptides, polymers, and biomolecules such as cellulose, were used as a training dataset, and these molecules were correctly predicted using the SIMS identification system. The importance of RF indicated that mass peaks representing these structures were detected in the ToF-SIMS spectra and that the materials were identified based on the essential chemical structures of a target molecule. Moreover, the ToF-SIMS spectra of Styrofoam-like Ocean plastic samples were correctly identified as polystyrene by the system. This study demonstrates the potential of our SIMS identification system to accurately identify unknown organic molecules from full ToF-SIMS spectra, offering a robust approach for expanding molecular identification in complex samples.
Time-of-flight secondary ion mass spectrometry (ToF-SIMS) data interpretation for organic materials is complicated because of various fragment ions produced from each molecule and the overlapping of certain mass peaks from different molecules. Fragmentation mechanisms in SIMS are complex because different sputtering and ionization processes can simultaneously occur. Therefore, a prediction system that can identify materials in a sample is required. A novel prediction system for peptides based on ToF-SIMS and amino-acid-based teaching information (labels) for supervised machine learning was developed. To develop the prediction system for general organic materials, the annotation of materials is crucial to creating effective labels for supervised learning. Peptides are composed of 20 amino acid residues, which can be used as labels. We previously developed a peptide prediction system using Random Forest, a supervised machine-learning method. However, only the amino acids contained in the target peptide were predicted, and the amino acid sequence was unable to be assumed. In this study, the amino acid sequence of the test peptide was determined by adding the information on two adjacent amino acids to the labels. Once the prediction system learned the target peptide spectra, the peptides in the newly obtained ToF-SIMS spectra could be identified. The new prediction system also provides useful information for the identification of unknown peptides. The prediction results indicate that two adjacent permutations of amino acids are effective pieces of teaching information for expressing the amino acid sequence of a peptide.
Methods that facilitate molecular identification and imaging are required to evaluate drug penetration into tissues. Time-of-flight secondary ion mass spectrometry (ToF-SIMS), which has high spatial resolution and allows 3D distribution imaging of organic materials, is suitable for this purpose. However, the complexity of ToF-SIMS data, which includes nonlinear factors, makes interpretation challenging. Therefore, in this study, ToF-SIMS data of a stratum corneum treated with diclofenac were analyzed using machine learning to enable the evaluation of drug distribution. Diclofenac-related mass peaks were identified using autoencoder results, and the degree of penetration was evaluated across 2-20th stripped tapes. In addition, the permeation pathway was clarified by comparing the secondary ion images of phosphatidylethanolamine (PhEA; a marker of the inside of the cell); cholesterol, which is abundant in cell membranes; and diclofenac. Based on the biomolecule-related ion images showing the penetration pathway of diclofenac applied to the skin, diclofenac penetrates both the extracellular space and inside cells.
Mass spectrometry imaging (MSI) is essential for visualizing drug distribution, metabolites, and significant biomolecules in pharmacokinetic studies. This study mainly focuses on imipramine, a tricyclic antidepressant that affects endogenous metabolite concentrations. The aim was to use atmospheric pressure matrix-assisted laser desorption/ionization (AP-MALDI)-MSI combined with different dimensionality reduction methods to examine the distribution and impact of imipramine on endogenous metabolites in the brains of treated wild-type mice. Brain sections from both control and imipramine-treated mice underwent AP-MALDI-MSI. Dimensionality reduction methods, including principal component analysis, multivariate curve resolution, and sparse autoencoder (SAE), were employed to extract valuable information from the MSI data. Only the SAE method identified phosphorylcholine (ChoP) as a potential marker distinguishing between the control and treated mice brains. Additionally, a significant decrease in ChoP accumulation was observed in the cerebellum, hypothalamus, thalamus, midbrain, caudate putamen, and striatum ventral regions of the treated mice brains. The application of dimensionality reduction methods, particularly the SAE method, to the AP-MALDI-MSI data is a novel approach for peak selection in AP-MALDI-MSI data analysis. This study revealed a significant decrease in ChoP in imipramine-treated mice brains.
Secondary ion mass spectrometry (SIMS) is a technique for chemical analysis and imaging of solid materials, with applications in many areas of science and technology. It involves bombarding a sample surface under high vacuum with energetic primary ions. The ejected secondary ions undergo mass-to-charge ratio (m/z) analysis and are detected. The resulting mass spectrum contains detailed surface chemical information with sub-monolayer sensitivity. Different experimental configurations provide chemically resolved depth distribution and 2D or 3D images. SIMS is complementary to other surface analysis techniques, such as X-ray photoelectron spectroscopy; chemical imaging techniques, for example, vibrational microspectroscopy methods such as Fourier transform infrared spectroscopy and Raman spectroscopy; and other mass spectrometry imaging techniques, including desorption electrospray ionization and matrix-assisted laser desorption ionization. Features of SIMS include high spatial resolution, high depth resolution and broad chemical sensitivity to all elements, isotopes and molecules up to several thousand mass units. This Primer describes the operating principles of SIMS and outlines how the instrument geometry and operational parameters enable different modes of operation and information to be obtained. Applications, including materials science, surface science, electronic devices, geosciences and life sciences, are explored, finishing with an outlook for the technique. Solid samples can be imaged and chemically analysed using secondary ion mass spectrometry. This Primer describes the secondary ion mass spectrometry experimental setup, in which a primary ion beam sputters secondary ions that are analysed and detected by a mass spectrometer, and explores applications in materials, geological and life sciences.
For understanding complex reactions in the electrodes of lith-ium-air batteries (LABs), time-of-flight secondary ion mass spectrometry (ToF-SIMS) is one of the most powerful methods because ToF-SIMS provides molecular information including organic and organic-inorganic complex materials. Although ToF-SIMS has been used to characterize various lithium-ion bat-tery electrodes, low-mass peaks, less than 100 Da, were mainly focused for obtaining the distribution images and depth profiles. Nonetheless, high-mass peaks are important for identifying the products and reactions that cause battery aging. The selection of the mass peaks specific to a particular sample among similar samples, such as a particular electrode and control electrodes, is generally very difficult by manual analysis because the ToF-SIMS spectra contain more than hundreds of mass peaks. This renders the detailed evaluation by TOF-SIMS for improved LABs challenging. Herein, we demonstrate that by applying a sparse autoencoder to the ToF-SIMS data of air electrodes con-taining a redox mediator that suppressed aging, it was possible to distinguish specific mass peaks corresponding to the elec-trodes with the redox mediator, initial electrodes before aging, and aged electrodes. The results indicate that the intensity of the mass peaks originating directly from the electrolyte material increases when the residual products on the air electrode are re-moved by the redox mediator. The analysis presented herein can be utilized to monitor the aging and degradation of these next-generation batteries.
Electron backscatter diffraction (EBSD) indexing based on Kikuchi diffraction patterns, which indicate the types and orientation of the crystal lattice, is effective for characterizing crystals. Most regions in a sample can be indexed due to simulation of diffraction patterns of possible crystal types, orientations, and angles. However, indexing some of the complex regions related to the grain boundaries, dislocations, and strain areas is difficult. Moreover, minor crystal structures are possibly omitted from the index results. To characterize all the regions, including such complicated boundaries, the analysis of raw data, including all Kikuchi patterns, is necessary. By analyzing all the Kikuchi patterns, significant information can be extracted from mixed crystal conditions. Stainless steel was used as the model sample in this study. As hydrogen diffusion in metals strongly depends on the crystal structure and grain boundaries, structural analysis is required to study hydrogen behavior in steel. In this study, all Kikuchi patterns at all pixels in a measurement area of stainless steel were analyzed simultaneously using unsupervised learning methods, such as principal component analysis and multivariate curve resolution, and the pixels of the measurement area were classified based on the Kikuchi patterns to investigate the grain boundaries and dislocations in detail.
飛行時間型二次イオン質量分析 (TOF-SIMS) は実測で得られる質量スペクトルが複雑で解釈が難しいため, TOF-SIMSデータの解析には主成分分析や多変量スペクトル分解などの多変量解析が用いられてきた.また,マトリックス効果などによってスペクトルパターンが変化するため,データベースの確立も難しい.本研究では,教師あり機械学習法による未知物質のスペクトル予測・ピーク同定を実現することを目的とし,未知物質の予測が可能となるデータ様式について検討した.教師あり学習に必要な試料情報を与えるラベルにSMILES記法によって文字列化した物質名の自動分割によって得られる化学構造を用いた.有機物・高分子・ペプチドから成るモデル試料のTOF-SIMSデータについて,ランダムフォレスト (Random Forest) による予測結果を評価した結果,自動分割した化学構造ラベルに基づいて正解率0.9以上で予測できることが示された.
fac-Tris(2-phenylpyridine) iridium [Ir(ppy)(3)] has been investigated by means of soft desorption/ionization induced by neutral SO2 clusters in combination with mass spectrometry. Desorption of intact Ir(ppy)(3 )was observed. Further analysis of the isotopic pattern revealed two forms of ionization, either by uptake of a proton or by electron abstraction. The relative contribution of the two processes depends on measurement time and H2O partial pressure, as well as preparation scheme and surface morphology of the samples.
Machine learning is a useful tool when extracting hidden information from complex measurement data obtained via surface analysis, as in secondary ion mass spectrometry. Flexible learning methods often require significant effort to adjust parameters, as these parameters may have a significant effect on results. However, machine learning methods enable the extraction of new information that cannot be found by manual analysis. This paper presents some examples of complex data analyses using conventional multivariate analysis methods based on linear combinations (principal component analysis and multivariate curve resolution), an unsupervised learning method based on artificial neural networks (sparse autoencoder), and a supervised learning method based on decision trees (random forest). To obtain reproducible and useful results from machine learning applications to surface analysis data, the preparation of data sets—including the selection of variables and the raw data conversion process—is crucial. Moreover, sufficient information representing analytical purposes, such as the chemical structures of unknown samples, material types, and physical or chemical properties of particular materials, must be contained in the data set for supervised learning.