Due to their high fat content, it is well known that walnuts are highly susceptible to oxidation and thermal degradation. These processes produce various compounds, including aromatics, aldehydes, acids, ketones, and alcohols, resulting in poor quality walnuts and shorter shelf life. The goal of this work was to demonstrate a high-level analytical workflow to follow and capture the kinetic volatile fingerprint of oxidized walnut oil through development of a method using headspace solid-phase microextraction sampling followed by comprehensive two-dimensional gas chromatography time-of-flight mass spectrometry (HS-SPME GC × GC-TOFMS) coupled with chemometric data analysis. Starting with freshly cold-pressed walnut oil, vials of the walnut oil underwent accelerated oxidation at 60 °C across a span of 6 weeks. Once HS-SPME GC × GC-TOFMS data was collected for each week of oxidation, tile-based Fisher ratio analysis and ANOVA were utilized to discover 59 class-distinguishing analytes across the 6-week experiment. Density-based spatial clustering of applications with noise visualized with principal component analysis revealed three clear kinetic profile class trends within the 59 analytes, aptly named: diminishing analytes, emerging analytes, and intermediate analytes. A very minor fourth category was also observed - unique analytes - consisting of kinetic trends dissimilar from the other three trends. Of the analytes discovered, previously unreported native compounds were discovered as part of the overall volatile chemical fingerprint of the walnut oil, providing fundamental metabolic insight to the oxidative and thermal degradation processes and their influencing factors. To illuminate the significance of these results, a partial least squares regression model to predict hexanal formation - a known oxidation product - was utilized to identify analyte responses correlated to oxidation across all kinetic profiles. Overall, the analytical workflow demonstrated herein confidently studied the walnut oil oxidation process by coupling high-resolution volatile chemical fingerprints of the walnut oil samples - captured via HS-SPME GC × GC-TOFMS analysis - with rigorous chemometric data analysis, ultimately providing a powerful approach for food science studies.
Accurate identification of all detectable analyte components in a single comprehensive two-dimensional (2D) gas chromatography time-of-flight mass spectrometry (GC × GC-TOFMS) chromatogram is a fundamental interest in the field. While commercial software tools intended for this purpose are available, the performance of these tools to generate an accurate peak table has generally not been validated. To address this, we developed a new algorithmic software approach called 2D mzCompare to generate accurate peak tables for GC × GC-TOFMS. Extending from our original method for one-dimensional GC-MS data, the 2D mzCompare algorithm discovers selective mass channels (m/z) for each analyte to resolve overlapping peaks and improve analyte identification, by leveraging the similarity in retention time and peak shape across m/z of the same analyte. The 2D mzCompare algorithm calculates the peak shape similarity between m/z at every modulation, followed by clustering and focusing steps, to generate a final peak table. To evaluate this software, we simulated realistic GC × GC-TOFMS data in the context of the statistical overlap theory (SOT), so the exact number and identities of analytes are known a priori. Utilizing an in-house mass spectrum library of similar compounds, GC × GC-TOFMS chromatograms were simulated with varying degrees of 2D chromatographic saturation (α2D). At low saturation factors (α2D = 0.01, 0.03, and 0.1), over 95% of the simulated components are found to be mathematically resolved singlets (pure analyte components) by 2D mzCompare. Meanwhile, approximately 62% were found at α2D = 1, exceeding predictions made by SOT. Application of 2D mzCompare computationally reduces the 2D peak widths, resulting in an ∼12-fold reduction in α2D. The results of this research are simultaneously two-fold. First, we provide a new algorithmic approach, 2D mzCompare, to resolve overlapped analytes in GC × GC-TOFMS data, and second, we validate the accuracy of the software performance using SOT.
To address the challenge of comparing entire chromatograms, three tile-based comparative analysis methods (Fisher ratio, pairwise, and relative variance ranking) were applied to data collected by comprehensive two-dimensional gas chromatography with vacuum ultraviolet spectroscopy detection (GC × GC-VUV). The tile-based scheme, originally developed for time-of-flight mass spectrometry (TOFMS) detection, was adapted herein to rank analytes on a hit list according to the metric for each method. We analyzed a GC × GC-VUV data set comprising 14 gas oils from five sample classes: straight-run (SR), light cycle oil (LCO), coker (CK), hydroconverted (HDC), and hydrotreated (HDT). Three replicates each of HDT and LCO gas oil GC × GC-VUV data were analyzed by Fisher ratio analysis with 459 out of 557 analyte hits exhibiting at least a 2-fold change, which were visualized by 2D chromatogram projections. Additionally, multivariate curve resolution-alternating least-squares was applied to obtain the pure VUV spectrum of each analyte, which was readily classified into five compound classes (saturates, olefins, mono-, di-, and triaromatics) and projected onto the 2D separation space to visualize the positions of these compound classes. Next, pairwise analysis, either between replicates of the same sample or between samples within the same sample class was studied. A concentration ratio >1.2 was determined as the 99% confidence level threshold for distinguishing analyte hits from background variation. Finally, relative variance analysis of all 24 gas oil chromatograms combined with principal component analysis enabled unsupervised differentiation of the five gas oil sample classes.
It is well established that partial least squares (PLS) regression is a powerful chemometric tool for linking complex fuel composition data to physicochemical properties when coupled with either gas chromatography (GC) or Fourier transform infrared spectroscopy (FT-IR). In this report, we investigate combining the analysis of high-speed GC with flame ionization detection (GC-FID) data with FT-IR data via data augmentation strategies with PLS modeling to optimize the prediction of viscosity, hydrogen content, heat of combustion, and density across a diverse set of 50 kerosene-based fuels. High-speed GC-FID separations were performed at 2-min and 5-fold slower at 10-min separation times, while comprehensive two-dimensional gas chromatography with time-of-flight mass spectrometry (GC × GC-TOFMS) served as a high-resolution method to benchmark the PLS modeling performance. PLS models constructed from GC × GC-TOFMS achieved normalized root mean square error of prediction (NRMSEP) values of 2.9% (viscosity), 6.7% (hydrogen content), 10.6% (heat of combustion), and 3.4% (density). GC-FID models with 10-min separations produced similar performance with NRMSEPs of 3.1%, 7.0%, 9.8%, and 3.2%, respectively, while the high-speed 2-min separations yielded 2.7%, 7.4%, 8.1%, and 5.7%. FT-IR models provided competitive performance, achieving NRMSEPs of 5.1% (viscosity), 3.4% (hydrogen content), 9.9% (heat of combustion), and 2.0% (density). Data augmentation of 2-min GC-FID separations and FT-IR spectra further improved the PLS models, yielding NRMSEPs of 2.7% (viscosity), 4.5% (hydrogen content), 6.9% (heat of combustion), and 3.6% (density). Overall, the results demonstrate that high-speed GC-FID and FT-IR spectroscopy provide synergistic analytical modalities with data augmentation for rapid and robust fuel property prediction.
Proton transfer reaction time-of-flight mass spectrometry (PTR-TOFMS) is a powerful tool for real-time analysis of volatile organic compounds (VOCs), including tetrachloroethylene (PCE). However, data analysis of large datasets of PTR-TOFMS data can be challenging to optimally extract chemical information. To address this challenge, we present a MATLAB-based workflow that processes PTR-TOFMS data directly from raw HDF5 files and analyzes all measurement cycles (i.e., spectra) in an untargeted manner. The workflow integrates mass spectrum preprocessing, alignment, and feature selection to maximize signal-to-noise ratio (S/N) and improve analyte discovery. Preprocessing included averaging, smoothing, and baseline correction, followed by correlation optimized warping (COW) alignment adapted to one-dimensional mass spectral data to correct m/z peak shifting across all spectra and samples. We implemented a tile-based F-ratio analysis to the one-dimensional (1D) PTR-TOFMS spectra to discover analytes correlated with PCE concentration. Samples with the five highest and five lowest PCE concentrations were separated into two classes, and 1D tile-based Fisher-ratio (F-ratio) analysis was followed by a 95% confidence interval t-test to generate a statistically filtered hit list. Principal component analysis (PCA) was used to visualize class separation, and performance was quantitatively assessed using the degree of class separation (DCS). PCE was ranked as the top hit (F-ratio = 129), and 20 of 487 discovered analyte hits passed the statistical significance threshold. Receiver operating characteristic analysis yielded an area under the curve (AUC) of 0.98, indicating effective discrimination of PCE-correlated analytes. The DCS between high- and low-PCE classes increased from 1.9 to 5.1, representing nearly a three-fold improvement when PCA was applied to only the top 20 F-ratio hits. Overall, this study demonstrates a workflow incorporating preprocessing and alignment across all mass spectra and samples and applies a tile-based automated F-ratio calculation framework to one-dimensional PTR-TOFMS data, enabling robust untargeted detection and quantification of analytes correlated with PCE.
We are developing high speed gas chromatographic (HSGC) instrumentation with an optimizable injection system, referred to herein as dynamic pressure gradient injection (DPGI). In the present study, we examine the effects of the DPGI pulse width and linear flow velocity on the resultant chromatographic peak widths and separation peak capacity. DPGI readily yields reproducible peak widths and retention times in a sub-second separation runtime regime over long periods of repeated injections. These repeated measurements facilitate a statistically rigorous analysis of the relationships between peak widths obtained and injection pulse width and/or linear flow velocity. Chromatographic performance was studied using a 1 m × 100 µm × 0.1 µm Rtx-5 chromatographic column at various linear flow velocities with hydrogen as the carrier gas, an isothermal temperature of 100 °C, with a test mixture of acetone, nonane, decane and undecane. At this column temperature, acetone is nominally unretained. For conditions where plate height is minimized (Hmin) at the so-called optimum linear flow velocity, uopt, and with the off-column band broadening approaching zero by optimizing DPGI performance, an Hmin of 77 µm was obtained. The chromatographic data corresponding to this Hmin included a minimum peak width-at-half height (w1/2) of 8±0.2ms for acetone, and a peak capacity (nc)of ∼30 for a separation runtime of 1.2 s. When all that is needed is the separation of a few key analytes as fast as possible, and if some peak capacity can be sacrificed, the fastest separation studied yielded a minimum peak width at half-height w1/2=5.5±0.09ms for acetone, and a nc of 10 with a separation runtime of 325ms.
Historically, tile-based Fisher ratio (F-ratio) analysis of comprehensive two-dimensional gas chromatography time-of-flight mass spectrometry (GC × GC-TOFMS) data was developed for analysts to use a supervised experimental design with defined sample classes to obtain a hit list to discover analytes that most significantly distinguish the sample classes at the top of the hit list. In this traditional application, a user-specified F-ratio threshold is used to discard most hits in order to focus on the top hits. To broaden the scope of tile-based F-ratio analysis, in the present study we explore the ability of the software to discover all analyte components that are detected in a set of samples, essentially taking full advantage of the tiling aspect of the software which uncovers all analytes that exhibit sufficient signal relative to the baseline noise across all samples to be deemed detectable and hence to produce an F-ratio. For this study a set of nine petroleum samples, i.e., two hydrobates (light naphthas), two reformates, four naphthas, and a “heavy” gasoline, are simultaneously analyzed and statistically compared via p-testing to blank chromatograms to produce one comprehensive hit list. The pin locations and signal areas at the top m/z F-ratio are used together with replicate blanks to generate a master peak table (MPT) that in turn is used to generate sample-specific peak tables (SSPT), one SSPT for each injection replicate of each petroleum sample (class), that are naturally retention-time aligned via the F-ratio software. The nine petroleum samples vary to a large extent in the identity and number of analytes present. Indeed, while a total of ∼715 analytes were found across all nine samples, only ∼260 of these analytes are fully shared across all sample classes. The number of analytes in the nine petroleum samples ranged from an average of 335 analytes for one of the hydrobates to 669 analytes for two of the naphthas. This workflow also facilitated generating simulated distillation curves for the nine petroleum samples to provide further insight.
Comprehensive two-dimensional gas chromatography (GC×GC)–mass spectrometry (MS) is currently the most powerful tool for analysing GC-amenable compounds. The technique is complemented by using MS as a third dimension, after two stages of chromatographic separation. In this Primer, various aspects of GC×GC–MS are explored, including basic principles and method optimization with both cryogenic and flow modulation. State-of-the-art instrumentation and data processing tools are discussed, including the advantages and limitations of GC×GC–MS compared with other GC techniques. A range of applications are explored, with an overall future outlook for the field. Information is provided on when and why GC×GC–MS should be used, how it can be fully exploited and potential advances that may occur in the next decade. Using two gas chromatography columns and a mass spectrometer, comprehensive two-dimensional gas chromatography–mass spectrometry (GC×GC–MS) is a powerful tool for separating and analysing gas-phase compounds. This Primer provides an overview of GC×GC–MS, including experimental set-up, analysis and applications in food science, environmental studies, petrochemicals and various -omics fields.
Malassezia yeasts are commensal microorganisms found in human and animal skin. Species of Malassezia have been connected to skin and opportunistic infections, where certain microenvironmental conditions are required in the host for the pathogenic processes to occur. We present the analysis of the volatile space of Malassezia pachydermatis grown at three pH values (5.7, 9.7, and 12.4) by comprehensive two-dimensional gas chromatography time-of-flight mass spectrometry (GC×GC-TOFMS). Since changes in pH also affect the growth media and the volatile organic compounds (VOCs) produced by it, media blanks at the three pHs were analyzed, with 5 replicates of each of the 6 samples. Following data collection, GC×GC-TOFMS chromatograms were analyzed by Fisher ratio software that found 566 analytes, out of which 288 were tentatively identified with a mass spectrum match value (MV) ≥ 800 based upon a NIST library search. A signal pattern for each of the 566 analytes was obtained by averaging the replicates, and two metrics (R and RSD) were calculated for each signal pattern. The R metric was defined to focus upon the differences between analyte signals of media blanks and M. pachydermatis by taking away the influence of pH changes, while the RSD metric was defined to evaluate only the influence of pH. Based on the R metric magnitude, the analytes were split into 3 categories: media analytes consumed by M. pachydermatis, analytes at similar concentration at a given pH in the media and M. pachydermatis, and analytes produced in M. pachydermatis only. Many of the M. pachydermatis produced analytes were already shown to be produced by other yeast species and shown to have biological significance when the pH is varied. Further, there is evidence of some bioconversions between the consumed analytes discovered versus the analytes produced. We also verified our classification results using a support vector machine (SVM) model, where cross-validation provided a very promising outcome with true positive rate (TPR) and true negative rate (TNR) both being over 0.95 and the error being below 0.03 (or 3%).
Partial least squares (PLS) regression is a valuable chemometric tool for property prediction when coupled with gas chromatography (GC). Since the separation run time and stationary phase selection are crucial for effective PLS modeling, we study these GC parameters on the prediction of viscosity, density and hydrogen content for 50 aerospace fuels. Due to the diversity of compounds in the fuels (primarily alkanes, cycloalkanes, and aromatics), we explore both polar and non-polar stationary phase columns. The robustness for the PLS models was evaluated by their normalized root mean square error of cross-validation (NRMSECV). PLS models built for viscosity across 1-min, 3-min, 7-min, and 10-min time window (TW) high-speed GC separations produced nearly the same NRMSECV with the polar column data with an average (standard deviation) of 4.41 % (0.34 %) versus the non-polar column data of 4.69 % (0.15 %). In contrast, while the NRMSECV of density modeling with the polar column data varied more than the viscosity models, averaging 7.54 % (0.67 %), the non-polar column data produced a significantly higher average NRMSECV of 10.06 % (0.35 %). Similarly, for hydrogen content, the NRMSECV with the polar column data averaged 9.50 % (0.87 %), which was significantly lower than the NRMSECV with the non-polar column data averaging 12.10 % (0.88 %). We also investigated the impact of smoothing the GC data on the corresponding PLS models. By applying varying degrees of smoothing, we can effectively obtain similar chromatographic peak patterns in a shorter TW. For example, a 10-min smoothed chromatogram appears like the 1-min separation with no smoothing but resulted in nearly the same NRMSECV. Overall, the fast separation with a 1-min TW produced robust PLS models for viscosity with either stationary phase column, whereas for density and hydrogen content the polar stationary phase column produced superior PLS models, thus with proper stationary phase selection, a fast separation run time could be readily applied with optimal PLS property modeling results.
The presence of flavor defects in coffee beans can negatively impact quality, the consumer experience, and commercial trade. Potato taste defect (PTD), a flavor defect specific to East African coffee, is often characterized by a musty, vegetable-like aroma. While previous work has correlated PTD with the presence of 2-isopropyl-3methoxypyrazine (IPMP), additional changes in the volatile profile of these beans can further amplify the distinct odor of this defect. The aim of this work was to develop a volatile fingerprint of PTD in roasted arabica coffee using headspace solid-phase microextraction coupled to comprehensive two-dimensional gas chromatography with time-of-flight mass spectrometry (HS-SPME-GC x GC-TOFMS) and chemometrics. Examination of the HSSPME-GC x GC-TOFMS data with tile-based Fisher ratio (F-ratio) analysis discovered 359 analytes that differentiated clean coffee samples from those impacted by severe PTD (p-value < 0.01). It was determined that 327 of the identified analytes were more prevalent in the clean coffee samples while 32 analytes, including IPMP, exhibited higher signals in the impacted coffee samples. Principal components analysis (PCA) of the F-ratio results demonstrated that the coffee samples clustered based on the presence of PTD. Partial least squares (PLS) regression modeling further demonstrated that the compounds discovered by F-ratio analysis were correlated with PTD by accurately predicting the concentration of IPMP in the samples. Investigation of the compounds highly weighted in both the PCA and PLS loadings suggest that the presence of microorganisms on coffee beans after antestia bug damage could be a potential pathway for PTD. This damage results in an overall decrease of analytes that are known to have positive sensory contributions to coffee aroma. Collectively, the volatile fingerprint shown herein illustrates that PTD alters the biochemical process in coffee beans.
Herein, two "orthogonal" characteristics of moisture damaged cacao beans (temporally dependent molding kinetics versus the time-independent geographical region of origin) are simultaneously analyzed in a comprehensive two-dimensional (2D) gas chromatography time-of-flight mass spectrometry (GC×GC-TOFMS) dataset using tile-based Fisher ratio (F-ratio) analysis. Cacao beans from six geographical regions were analyzed once a day for six days following the initiation of moisture damage to trigger the molding process. Thus, there are two "extremes" to the experimental sample class design: six time points for the molding kinetics versus the six geographical regions of origin, resulting in a 6 × 6 element signal array referred to as a composite chemical fingerprint (CCF) for each analyte. Usually, this study would involve initial generation of two separate hit lists using F-ratio analysis, one hit list from inputting the data with the six time point classes, then another hit list from inputting the dataset from the perspective of geographic region of origin. However, analysis of two separate hit lists with the intent to distill them down to one hit list is extremely time-consuming and fraught with shortcomings due to the challenges associated with attempting to match analytes across two hit lists. To address this challenge, tile-based F-ratio analysis is "orthogonally applied" to each analyte CCF to simultaneously determine two F-ratios at the chromatographic 2D location (F-ratiokinetic and F-ratioregion) for each hit, by ranking a single hit list using the higher of the two F-ratios resulting in the discovery of 591 analytes. Further, using a pseudo-null distribution approach, at the 99.9% threshold over 400 analytes were deemed suitable for PCA classification. Using a more stringent 99.999% threshold, over 100 analytes were explored more deeply using PARAFAC to provide a purified mass spectrum.
Chemometric decomposition methods like multivariate curve resolution-alternating least squares (MCR-ALS) are often employed in gas chromatography-mass spectrometry (GC-MS) to improve analyte identification and quantitation. However, these methods can perform poorly for analytes with a low chromatographic resolution (Rs) and a high degree of spectral contamination from noise and background interferences. Thus, we propose a novel computational algorithm, termed mzCompare, to improve analyte identification and quantitation when coupled to MCR-ALS. The mzCompare method utilizes an underlying requirement that the retention time and peak shape between mass channels (m/z) of the same analyte should be similar. By discovering the selective m/z for a given analyte in a chromatogram, a pure elution profile can be generated and used as an equality constraint in MCR-ALS. The performance of the mzCompare methodology is demonstrated with both experimental and simulated chromatograms. Experimentally, unresolved analytes with a Rs as low as 0.05 could be confidently identified with mzCompare assisted MCR-ALS. Furthermore, application of the mzCompare algorithm to a complex aerospace fuel resulted in the discovery of 335 analytes, a 44 % increase compared to conventional peak detection methods. GC-MS simulations of target-interferent analyte pairs demonstrated that the performance of MCR-ALS deteriorated below a Rs of ∼0.25. However, mzCompare assisted MCR-ALS showed excellent identification and acceptable quantitative accuracy at a Rs of ∼0.02. These results show that the mzCompare algorithm can help analysts overcome modeling ambiguities resulting from the chemometric multiplex disadvantage.
Comprehensive two-dimensional gas chromatography coupled with mass spectrometry (GC × GC–MS) is now an integral analytical technique for the characterization of volatile and semivolatile compounds. With the rise of this technique in various application fields, the size and complexity of data sets have also grown, making manual interpretation of GC × GC–MS data sets cumbersome and tedious. However, for these comparative analysis studies, chemometric methods can efficiently discover chemical differences between samples. This chapter discusses the recent developments in chemometric analysis of GC × GC–MS data. As the application of these computational approaches is dependent on the quality of the data, instrumentation and data preprocessing methods are outlined. Key principles of both unsupervised and supervised nontargeted methods and their application in a pixel-based, peak table-based, and tile-based fashion are also reviewed. Finally, this chapter explores the application of targeted methods during a comparative analysis workflow to improve analyte identification and quantification.
In this study, we introduce a new nontargeted tile-based supervised analysis method that combines the four-grid tiling scheme previously established for the Fisher ratio (F-ratio) analysis (FRA) with the estimation of tile hit importance using the machine learning (ML) algorithm Random Forest (RF). This approach is termed tile-based RF analysis. As opposed to the standard tile-based F-ratio analysis, the RF approach can be extended to the analysis of unbalanced data sets, i.e., different numbers of samples per class. Tile-based RF computes out-of-bag (oob) tile hit importance estimates for every summed chromatographic signal within each tile on a per-mass channel basis (m/z). These estimates are then used to rank tile hits in a descending order of importance. In the present investigation, the RF approach was applied for a two-class comparison of stool samples collected from omnivore (O) subjects and stored using two different storage conditions: liquid (Liq) and lyophilized (Lyo). Two final hit lists were generated using balanced (8 vs Eight comparison) and unbalanced (8 vs Nine comparison) data sets and compared to the hit list generated by the standard F-ratio analysis. Similar class-distinguishing analytes (p < 0.01) were discovered by both methods. However, while the FRA discovered a more comprehensive hit list (65 hits), the RF approach strictly discovered hits (31 hits for the balanced data set comparison and 29 hits for the unbalanced data set comparison) with concentration ratios, [OLiq]/[OLyo], greater than 2 (or less than 0.5). This difference is attributed to the more stringent feature selection process used by the RF algorithm. Moreover, our findings suggest that the RF approach is a promising method for identifying class-distinguishing analytes in settings characterized by both high between-class variance and high within-class variance, making it an advantageous method in the study of complex biological matrices.
We examine and then optimize alignment of chromatograms collected on nominally identical columns using retention time locking (RTL), an instrumental alignment tool, and software-based alignment using correlation optimized warping (COW). For this purpose, three samples are constructed by spiking two sets of analytes into a base test mixture. The three samples are analyzed by high-speed gas chromatography with four nominally identical columns and identical separation conditions. The data is first analyzed without alignment, then using COW alone, then RTL alone, and finally with RTL followed by COW to correct the severe column-to-column misalignment. Principal component analysis (PCA) is used to investigate how well each alignment method clustered the chromatograms into the three sample classes via a scores plot without being compromised by the specific column(s) used. The degree-of-class separation (DCS) is used as a classification metric, measured as the Euclidian distance between the centroids of two clusters in PC space in the scores plot, normalized by their pooled variance. With no alignment, the average DCS between sample classes (DCSsam) was 3.0, while the average DCS between the four nominally identical columns, i.e., column classes (DCScol) was 76.1 (ideally the DCScol should be 0), indicating the chromatograms were initially classified by the columns used. Using either COW or RTL alone also produced unsatisfactory results, with COW alone incorrectly aligning many peaks, leading to a DCSsam of only 1.9 and DCScol of 1.7, while RTL alone provided a DCSsam of 4.7 and DCScol of 4.2. Finally, using RTL followed by COW alignment, DCSsam increased to 32.5, indicating successful classification by chemical differences between sample classes, while the DCScol decreased to 0.4, indicating virtually no classification due to column-to-column differences, as desired. Thus, RTL provided a "first-order" correction of the initial retention mismatch observed for the nominally identical columns, while additional alignment via COW was required to optimize sample classification by PCA.
Nontargeted analyses of low-concentration analytes in the information-rich data collected by liquid chromatography with high-resolution mass spectrometry detection can be challenging to accomplish in an efficient and comprehensive manner. The aim of this study is to demonstrate a workflow involving targeted parameter optimization for entire chromatograms using region of interest (ROI) data compression uncoupled from a subsequent tile-based Fisher ratio (F-ratio) analysis, a supervised discovery-based method, for the discovery of low-concentration analytes. Soil samples spiked with 18 pesticides at nominal concentrations ranging from 0.1 to 50 ppb for a total of six sample classes served as challenging samples to demonstrate the overall workflow. Optimization of two parameters proved to be the most critical for ROI data compression: the signal threshold parameter and the admissible mass deviation parameter. The parameter optimization method workflow we introduce is based upon spiking known analytes into a representative sample and determining the number of detectable spikes and the Δppm for various combinations of the signal threshold and admissible mass deviation, where Δppm is the absolute value of the difference between the theoretical m/z and the ROI m/z. Once optimal parameters are determined providing the lowest average Δppm and the greatest number of detectable analytes, the optimized parameters can be utilized for the intended analysis. Herein, tile-based F-ratio analysis was performed on the ROI compressed data of all spiked soil samples first by applying ROI parameters recommended in the literature, referred to herein as the initial ROI parameters, and finally by the combination of the two optimized parameters. Using the initial ROI parameters, three pesticides were discovered, whereas all 18 spiked pesticides were discovered by optimizing both ROI parameters.
ADVERTISEMENT RETURN TO ISSUEPREVReviewNEXTRecent Advances in GC×GC and Chemometrics to Address Emerging Challenges in Nontargeted AnalysisTimothy J. TrinkleinTimothy J. TrinkleinDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United StatesMore by Timothy J. TrinkleinView Biographyhttps://orcid.org/0000-0003-3475-5981, Caitlin N. CainCaitlin N. CainDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United StatesMore by Caitlin N. CainView Biographyhttps://orcid.org/0000-0001-8367-5799, Grant S. OchoaGrant S. OchoaDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United StatesMore by Grant S. OchoaView Biography, Sonia SchöneichSonia SchöneichDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United StatesMore by Sonia SchöneichView Biography, Lina MikaliunaiteLina MikaliunaiteDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United StatesMore by Lina MikaliunaiteView Biographyhttps://orcid.org/0000-0002-8546-1927, and Robert E. Synovec*Robert E. SynovecDepartment of Chemistry, University of Washington, Box 351700, Seattle, Washington 98195-1700, United States*Phone: +1-206-685-2328. Fax: +1-206-685-8665. E-mail: [email protected]More by Robert E. SynovecView BiographyCite this: Anal. Chem. 2023, 95, 1, 264–286Publication Date (Web):January 10, 2023Publication History Received26 September 2022Published online10 January 2023Published inissue 10 January 2023https://doi.org/10.1021/acs.analchem.2c04235Copyright © 2023 American Chemical SocietyRIGHTS & PERMISSIONSArticle Views941Altmetric-Citations3LEARN ABOUT THESE METRICSArticle Views are the COUNTER-compliant sum of full text article downloads since November 2008 (both PDF and HTML) across all institutions and individuals. These metrics are regularly updated to reflect usage leading up to the last few days.Citations are the number of other articles citing this article, calculated by Crossref and updated daily. Find more information about Crossref citation counts.The Altmetric Attention Score is a quantitative measure of the attention that a research article has received online. Clicking on the donut icon will load a page at altmetric.com with additional details about the score and the social media presence for the given article. Find more information on the Altmetric Attention Score and how the score is calculated. Share Add toView InAdd Full Text with ReferenceAdd Description ExportRISCitationCitation and abstractCitation and referencesMore Options Share onFacebookTwitterWechatLinked InReddit Read OnlinePDF (16 MB) Get e-AlertscloseSUBJECTS:Beverages,Chemometrics,Chromatography,Mathematical methods,Software Get e-Alerts
Chemometric methods like partial least squares (PLS) regression are valuable for correlating sample-based differences hidden in comprehensive two-dimensional gas chromatography (GC x GC) data to indepen-dently measured physicochemical properties. Herein, this work establishes the first implementation of tile-based variance ranking as a selective data reduction methodology to improve PLS modeling perfor-mance of 58 diverse aerospace fuels. Tile-based variance ranking discovered a total of 521 analytes with a square of the relative standard deviation (RSD2) in signal between 0.07 to 22.84. The goodness-of-fit for the models were determined by their normalized root-mean-square error of cross-validation (NRMSECV) and normalized root-mean-square error of prediction (NRMSEP). PLS models developed for viscosity, hy-drogen content, and heat of combustion using all 521 features discovered by tile-based variance ranking had a respective NRMSECV (NRMSEP) equal to 10.5 % (10.2 %), 8.3 % (7.6 %), and 13.1 % (13.5 %). In con-trast, use of a single-grid binning scheme, a common data reduction strategy for PLS analysis, resulted in less accurate models for viscosity (NRMSECV = 14.2 %; NRMSEP = 14.3 %), hydrogen content (NRM-SECV = 12.1 %; NRMSEP = 11.0 %), and heat of combustion (NRMSECV = 14.4 %; NRMSEP = 13.6 %). Further, the features discovered by tile-based variance ranking can be optimized for each PLS model with RReliefF analysis, a machine learning algorithm. RReliefF feature optimization selected 48, 125, and 172 analytes out of the original 521 discovered by tile-based variance ranking to model viscosity, hydrogen content, and heat of combustion, respectively. The RReliefF optimized features developed highly accu-rate property-composition models for viscosity (NRMSECV = 7.9 %; NRMSEP = 5.8 %), hydrogen content (NRMSECV = 7.0 %; NRMSEP = 4.9 %), heat of combustion (NRMSECV = 7.9 %; NRMSEP = 8.4 %). This work also demonstrates that processing the chromatograms with a tile-based approach allows the analyst to directly identify the analytes of importance in a PLS model. Coupling tile-based feature selection with PLS analysis allows for deeper understanding in any property-composition study.(c) 2023 Elsevier B.V. All rights reserved.
Comprehensive three-dimensional (3D) gas chromatography with time-of-flight mass spectrometry (GC3-TOFMS) is a promising instrumental platform for the separation of volatiles and semi-volatiles due to its increased peak capacity and selectivity relative to comprehensive two-dimensional gas chromatography with TOFMS (GC×GC-TOFMS). Given the recent advances in GC3-TOFMS instrumentation, new data analysis methods are now required to analyze its complex data structure efficiently and effectively. This report highlights the development of a cuboid-based Fisher ratio (F-ratio) analysis for supervised, non-targeted studies. This approach builds upon the previously reported tile-based F-ratio software for GC×GC-TOFMS data. Cuboid-based F-ratio analysis is enabled by constructing 3D cuboids within the GC3-TOFMS chromatogram and calculating F-ratios for every cuboid on a per-mass channel basis. This methodology is evaluated using a GC3-TOFMS data set of jet fuel spiked with both non-native and native components. The neat and spiked jet fuels were collected on a total-transfer (100 % duty cycle) GC3-TOFMS instrument, employing thermal modulation between the first (1D) and second dimension (2D) columns and dynamic pressure gradient modulation between the 2D and third dimension (3D) columns. In total, cuboid-based F-ratio analysis discovered 32 spiked analytes in the top 50 hits at concentration ratios as low as 1.1. In contrast, tile-based F-ratio analysis of the corresponding GC×GC-TOFMS data only discovered 28 of the spiked analytes total, with only 25 of them in the top 50 hits. Along with discovering more analytes, cuboid-based F-ratio analysis of GC3-TOFMS data resulted in fewer false positives. The increased discoverability is due to the added peak capacity and selectivity provided by the 3D column with GC3-TOFMS resulting in improved chromatographic resolution.