
ABSTRACT ATR‐FTIR spectroscopy combined with chemometric modelling is increasingly used for asphalt binder classification, but model reliability can be overestimated when repeated spectra, related parent samples, and instrument/source differences are not controlled during validation. This study evaluates ATR‐FTIR–based asphalt binder classification models under sample‐level and cross‐source conditions using a multi‐source spectral database organized into measurement records, parent samples, and final split units. Binary classification of base/non‐SBS and SBS‐modified binders was used as the case task. Candidate preprocessing‐model pipelines were evaluated at the parent sample–level, and two cross‐source validation directions were used to test stability across source conditions. Random spectrum‐level splitting produced optimistic performance estimates. SNV+LinearSVM achieved the highest internal main BA (0.9482), but its minimum cross‐source BA decreased to 0.5621. In contrast, the selected SNV+PLS‐DA–style pipeline achieved comparable internal performance (main BA = 0.9428) while maintaining a higher minimum‐direction BA (0.9554) and a smaller direction gap BA (0.0237). Diagnostic experiments showed that cross‐source stability depends on both sufficient training sample support and adequate source representation. Latent‐variable analysis, coefficient/VIP spectra, and window masking indicated that the model used multiple spectral regions rather than a single SBS marker peak. Source‐predictability analysis further showed that source effects were attenuated but not eliminated. The results support a sample‐level and cross‐source evaluation strategy for reproducible ATR‐FTIR chemometric classification.
ABSTRACT During spectral detection operations, when the streaming samples under test differ significantly from the modeling samples, the prediction error of the spectral model tends to increase. In such cases, it becomes necessary to incorporate novel samples into the modeling set for model updating. A common approach to assess sample novelty involves calculating the spectral residual, typically using the Q residual. However, existing methods generally employ a single distance scalar to evaluate the novelty of multidimensional spectra. This approach suffers from the issue that residuals at different positions are aggregated into a single value, which is overly integrative. Consequently, variations at different locations may lead to the same novelty value. Therefore, a window‐based method that evaluates data in segments is more reasonable. The proposed method segments the residual dimensionality using a windowing approach, discretizes the residual values, and classifies and counts the patterns of residual variations within each window. The slope entropy derived from the residual is then used to represent sample novelty. The effectiveness of the proposed method is demonstrated through both simulated and real‐world data.
ABSTRACT Accurate molecular identification from mass spectra is essential in analytical workflows, yet conventional library search typically returns the closest match even when the true compound is absent, leading to overconfident false positives. We propose a probabilistic framework that supports both reliable identification of compounds represented in a reference library and model‐based flagging of spectra that may not be adequately represented in the reference library. Spectra are modeled within a Bayesian nonparametric method that does not predefine the number of clusters; instead, the model can allocate new clusters when incoming spectra are insufficiently explained by existing references. This property provides a statistical mechanism for flagging potentially unseen compounds while maintaining coherent grouping of known ones. Experiments on large‐scale electron ionization libraries demonstrate stable performance under diverse noise conditions, consistent grouping of replicate spectra, and the tendency of spectra absent from the reference database to form new clusters. The framework supports reliable identification while reducing forced matches for spectra that are poorly represented in the reference data.
ABSTRACT The discretized snake optimization algorithm was first proposed as a variable selection method to reduce irrelevant variables and enhance the prediction accuracy of complex samples. In discretized snake optimization (SO), the positions of the snakes were updated, and three transfer functions, V‐shaped, arctangent and sigmoid functions, were introduced and compared for discretization. The partial least squares (PLS) model was built using the spectral variables selected by discretized SO. The performance of snake population, transfer functions in SO, and distribution of selected variables for different methods are investigated. To verify the feasibility of SO‐PLS, the predictive accuracy of SO‐PLS was compared with full‐spectrum PLS, uninformative variable elimination‐PLS (UVE‐PLS), Monte Carlo UVE‐PLS (MCUVE‐PLS), randomization test‐PLS (RT‐PLS), grey wolf optimizer‐PLS (GWO‐PLS), and whale optimization algorithm‐PLS (WOA‐PLS) models on four complex sample datasets. The results indicate that the V‐shaped function is the best transfer function. Compared with the other variable selection methods, SO‐PLS uses the least number of variables and gets the best prediction accuracy. Furthermore, compared to PLS, the SO‐PLS model reduced the root mean squared error of prediction (RMSEP) by 52%, 43%, 38%, and 23% for predicting protein, sugar, alcohol, and fat in wheat, orange juice, wine, and cocoa bean dataset, respectively. The conresponding correlation coefficients ( R ) increased from 0.8942, 0.7375, 0.9984, and 0.8121 to 0.9782, 0.8935, 0.9996, and 0.8920, respectively.
ABSTRACT In multi‐block data, the dominant sources of variation are not always most relevant to a response of interest, meaning that purely exploratory decompositions may fail to recover subtle but important response‐associated structure. We introduce PESCAR, a supervised extension of Penalised Exponential Simultaneous Component Analysis (PESCA) that incorporates response information directly into the estimation of common, local and distinct (CLD) structure across multiple data blocks. This allows simultaneous multiblock decomposition and response variable‐influenced recovery of latent structure. Through simulation studies, we show that PESCAR can detect weak response‐related components across a range of settings, including different noise levels and model‐rank mis‐specification. Applied to a real multi‐omics dataset, PESCAR recovers biologically meaningful response‐associated patterns and retains interpretable block structure. We further demonstrate that sparsity in the fitted loading matrices admits a hypergraph‐based interpretability layer, summarising overlapping support patterns across components and blocks. These results show that direct incorporation of response information into multiblock decomposition can improve detection of subtle relevant signal and facilitate interpretation in complex systems.
Atmospheric pollution presents a significant global health burden, yet the establishment and maintenance of extensive 'gold standard' monitoring networks remain prohibitively expensive. We explore the potential of machine learning to create virtual monitoring sites for predicting and reconstructing urban air pollutant concentrations, offering a cost-effective alternative. Partial least squares (PLS) and support vector regression (SVR) models were developed and validated for daily concentrations of nitrogen dioxide (NO2) and particulate matter (PM2.5 and PM10) within a well-monitored central London area. Our methodology involved comprehensive data pre-processing, including normalization, and rigorous model validation using RMSE, MSE, MAE, bias and R 2 metrics. A key finding was the superior predictive performance of PLS models incorporating meteorological data alongside pollutant concentrations, compared to models including land-use data. Specifically, these PLS models exhibited lower error metrics across all pollutants, with NO2 and PM2.5 predictions showing R 2 = 0.891 and R 2 = 0.688, respectively, in the transformed space. Upon back-transformation to original units (mu g/m3), the PLS models maintained high performance, particularly for PM2.5 (RMSEP = 2.261, R 2 = 0.796). Whereas SVR models occasionally showed slightly higher R 2 values in the transformed space, their error metrics were consistently larger than those of PLS models, especially for PM2.5. This research underscores the effectiveness of PLS for accurate urban air pollutant prediction and highlights the critical influence of meteorological data over land-use data in densely built-up areas. The Air-01 London dataset, a compilation of air pollutant, meteorological and land-use data created for this study, has been made openly available to facilitate further research.
Near-infrared spectroscopy (NIRS) is a non-destructive analytical technique. It measures absorption signals generated by overtone and combination vibrations of chemical bonds in hydrogen-containing functional groups. Owing to its rapid and efficient analytical capability, NIRS has been widely used for quality evaluation of agricultural products and traditional Chinese medicinal materials. However, NIRS is often limited by weak absorption intensity and severe spectral band overlap. NIRS data are usually high-dimensional, highly collinear, and information-redundant. These features increase the difficulty of model development and may reduce analytical accuracy. This study developed a dynamic entropy-guided Mixup data augmentation method to enhance feature representation and improve classification accuracy in small-sample near-infrared spectral analysis. The proposed method was further applied to the geographical origin identification of Euryales Semen (ES, Euryale ferox Salisb.). Information entropy was used to quantify sample uncertainty and identify high-uncertainty samples near classification boundaries. Based on these samples, a targeted Mixup data augmentation strategy was developed and combined with deep learning classifiers to improve the learning capacity and generalization performance of models for small-sample spectral data. The results demonstrate that the overall performance of three deep learning models, namely multilayer perceptron (MLP), one-dimensional convolutional neural network (1D-CNN), and long short-term memory (LSTM), was improved to varying degrees after applying data augmentation strategies for the geographical origin classification of ES. Among them, the proposed dynamic entropy-guided Mixup (DEG-Mixup) method achieved the best performance, with test accuracies of 86.8%, 83.27%, and 83.24% for MLP, 1D-CNN, and LSTM, respectively. Compared with models without augmentation, the accuracies were improved by 10.26%, 10.88%, and 12.47%, respectively. Compared with conventional Mixup, the improvements were 7.20%, 7.79%, and 8.46%, respectively. These results indicate that DEG-Mixup effectively enhances the discriminative ability and generalization performance of models for small-sample spectral data, enabling accurate geographical origin identification of ES. This study established a rapid, effective, and reliable near-infrared spectral classification strategy based on entropy-guided data augmentation. The proposed strategy provides a new methodological reference for the geographical origin identification and quality evaluation of traditional Chinese medicinal materials under small-sample conditions, and shows promising potential for broader spectral analysis tasks.
Virtual reality has become a powerful tool for analyzing complex chemical data, allowing researchers to make better decisions than automated algorithms alone. However, these systems rely almost entirely on vision, excluding blind users, some of whom demonstrate strong abilities in spatial reasoning and pattern recognition, though these abilities vary considerably across individuals. We developed a haptic-first VR system that makes immersive data analysis accessible to blind users through touch-based feedback via haptic gloves combined with spatial audio cues. The system translates visual data representations into tactile experiences where users can physically explore and manipulate data objects. We encoded hundreds of sample similarity measures into glyph features including vibration patterns, force feedback, and musical tones, enabling users to perform sophisticated analytical tasks like detecting outliers, identifying clusters, and classifying samples. Our pilot feasibility study with four blind university students explored a real chemical dataset with known structure, demonstrating that users could engage with and reason about chemometric data through haptic and audio modalities. Audio feedback proved consistently effective across all users, while haptic feedback varied in usefulness depending on the specific tactile channel and individual preferences. Navigation emerged as the main challenge, although participants demonstrated genuine analytical engagement and could articulate their decision-making process. This work provides proof-of-concept evidence and design insights indicating that complex chemical analysis can be explored through coordinated touch and sound, identifying navigation and calibration challenges that must be resolved before practical deployment.
The relationship between the Mahalanobis distance and the χ 2 and Hotelling T 2 distributions is described. The methods are illustrated by two case studies, a 40 × 2 simulation and a 54 × 34 dataset consisting of the NMR of metabolic extracts from maize harvested at 8.5°C. It is shown that if the number of samples is close to but still greater than the number of variables, p values calculated using T 2 as usually defined are higher than should be normally expected. This is interpreted as a problem caused by the difficulty of determining the number of degrees of freedom in multivariate matrices. It is recommended to use the χ 2 distribution in preference when calculating multivariate p values for probabilities that a sample belongs to a predefined distribution. When the sample size is large relative to the number of variables, there is little point in using T 2 , and when small, T 2 can predict misleading p values due to problems determining the number of degrees of freedom.
Near-infrared (NIR) spectra exhibit strong local collinearity and smooth absorption patterns, requiring variable selection to mitigate overfitting in multivariate calibration. Point-wise methods yield scattered subsets that do not account for spectral ordering, whereas fixed-window approaches lack the adaptability to capture the varying widths of physical absorption bands. To address these limitations, we propose Fused-BiPLS, a two-stage wavelength-interval selection framework that decouples spectral segmentation from predictive coefficient estimation. First, a fused lasso penalty is applied to adjacent coefficient differences and is solved along its exact solution path. Guided by Mallows' criterion, this yields a piecewise-constant coefficient profile that forms data-driven, contiguous segments, replacing manually defined interval widths. Second, backward interval partial least squares (BiPLS) identifies a subset of these intervals without further shrinkage. Evaluated across six benchmark NIR datasets spanning agricultural, petrochemical, biological, and pharmaceutical materials, Fused-BiPLS achieved prediction errors comparable to or lower than those of established chemometric models. The method yielded medium to large positive effect sizes () against most baselines on the Diesel, Soil, and Meat datasets. The adaptive segmentation isolated contiguous blocks corresponding to known molecular overtones, demonstrating stability across sampling variations. Although the exact solution path incurs a higher off-line calibration cost, the resulting interval-based models are computationally efficient during inference. Because the selected contiguous intervals align directly with the bandwidth specifications of optical filters, the framework supports the configuration of compact, portable NIR devices.
Gaussian mixture regression (GMR) is a useful framework for both forward prediction and direct inverse analysis because it models the joint probability distribution of input variables x and output variables y. However, when full covariance matrices are used, the number of fitting parameters increases rapidly with dimensionality, making estimation unstable when only a limited number of paired (x, y) samples are available. In this study, two semi-supervised extensions of GMR are proposed to exploit abundant unlabeled x data. The first method, semi-supervised GMR with x-only GMM pretraining (ssGMR-xGMM), builds a Gaussian mixture model (GMM) using only x samples and uses the obtained parameters to initialize the joint GMM for GMR. The second method, semi-supervised GMR with x-only GMM pretraining and x-mean anchoring (ssGMR-xGMM-xMA), further fixes the x-side mean vectors during training to preserve the mixture structure learned from unlabeled x data. The proposed methods were evaluated using numerical simulation data and real molecular and spectral datasets, including boiling point, aqueous solubility, pharmacological activity, environmental toxicity, and tablet spectral datasets. Compared with conventional GMR, both proposed methods improved predictive performance, and ssGMR-xGMM-xMA showed the best overall performance. Under the transductive semi-supervised setting examined in this study, these results indicate that leveraging unlabeled x data can be effective for stabilizing parameter estimation and improving GMR accuracy in data-scarce settings.
In classification tasks, predicted class probabilities are often overconfident for samples outside the training distribution, which can lead to unsafe decision-making. kNNPC, a k-nearest neighbor algorithm per class, is proposed to mitigate such overconfidence under extrapolation or distribution shift. For a given test sample, kNNPC adjusts the predicted class-probability vector so that it becomes closer to a maximum-uncertainty distribution when the sample is dissimilar to the training data in the feature space. Importantly, kNNPC is not intended to identify extrapolation regions explicitly; rather, it provides a conservative probability output that serves as a practical safeguard against overconfident predictions. kNNPC was evaluated on benchmark and real-world datasets, including the Iris dataset, a toxicity classification dataset, and a superconducting critical-temperature dataset, and compares with standard probabilistic classifiers. The results indicate that kNNPC can effectively suppress overconfident probabilities for out-of-domain samples while maintaining comparable predictive performance. The Python code for the proposed method is available at https://github.com/hkaneko1985/dcekit.
Precise and fast estimation of 1,8-cineole levels in large cardamom is essential for proper quality assessment and improving value chain efficiency in the essential oil sector. This work introduces an adaptive-band convolutional transformer neural network (AdBand-CTNet) designed to forecast 1,8-cineole content using near-infrared (NIR) spectral measurements. The approach incorporates a differentiable adaptive band selection (ABS) module combined with an attention scheme, allowing the model to learn the most informative wavelength zones and the regression parameters simultaneously in a fully integrated training process. In contrast to traditional feature selection strategies or fixed-width spectral partitioning, the proposed architecture dynamically optimizes both the center and span of every band through gradient updates, producing concise and interpretable spectral features. Samples from seven cultivars of large cardamom collected across multiple agroclimatic areas of West Bengal and Sikkim, India, were studied. Reference values for 1,8-cineole were determined using gas chromatography-mass spectrometry (GC-MS), and NIR spectra were captured between 900 and 1700 nm. Over 100 randomized trials, the AdBand-CTNet reached a mean absolute error of 0.0633, a root mean squared error of 0.1301, and an R 2 of 0.9969 on the test data, demonstrating its outstanding precision and robustness. The introduced method removes unnecessary wavelength segments while still maintaining excellent prediction capability. Its flexible and explainable architecture makes it well-suited for industrial scenarios that require trustworthy and accurate chemical composition assessment, offering a more economical option compared with standard chromatographic methods.
Traditional electroanalytical methods remain a cornerstone in electrochemical biosensor optimizations. In the particular case of chronoamperometry, scalability, cost, and time-consuming practices challenge to a certain extent its use in multivariate biosensor design workflows. In this study, we demonstrate that spectrophotometric enzymatic assays can be employed as rapid and reliable indicators of the electrochemical response of glucose biosensors, significantly streamlining statistical optimization processes. To this end, second-generation glucose biosensors were fabricated by co-depositing flavin adenine dinucleotide-dependent glucose dehydrogenase (FAD-GDH) and polydopamine onto graphite electrodes. A two-stage Design of Experiments (DoE) approach was implemented to systematically evaluate the effects of dopamine and enzyme concentrations. During the exploration step, UV-Vis measurements of ferricyanide reduction kinetics (ferricyanide being a redox mediator) were used as a proxy for the enzymatic activity, which was compared with the electrochemical sensitivity values obtained via chronoamperometry in the optimization step. A strong correlation between both responses was found, as evidenced by similar predictions of the optimal sensor fabrication conditions, confirming that simple spectrophotometric enzymatic assays can reliably reflect the trends in the biosensor performance. These findings suggest that spectrophotometric measurements offer a rapid, cost-effective and parallelizable alternative to traditional amperometric workflows, particularly during early stages in multivariate optimizations, implying minimal experimental workload.
This study presents a comprehensive evaluation of various computational models for glioblastoma cell classification using Raman spectroscopy data. We compare traditional machine learning methods such as support vector machines, boosting and random forests with more modern methods including convolutional neural networks, vision transformers and the broad learning system. We find that convolutional neural networks are the most successful model for this data set and perform best when some feature normalisation is performed but without preprocessing methods such as background drift removal. We also find that data augmentation did not improve performance which is contrary to other published work in the area.
Electronic-nose screening models degrade after deployment because sensor drift and environmental variability change both separability and confidence calibration, yet field datasets rarely allow controlled severity analysis. We study calibration-aware selective classification on a synthetic 16-channel MOX benchmark that parameterizes humidity, temperature, aging, device variability, and noise to create mild, moderate, and severe target domains, and includes calibration anchors. The proposed CAN-OCS pipeline combines anchor normalization, nuisance suppression, CORAL covariance alignment using unlabeled target batches, and an uncertainty gate that defers low-confidence samples to confirmatory testing while reporting coverage. At matched coverage, CAN-OCS outperformed a standardized SVM baseline with severity-dependent gains (Holm-adjusted p = 0.0146 in the severe domain; borderline in moderate; not significant in mild). On a public one-year drift dataset (24 temporal splits), fixed thresholding caused strong coverage contraction (median 0.439). Results emphasize that accepted-sample accuracy must be interpreted jointly with coverage and anchor validity.
Understanding the biochemical basis of bacterial inhibition from spectral signatures requires modeling approaches capable of handling high-dimensional, collinear, and multiresponse data. In this study, we proposed a novel hybrid framework that integrates partial least squares regression for multiple responses (PLS2) within the multivariate adaptive regression splines (MARS) model, referred to as the MARS-PLS2 approach. Unlike conventional MARS, which applies linear regression at terminal nodes, the proposed framework leverages PLS2 to simultaneously model five inhibition responses (Escherichia coli, Bacillus subtilis, MR-Staphylococcus aureus, Klebsiella pneumoniae, and S. aureus), enabling an efficient representation of shared spectral-biochemical variations. A sparsity-driven feature selection strategy identifies a subset of informative wavenumbers associated with key vibrational modes, including C-H stretching and bending, O-H, C-N, C=C, and N-H stretching, which are biochemically relevant to membrane dynamics, protein interactions, and cellular response mechanisms. The proposed MARS-PLS2 framework substantially enhances prediction accuracy and model interpretability. For example, for E. coli, the MARS-PLS2 model achieves a high , with a significantly lower RMSE of 1.655 and MAE of 1.201, compared with the standard PLS2 model (, RMSE = 6.094, MAE = 5.655) and standard MARS model (, RMSE = 18.508, MAE = 10.412). Similarly, for B. subtilis, the MARS-PLS2 framework yields , RMSE = 1.466, and MAE = 1.124, considerably outperforming the standard PLS2 results (, RMSE = 5.383, MAE = 4.071) and standard MARS (, RMSE = 13.745, MAE = 9.130). Furthermore, validation using permutation testing and bootstrap analysis confirms the robustness and stability of the proposed model, with consistently high values and low prediction errors across resampled datasets, indicating that the observed performance is not due to random variation. Overall, the proposed MARS-PLS2 framework provides a robust, interpretable, and accurate approach for linking spectral features to bacterial inhibition responses, offering both improved predictive performance and meaningful biochemical insight. The framework can be extended to broader chemometric applications, including rapid diagnostics and antimicrobial research.
Ensuring authenticity of edible oils is vital for consumer trust, regulatory compliance and permissibility assurance, particularly for Muslim consumers. This study introduces a spectral biosensing strategy using Fourier transform infrared spectroscopy coupled with attenuated total reflectance (FTIR-ATR) and chemometric modelling to detect lard adulteration in palm oil (PO) under thermal stress. PO samples spiked with 1%-50% v/v lard were heated from 25 degrees C to 200 degrees C for 30 min, simulating industrial and culinary conditions. FTIR-ATR spectra showed distinct shifts in carbonyl and fingerprint regions due to lard incorporation and heat-induced lipid degradation. Discriminant analysis (DA) achieved 100% classification accuracy across the full spectrum (4000-650 cm-1), demonstrating strong discriminatory power. Partial least squares-discriminant analysis (PLS-DA) identified the fingerprint region (1000-650 cm-1) as most diagnostic, yielding robust performance with R 2 Y = 0.895, R 2 X = 1.000, Q 2 = 0.893 and 100% correct classification in both training and validation datasets. Principal component analysis (PCA) revealed clear clustering of pure and adulterated samples, even under severe thermal conditions. Moreover, using the optimised FTIR-ATR/PLS model, the lard adulteration in thermally treated PO could be reliably detected at levels as low as 1% v/v with LOD and LOQ ranges of 0.01%-1.14% and 0.02%-3.34% v/v, respectively. These findings position FTIR-ATR with multivariate chemometrics as a rapid, nondestructive and thermally resilient platform for lard detection in PO. The approach extends to broader food quality and safety monitoring in real-world processing scenarios.