Rice is one of the world's most important staple crops, but its large varietal diversity makes authenticity control and fraud prevention difficult. This study investigated whether near-infrared hyperspectral imaging (NIR-HSI) combined with chemometrics could support rice variety discrimination through morphological and chemical information. A total of 47 rice varieties, 36 of which Italian and 11 from different foreign countries, were analysed, with 115 grains per variety, and the resulting image data were explored by principal component analysis and classified using partial least squares discriminant analysis (PLS-DA) and hierarchical modelling. Morphological features gave the best performance: the five grain-shape classes were classified with 94.1 % NER and up to 96.0 % class accuracy, while hierarchical morphology-based models further improved discrimination, separating 11-12 groups with accuracies above 93 % and reaching 99.6 % in the best case. NIR spectra were less informative for variety discrimination: amylose-based PLS-DA achieved 86.7 % NER and 91.0 % accuracy, whereas hierarchical spectral models showed high specificity but low sensitivities for several classes. Data fusion of morphological and spectral variables did not improve the predictive performance beyond morphology alone, suggesting limited complementarity between the two information sources. Overall, the results demonstrate that NIR-hyperspectral imaging is a rapid, non-destructive, and promising approach for rice authentication, with morphological descriptors providing the most robust basis for classification. These findings support the use of chemometric imaging strategies to strengthen quality control and help prevent food fraud in the rice supply chain.
Background: Liquid chromatography-ion mobility-tandem mass spectrometry (LC-IMS-MS/MS) enables multidimensional characterization of complex lipid mixtures. However, the resulting high-dimensional datasets pose significant challenges for reliable component resolution and interpretation. In this work, the Regions-Of-Interest-based Multivariate Curve Resolution strategy (ROIMCR) is proposed for the integrated analysis of LC-IMS-MS/MS data. It explicitly preserves ion mobility information in its native profile form rather than reducing it to a single collision cross section (CCS) value. The proposed approach simultaneously resolves chromatographic elution profiles, ion mobility profiles, and mass spectra profiles (MS1 and MS2). Bilinear or trilinear models are used to enforce chemically consistent solutions across the elution time, m/z, and ion mobility dimensions, yielding interpretable component profiles while maintaining satisfactory model performance. Results: This study has proven the applicability and chemical interpretability of the trilinear model in resolving complex lipid mixtures using LC-IMS-MS/MS three-way data. Compound annotation and identification were achieved through direct comparison of the ROIMCR-resolved retention time, arrival time, and tandem mass spectra with those of authentic standards or reference data libraries. The applicability of the method was demonstrated using two data examples. The first one is used to validate the proposed approach using a simulated mixture of four analytes. The second example shows the application of the proposed method in the LC-IMS-MS/MS experimental analysis of a complex mixture of lipids. Overall, the proposed three-way MCR framework provides an accurate, interpretable, and scalable solution for high-dimensional LC-IMS-MS/MS data analysis.
While several techniques exist to study the secondary structure of proteins, mid-IR and electronic circular dichroism spectroscopy are complementary techniques that stand out for liquid measurements due to their non-destructive nature, broad applicability and ease of use. CD spectroscopy however is only used to measure low protein concentrations and suffers from signal interferences from common buffers. With advances in laser-technology, the achievable pathlengths used in IR-spectroscopy have increased with the use of external cavity quantum cascade lasers (EC-QCL) as light sources. This has not only enabled protein denaturation to be studied without aggregates clogging the measurement cell, as was the case with shorter pathlengths, but also expanded the concentration range that can be measured with the technique. Yet, the concentration range where both IR and CD measurements can be made has only a small overlap.In this study, the IR and CD spectra obtained during the thermal denaturation of α-chymotrpysin (α-CT) in the overlapping concentration range of the two spectroscopy techniques, between 10mgmL-1 to 40mgmL-1, are used to extract more information about the denaturation process of the protein and its concentration dependence. To this end, the protein denaturation spectra from both methods between 20°C to 80°C are fused and analyzed using multivariate curve resolution - alternating least squares (MCR-ALS), a bilinear unmixing technique that provides thermal profiles and pure fingerprints of the protein conformations involved in the denaturation process. However, in these measurements, a large and information-dense part of the CD spectra found at the lowest UV wavelengths cannot be used due to the saturated signal linked to high protein concentration levels. A modified version of MCR-ALS is hence introduced to overcome this issue, which allows all available information to be used without sacrificing the quality of the fit. The prowess of the adapted MCR-ALS technique as a tool for data fusion was demonstrated by not only providing a better fit than the individual models (CD and IR spectra analyzed separately), but also by aiding in squeezing out information about a third protein conformation that forms during the denaturation of α-CT.
Acute mesenteric ischemia (AMI) is a life-threatening, rapidly progressive disease. Conventional diagnostic methods have their own limitations. Therefore, a non-invasive, real-time and efficient diagnostic method is needed. Hyperspectral imaging (HSI) has the potential to differentiate between normal and necrotic intestinal tissue as a non-invasive, real-time imaging modality. In this study, hyperspectral images of intestinal rabbit tissue were analyzed over time with multivariate curve resolution (MCR) to identify and visualize the evolution of necrotic segments. VIS-NIR reflectance spectral data obtained by HSI showed characteristic peaks that distinguish normal and necrotic tissues. Ultimately, three-dimensional data, spectral, spatial, and temporal, were simultaneously analyzed in this study. The results demonstrated that combining MCR with HSI can accurately detect intestinal necrosis and analyze its progression over time. This study provides a robust foundation for improving AMI diagnosis and offers a new perspective for non-invasive evaluation of intestinal viability, with significant implications for clinical decision-making and patient outcomes.
Hyperspectral imaging (HSI) is a very complete analytical measurement that encloses rich spatial and chemical information. This double side enables HSI to outperform classical spectroscopic measurements and vision systems based only on color information. However, HSI requires powerful data analysis tools for interpretation and to facilitate its implementation in process analytical technology (PAT) contexts. This work is a brief perspective that shows a diversity of PAT challenges that can be solved with the combined use of HSI and dedicated chemometric procedures.
Mid-infrared (IR) spectroscopy and electronic circular dichroism (ECD) spectroscopy are well-established, complementary techniques that can be used to study protein structure and denaturation. Attempts have been made to combine them in the past using chemometrics. In this work we try to combine the techniques using multivariate curve resolution- alternating least squares (MCR-ALS). To this end, we simulate IR and ECD spectra to first demonstrate the effect of concentration on these spectra and subsequently analyze them, first separately and then together, using MCR-ALS. The results show that a modified version of MCR-ALS, capable of dealing with missing data while still being able to use the benefits of having information from two different techniques would be the correct choice for the given problem, especially when layers of complexity are added with measured data in the future.
Multiscale and multimodal image fusion is a challenge derived from the diversity of chemical and spatial information provided by the current hyperspectral image platforms. Efficient image fusion approaches are essential to exploit the complementary chemical information across different zoom scales. Most current image fusion algorithms tend to work by equalizing the spatial characteristics of the platforms to be combined, i.e., downsampling pixel size and cropping noncommon scanned sample areas if required. In this work, a new image unmixing algorithm based on a flexible mathematical framework is proposed to enable working with all available image information while preserving the original spatial properties of every imaging measurement. The algorithm is tested on a challenging image fusion scenario of fluorescence and Raman images collected on labeled HeLa cells. The system is relevant from an analytical point of view, since smart fluorescence labeling allows profiting from the excellent morphological information without causing interferences in the rich chemical information furnished by Raman. From a data handling perspective, it offers a challenging multiscale problem, where the fast fluorescence imaging acquisition allows recording full cell images, and the slower Raman image acquisition is focused on scanning only relevant small regions of the cells analyzed. By applying the image fusion algorithm proposed, an improved morphological and chemical characterization of cell constituents in the full cell area is obtained despite the different spatial scales used in the original imaging measurements.
Nanofiltration (NF) membranes are essential in wastewater treatment, battery industries, and brine management for selectively removing multivalent ions. However, fouling reduces their lifespan and necessitates harsh cleaning. The layer-by-layer (LBL) technique addresses this by modifying surface properties, enhancing rejection of divalent cations, such as magnesium and calcium, and minimizing fouling. This study evaluated a semiaromatic-based polyamide NF membrane (Fortilife-XN) modified using a LBL technique. Surface properties such as contact angle (CA), roughness, morphology, uniformity, and thickness were analyzed before and after modification. FTIR characterization revealed the membrane's structure, comprising a polyethylene terephthalate (PET) support, a polysulfone (PS) substrate, and a polyamide (PA) active surface. Poly-(sodium 4-styrenesulfonate) (PSS) was used as the polyanion, while poly-(diallyldimethylammonium chloride) (PDADMAC) and poly-(allylamine hydrochloride) (PAH) served as strong and weak polycations, respectively. Modifications in varying bilayers (1.5-6.5 BLs) with a positive terminal half-layer introduced peaks at 1034 and 923 cm-1, corresponding to sulfonate and C-N bonds. CA and roughness analysis showed that (PDADMAC/PSS)-4.5 (CA 20°, roughness 77 nm) and (PAH/PSS)-1.5 (CA 61°, roughness 79 nm) achieved superior wettability and roughness, confirmed by initial NF testing with a permeability of around 10 L/(m2 h bar). Ellipsometry, using Cauchy and Sellmeier models, measured multilayer thickness, estimating bilayers at ∼2 nm. Raman imaging visualized cross-sectional modifications, distinguishing raw and modified layers. Surface imaging revealed more uniform PAH deposition, while PDADMAC showed higher swelling with increased layers due to stronger affinity. These combined analytical techniques provided insights into the impact of LBL modification on membrane morphology and properties, aiding performance optimization. All membranes were preliminarily tested with a mixed salt solution (NaCl, CaSO4, and MgSO4). The average selectivity of Na/Ca for the bare membrane was 2.6 ± 0.4, increasing by 108% for (PDADMAC/PSS)-5.5 and 134% for PAH 6.5. For Na/Mg, the selectivity was 4.5 ± 0.3, rising by 49% for (PDADMAC/PSS)-5.5 and 131% for PAH 6.5.
Dealing with missing data poses a challenge in Principal Component Analysis (PCA) since the most common algorithms are not designed to handle them. Several approaches have been proposed to solve the missing value problem in PCA, such as Imputation based on SVD (I-SVD), where missing entries are filled by imputation and updated in every iteration until convergence of the PCA model, and the adaptation of the Nonlinear Iterative Partial Least Squares (NIPALS) algorithm, able to work skipping the missing entries during the least-squares estimation of scores and loadings. However, some limitations have been reported for both approaches. On the one hand, convergence of the I-SVD algorithm can be very slow for data sets with a high percentage of missing data. On the other hand, the orthogonality properties among scores and loadings might be lost when using NIPALS.To solve these issues and perform PCA of data sets with missing values without the need of imputation steps, a novel algorithm called Orthogonalized-Alternating Least Squares (O-ALS) is proposed. The O-ALS algorithm is an alternating least-squares algorithm that estimates the scores and loadings subject to the Gram-Schmidt orthogonalization constraint. The way to estimate scores and loadings is adapted to work only with the available information.In this study, the performance of O-ALS is tested and compared with NIPALS and I-SVD in simulated data sets and in a real case study. The results show that O-ALS is an accurate and fast algorithm to analyze data with any percentage and distribution pattern of missing entries, being able to provide correct scores and loadings in cases where I-SVD and NIPALS do not perform satisfactorily.
Multivariate Curve Resolution (MCR) deals with the mixture analysis problem by decomposing a data set with mixed information into a bilinear model of pure component contributions. Multiset analysis deals with fused data blocks linked to related experiments and/or techniques. Nevertheless, experiments and techniques often show differences that lead, when concatenated, to incomplete multisets with missing blocks of information. Incomplete multisets aim at incorporating all available information in the initial blocks of measurements but require adapted algorithms to be properly handled. This work presents the evolution of the different perspectives adopted to analyze incomplete multisets with advantages and drawbacks. Finally, a new methodology is proposed that adapts to any data configuration with missing entries without the need to perform data imputation or multiple factorizations. The new method adapts very well to analytical applications where the blocks of information to be fused are not acquired in equivalent experimental conditions.
The goal of native mass spectrometry is to obtain information on non-covalent interactions in solution through mass spectrometry measurements in the gas phase. Characterizing intramolecular folding re-quires using structural probing techniques such as ion mobility spectrometry. However, inferring solu-tion structures of nucleic acids is difficult because the low-charge state ions produced from aqueous solutions at physiological ionic strength get compacted during electrospray. Here we explored whether native supercharging could produce higher charge states that would better reflect solution folding, and whether the voltage required for collision-induced unfolding (CIU) could reflect preserved intramolecu-lar hydrogen bonds. We studied pH-responsive i-motif structures with different loops, and unstructured controls. We also implemented a multivariate curve resolution procedure to extract physically meaning-ful pure components from the CIU data and reconstruct unfolding curves. We found that the relative unfolding voltages reflect to some extent, but not always unambiguously, the number of intramolecular hydrogen bonds that were present in solution. Reaching phosphate charging densities over 0.25 makes it easier to discriminate between structures, and the use of native supercharging agents is thus essential. We also uncovered several caveats in data interpretation: (1) when different structures (for example the i-motif with and without hairpin) unfold via different pathways, the unfolding voltages do not necessari-ly reflect the number of hydrogen bonds, (2) unstructured controls also undergo unfolding, and the base composition influences the unfolding voltage, (3) changing the solution pH also unexpectedly changed the unfolding voltage, and (4) the ion mobility patterns become more complicated when two structures are present simultaneously, such as an i-motif and a harpin, because of opposite effects on the collision cross section upon activation.
Trilinearity is a property of some chemical data that leads to unique decompositions when curve resolution or multiway decomposition methods are used. Curve resolution algorithms, such as Multivariate Curve Resolution-Alternating Least Squares (MCR-ALS), can provide trilinear models by implementing the trilinearity condition as a constraint. However, some trilinear analytical measurements, such as excitation-emission matrix (EEM) measurements, usually exhibit systematic patterns of missing data due to the nature of the technique, which imply a challenge to the classical implementation of the trilinearity constraint. In this instance, extrapolation or imputation methodologies may not provide optimal results. Recently, a novel algorithmic strategy to constrain trilinearity in MCR-ALS in the presence of missing data was developed. This strategy relies on the sequential imposition of a classical trilinearity restriction on different submatrices of the original investigated dataset, but, although effective, was found to be particularly slow and requires a proper submatrix selection criterion. In this paper, a much simpler implementation of the trilinearity constraint in MCR-ALS capable of handling systematic patterns of missing data and based on the principles of the Nonlinear Iterative Partial Least Squares (NIPALS) algorithm is proposed. This novel approach preserves the trilinearity of the retrieved component profiles without requiring data imputation or subset selection steps and, as with all other constraints designed for MCR-ALS, offers the flexibility to be applied component-wise or data block-wise, providing hybrid bilinear/trilinear models. Furthermore, it can be easily extended to cope with any trilinear or higher-order dataset with whatever pattern of missing values.
Hyperspectral image analysis offers many contexts where the treatment of second-order data is mandatory. To mention a few examples, working with sets of related images for quantitative analysis purposes, process monitoring, or fusing images acquired on the same sample with different spectroscopic platforms requires the use of data analysis tools that may handle simultaneously related blocks of information. Unmixing analysis in hyperspectral imaging is one of the most common tasks performed, as a complete spatial and chemical characterization of the compounds in a scanned sample needs retrieving their pure distribution maps and spectral signatures. Among the different algorithms, multivariate curve resolution-alternating least squares (MCR-ALS) offers a particularly flexible framework that allows combining images that may present significant differences in spatial and spectroscopic properties. Through challenging data configurations and the combined use of bilinear and multilinear models, a large diversity of image fusion scenarios can be properly addressed. This chapter gradually shows the image requirements and the chemometric solutions available for every problem.
Time-resolved fluorescence spectroscopy plays a crucial role when studying dynamic properties of complex photochemical systems. Nevertheless, the analysis of measured time decays and the extraction of exponential lifetimes often requires either the experimental assessment or the modeling of the instrument response function (IRF). However, the intrinsic nature of the IRF in the measurement process, which may vary across measurements due to chemical and instrumental factors, jeopardizes the results obtained by reconvolution approaches. In this paper, we introduce a novel methodology, called blind instrument response function identification (BIRFI), which enables the direct estimation of the IRF from the collected data. It capitalizes on the properties of single exponential signals to transform a deconvolution problem into a well-posed system identification problem. To delve into the specifics, we provide a step-by-step description of the BIRFI method and a protocol for its application to fluorescence decays. The performance of BIRFI is evaluated using simulated and time-correlated single-photon counting data. Our results demonstrate that the BIRFI methodology allows an accurate recovery of the IRF, yielding comparable or even superior results compared with those obtained with experimental IRFs when they are used for reconvolution by parametric model fitting.
Time -resolved fluorescence spectroscopy plays a crucial role when studying dynamic properties of complex photochemical systems. Nevertheless, the analysis of measured time decays and the extraction of exponential lifetimes often requires either the experimental assessment or the modeling of the instrument response function (IRF). However, the intrinsic nature of the IRF in the measurement process, which may vary across measurements due to chemical and instrumental factors, jeopardizes the results obtained by reconvolution approaches. In this paper, we introduce a novel methodology, called blind instrument response function identi fication (BIRFI), which enables the direct estimation of the IRF from the collected data. It capitalizes on the properties of single exponential signals to transform a deconvolution problem into a well -posed system identi fication problem. To delve into the speci fics, we provide a step-by-step description of the BIRFI method and a protocol for its application to fluorescence decays. The performance of BIRFI is evaluated using simulated and time -correlated single -photon counting data. Our results demonstrate that the BIRFI methodology allows an accurate recovery of the IRF, yielding comparable or even superior results compared with those obtained with experimental IRFs when they are used for reconvolution by parametric model fitting.
Emission (3D) and excitation-emission (4D) fluorescence images allow covering wide excitation and emission spectral ranges and, hence, provide very complete information for a good characterization and location of fluorophores in samples. However, when the acquisition time of the image is too long, degradation of the fluorescence signal of compounds and sample photodamage can occur due to photobleaching. This phenomenon is due to the long exposure time of the sample to the light source and can hinder the detection and the proper characterization of the fluorophores in samples. The main purpose of this research is providing a methodology to obtain and interpret the information of fluorescence images for the characterization of samples without suffering the consequences of photobleaching. Such a goal implies a first thorough knowledge of the photobleaching phenomenon to adapt the fluorescence imaging measurement for an optimal characterization of the fluorophores present in samples. The proposed approach relies first on a study of time-series of 3D or 4D fluorescence images to characterize spatially and spectroscopically the fluorophores present in the samples and their photobleaching behaviour. Since photobleaching is fluorophore-dependent, the unmixing algorithm Multivariate Curve Resolution-Alternating Least Squares (MCR-ALS) is applied to the set of fluorescence images acquired as a function of time to understand the specific behaviour of every fluorophore. The characteristics of the photobleaching phenomenon and the nature of the fluorescence measurement offer a challenging scenario to look for adapted implementations of trilinear and quadrilinear models within the MCR framework. From the results obtained, appropriate instrumental settings are adopted for an image acquisition that allows the correct spatial and spectroscopic characterization of fluorophores in samples. To test the potential of this methodology, the characterization of thin cross-sections of the Oryza sativa (commonly called rice) root have been studied due to the co-occurrence of several natural fluorophores in vegetal tissues.
Chemical analyses based on digital images have been widely investigated due to their non-invasive character and simplicity on the measurement. Based on a bibliographic survey, it is possible to notice the importance of controlling instrumental and structural parameters for image acquisition, which can compromise the repeatability and reproducibility of the analyses. However, despite the remarkable progress in literature, the role of many of these parameters has not yet been thoroughly unraveled. Another, more practical, limitation concerning image acquisition is related to the high cost involved in accessing more robust instruments. To circumvent these limitations, a low-cost prototype was built to evaluate conditions that directly influence the quality of the digital images. Thus, an image acquisition system based on a Raspberry Pi® 4 4 GB (a single board computer) coupled with electronic components (luminosity and temperature sensor) and a Pi Camera® 5MP was designed. To control, integrate and automate the process of acquisition and analysis of digital images, an algorithm in Python programming language was developed. To analyze the images obtained, multivariate tools were used for exploratory analyses. And the results were validated (repeatability, reproducibility and robustness) by evaluating the similarities of the colors (Euclidean distance of the predominant intensity between the colorgrams and the comparison of the cosines between the histograms) and compared to those obtained by spectrophotometry in the visible region.
Industrially tempered dark chocolate processed under different conditions was analyzed using Attenuated Total Reflection-Fourier Transform-Infrared spectroscopy (ATR-FT-IR). Spectra were collected during the cooling process right after tempering using two distinct machines and machine operational modes and over short and long storage (1 day and two months, respectively). To interpret the spectra collected, Multivariate Curve Resolution-Alternating Least Squares (MCR-ALS) analysis was used. MCR-ALS incorporated a dedicated implementation of hard modeling constraints for the elucidation of the spectral and concentration profiles of chocolate contributions during the monitored process with the aim of reducing ambiguity in the data interpretation. Thus, the spectral signature of the amorphous form of chocolate was constrained to follow a Gaussian shape, in line with previously reported works and with measurements observed at the beginning of the cooling process. Additionally, hard modeling constraints were applied to the concentration profiles linked to the cooling process by implementing the Avrami kinetic model, often postulated to describe thermally induced fat crystal transformations. The application of the hard modeling constraints eased the generation of interpretable solutions and resulted in concentration profiles and spectral signatures that were defining much better the crystal state of the chocolate samples. A two-component system explained the studied procedures, with a contribution S1 representing the highly crystalline state of the chocolate fats and S2 representative of the amorphous/less stable state. In general, we were able to track the transition from a less ordered to a highly ordered crystal state of the fats during the cooling stage of chocolate. After short storage, the chocolates had contributions of both S1 and S2 profiles. After long storage, the S1 crystalline form was the most dominant interpreted as the reorganization of the triglyceride acyl chains towards a thermodynamically favorable state. The latter was observed for all analyzed chocolates, implying that, regardless the tempering process or temper regime tested, the fats in chocolates reach the same high crystalline state after maturing.
L.M.C. Buydens合作论文数Department of Analytical Chemistry, University of Nijmegen, Toernooiveld 1, 6525 ED Nijmegen, Netherlands4