Researchers interested in developing new multivariate statistical methods often need to be able to generate multivariate datasets with specific characteristics to test the effectiveness of their data analysis algorithms under specific conditions. In this paper, we present a family of methods for generating multivariate centred datasets by simultaneously controlling features of the cross-product matrices X inverted perpendicular X and XX inverted perpendicular. This provides an interesting trade-off to control for the variance structure in the data, important for the family of algorithms that operate on the data matrix, like, e.g., Principal Component Analysis, and control for the distances among objects, important for algorithms that operate on the distance matrix, like Multidimensional Scaling. The proposed methods form a general framework that can be understood as a jigsaw puzzle, joining pieces obtained from the spectral decomposition of a target covariance matrix and the singular value decomposition of a target data matrix. These methods have in common that they are derived from a two-sided orthogonal Procrustes problem.
The traditional Six Sigma toolkit presents significant limitations for problem solving in Industry 4.0 settings. To address these limitations, it can be extended with latent variable-based multivariate statistical techniques such as Principal Component Analysis (PCA) and Partial Least Squares (PLS), in what has been referred to as multivariate Six Sigma. In this work, this approach is applied to address vibration performance issues in the caliper, a key component of a car's braking system. By appropriately integrating these techniques into the five-step DMAIC cycle, the root cause was satisfactorily identified, an effective corrective action was implemented, and the objectives of the project were successfully achieved. This case study provides further evidence of the practical utility of multivariate Six Sigma and demonstrates its potential for broader adoption in industrial projects.
We present a novel Latent Space-based Multivariate Capability Index (LSb-MCpk) aligned with the Quality by Design initiative and used as a criterion for ranking and selecting suppliers for a particular raw material used in a manufacturing process. The novelty of this new index is that, contrary to other multivariate capability indexes that are defined either in the raw material space or in the Critical Quality Attributes (CQAs) space of the product manufactured, this new LSb-MCpk is defined in the latent space connecting both spaces. This endows the new index with a clear advantage over classical ones as it quantifies the capacity of each raw material supplier of providing assurance of quality with a certain confidence level for the CQAs of the manufactured product before manufacturing a single unit of the product. All we need is a rich database with historical information of several raw material properties along with the CQAs. Besides, we present a novel methodology to carry out the diagnosis for assignable causes when a supplier does not score a good capability index. The proposed LSb-MCpk is based on Partial Least Squares (PLS) regression, and it is illustrated using data from both an industrial and a simulation study.
The concepts of null space and orthogonal space have been developed in independent contexts and with different purposes: the former arises in the inversion of partial least-squares (PLS) regression models, and the latter in orthogonal PLS (O-PLS) modeling. In this study, we bridge PLS model inversion and O-PLS modeling by mathematically proving that the null space and the orthogonal space are the same space. We also provide a graphical interpretation of the equivalence between the two spaces, using both a simulated and a real case study.
This work was centred on developing an objective, reproducible and non-destructive methodology to predict the lethality of C. elegans populations contained in liquid culture mediums, addressing the handicaps presented for imaging analysis in those media types from a numerical point of view, applying chemometric and machine learning procedures on imaging data obtained with a basic image device and processing. The experiment was carried out by taking videos from nematode populations exposed to different conditions of three stressors (hydrogen peroxide, heat and UV radiation). The processed video datasets were used as predictors for different configurations in regression methods. The dimensionality reduction approach improved the prediction capacity of the imaging information compared to the raw dataset. Moreover, the best result was achieved with a super learner model, demonstrating the synergistic effect of combining results from models with lower prediction capacity to develop a meta-model with high prediction capabilities.
Process optimization and innovation are essential in a competitive and digitalized industry driven by the Quality-by-Design paradigm. This requires building a causal model that explains how variations in the inputs relate to variations in the outputs. Traditionally, deterministic models are preferred for this purpose, but these are often unfeasible due to limited knowledge and high development costs. As a result, data-driven models are increasingly used. To maintain causality in these models, independent input variation is necessary-typically achieved through Design of Experiments (DOE). However, in Industry 4.0 contexts, performing DOE is challenging due to complex variable correlations and the high number of factors involved. Although large volumes of production data are available, they often lack the independence required for causal inference, making traditional statistical and machine learning models ineffective for optimization. Consequently, there is growing interest in developing causal models from historical data. This paper explores two promising methods: retrospective DOE and causal latent space-based modeling.
The sequential multi-block partial least squares (SMB-PLS) is proposed for implementing a multivariate statistical process control scheme. This is of interest when the system is composed of several blocks following a sequential order and presenting correlated information, for instance, a raw material properties block followed by a process variables block that is manipulated according to raw material properties. The SMB-PLS uses orthogonalization to separate correlated information between blocks from orthogonal variations. This allows monitoring the system in different stages considering only the remaining orthogonal part in each block. Thus, the SMB-PLS increases the interpretability and process understanding in the model building (Phase I), since it provides a deep insight about the nature of the system variations. Besides, it prevents any special cause from propagating to subsequent blocks enabling their use in the model exploitation (Phase II). The methodology is applied to a real case study from a food manufacturing process.
Currently, magnetic resonance imaging is the most sensitive imaging technique for detecting cancerous processes in early stages. As for breast cancer, due to the tubular structure of the tissue, being formed by ducts, anisotropic diffusion should be considered instead of the general isotropic diffusion. Anisotropic diffusion is studied by applying a technique called Diffusion Tensor Imaging (DTI), where the diffusion gradient is applied by changing the magnetic field in several spatial directions.To date, the application of Multivariate Curve Resolution (MCR) models in diffusion sequences has demonstrated its ability to develop cancer biomarkers of easy clinical interpretation in the case of isotropic tissues, such as the prostate. But so far, it has never been applied in the case of anisotropic tissues, as the breast.Therefore, the main objective of this work is to obtain easy-to-interpret imaging biomarkers useful for early breast cancer diagnosis from diffusion magnetic resonance imaging based on the Diffusion Tensor using multivariate curve resolution (MCR) models. A classification model to identify healthy and tumor affected pixels is also proposed.
Functional MRI is, currently, the most sensitive technique in breast cancer for detecting early tumors, and perfusion (DCE-MRI) has become the most important sequence to depict and characterize angiogenesis and neovascularization. In this work, we propose the use of new biomarkers that are related to clear physiological phenomena, obtained from MCR-ALS as an alternative to curve-based pseudo-biomarkers and pharmacokinetics models. In order to provide a discrimination and prediction model between healthy tissue and cancer, we propose using PLS-DA with double cross-validation (2CV) and variable selection, repeated several times and obtaining excellent average results for the performance indexes (f-score: 0.9149, MCC: 0.8538, AUROC: 0.8794). After selecting the optimal prediction model, a unique probabilistic map called “virtual biopsy” that shows in different colors the probability that each pixel of the image has a tumor behavior is obtained, helping the specialist with the identification and characterization of breast tumors with only one easy-to-interpret biomarker map.
High-dimensional and multivariate data sets often contain missing data and/or cellwise/rowwise outliers. Whereas several solutions have been proposed to deal with each one of these issues independently, the number of suitable techniques that simultaneously confront these phenomena is drastically reduced. In this paper, we introduce RadarTSR, a Robust Adaptation for Data with Anomalous Rows and/or cells of the Trimmed Scores Regression method, which is based on a Principal Component Analysis (PCA). RadarTSR detects cellwise and rowwise outliers, imputes missing data without the harmful effect of outliers, and, if grouped rowwise outliers are detected, RadarTSR imputes them with their own model. The performance of RadarTSR is compared to the MacroPCA algorithm; as far as we are concerned, the only proposal that deals with missing data and contemplates these two different types of outliers. Several simulated and real data sets are used. The RadarTSR code is available in Matlab.
The study introduces three novel strategies for incorporating capabilities for dynamic modelling into multiblock regression methods by integrating sequentially orthogonalised partial least squares (SO-PLS) with different dynamic modelling techniques. The study evaluates these strategies using synthetic datasets and an industrial example, comparing their performance in predictive ability, identification of process dynamics, and quantification of block contributions. Results suggest that these approaches can effectively model the dynamics with performance comparable to state-of-the-art methods, providing, at the same time, insight into the dynamic order and block contributions. One of the strategies, sequentially orthogonalised dynamic augmented (SODA)-PLS, shows promise by ensuring that redundant information in the time dimension is not included, resulting in simpler and more easily interpretable dynamic models. These multiblock dynamic regression strategies have potential applications for improved process understanding in industrial settings, especially where multiple data sources and inherent time dynamics are present.
We present here a novel methodology to analyze data from two-level factorial experimental designs, with or without missing runs, with just one method: partial least squares regression with one response variable (PLS1, hereinafter PLS). This property is very attractive for practitioners because, to the best of our knowledge, no other statistical tool has comparable versatility. In the case of a full and fractional factorial design, the one-PLS component model yields the same analytical solution as multiple linear regression (MLR), not only in the estimation of the effects but also in their statistical significance. When having missing runs in the factorial design, PLS is of particular interest as it is a powerful tool when dealing with complex correlation structures, as opposed to MLR. Thus, we challenge the widely held view that PLS is useful only when dealing with nonexperimental design (i.e., correlated observational data). The methodology is illustrated by two illustrative examples and synthesized by an easy-to-follow route map useful for practitioners.
Digital sensors and machine learning enable efficiency improvements in production processes, through process monitoring, anomaly detection, soft sensing, and process control. However, the development of such solutions requires several data preprocessing steps. In continuous processes, a crucial part of the data preparation is adjusting for time delays between different sensors. This is necessary to ensure that each sensor measurement relate to the same volume of materials going through various processing steps.This study provides an overview of data-driven methods for estimating time lags between sensors in continuous processes. The methods are assessed in a large simulation study, on data sets with different sample sizes, model complexities and autocorrelation functions. Our results shows that most methods work well if the relationships are close to linear, but more flexible metrics like distance correlation and maximum information coefficient are needed in more complex systems. Finally, we present a real industrial example to illustrate some real-world aspects of the variable time delay estimation process.
One of the most common sources of information in Synthetic Biology is the data coming from plate reader fluorescence measurements. These experiments provide a measure of the light emitted by a certain fluorescent molecule, such as the Green Fluorescent Protein (GFP). However, these measurements are generally expressed in arbitrary units and are affected by the measurement device gain. This limits the range of measurements in a single experiment and hampers the comparison of results among experiments. In this work, we describe PLATERO, a calibration protocol to express fluorescence measures in concentration units of a reference fluorophore. The protocol removes the gain effect of the measurement device on the acquired data. In addition, the fluorescence intensity values are transformed into units of concentration using a Fluorescein calibration model. Both steps are expressed in a single mathematical expression that returns normalized, gain-independent, and comparable data, even if the acquisition was done at different device gain levels. Most important, the PLATERO embeds a Linearity and Bias Analysis that provides an assessment of the uncertainty of the model estimations, and a Reproducibility and Repeatability analysis that evaluates the sources of variability originating from the measurements and the equipment. All the functions used to build the model, exploit it with new data, and perform the uncertainty and variability assessment are available in an open access repository.
The Sequential Multi-Block Partial Least Squares (SMB-PLS) model inversion is applied for defining analytically the multivariate raw material region providing assurance of quality with a certain confidence level for the critical to quality attributes (CQA). The SMB-PLS algorithm does identify the variation in process conditions uncorrelated with raw material properties and known disturbances, which is crucial to implement an effective process control system attenuating most raw material variations. This allows expanding the specification region and, hence, one may potentially be able to accept lower cost raw materials that will yield products with perfectly satisfactory quality properties. The methodology can be used with historical/happenstance data, typical in Industry 4.0. This is illustrated using simulated data from an industrial case study.
Background: The implementation of process analytical technologies (PAT) has gained attention since 2004 when its formal introduction through the U.S. Food and Drug Administration was introduced. Manufacturers that need to evaluate the employment of new monitoring systems could face different challenges: identification of suitable sensors, verification of data meaning, evaluation of several statistical strategies to obtain insights about data and achieve process understanding and finally, the actual possibilities for monitoring. Kefir fermentations were chosen as an example because of the chemical and physical transformations that occurred during the process, which could be common to several other fermentation processes. In order to pave the way for monitoring establish the information contained in the data and find the right tools for extracting them is of extreme importance. Strategies to identify different experimental conditions in the spectra acquired with a miniaturized NIR (1350-2550 nm) during process occurrence were addressed.Results: The study aims to offer insights into good practices and steps to pave the way for process monitoring with handheld NIR data. The main aspects of interest for batch processes in preliminary evaluations were investigated and discussed. On the one hand, process understanding and, on the other, the possibilities for process monitoring and endpoint determination were examined. The combination of different statistical tools allowed the extraction of information from the data and the identification of the link between them and the chemical and physical changes during the process. In addition, insights into the spectra characteristics in the studied spectroscopic range for kefir fermentation were reported.Significance: The capabilities for miniaturized NIR spectra to represent and statistical strategies to characterize different experimental conditions in a real case fermentation occurrence were proved. The strengths and limi-tations of some of the common approaches to catch changes in fermentation condition were highlighted. For the various statistical approaches, the chances offered in the research and development stages and to set the scene for monitoring and end-point detection were explored.
Bladder cancer (BC) is the sixth leading cause of death by cancer. Depending on the invasiveness of tumors, patients with BC will undergo surgery and surveillance lifelong, owing the high rate of recurrence and progression. In this context, the development of strategies to support non-invasive BC diagnosis is focusing attention. Voltammetric electronic tongue (VET) has been demonstrated to be of use in the analysis of biofluids. Here, we present the implementation of a VET to study 207 urines to discriminate BC and non-BC for diagnosis and surveillance to detect recurrences. Special attention has been paid to the experimental setup to improve reproducibility in the measurements. PLSDA analysis together with variable selection provided a model with high sensitivity, specificity, and area under the ROC curve AUC (0.844, 0.882, and 0.917, respectively). These results pave the way for the development of non-invasive low-cost and easy-to-use strategies to support BC diagnosis and follow-up.
The clinical course of COVID-19 is highly variable. It is therefore essential to predict as early and accurately as possible the severity level of the disease in a COVID-19 patient who is admitted to the hospital. This means identifying the contributing factors of mortality and developing an easy-to-use score that could enable a fast assessment of the mortality risk using only information recorded at the hospitalization. A large database of adult patients with a confirmed diagnosis of COVID-19 (n = 15,628; with 2,846 deceased) admitted to Spanish hospitals between December 2019 and July 2020 was analyzed. By means of multiple machine learning algorithms, we developed models that could accurately predict their mortality. We used the information about classifiers' performance metrics and about importance and coherence among the predictors to define a mortality score that can be easily calculated using a minimal number of mortality predictors and yielded accurate estimates of the patient severity status. The optimal predictive model encompassed five predictors (age, oxygen saturation, platelets, lactate dehydrogenase, and creatinine) and yielded a satisfactory classification of survived and deceased patients (area under the curve: 0.8454 with validation set). These five predictors were additionally used to define a mortality score for COVID-19 patients at their hospitalization. This score is not only easy to calculate but also to interpret since it ranges from zero to eight, along with a linear increase in the mortality risk from 0% to 80%. A simple risk score based on five commonly available clinical variables of adult COVID-19 patients admitted to hospital is able to accurately discriminate their mortality probability, and its interpretation is straightforward and useful.