This chapter describes the transformation of a young physical organic chemist (SW, 1964), from a believer in first principles models to a middle-aged chemometrician (SW, 1974) promoting empirical and semiempirical "data driven, soft, analogy" models for the design of experiments and the analysis of the resulting data. This transformation was marked by a number of influential events, each tipping the balance towards the data driven, soft, analogy models until the point of no return in 1974. On June 10, 1974, Bruce and I together with our research groups joined forces formed the Chemometrics Society (later renamed to the International Chemometrics Society), and we took off into multidimensional space. This review of my personal scientific history, inspired and encouraged by Bruce, is illustrated by examples of method development driven by necessity to solve specific problems and leading to data driven soft models, which, at least in my own eyes, were superior to the classical first principles approaches to the same problems. Bruce and I met at numerous conferences between 1975 and 1990, but after that, Bruce and I gradually slid out of the academic world, and now Bruce has taken his final step.
A personal view is given about the gradual development of projection methods—also called bilinear, latent variable, and more—and their use in chemometrics. We start with the principal components analysis (PCA) being the basis for more elaborate methods for more complex problems such as soft independent modeling of class analogy, partial least squares (PLS), hierarchical PCA and PLS, PLS‐discriminant analysis, Orthogonal projection to latent structures (OPLS), OPLS‐discriminant analysis and more.From its start around 1970, this development was strongly influenced by Bruce Kowalski and his group in Seattle, and his realization that the multidimensional data profiles emerging from spectrometers, chromatographs, and other electronic instruments, contained interesting information that was not recognized by the current one variable at a time approaches to chemical data analysis.This led to the adoption of what in statistics is called the data analytical approach, often called also the data driven approach, soft modeling, and more. This approach combined with PCA and later PLS, turned out to work very well in the analysis of chemical data. This because of the close correspondence between, on the one hand, the matrix decomposition at the heart of PCA and PLS and, on the other hand, the analogy concept on which so much of chemical theory and experimentation are based. This extends to numerical and conceptual stability and good approximation properties of these models.The development is informally summarized and described and illustrated by a few examples and anecdotes. Copyright © 2014 John Wiley & Sons, Ltd.
The amount of data measured during a typical manufacturing process is immense. To efficiently utilize these data without becoming overwhelmed with confusing and often conflicting information is difficult to impossible when using traditional univariate methods. Multivariate data mining methods can be used to examine large data sets by extracting relationships between variables to highlight variable correlations and deviations. Specifically, PLS-trees can be used to quickly identify significant clusters in large datasets and to highlight the differences within the groups.
The period 1965 to 1990 saw a dramatic change in Chemistry, from ”wet” to instrument based ”electronic” Chemistry. The former was based on weighing and titration for quantitative results, and reagents and indicators for qualitative information. The latter uses instrumentation such as chromatographs and spectrometers providing quantitative data about samples and reactions, but no immediate information. Computerized data analysis was now needed for the conversion of the data to interpretable results, and the emerging field of ”chemical data analysis” was called ”Chemometrics” in analogy to Biometrics, Econometrics, Psychometrics, etc. In its infancy, Chemometrics was focussed on classification (”supervised pattern recognition”) and the graphical overview of data, but gradually quantitative relationships such as multiple regression were becoming the dominating interest (including graphical aspects thereof). Chemical data sets often have many variables and rather few observations (cases, samples), which made existing data analytical and statistical methods difficult to apply without recourse to mutilating variable selection. Approaches for the analysis of data sets with many variables and few observations were clearly needed. Against everybody’s advice, such approaches were developed in Chemometrics, first for classification (the SIMCA method and similar), and thereafter for regression problems (PLS-regression and extensions). The development continued in various directions, particularly towards interesting applications in biology and engineering, and the interpretability of resulting models, so that today Chemometrics provides a fairly adequate data analytical toolbox covering common applications of ”chemical data analysis”. Thinking back at these happy years, I try to contemplate what we may learn from this development, and of course I become distracted by sentimental memories of funny moments, good friends, and unexpected revelations. The talk provides some glimpses thereof, as well as some reflections on possible cross-fertilizations between the ”multi-block PLS” and the Chemometrics communities.
Data mining by means of projection methods such as PLS (projection to latent structures), and their extensions is discussed. The most common data analytical questions in data mining are covered, and illustrated with examples. (a) Clustering, i.e., finding and interpreting “natural” groups in the data (b) Classification and identification, e.g., biologically active compounds vs inactive (c) Quantitative relationships between different sets of variables, e.g., finding variables related to quality of a product, or related to time, seasonal or/and geographical change Sub-problems occurring in both (a) to (c) are discussed. (1) Identification of outliers and their aberrant data profiles (2) Finding the dominating variables and their joint relationships (3) Making predictions for new samples The use of graphics for the contextual interpretation of results is emphasized. With many variables and few observations (samples) – a common situation in data mining – the risk to obtain spurious models is substantial. Spurious models look great for the training set data, but give miserable predictions for new samples. Hence, the validation of the data analytical results is essential, and approaches for that are discussed.
We introduce a new measure for the importance of predictor variables, X, for the separation of two groups (classes) of observations. The measure is a Graphical Index of Separation (GIOS), and is, for each predictor, determined from the distribution of all possible pairs of observations with one from each group. GIOS is quantitative, intuitively simple and easy to interpret. The GIOS is straightforward to visualize in bivariate plots, and line or bar plots for larger number of variables. The approach applies both to discriminant analyses such as LDA, SIMCA, PLS-DA, OPLS-DA and to quantitative modeling such as MLR, PLS and OPLS. In the latter case, the observations are first divided into two groups based on their response values, Y. The GIOS approach is illustrated by PLS-DA/OPLS-DA and SIMCA-classification of a number of multivariate data sets with few and many variables relative to the number of observations. Copyright (C) 2010 John Wiley & Sons, Ltd.
A hierarchical clustering approach based on a set of PLS models is presented. Called PLS‐Trees®, this approach is analogous to classification and regression trees (CART), but uses the scores of PLS regression models as the basis for splitting the clusters, instead of the individual X‐variables. The split of one cluster into two is made along the sorted first X‐score (t1) of a PLS model of the cluster, but may potentially be made along a direction corresponding to a combination of scores. The position of the split is selected according to the improvement of a weighted combination of (a) the variance of the X‐score, (b) the variance of Y and (c) a penalty function discouraging an unbalanced split with very different numbers of observations. Cross‐validation is used to terminate the branches of the tree, and to determine the number of components of each cluster PLS model. Some obvious extensions of the approach to OPLS‐Trees and trees based on hierarchical PLS or OPLS models with the variables divided in blocks depending on their type, are also mentioned. The possibility to greatly reduce the number of variables in each PLS model on the basis of their PLS w‐coefficients is also pointed out. The approach is illustrated by means of three examples. The first two examples are quantitative structure‐activity relationship (QSAR) data sets, while the third is based on hyper‐spectral images of liver tissue for identifying different sources of variability in the liver samples. Copyright © 2009 John Wiley & Sons, Ltd.
Pell, Ramos and Manne (PRM) in a recent article in this journal claim that the 'conventional' PLS algorithm with orthogonal scores has an inherent inconsistency in that it uses different model spaces for calculating the prediction model coefficients and for calculating the X-space model and it's residuals [1]. We disagree with PRM. All PLS model scores, residuals, coefficients, etc., obtained by the conventional PLS algorithm do come from the same underlying latent variable (LV) model, and not from different models or model spaces as PRM suggest. PRM have simply posed a different model with different assumptions and obtained slightly different results, as should have been expected. Copyright (C) 2008 John Wiley & Sons, Ltd.
The amount of data measured during a typical manufacturing process is immense. To efficiently utilize this data without becoming overwhelmed with confusing, and often conflicting information is difficult to impossible when using traditional Univariate methods. Batch processes in particular - where there is a start and a stop for a given product - require special handling so time-dependent process evolution information is not lost.Measured variables are typically used for process control, but they are also useful for the overview of the process (process monitoring). for fault (Upset) detection, comparing differences in processing equipment (equipment matching), predicting product or process properties (soft sensors or virtual metrology). and for improved process understanding. The use of multivariate analysis to accomplish these objectives Is discussed and illustrated with a batch processing example from the semiconductor industry.
Batch-wise manufacturing (applied in fermentation, cell culturing, and chemical synthesis), gives rise to three-way data arrays when several process variables are measured on the process at regular intervals. The same data structure results from in pharmacokinetics and metabonomics, when data profiles of are taken from individuals at specified intervals. The modeling approaches of three-way batch data for the purpose of understanding, fault detection, control, and prediction, fall in two broad categories, using summarizing variables such as discrete features from the trajectories – landmark points (e.g., peak temp., slopes, times in various phases, etc.), and then forming a batchwise X matrix from these and analyzing by regular PCA/PLS; unfolding the three way array of batch data into a two way matrix (can be done in several ways), followed by PCA/PLS of the two way array to extract an efficient feature set – i.e. latent variables. The established approaches of batch data analysis are reviewed and illustrated by three examples, of yeast production, nylon manufacturing, and of a drying process step.
To find relationships in large data sets, such data sets are almost always divided into groups containing fairly homogeneous (non-grouped) data. Hence, to make multivariate data analysis applicable in data mining, and other analyses of large and complex data sets, we need one or several clustering algorithms, preferably combined with an embedded multivariate regression step (i.e. PLS, OPLS® or O2PLS®). This multivariate clustering algorithm must work well for large data sets with potentially very many and collinear variables, missing data, noise, and other common complications.