Data tables for machine learning and structure-activity relationship modelling (QSAR) are often naturally organized in blocks of data, where multiple molecular representations or sets of descriptors form the blocks. Multi-block Orthogonal Component Analysis (MOCA), a new analytical tool, can be used to explore such data structures in a single model, identifying principal components that are unique to a single block or joint over multiple blocks. We applied MOCA to two sets of 550 and 300 molecules and up to 9213 molecular descriptors organized in 11 blocks. The MOCA models reveal relationships between the blocks and overarching trends across the whole dataset. Based on the MOCA joint components, we propose a quantitative metric for the redundancy of blocks, useful for a priori block-wise feature selection or evaluation of new molecular representations. The second data set includes 7 ecotoxicological study endpoints for crop protection chemicals, for which we (re-)discovered some general trends and linked them to molecular properties. Using a single MOCA model we estimated the predictive potential of each block and the model-ability of the target block.
This data table contains the molecules and their calculated molecular descriptors for the Pesticides data set as described in the journal article.
Recently a new parameter to infer variable importance in orthogonal projections to latent structures (OPLS) was presented. Called OPLS-VIP (variable influence on projection), this parameter is here applied in multivariate time series analysis to achieve an improved diagnosis of process dynamics. To this end, OPLS-VIP has been tested in three real-world industrial data sets; the first data set corresponds to a pulp manufacturing process using a continuous digester, the second one involves data from an industrial heater that experienced problems, and the third data set contains measures of the chemical oxygen demand into the effluent of a newsprint mill. The outcomes obtained using OPLS-VIP are benchmarked against classical PLS-VIP results. It is demonstrated how OPLS-VIP provides a better diagnosis and understanding of the time series behavior than PLS-VIP.
A personal view is given about the gradual development of projection methods—also called bilinear, latent variable, and more—and their use in chemometrics. We start with the principal components analysis (PCA) being the basis for more elaborate methods for more complex problems such as soft independent modeling of class analogy, partial least squares (PLS), hierarchical PCA and PLS, PLS‐discriminant analysis, Orthogonal projection to latent structures (OPLS), OPLS‐discriminant analysis and more.From its start around 1970, this development was strongly influenced by Bruce Kowalski and his group in Seattle, and his realization that the multidimensional data profiles emerging from spectrometers, chromatographs, and other electronic instruments, contained interesting information that was not recognized by the current one variable at a time approaches to chemical data analysis.This led to the adoption of what in statistics is called the data analytical approach, often called also the data driven approach, soft modeling, and more. This approach combined with PCA and later PLS, turned out to work very well in the analysis of chemical data. This because of the close correspondence between, on the one hand, the matrix decomposition at the heart of PCA and PLS and, on the other hand, the analogy concept on which so much of chemical theory and experimentation are based. This extends to numerical and conceptual stability and good approximation properties of these models.The development is informally summarized and described and illustrated by a few examples and anecdotes. Copyright © 2014 John Wiley & Sons, Ltd.
A new approach for variable influence on projection (VIP) is described, which takes full advantage of the orthogonal projections to latent structures (OPLS) model formalism for enhanced model interpretability. This means that it will include not only the predictive components in OPLS but also the orthogonal components. Four variants of variable influence on projection (VIP) adapted to OPLS have been developed, tested and compared using three different data sets, one synthetic with known properties and two real‐world cases. Copyright © 2014 John Wiley & Sons, Ltd.
In batch statistical process control (BSPC), data from a number of "good" batches are used to model the evolution (trajectory) of the process and they also define model control limits, against which new batches may be compared. The benchmark methods used in BSPC include partial least squares (PLS) and principal component analysis (PCA). In this paper, we have used orthogonal projections to latent structures (OPLS) in BSPC and compared the results with PLS and PCA. The experimental study used was a batch hydrogenation reaction of nitrobenzene to aniline characterized by both UV spectroscopy and process data. The key idea is that OPLS is able to separate the variation in data that is correlated to the process evolution (also known as 'batch maturity index') from the variation that is uncorrelated to process evolution. This separation of different types of variations can generate different batch trajectories and hence lead to different established model control limits to detect process deviations. The results demonstrate that OPLS was able to detect all process deviations and provided a good process understanding of the root causes for these deviations. PCA and PLS on the other hand were shown to provide different interpretations for several of these process deviations, or in some cases they were unable to detect actual process deviations. Hence, the use of OPLS in BSPC can lead to better fault detection and root cause analysis as compared to existing benchmark methods and may therefore be used to complement the existing toolbox.
Partial least squares (PLS) regression is a flexible data analytical approach, which can be made even more versatile and useful by various modifications. In this article we describe the extension into orthogonal PLS modeling, in terms of two new methods, called OPLS and O2PLS, with similar prediction capacity but improved model interpretation.
This paper presents an extension to the recently published OnPLS data analysis method. Bi‐modal OnPLS allows for arbitrary block relationships in both columns and rows and is able to extract orthogonal variation in both columns and rows without bias towards any particular direction or matrix: the method is fully symmetric with regard to both rows and columns. Bi‐modal OnPLS extracts a minimal number of globally predictive score vectors that exhibit maximal covariance and correlation in the column space and a corresponding set of predictive loading vectors that exhibit maximal correlation in the row space. The method also extracts orthogonal variation (i.e. variation that is not related to all other matrices) in both columns and rows. The method was applied to two synthetic datasets and one real data set regarding sensory information and consumer likings of dairy products. It was shown that Bi‐modal OnPLS greatly improves the intercorrelations between both loadings and scores while still finding the correct variation. This facilitates interpretation of the predictive components and makes it possible to study the orthogonal variation in the data. Copyright © 2012 John Wiley & Sons, Ltd.
Data mining by means of projection methods such as PLS (projection to latent structures), and their extensions is discussed. The most common data analytical questions in data mining are covered, and illustrated with examples. (a) Clustering, i.e., finding and interpreting “natural” groups in the data (b) Classification and identification, e.g., biologically active compounds vs inactive (c) Quantitative relationships between different sets of variables, e.g., finding variables related to quality of a product, or related to time, seasonal or/and geographical change Sub-problems occurring in both (a) to (c) are discussed. (1) Identification of outliers and their aberrant data profiles (2) Finding the dominating variables and their joint relationships (3) Making predictions for new samples The use of graphics for the contextual interpretation of results is emphasized. With many variables and few observations (samples) – a common situation in data mining – the risk to obtain spurious models is substantial. Spurious models look great for the training set data, but give miserable predictions for new samples. Hence, the validation of the data analytical results is essential, and approaches for that are discussed.
A new set of amino acid descriptor scales was recently introduced. The new scales relate to amino acid side chain (i) rigidity and (ii) flexibility. These two properties were found to be orthogonal to each other. In this study, (i) the understanding of their meaning is improved, (ii) their utility further corroborated, and (iii) the scope of their applicability is broadened. Notably, the rigidity and flexibility scales are benchmarked against previous amino acid description precedence and found to contribute (1) new information, and (2) direct interpretation. The suggested description extensions were tested using empirical data from peptide description and in quantitative structure–activity relationships (QSAR).
Awards and medals are an important mechanism by which scientific societies increase awareness and recognition of important fields of study. To honor researchers who have made an important impact in the field of chemometrics, the Chemometrics Division of the Swedish Chemical Society founded The Herman Wold medal in 1995. The Herman Wold medal is struck in pure gold and highlights important advances in chemometrics. It is named after the Swedish statistician Herman Wold (1908–1992) who worked in the field of time series analysis and econometrics at Uppsala University. The recipients of the medal are persons who have contributed significantly to the development and proliferation of chemometrics in research, development and production. In addition, this should be done “in the spirit of Herman Wold.” It has become a tradition that the winner of the Herman Wold medal is announced at a meeting in the biannual series of Scandinavian Symposium on Chemometrics conferences (SSC). The first Herman Wold medal recipient was Svante Wold (SSC4, Lund, Sweden, 1995), followed by Agnar Höskuldsson (SSC5, Lahti, Finland, 1997), Harald Martens (SSC6, Porsgrunn, Norway, 1999), John MacGregor (SSC7, Copenhagen, Denmark, 2001), Rolf Carlson (SSC8, Mariehamn, Åland, Finland, 2003), and Olav Kvalheim (SSC9, Reykjavik, Iceland, 2005). At SSC10 in Lappeenranta 2007 Professor Pentti Minkkinen, Lappeenranta University of Technology, Finland, was awarded with the seventh Herman Wold medal for his work in sampling strategies. As an associate expert in two United Nations' Mineral exploration projects in Turkey and Egypt in 1973–1975, professor Minkkinen became interested in analytical quality control. In 1976 he was appointed as an Associate Professor and later on full Professor in Inorganic and Analytical Chemistry at Lappeenranta University of Technology, Finland. After retiring at the end of 2007, he still remains scientifically active, especially in the field of sampling. He was also a member of EURACHEM Working Group for Uncertainty arising from Sampling 2004–2007and an Associate Member of the Division of Analytical Chemistry, International Union of Pure and Applied Chemistry (IUPAC) 2006–2009. Professor Minkkinen has over 80 original research articles in peer review journals, of which 15 are on sampling. The eighth Herman Wold medal was received by Professor Michael Sjöström, Umeå University, Sweden in 2008. This was announced at the Euro-QSAR meeting in Uppsala in September 2008. Professor Sjöström was the first graduate student at the Research Group for Chemometrics, Umeå University, Sweden. He got his Ph.D. there in 1976 with a thesis on the theory and statistics of the Hammett equation and other extra-thermodynamical relationships. This was the beginning of the Umeå Chemometrics group, which then led to the SIMCA method and later to PLS for quantitative pattern recognition. After his dissertation he spent a post-doc year with Bruce Kowalski in University of Washington in Seattle. He then returned to Umeå, where he spent a fruitful chemometrics/physical organics career on research and education in the areas of structure-reactivity and structure activity (QSAR) relationships, solvent effects (the Kamlet-Taft controversy), sequence activity relationships, PLS, PLS-discriminant analysis, and similar problems. Professor Sjöström has over 175 scientific articles in international journals. Sjöström has been a central person in Chemometrics, and has been part of the Umeå group since its beginning around 1968. His kind but clever personality very much contributed to its success. Michael is also one of the founders of the International Chemometrics Society. Associate Professor Johan Trygg, Umeå University, Sweden, received his Herman Wold medal at the SSC11 in Loen/Stryn, Norway, 2009, joining the group of well-known and legendary professors at an early age. He was rewarded for his creative development of new methods for modeling and interpretation of biological and medical data by OPLS (orthogonal projections to latent structures). OPLS is already in use by more than 150 Swedish companies, 50 international institutions and the 10 largest pharmaceutical companies in the world. It has also become a standard method in the rapidly growing field of “omics” research. Professor Trygg has over 90 publications in international Journals. Johan Trygg is also now the Chair of the Chemometrics Division of the Swedish Chemical Society. The editorial group behind this double issue also recently had a chance to interview Professor Svante Wold, the son of Herman Wold, who together with Michael Sjöström and Rolf Carlson shaped the “Umeå-school” of chemometrics research and education: “I am impressed by the broad scope and the investigative skills presented by the group of articles compiled in this issue,” says Svante, and continues “I really enjoy seeing so many excellent studies being produced in the spirit of the work done by Herman.” Svante has formally retired from Umeå University and today runs his own company (NNS Consulting) together with his wife Nouna. He has also contributed one article to this special issue: “I will always continue to do chemometrics research, at my own pace, and try to propose simple, reliable and interpretable solutions to important data analytical issue,” says Svante. With this statement, he actually summarizes the gist of the many papers featured in this issue, namely to advance the chemometrics toolbox to include ever improving methods for the benefit of the people who analyze our continually increasing masses of data. This special issue of the Journal of Chemometrics is to honor the Herman Wold medal winners 2007–2009: Professor Pentti Minkkinen, Professor Michael Sjöström, and Associate Professor Johan Trygg. The editorial group is very grateful to all the contributors and we are proud of the final result.
We introduce a new measure for the importance of predictor variables, X, for the separation of two groups (classes) of observations. The measure is a Graphical Index of Separation (GIOS), and is, for each predictor, determined from the distribution of all possible pairs of observations with one from each group. GIOS is quantitative, intuitively simple and easy to interpret. The GIOS is straightforward to visualize in bivariate plots, and line or bar plots for larger number of variables. The approach applies both to discriminant analyses such as LDA, SIMCA, PLS-DA, OPLS-DA and to quantitative modeling such as MLR, PLS and OPLS. In the latter case, the observations are first divided into two groups based on their response values, Y. The GIOS approach is illustrated by PLS-DA/OPLS-DA and SIMCA-classification of a number of multivariate data sets with few and many variables relative to the number of observations. Copyright (C) 2010 John Wiley & Sons, Ltd.
A hierarchical clustering approach based on a set of PLS models is presented. Called PLS‐Trees®, this approach is analogous to classification and regression trees (CART), but uses the scores of PLS regression models as the basis for splitting the clusters, instead of the individual X‐variables. The split of one cluster into two is made along the sorted first X‐score (t1) of a PLS model of the cluster, but may potentially be made along a direction corresponding to a combination of scores. The position of the split is selected according to the improvement of a weighted combination of (a) the variance of the X‐score, (b) the variance of Y and (c) a penalty function discouraging an unbalanced split with very different numbers of observations. Cross‐validation is used to terminate the branches of the tree, and to determine the number of components of each cluster PLS model. Some obvious extensions of the approach to OPLS‐Trees and trees based on hierarchical PLS or OPLS models with the variables divided in blocks depending on their type, are also mentioned. The possibility to greatly reduce the number of variables in each PLS model on the basis of their PLS w‐coefficients is also pointed out. The approach is illustrated by means of three examples. The first two examples are quantitative structure‐activity relationship (QSAR) data sets, while the third is based on hyper‐spectral images of liver tissue for identifying different sources of variability in the liver samples. Copyright © 2009 John Wiley & Sons, Ltd.
To find relationships in large data sets, such data sets are almost always divided into groups containing fairly homogeneous (non-grouped) data. Hence, to make multivariate data analysis applicable in data mining, and other analyses of large and complex data sets, we need one or several clustering algorithms, preferably combined with an embedded multivariate regression step (i.e. PLS, OPLS® or O2PLS®). This multivariate clustering algorithm must work well for large data sets with potentially very many and collinear variables, missing data, noise, and other common complications.