To find relationships in large data sets, such data sets are almost always divided into groups containing fairly homogeneous (non-grouped) data. Hence, to make multivariate data analysis applicable in data mining, and other analyses of large and complex data sets, we need one or several clustering algorithms, preferably combined with an embedded multivariate regression step (i.e. PLS, OPLS® or O2PLS®). This multivariate clustering algorithm must work well for large data sets with potentially very many and collinear variables, missing data, noise, and other common complications.