Many of the datasets encountered in statistics are two-dimensional in nature and can be represented by a matrix. Classical clustering procedures seek to construct separately an optimal partition of rows or, sometimes, of columns. In contrast, co-clustering methods cluster the rows and the columns simultaneously and organize the data into homogeneous blocks (after suitable permutations). Methods of this kind have practical importance in a wide variety of applications such as document clustering, where data are typically organized in two-way contingency tables. Our goal is to offer coherent frameworks for understanding some existing criteria and algorithms for co-clustering contingency tables, and to propose new ones. We look at two different frameworks for the problem of co-clustering. The first involves minimizing an objective function based on measures of association and in particular on phi-squared and mutual information. The second uses a model-based co-clustering approach, and we consider two models: the block model and the latent block model. We establish connections between different approaches, criteria and algorithms, and we highlight a number of implicit assumptions in some commonly used algorithms. Our contribution is illustrated by numerical experiments on simulated and real-case datasets that show the relevance of the presented methods in the document clustering field.
The predictive analysis of the operating state of complex dynamic systems remains a challenge in numerous applications, particularly in the transport field. Its main objective is the estimation and prediction of the state of health of these systems from data usually available in the form of multivariate time series and emanating from multiple sensors. A classic approach to tackle this problem consists in assuming that the state of health switches between a finite set of operating states. In this case, supervised and unsupervised classification methods can be exploited, as well as finite state space dynamic models (eg. hidden Markov models). This contribution addresses the same issue, but from a different point of view: the unknown dynamics of the state of health is searched into a continuous low-dimensional space. The resulting model is a dynamic factor analytic model which can be seen as a specific state-space models. It can also be exploited for visualizing the evolution of the system's operating state over time. This article will describe the implementation of this model for the estimation, forecasting and visualization of the state of health of a specific railway component: the switch mechanism.
Simultaneous clustering of rows and columns, usually designated by bi-clustering, co-clustering or block clustering, is an important technique in two way data analysis. A new standard and efficient approach has been recently proposed based on the latent block model (Govaert and Nadif 2003) which takes into account the block clustering problem on both the individual and variable sets. This article presents our R package blockcluster for co-clustering of binary, contingency and continuous data based on these very models. In this document, we will give a brief review of the model-based block clustering methods, and we will show how the R package blockcluster can be used for co-clustering.
A new approach is introduced in this article for describing and visualizing time series of curves, where each curve has the particularity of being subject to changes in regime. For this purpose, the curves are represented by a regression model including a latent segmentation, and their temporal evolution is modeled through a Gaussian random walk over low-dimensional factors of the regression coefficients. The resulting model is nothing else than a particular state-space model involving discrete and continuous latent variables, whose parameters are estimated across a sequence of curves through a dedicated variational Expectation-Maximization algorithm. The experimental study conducted on simulated data and real time series of curves has shown encouraging results in terms of visualization of their temporal evolution and forecasting.
Dans le cadre du diagnostic de systemes complexes, ou les donnees sont generalement collectees sous la forme de signaux, cet article aborde la problematique de la classification non supervisee de donnees dont les classes evoluent de maniere non stationnaire. Un modele de melange dont les parametres sont modelises de maniere stochastique est propose dans ce contexte, ainsi qu'un algorithme EM variationnel pour l'estimation des parametres de ce modele. Une etude experimentale est menee sur des donnees simulees.
Co-clustering leads to parsimony in data visualisation with a number of parameters dramatically reduced in comparison to the dimensions of the data sample. Herein, we propose a new generalized approach for nonlinear mapping by a re-parameterization of the latent block mixture model. The densities modeling the blocks are in an exponential family such that the Gaussian, Bernoulli and Poisson laws are particular cases. The inference of the parameters is derived from the block expectation–maximization algorithm with a Newton–Raphson procedure at the maximization step. Empirical experiments with textual data validate the interest of our generalized model.
A type of learning problem is considered, in which the class of training examples is only partially specified. Two approaches to such problems are described: the maximum likelihood approach, in which a probabilistic model relating the imprecise label to the true class is postulated, and the Transferable Belief Model approach, which relies on a non probabilistic formalism for representing and manipulating imprecise information. These two methods are compared experimentally using simulated data sets.
Mixmod is a well-established software package for fitting mixture models of multivariate Gaussian or multinomial probability distribution functions to a given dataset with either a clustering, a density estimation or a discriminant analysis purpose. The Rmixmod S4 package provides an interface from the R statistical computing environment to the C++ core library of Mixmod (mixmodLib). In this article, we give an overview of the model-based clustering and classification methods implemented, and we show how the R package Rmixmod can be used for clustering and discriminant analysis.
Semi-supervised classification can help to improve generative classifiers by taking into account the information provided by the unlabeled data points, especially when there are far more unlabeled data than labeled data. The aim is to select a generative classification model using both unlabeled and labeled data. A predictive deviance criterion, AICcond, aiming to select a parsimonious and relevant generative classifier in the semi-supervised context is proposed. In contrast to standard information criteria such as AIC and BIC, AICcond is focused on the classification task, since it attempts to measure the predictive power of a generative model by approximating its predictive deviance. However, it avoids the computational cost of cross-validation criteria, which make repeated use of the EM algorithm. AICcond is proved to have consistency properties that ensure its parsimony when compared with the Bayesian Entropy Criterion (BEC), whose focus is similar to that of AICcond. Numerical experiments on both simulated and real data sets show that the behavior of AICcond as regards the selection of variables and models, is encouraging when it is compared to the competing criteria.
A non linear regression approach which consists of a specific regression model incorporating a latent process, allowing various polynomial regression models to be activated preferentially and smoothly, is introduced in this paper. The model parameters are estimated by maximum likelihood performed via a dedicated expecation-maximization (EM) algorithm. An experimental study using simulated and real data sets reveals good performances of the proposed approach.
A new approach for signal parametrization, which consists of a specific regression model incorporating a discrete hidden logistic pro- cess, is proposed. The model parameters are estimated by the maximum likelihood method performed by a dedicated Expectation Maximization (EM) algorithm. The parameters of the hidden logistic process, in the inner loop of the EM algorithm, are estimated using a multi-class Itera- tive Reweighted Least-Squares (IRLS) algorithm. An experimental study using simulated and real data reveals good performances of the proposed approach.
Cet article propose une methode de regression non lineaire qui s'appuie sur un modele integrant un processus latent qui permet d'activer preferentiellement un modele de regression polynomial parmi K modeles. L'utilisation d'une fonction logistique comme loi conditionnelle des variables latentes assure une souplesse de transition (lente ou rapide) entre les differents polynomes, ce qui permet d'obtenir une modelisation correcte de non linearites. L'estimation des parametres du modele propose est effectuee par un algorithme EM dedie. Une etude experimentale menee sur des donnees simulees revele de bonnes performances de la methode proposee en termes de precision d'estimation, comparee a la methode de regression polynomiale par morceaux.
Résumé. L’analyse exploratoire de données multidimensionnelles e st un problème complexe. Nous proposons d’extraire certains invari ants topologiques appelés nombre de Betti, pour synthétiser la topologie de la st ructure sous-jacente aux données. Nous définissons un modèle génératif basé sur le complexe simplicial de Delaunay dont nous estimons les paramètres par l’ optimisation du critère d’information Bayésien (BIC). Ce Complexe Simplic ial Génératif nous permet d’extraire les nombres de Betti de données jouets et d ’images d’objets en rotation. Comparé à la technique géométrique des Witness Complex, le CSG apparait plus robuste aux données bruitées.
This paper addresses the problem of temporal data clustering using a dynamic Gaussian mixture model whose means are considered as latent variables distributed according to random walks. Its final objective is to track the dynamic evolution of some critical railway components using data acquired through embedded sensors. The parameters of the proposed algorithm are estimated by maximum likelihood via the Expectation-Maximization algorithm. In contrast to other approaches as the maximum a posteriori estimation in which the covariance matrices of the random walks have to be fixed by the user, the results of the simulations show the ability of the proposed algorithm to correctly estimate these covariances while keeping a low clustering error rate.
Studies based on human mobility, including Bicycle Sharing System analysis, has expanded over the past few years. They aim to give insight of the underlying urban phenomena linked to city dynamics. This paper presents a generative count-series model using adapted Poisson mixtures to automatically analyse and find temporal-based clusters over the Velib' origin-destination flow-data. Such an approach may provide latent factors that reveal how regions of different usage interact over the time. More generally, the proposed methodology can be used to cluster edges of temporal valued-graph with respect to their temporal profiles.
This paper proposes a method of segmenting temporal data into ordered classes. It is based on mixture models and a discrete latent process, which enables to successively activates the classes. The classification can be performed by maximizing the likelihood via the EM algorithm or by simultaneously optimizing the model parameters and the partition by the CEM algorithm. These two algorithms can be seen as alternatives to Fisher's algorithm, which improve its computing time.
Christophe Biernacki合作论文数UFR de Mathematiques - UMR CNRS 8524 (laboratoire Painleve), Universite des Sciences et Technologies de Lille (Lille 1)9