This paper addresses model order selection under large-dimensional, correlated, non-Gaussian noise. Sources are assumed to be embedded in additive Complex Elliptically Symmetric (CES) noise with an unknown Toeplitz-structured scatter matrix. We propose a two-stage robust framework: (i) a noise-whitening step based on a Toeplitz-rectified M-estimator of the scatter matrix, and (ii) signal subspace rank inference via large-dimensional Random Matrix Theory (RMT). Almost sure consistency of the proposed estimators is established, together with explicit RMT eigenvalue upper bounds separating signal from noise components, in the regime where the observation dimension m and the sample size N grow proportionally. Three estimation branches are derived, based respectively on the sample covariance matrix (SCM), Maronna's M-estimator, and the distribution-free Tyler M-estimator for whitening. The methodology is validated on synthetic data, real hyperspectral images, EEG recordings, and financial data, with significant gains over AIC and unwhitened methods.
This paper deals with Elliptical Wishart distributions - which generalize the Wishart distribution - in the context of signal processing and machine learning. Two algorithms to compute the maximum likelihood estimator (MLE) are proposed: a fixed point algorithm and a Riemannian optimization method based on the derived information geometry of Elliptical Wishart distributions. The existence and uniqueness of the MLE are characterized as well as the convergence of both estimation algorithms. Statistical properties of the MLE are also investigated such as consistency, asymptotic normality and an intrinsic version of Fisher efficiency. On the statistical learning side, novel classification and clustering methods are designed. For the t-Wishart distribution, the performance of the MLE and statistical learning algorithms are evaluated on both simulated and real EEG and hyperspectral data, showcasing the interest of our proposed methods.
Detection and identification of emitters provide vital information for defensive strategies in electronic intelligence. Based on a received signal containing pulses from an unknown number of emitters, this paper introduces an unsupervised methodology for deinterleaving RADAR signals based on a combination of clustering algorithms and optimal transport distances. The first step involves separating the pulses with a clustering algorithm under the constraint that the pulses of two different emitters cannot belong to the same cluster. Then, as the emitters exhibit complex behavior and can be represented by several clusters, we propose a hierarchical clustering algorithm based on an optimal transport distance to merge these clusters. A variant is also developed, capable of handling more complex signals. Finally, the proposed methodology is evaluated on simulated data provided through a realistic simulator. Results show that the proposed methods are capable of deinterleaving complex RADAR signals.
The multivariate generalized Gaussian distribution (MGGD), also known as the multivariate exponential power (MEP) distribution, is widely used in signal and image processing. However, estimating MGGD parameters, which is required in practical applications, still faces specific theoretical challenges. In particular, establishing convergence properties for the standard fixed-point approach when both the distribution mean and the scatter (or the precision) matrix are unknown is still an open problem. In robust estimation, imposing classical constraints on the precision matrix, such as sparsity, has been limited by the non-convexity of the resulting cost function. This paper tackles these issues from an optimization viewpoint by proposing a convex formulation with well-established convergence properties. We embed our analysis in a noisy scenario where robustness is induced by modelling multiplicative perturbations. The resulting framework is flexible as it combines a variety of regularizations for the precision matrix, the mean and model perturbations. This paper presents proof of the desired theoretical properties, specifies the conditions preserving these properties for different regularization choices and designs a general proximal primal-dual optimization strategy. The experiments show a more accurate precision and covariance matrix estimation with similar performance for the mean vector parameter compared to Tyler's M-estimator. In a high-dimensional setting, the proposed method outperforms the classical GLASSO, one of its robust extensions, and the regularized Tyler's estimator.
Though very popular, it is well known that the Expectation-Maximisation (EM) algorithm for the Gaussian mixture model performs poorly for non-Gaussian distributions or in the presence of outliers or noise. In this paper, we propose a Flexible EM-like Clustering Algorithm (FEMCA): a new clustering algorithm following an EM procedure is designed. It is based on both estimations of cluster centers and covariances. In addition, using a semi-parametric paradigm, the method estimates an unknown scale parameter per data point. This allows the algorithm to accommodate heavier tail distributions, noise, and outliers without significantly losing efficiency in various classical scenarios. We first present the general underlying model for independent, but not necessarily identically distributed, samples of elliptical distributions. We then derive and analyze the proposed algorithm in this context, showing in particular important distribution-free properties of the underlying data distributions. The algorithm convergence and accuracy properties are analyzed by considering the first synthetic data. Finally, we show that FEMCA outperforms other classical unsupervised methods of the literature, such as k-means, EM for Gaussian mixture models, and its recent modifications or spectral clustering when applied to real data sets as MNIST, NORB, and 20newsgroups .
This paper deals with the Elliptical Wishart and Inverse Elliptical Wishart distributions, which play a major role when handling covariance matrices. Similarly to multivariate elliptical distributions, these form a large family of covariance distributions, encompassing, e.g., the Wishart or t-Wishart ones. Our first major contribution is to derive a stochastic representation for Elliptical Wishart and Inverse Elliptical Wishart matrices. This later enables us to obtain various key statistical properties of Elliptical Wishart and Inverse Elliptical Wishart distributions such as expectations, variances, and Kronecker moments up to any orders. The stochastic representation also allows us to provide an efficient method to generate random matrices from Elliptical Wishart and Inverse Elliptical Wishart distributions. Finally, the practical interest of Elliptical Wishart distributions - in particular the t-Wishart one - is demonstrated through a fitting experiment on real electroencephalographic data. This showcases their effectiveness in accurately modeling real covariance matrices.
The Fermat distance has been recently established as a useful tool for machine learning tasks when a natural distance is not directly available to the practitioner or to improve the results given by Euclidean distances by exploding the geometrical and statistical properties of the dataset. This distance depends on a parameter $\alpha$ that greatly impacts the performance of subsequent tasks. Ideally, the value of $\alpha$ should be large enough to navigate the geometric intricacies inherent to the problem. At the same, it should remain restrained enough to sidestep any deleterious ramifications stemming from noise during the process of distance estimation. We study both theoretically and through simulations how to select this parameter.
In this study, we consider the realm of covariance matrices in machine learning, particularly focusing on computing Fréchet means on the manifold of symmetric positive definite matrices, commonly referred to as Karcher or geometric means. Such means are leveraged in numerous machine-learning tasks. Relying on advanced statistical tools, we introduce a random matrix theory-based method that estimates Fréchet means, which is particularly beneficial when dealing with low sample support and a high number of matrices to average. Our experimental evaluation, involving both synthetic and real-world EEG and hyperspectral datasets, shows that we largely outperform state-of-the-art methods.
Sparse principal component analysis (PCA) aims at mapping large dimensional data to a linear subspace of lower dimension. By imposing loading vectors to be sparse, it performs the double duty of dimension reduction and variable selection. Sparse PCA algorithms are usually expressed as a trade-off between explained variance and sparsity of the loading vectors (i.e., number of selected variables). As a high explained variance is not necessarily synonymous with relevant information, these methods are prone to select irrelevant variables. To overcome this issue, we propose an alternative formulation of sparse PCA driven by the false discovery rate (FDR). We then leverage the Terminating-Random Experiments (T-Rex) selector to automatically determine an FDR-controlled support of the loading vectors. A major advantage of the resulting T-Rex PCA is that no sparsity parameter tuning is required. Numerical experiments and a stock market data example demonstrate a significant performance improvement.
Gradient-Boosted Decision Trees (GBDT) stand out as a powerful Machine Learning tool in tackling classification and regression tasks. Despite its effectiveness, GBDT, like other ensemble methods, suffers from a lack of explainability. Understanding these models' inner workings is crucial for comprehensively grasping their decision-making processes. In this study, we propose a method to enhance the explainability of GBDT, focusing on identifying specific training data points termed “comparable samples,” which play a pivotal role in the model's predictions. Inspired by the Frank-Wolfe algorithm, we introduce Explainable Gradient Boosting (ExpGB), which aims to shed light on the relationships between input data attributes and model predictions. ExpGB operates by ranking training samples based on their decomposition coefficients within the algorithm's output. Higher weights assigned to particular training samples indicate a closer resemblance to the sample being analyzed. To validate the efficiency of our approach, we conduct a comparative analysis across three diverse datasets, contrasting ExpGB with traditional GBDT algorithms. Through this analysis, we evaluate the quality of estimation provided by ExpGB, thereby enhancing our understanding of GBDT's workings.
Though very popular, it is well known that the Expectation-Maximisation (EM) algorithm for the Gaussian mixture model performs poorly for non-Gaussian distributions or in the presence of outliers or noise. In this paper, we propose a Flexible EM-like Clustering Algorithm (FEMCA): a new clustering algorithm following an EM procedure is designed. It is based on both estimations of cluster centers and covariances. In addition, using a semi-parametric paradigm, the method estimates an unknown scale parameter per data point. This allows the algorithm to accommodate heavier tail distributions, noise, and outliers without significantly losing efficiency in various classical scenarios. We first present the general underlying model for independent, but not necessarily identically distributed, samples of elliptical distributions. We then derive and analyze the proposed algorithm in this context, showing in particular important distribution-free properties of the underlying data distributions. The algorithm convergence and accuracy properties are analyzed by considering the first synthetic data. Finally, we show that FEMCA outperforms other classical unsupervised methods of the literature, such as k-means, EM for Gaussian mixture models, and its recent modifications or spectral clustering when applied to real data sets as MNIST, NORB, and 20newsgroups .
Linear and Quadratic Discriminant Analysis (LDA and QDA) are well-known classical methods but can heavily suffer from non-Gaussian distributions and/or contaminated datasets, mainly because of the underlying Gaussian assumption that is not robust. This paper studies the robustness to scale changes in the data of a new discriminant analysis technique where each data point is drawn by its own arbitrary Elliptically Symmetrical (ES) distribution and its own arbitrary scale parameter. Such a model allows for possibly very heterogeneous, independent but non-identically distributed samples. The new decision rule derived is simple, fast, and robust to scale changes in the data compared to other state-of-the-art method
This paper provides a new classification method of covariance matrices exploiting the $t$ -Wishart distribution, which generalizes the Wishart distribution. Compared to the Wishart distribution, it is more robust to aberrant covariance matrices and more flexible to distribution mismatch. Following recent developments on this matrix-variate distribution, the proposed classifier is obtained by leveraging the Discriminant Analysis framework and providing original decision rules. The practical interest of our approach is shown thanks to numerical experiments on real data. More precisely, the proposed classifier yields the best results on two standard electroencephalography datasets compared to the best state-of-the-art minimum distance-to-mean (MDM) classifiers.
Although linear and quadratic discriminant analysis are widely recognized classical methods, they can encounter significant challenges when dealing with non-Gaussian distributions or contaminated datasets. This is primarily due to their reliance on the Gaussian assumption, which lacks robustness. We first explain and review the classical methods to address this limitation and then present a novel approach that overcomes these issues. In this new approach, the model considered is an arbitrary Elliptically Symmetrical (ES) distribution per cluster with its own arbitrary scale parameter. This flexible model allows for potentially diverse and independent samples that may not follow identical distributions. By deriving a new decision rule, we demonstrate that maximum-likelihood parameter estimation and classification are simple, efficient, and robust compared to state-of-the-art methods.
We propose estimating the scale parameter (mean of the eigenvalues) of the scatter matrix of an unspecified elliptically symmetric distribution using weights obtained by solving Tyler's M-estimator of the scatter matrix. The proposed Tyler's weights-based estimate (TWE) of scale is then used to construct an affine equivariant Tyler's M-estimator as a weighted sample covariance matrix using normalized Tyler's weights. We then develop a unified framework for estimating the unknown tail parameter of the elliptical distribution (such as the degrees of freedom (d.o.f.) $\nu$ of the multivariate $t$ (MVT) distribution). Using the proposed TWE of scale, a new robust estimate of the d.o.f. parameter of MVT distribution is proposed with excellent performance in heavy-tailed scenarios, outperforming other competing methods. R-package is available that implements the proposed method.
This work deals with elliptical Wishart distributions on the set of symmetric positive definite matrices.It contains two major contributions.First, the information geometry associated with elliptical Wishart distributions is derived.Second, this geometry is leveraged to propose Riemannian-optimization-based maximum likelihood estimators of any elliptical Wishart distribution.Particular attention is given to two specific distributions: the tand Kotz Wishart ones.The performance of the proposed methods is assessed through numerical experiments on simulated data.
The article proposes and theoretically analyses a \emph{computationally efficient} multi-task learning (MTL) extension of popular principal component analysis (PCA)-based supervised learning schemes \cite{barshan2011supervised,bair2006prediction}. The analysis reveals that (i) by default learning may dramatically fail by suffering from \emph{negative transfer}, but that (ii) simple counter-measures on data labels avert negative transfer and necessarily result in improved performances. Supporting experiments on synthetic and real data benchmarks show that the proposed method achieves comparable performance with state-of-the-art MTL methods but at a \emph{significantly reduced computational cost}.
In this article, covariance matrix estimation of compound-Gaussian vectors with texture-correlation (spatial correlation for the adaptive radar detectors) is examined. The texture parameters are treated as hidden random parameters whose statistical description is given by a Markov chain. States of the chain represent the value of texture coefficient and the transition probabilities establish the correlation in the texture sequence. An expectation–maximization (EM) method-based covariance matrix estimation solution is given for both noiseless and noisy snapshots. An extension to the practically important case of persymmetric covariance matrices is developed and possible extensions to other structured covariance matrices are described. The numerical results indicate that the benefit of utilizing spatial correlation in the covariance matrix estimation can be significant especially when the total number of snapshots in the secondary data is small. From applications viewpoint, the suggested model is well suited for the adaptive target detection in sea clutter, where some spatial correlation between range cells has been experimentally observed. The performance improvements of the suggested approach for small number of snapshots can be particularly important in this application area.
Expectation-Maximization (EM) algorithm is a widely used iterative algorithm for computing (local) maximum likelihood estimate (MLE). It can be used in an extensive range of problems, including the clustering of data based on the Gaussian mixture model (GMM). Numerical instability and convergence problems may arise in situations where the sample size is not much larger than the data dimensionality. In such low sample support (LSS) settings, the covariance matrix update in the EM-GMM algorithm may become singular or poorly conditioned, causing the algorithm to crash. On the other hand, in many signal processing problems, a priori information can be available indicating certain structures for different cluster covariance matrices. In this paper, we present a regularized EM algorithm for GMM-s that can make efficient use of such prior knowledge as well as cope with LSS situations. The method aims to maximize a penalized GMM likelihood where regularized estimation may be used to ensure positive definiteness of covariance matrix updates and shrink the estimators towards some structured target covariance matrices. We show that the theoretical guarantees of convergence hold, leading to better performing EM algorithm for structured covariance matrix models or with low sample settings.