In the era of information explosion, clustering analysis of multi-view data plays a crucial role in revealing the intrinsic structures of data. Despite the advancements in existing multi-view clustering methods for processing complex data, they often overlook the weight differences among various views and the diversity between clusters. To address the issues, the paper introduces a novel multi-view clustering approach termed weight consistency and cluster diversity based concept factorization for multi-view clustering (MVCF-WD). Specifically, the proposed method automatically learns the weights of the views, and incorporates a cluster diversity term to enhance the discriminability of clusters. Furthermore, to solve the formulated optimization model, an iterative optimization algorithm based on multiplication rules is developed and the convergence is analyzed. Extensive experiments conducted across seven datasets compared with ten state-of-the-art clustering algorithms demonstrate the superior clustering performance of the proposed method.
Graph-based multiview clustering (MVC) approaches have demonstrated impressive performance by leveraging the consistency properties of multiview data in an unsupervised manner. However, existing methods for graph learning heavily rely on either Euclidean structures or the manifold topological structures derived from fixed view-specific graphs. Unfortunately, these approaches may not accurately reflect the consensus topological structure in a multiview setting. To address this limitation and enhance the intrinsic graph learning process, an adaptive exploration of a more appropriate consistency topological structure is required. Toward this end, we propose a novel approach called collaborative topological graph learning (CTGL) for MVC. The key idea is to adaptively discover the consistent topological structure to guide intrinsic graph learning. We achieve this by introducing an auxiliary consistency graph that formulates the topological relevance learning function. However, estimating the auxiliary consistency graph is not straightforward, as it is based on the learned view-specific graphs and requires prior availability. To overcome this challenge, we develop a collaborative learning strategy that simultaneously learns both the auxiliary consistency graph and view-specific graphs using tensor learning techniques. This strategy enables the adaptive exploration of the consistency topological structure during graph learning, resulting in more accurate clustering outcomes. Extensive experiments are provided to show the effectiveness of the proposed method. The source code can be found at https://github.com/CLiu272/CTGL.
Unbalanced incomplete multiview data are widely generated in engineering areas due to sensor failures, data acquisition limitations, etc. However, current research works are rarely focused on unbalanced incomplete multiview unsupervised feature selection (MUFS). To address this issue, this article proposes an MUFS method called unbalanced incomplete multiview unsupervised feature selection with low-redundancy constraint in low-dimensional space (UIMUFSLR). Specifically, the proposed method mitigates the impact of missing samples by learning a unified graph with assigning weights of samples adaptively. In addition, a novel regularization is designed by utilizing the inner product of selected features to obtain low redundancy. An iterative optimization algorithm is devised for UIMUFSLR, accompanied by a comprehensive analysis of its convergence behavior and computational complexity. Experimental results demonstrate the competitiveness of UIMUFSLR in handling unbalanced incomplete multiview data on seven public datasets.
Non-negative Matrix Factorization (NMF) has been an ideal tool for machine learning. Non-negative Matrix Tri-Factorization (NMTF) is a generalization of NMF that incorporates a third non-negative factorization matrix, and has shown impressive clustering performance by imposing simultaneous orthogonality constraints on both sample and feature spaces. However, the performance of NMTF dramatically degrades if the data are contaminated with noises and outliers. Furthermore, the high-order geometric information is rarely considered. In this paper, a Robust NMTF with Dual Hyper-graph regularization (namely RDHNMTF) is introduced. Firstly, to enhance the robustness of NMTF, an improvement is made by utilizing the l(2,1)-norm to evaluate the reconstruction error. Secondly, a dual hyper-graph is established to uncover the higher-order inherent information within sample space and feature spaces for clustering. Furthermore, an alternating iteration algorithm is devised, and its convergence is thoroughly analyzed. Additionally, computational complexity is analyzed among comparison algorithms. The effectiveness of RDHNMTF is verified by benchmarking against ten cuttina-edae alaorithms across seven datasets corrupted with four types of noise.
Recent advancements in multi-view unsupervised feature selection (MUFS) have been notable, yet two primary challenges persist. First, real-world datasets frequently consist of unbalanced incomplete multi-view data, a scenario not adequately addressed by current MUFS methodologies. Second, the inherent complexity and heterogeneity of multi-view data often introduce significant noise, an aspect largely neglected by existing approaches, compromising their noise robustness. To tackle these issues, this paper introduces a Tensor-Based Error Robust Unbalanced Incomplete Multi-view Unsupervised Feature Selection (TERUIMUFS) strategy. The proposed MUFS framework specifically caters to unbalanced incomplete multi-view data, incorporating self-representation learning with a tensor low-rank constraint and sample diversity learning. This approach not only mitigates errors in the self-representation process but also corrects errors in the self-representation tensor, significantly enhancing the model’s resilience to noise. Furthermore, graph learning serves as a pivotal link between MUFS and self-representation learning. An innovative iterative optimization algorithm is developed for TERUIMUFS, complete with a thorough analysis of its convergence and computational complexity. Experimental results demonstrate TERUIMUFS’s effectiveness and competitiveness in addressing unbalanced incomplete multi-view unsupervised feature selection (UIMUFS), marking a significant advancement in the field.
In this work, we proposed an advanced deep clustering approach that leverages a pre-trained image encoder from CLIP to enhance image clustering by refining the learned feature representations. Specifically, this approach first designs a consistency loss mechanism to ensure the alignment of pseudo-labels generated by the clustering head with feature representations derived from the instance head. This mechanism facilitates the acquisition of more reliable and coherent feature representations. Second, in lieu of conventional strong data augmentation techniques, this approach employs a random masking strategy to enhance the diversity of the image datasets, enabling the model to prioritize finer image details and potentially improve its generalization ability. The efficacy and superior performance of this framework have been demonstrated through extensive experimentation across six public image datasets: STL-10, CIFAR-10, CIFAR-100, ImageNet-Dog, ImageNet-10 and Tiny-ImageNet. Notably, the proposed method achieves impressive results on the CIFAR-100 dataset, surpassing existing techniques by up to 11% in Normalized Mutual Information (NMI). It records scores of 0.562 for Accuracy, 0.586 for NMI, and 0.413 for Adjusted Rand Index. Ablation studies further underscore the individual contributions of the feature model, consistency loss, and random masking components to the framework's overall performance. The code can be available at https://github.com/MMengjuan-Li/Contrastive-Learning-and-CLIP-for-Clustering.
As a core branch of financial forecasting, stock forecasting plays a crucial role for financial analysts, investors, and policymakers in managing risks and optimizing investment strategies, significantly enhancing the efficiency and effectiveness of economic decision-making. With the rapid development of information technology and computer science, data-driven neural network technologies have increasingly become the mainstream method for stock forecasting. Although recent review studies have provided a basic introduction to deep learning methods, they still lack detailed discussion on network architecture design and innovative details. Additionally, the latest research on emerging large language models and neural network structures has yet to be included in existing review literature. In light of this, this paper comprehensively reviews the literature on data-driven neural networks in the field of stock forecasting from 2015 to 2023, discussing various classic and innovative neural network structures, including Recurrent Neural Networks (RNNs), Convolutional Neural Networks (CNNs), Transformers, Graph Neural Networks (GNNs), Generative Adversarial Networks (GANs), and Large Language Models (LLMs). It analyzes the application and achievements of these models in stock market forecasting. Moreover, the article also outlines the commonly used datasets and various evaluation metrics in the field of stock forecasting, further exploring unresolved issues and potential future research directions, aiming to provide clear guidance and reference for researchers in stock forecasting.
Multi-view clustering methods based on deep matrix factorization play a vital role in data analysis within the healthcare sector. However, existing methods predominantly conduct deep matrix factorization in the original data space, which is not conducive to addressing non-linear and complex data patterns. To address this issue, the Multi-kernel based Multi-view Deep Non-negative Matrix Factorization with Optimal Consensus Graph (OGMKMDNMF) is introduced. This approach utilizes deep non-negative matrix factorization after projecting the data matrix into a high-dimensional kernel space. Additionally, it employs optimal consensus graph to alleviate the detrimental effects arising from misassigned nearest neighbors during the construction of similarity matrix. An innovative iterative optimization algorithm is developed for OGMKMDNMF. The experimental results demonstrate the effectiveness and competitive advantage of OGMKMDNMF in addressing multi-view healthcare data clustering tasks.
Incomplete multi-view clustering addresses scenarios where data completeness cannot be guaranteed, diverging from traditional methods that assume fully observed features. Existing approaches often overlook high-order correlations present in multiple similarity graphs, and suffer from inefficiencies due to iterative optimization procedures. To overcome these limitations, we propose a graph-based model leveraging graph propagation to effectively handle incomplete data. The proposed method translates incomplete instances into incomplete graphs, and infers missing entries through a graph propagation strategy, ensuring the inferred data is meaningful and contextually relevant. Specifically, a self-guided graph is constructed to capture global relationships, while partial graphs represent view-specific similarities. The self-guided graph is first completed through self-guided graph propagation, which subsequently aids in the propagation of the partial graphs. The key contribution of graph propagation is to propagate information from complete data to incomplete data. Furthermore, the high-order correlation across multiple views is captured by low-rank tensor learning. To enhance computational efficiency, the optimization procedure is decoupled and implemented in a stepwise manner, eliminating the need for iterative updates. Extensive experiments validate the robustness of the proposed method, demonstrating superior performance compared to state-of-the-art methods, even when all instances are incomplete.
As the advancement of sensing technology and cyber-physical equipment, the complexity and dimensionality of the data from consumer internet of things have increased dramatically. Processing the complex and multi-view data to analyze the consumption characteristics becomes a significant issue. Deep matrix factorization (DMF) is a widely spread method with the ability to learn intrinsic components in data based on its multiple-layers structures. However, the multiple-layers structures may magnify noises in data and the standard DMF cannot reflect the manifold and low-rank structures on the coefficient matrix well. In the paper, a robust deep matrix factorization with low-rank and hypergraph learning (RDMFLRH) model is proposed for multi-view data processing. Firstly, an arc tangent loss function with less sensitivity to noises and outlines is introduced to enhance robustness. Additionally, low-rank and hypergraph learning are adopted to extract the relationships between clusters and maintain the binary manifold structures from the original space. Then, an optimization algorithm based on the multiplicative update rule is developed to solve the proposed model with convergence proven theoretically and experimentally. Finally, abundant experiments on real-world and noisy datasets validate the superiority of the proposed method, improving around 1% to 10% performance compared with the comparison state-of-the-art methods.
Incomplete multi-view clustering (IMVC) aims to address the clustering problem of multi-view data with partially missing samples and has received widespread attention in recent years. Most existing IMVC methods still have the following issues that require to be further addressed. They focus solely on the first-order correlation information among samples, neglecting the more intricate high-order connections. Additionally, these methods always overlook the noise or inaccuracies in the self-representation matrix. To address above issues, a novel method named Robust Mixed-order Graph Learning (RMoGL) is proposed for IMVC. Specifically, to enhance the robustness to noise, the self-representation matrices are separated into clean graphs and noise graphs. To capture complex high-order relationships among samples, the dynamic high-order similarity graphs are innovatively constructed from the recovered data. The clean graphs are endowed with mixed-order information and tend towards to obtain a consensus graph via a self-weighted manner. An efficient algorithm based on Alternating Direction Method of Multipliers (ADMM) is designed to solve the proposed RMoGL, and superior performance is demonstrated by compared with nine state-of-the-art methods across eight datasets. The source code of this work is available at https://github.com/guowei1314/RMoGL.
This article investigates a class of systems of nonlinear equations (SNEs). Three distributed neurodynamic models (DNMs), namely a two-layer model (DNM-I) and two single-layer models (DNM-II and DNM-III), are proposed to search for such a system's exact solution or a solution in the sense of least-squares. Combining a dynamic positive definite matrix with the primal-dual method, DNM-I is designed and it is proved to be globally convergent. To obtain a concise model, based on the dynamic positive definite matrix, time-varying gain, and activation function, DNM-II is developed and it enjoys global convergence. To inherit DNM-II's concise structure and improved convergence, DNM-III is proposed with the aid of time-varying gain and activation function, and this model possesses global fixed-time consensus and convergence. For the smooth case, DNM-III's globally exponential convergence is demonstrated under the Polyak-Łojasiewicz (PL) condition. Moreover, for the nonsmooth case, DNM-III's globally finite-time convergence is proved under the Kurdyka-Łojasiewicz (KL) condition. Finally, the proposed DNMs are applied to tackle quadratic programming (QP), and some numerical examples are provided to illustrate the effectiveness and advantages of the proposed models.
Deep matrix factorization (DMF) has the capability to discover hierarchical structures within raw data by factorizing matrices layer by layer, allowing it to utilize latent information for superior clustering performance. However, DMF-based approaches face limitations when dealing with complex and nonlinear raw data. To address this issue, Auto-weighted Multi-view Deep Nonnegative Matrix Factorization with Multi-kernel Learning (MvMKDNMF) is proposed by incorporating multi-kernel learning into deep nonnegative matrix factorization. Specifically, samples are mapped into the kernel space which is a convex combination of several predefined kernels, free from selecting kernels manually. Furthermore, to preserve the local manifold structure of samples, a graph regularization is embedded in each view and the weights are assigned adaptively to different views. An alternate iteration algorithm is designed to solve the proposed model, and the convergence and computational complexity are also analyzed. Comparative experiments are conducted across nine multi-view datasets against seven state-of-the-art clustering methods showing the superior performances of the proposed MvMKDNMF.
Tensor based subspace clustering, a method aimed at partitioning multi-view data into distinct clusters through the tensor low-rank representation, has gained widespread popularity. Despite the effectiveness of tensor singular value decomposition (t-SVD) based nuclear norm in extracting high-dimensional information from multiple views, certain limitations persist. Firstly, while the commonly used l2,1 norm enhances clustering robustness, it remains susceptible to the impact of high-intensity noises. Secondly, the t-SVD based nuclear norm provides a biased estimation of tensor rank. Thirdly, existing methods either fail to preserve manifold structures from the original space or only extract binary relationships between data points. To address these challenges and enhance clustering performance in both high-intensity noisy and real-world conditions, an error-robust multi-view subspace clustering with nonconvex low-rank tensor approximation and hyper-Laplacian graph embedding (EMSC-NLTHG) method is proposed. Specifically, Cauchy loss function which reduces sensitivity to larger noises and outliers is introduced. Based on the Cauchy loss function, we develop the Cauchy pseudo norm to represent the reconstruction error in clustering to enhance the robustness. Moreover, rather than uniformly weighting all singular values in the t-SVD nuclear norm, a nonconvex tensor nuclear norm is adopted to add weights adaptively and approximate the true tensor rank. Additionally, hypergraph is embedded in the representation matrices to preserve local manifold structures and extract the complex multivariate relationships between data points. Extensive experiments demonstrate that the proposed EMSC-NLTHG method consistently outperforms state-of-the-art techniques by approximately 18% to 40% on datasets corrupted by high-intensity noises and around 1% to 27% on eight real-world datasets across six popular evaluation metrics.
Multi-view data processing is an effective tool to differentiate the levels of consumers on electronics. Recently, the graph based multi-view clustering methods have attracted widespread attention because they can obtain the relationships of multi-view data points efficiently. However, there exist several shortcomings on most existing graph based clustering methods. Firstly, the mostly adopted Euclidean distance can not extract the nonlinear manifold structure. Secondly, graph based methods are mainly hard clustering methods, which means that each data point belongs to only the one cluster exactly. Thirdly, the high-dimension information between multiple views are not taken into account. Thus, a low-rank tensor regularized graph fuzzy learning (LRTGFL) method for multi-view data processing is proposed. In LRTGFL, Jensen-Shannon divergence is adopted to replace the Euclidean distance for obtaining more completely nonlinear structures. In addition, fuzzy learning is adopted to make graph clustering be a soft clustering method. Furthermore, a tensor nuclear norm based on the tensor singular value decomposition (t-SVD) is adopted to take advantage of the high-dimension information. Then, alternating direction method of multipliers (ADMM) is adopted to solve the LRTGFL model. Finally, the effectiveness and superiority of LRTGFL are demonstrated by comparing with various state-of-the-art algorithms on eight real-world datasets.
Classifying fruits and vegetables is a challenging task for traditional machine learning models, particularly convolutional neural networks (CNNs), which struggle to differentiate between similar-looking items. This differentiation is crucial in agriculture for efficient sorting and quality control. This study aimed to enhance the performance of classification models using a dataset of 3,115 images across 36 fruit and vegetable classes from Kaggle. The research explored four architectures: CNN, MobileNet, DenseNet, and Xception. Fine-tuning each model for the dataset, the CNN achieved a baseline accuracy of 96
Due to the scarcity of data labels, unsupervised feature selection has received a lot of attention in recent years. While many unsupervised feature selection methods are capable of selecting relevant features, they often fail to comprehensively consider the impact of both local and global information of the data on feature selection, nor can they effectively handle the complex nonlinear relationships commonly found in real-world data. As a result, suboptimal feature subsets are often selected. In this paper, inspired by the Uniform Manifold Approximation and Projection (UMAP) manifold learning technique and the nonlinear sparse learning method based on Feature-Wise Kernelized Lasso, we propose a novel unsupervised feature selection method called Multi-Cluster Unsupervised Nonlinear Feature Selection based on UMAP and block HSIC Lasso (MUNFS). MUNFS greatly improves the representation of high-dimensional data during dimensionality reduction and effectively handles complex nonlinear relationships in such data. Specifically, by capturing the intrinsic topology of the data, MUNFS accurately preserves the local structure of the data while keeping as much of the global structure as possible. Furthermore, the kernel-based Hilbert–Schmidt Independence Criterion (HSIC) may measure the nonlinear dependency between the features and the target variables, while applying the l1 regularization term in feature selection to achieve sparsity. This allows for a more precise assessment of the significance of each feature. Extensive experimental results on five benchmark datasets and eight hyperspectral datasets demonstrate that the MUNFS method performs much better than several other feature selection methods.
Incomplete Multi-View Clustering (IMVC) is a promising topic in multimedia as it breaks the data completeness assumption. Most existing methods solve IMVC from the perspective of graph learning. In contrast, self-representation learning enjoys a superior ability to explore relationships among samples. However, only a few works have explored the potentiality of self-representation learning in IMVC. These self-representation methods infer missing entries from the perspective of whole samples, resulting in redundant information. In addition, designing an effective strategy to retain salient features while eliminating noise is rarely considered in IMVC. To tackle these issues, we propose a novel self-representation learning method with missing sample recovery and enhanced low-rank tensor regularization. Specifically, the missing samples are inferred by leveraging the local structure of each view, which is constructed from available samples at the feature level. Then an enhanced tensor norm, referred to as Logarithm-p norm is devised, which can obtain an accurate cross-view description by adaptive weights. Our proposed method achieves exact subspace representation in IMVC by leveraging high-order correlations and inferring missing information at the feature level. Extensive experiments on several widely used multi-view datasets demonstrate the effectiveness of the proposed method.
Incomplete multi-view clustering (IMVC) presents a significant challenge due to the need for effectively exploring complementary and consistent information within the context of missing views. One promising strategy to tackle this challenge is to recover missing views by inferring the missing samples. However, such approaches often fail to fully utilize discriminative structural information or adequately address consistency, as it requires such information to be known or learnable in advance, which contradicts the incomplete data setting. In this study, we propose a novel approach called Latent Structure-Aware view recovery (LaSA) for the IMVC task. Our objective is to recover missing views through discriminative latent representations by leveraging structural information. Specifically, our method offers a unified closed-form formulation that simultaneously performs missing data inference and latent representation learning, using a learned intrinsic graph as structural information. This formulation, incorporating graph structure information, enhances the inference of missing data while facilitating discriminative feature learning. Even when intrinsic graph is initially unknown due to incomplete data, our formulation allows for effective view recovery and intrinsic graph learning through an iterative optimization process. To further enhance performance, we introduce an iterative consistency diffusion process, which effectively leverages the consistency and complementary information across multiple views. Extensive experiments demonstrate the effectiveness of the proposed method compared to state-of-the-art approaches.
In this work, we study a more realistic challenging scenario in multiview clustering (MVC), referred to as incomplete MVC (IMVC) where some instances in certain views are missing. The key to IMVC is how to adequately exploit complementary and consistency information under the incompleteness of data. However, most existing methods address the incompleteness problem at the instance level and they require sufficient information to perform data recovery. In this work, we develop a new approach to facilitate IMVC based on the graph propagation perspective. Specifically, a partial graph is used to describe the similarity of samples for incomplete views, such that the issue of missing instances can be translated into the missing entries of the partial graph. In this way, a common graph can be adaptively learned to self-guide the propagation process by exploiting the consistency information, and the propagated graph of each view is in turn used to refine the common self-guided graph in an iterative manner. Thus, the associated missing entries can be inferred through graph propagation by exploiting the consistency information across all views. On the other hand, existing approaches focus on the consistency structure only, and the complementary information has not been sufficiently exploited due to the data incompleteness issue. By contrast, under the proposed graph propagation framework, an exclusive regularization term can be naturally adopted to exploit the complementary information in our method. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods. The source code of our method is available at the https://github.com/CLiu272/TNNLS-PGP.
Hau-San Wong (黄厚生)合作论文数Department of Computer Science, City University of Hong Kong3