With the rapid growth of activities on the web, large amounts of interaction data on multimedia platforms are easily accessible, including e-commerce, music sharing, and social media. By discovering various interests of users, recommender systems can improve user satisfaction without accessing overwhelming personal information. Compared to graph-based models, hypergraph-based collaborative filtering has the ability to model higher-order relations besides pair-wise relations among users and items, where the hypergraph structures are mainly obtained from specialized data or external knowledge. However, the above well-constructed hypergraph structures are often not readily available in every situation. To this end, we first propose a novel framework named HGRec, which can enhance recommendation via automatic hypergraph generation. By exploiting the clustering mechanism based on the user/item similarity, we group users and items without additional knowledge for hypergraph structure learning and design a cross-view recommendation module to alleviate the combinatorial gaps between the representations of the local ordinary graph and the global hypergraph. Furthermore, we devise a sparse optimization strategy to ensure the effectiveness of hypergraph structures, where a novel integration of the $\ell _{2,1}$ -norm and optimal transport framework is designed for hypergraph generation. We term the model HGRec with sparse optimization strategy as HGRec++. Extensive experiments on public multi-domain datasets demonstrate the superiority brought by our HGRec++, which gains average 8.1 $\%$ and 9.8 $\%$ improvement over state-of-the-art baselines regarding Recall and NDCG metrics, respectively.
Graph convolutional network (GCN) with the powerful capacity to explore graph-structural data has gained noticeable success in recent years. Nonetheless, most of the existing GCN-based models suffer from the notorious over-smoothing issue, owing to which shallow networks are extensively adopted. This may be problematic for complex graph datasets because a deeper GCN should be beneficial to propagating information across remote neighbors. Recent works have devoted effort to addressing over-smoothing problems, including establishing residual connection structure or fusing predictions from multilayer models. Because of the indistinguishable embeddings from deep layers, it is reasonable to generate more reliable predictions before conducting the combination of outputs from various layers. In light of this, we propose an alternating graph-regularized neural network (AGNN) composed of graph convolutional layer (GCL) and graph embedding layer (GEL). GEL is derived from the graph-regularized optimization containing Laplacian embedding term, which can alleviate the over-smoothing problem by periodic projection from the low-order feature space onto the high-order space. With more distinguishable features of distinct layers, an improved Adaboost strategy is utilized to aggregate outputs from each layer, which explores integrated embeddings of multi-hop neighbors. The proposed model is evaluated via a large number of experiments including performance comparison with some multilayer or multi-order graph neural networks, which reveals the superior performance improvement of AGNN compared with the state-of-the-art models.
Transformers designed for natural language processing have originally been explored for computer vision in recent research. Various Vision Transformers (ViTs) play an increasingly important role in the field of image tasks such as computer vision, multimodal fusion and multimedia analysis. However, to obtain promising performance, most existing ViTs usually rely on artificially filtered high-quality images, which may suffer from inherent noise risk. Generally, such well-constructed images are not always available in every situation. To this end, we propose a Robust ViT (RViT) to focus on the relevant and robust representation learning for image classification tasks. Specifically, we first develop a novel Denoising VTUnet module, where we conceptualize the nonrobust noise as the uncertainty under the variational conditions. Furthermore, we design a fusion transformer backbone with a tailored fusion attention mechanism to perform image classification based on the extracted robust representations effectively. To demonstrate the superiority of our model, the compared experiments are conducted on several popular datasets. Benefiting from the sequence regularity of the Transformer and captured robust feature, the proposed method exceeds compared Transformer-based models with superior performance in visual tasks.
Multimedia recommender systems (MRS) become prevalent due to their rich multimodal data (e.g., visual and textual content). Recent advancements have leveraged Graph Neural Networks (GNNs) to integrate these data, they often fall short in capturing the complex high-order relations within multimodal data, but readily hypergraph structures are not always available. To this end, we introduce the HMRec framework, a novel approach in Heterogeneous Hypergraph Structure Learning tailored for MRS. Specifically, we formulate the construction of a heterogeneous hypergraph as determining item associations across modalities, and introduce an adaptive hypergraph convolution mechanism for differentially weighting multimodal hyperedges. Furthermore, we propose an enhanced multimedia recommendation module, which introduces a contrastive fusion mechanism to effectively integrate graph-view, hypergraph-view, and ID-specific embeddings. Extensive experiments on real-world multimodal datasets show the superiority of our proposed HMRec framework in offering great potential for multimedia recommendations over the state-of-the-art baselines regarding the Recall and NDCG metrics.
Recently, researchers have focused on utilizing given heterogeneous features to explore obvious discrimination information for clustering. Most of the current work exploits consistency using some fusion metrics, but the complementarity of multi-view features is not well leveraged. In this paper, we propose an efficient consistent contrastive representation network (CCR-Net) for multi-view clustering, which provides a generalized framework for multi-view learning tasks. First, the proposed model explores the complementarity by a designed contrastive fusion module to learn a shared fusion weight. Second, the proposed method utilizes a consistent representation module to ensure consistency and obtains a consistent graph. Furthermore, we also extend the proposed method to incomplete multi-view scenarios. The designed contrastive fusion module utilizes the complementarity of multiple views to fill in the missing view graphs. Moreover, the consistent feature representation module adds a maxpooling layer on CCR-Net to explore a shared local structure and extract a latent low-dimensional embedding. Finally, the proposed method presents end-to-end training and flexible task interfaces for multi-view learning. Comprehensive evaluations on challenging multi-view tasks demonstrate that the proposed method achieves outstanding performance.
With the rapid growth of multimedia-sharing platforms (e.g. Twitter and TikTok), multimedia recommender systems have become fundamental for helping users alleviate information overload and discover items of interest. Existing multimedia recommendation methods often incorporate various auxiliary modalities (e.g., visual, textual, and acoustic) to describe item characteristics and improve task performance. However, these methods usually assume that each item is associated with complete modalities, ignoring the prevalence of missing modality issues in real-world scenarios. To deal with the challenge of missing modalities, in this paper, we propose a novel framework of Contrastive Intra- and Inter-Modality Generation (CI2MG) for enhancing incomplete multimedia recommendation. We first develop a contrastive intra- and inter-modality generation module for the missing modalities, where the intra-modality representation is updated through clustering-based hypergraph convolution and inter-modality representation is obtained by optimal transport between different modalities. To tackle the challenge of insufficient and incomplete supervision labels during intra- and inter-modality generation, a modality-aware contrastive learning paradigm is introduced based on an augmentation between the intra-modality view and inter-modality view. Furthermore, to learn task-related representations from the generative modalities and further improve the performance of recommendation, we design an enhanced multimedia recommendation module to alleviate the influences driven by task-irrelevant noise. Extensive experiments on real-world datasets show the superiority of our proposed CI2MG framework in offering great potential for personalized multimedia recommendation over the state-of-the-art baselines regarding Recall, NDCG, and Precision metrics.
As real-world data become increasingly heterogeneous, multi-view semi-supervised learning has garnered widespread attention. Although existing studies have made efforts towards this and achieved decent performance, they are restricted to shallow models and how to mine deeper information from multiple views remains to be investigated. As a recently emerged neural network, Graph Convolutional Network (GCN) exploits graph structure to propagate label signals and has achieved encouraging performance, and it has been widely employed in various fields. Nonetheless, research on solving multi-view learning problems via GCN is limited and lacks interpretability. To address this gap, in this paper we propose a framework termed Interpretable Multi-view Graph Convolutional Network (IMvGCN). We first combine the reconstruction error and Laplacian embedding to formulate a multi-view learning problem that explores the original space from feature and topology perspectives. In light of a series of derivations, we establish a potential connection between GCN and multi-view learning, which holds significance for both domains. Furthermore, we propose an orthogonal normalization method to guarantee the mathematical connection, which solves the intractable problem of orthogonal constraints in deep learning. In addition, the proposed framework is applied to the multi-view semi-supervised learning task. Comprehensive experiments demonstrate the superiority of our proposed method over other state-of-the-art methods.
For a multi-view learning task, it is crucial to assign appropriate weights to each view in order to learn complementary and consistent information across different views. In the field of multi-view clustering, most existing methods have been able to handle the weights of different views. However, these algorithms face the problem of unacceptable time complexity when dealing with large-scale datasets, and the learned similarity matrix fails to satisfy the graph regularization. In this paper, we propose an auto-weight learning method called multi-view clustering with graph regularized optimal transport. First, an anchor-based method is employed to overcome the problem of heavy time complexity when processing large-scale datasets, and it is able to automatically learn an appropriate weight for each view. Second, by introducing optimal transport we learn a regularized doubly-stochastic similarity matrix applicable to multi-view clustering tasks. Third, the optimal regularized anchor graph can be classified into specific clusters by adding a rank constraint. Finally, an effective optimization method is designed to optimize the formulated problem. Comprehensive experiments on multiple real-world datasets demonstrate that the proposed algorithm achieves superior performance to other state-of-the-arts algorithms.