Many representation learning methods have gradually emerged to better exploit the properties of multi-view data. However, these existing methods still have the following areas to be improved: 1) Most of them overlook the ex-ante interpretability of the model, which renders the model more complex and more difficult for people to understand; 2) They underutilize the potential of the bi-topological spaces, which bring additional structural information to the representation learning process. This lack is detrimental when dealing with data that exhibits topological properties or has complex geometrical relationships between different views. Therefore, to address the above challenges, we propose an optimization-oriented multi-view representation learning framework in implicit bi-topological spaces. On one hand, we construct an intrinsically interpretability end-to-end white-box model that directly conducts the representation learning procedure while improving the transparency of the model. On the other hand, the integration of bi-topological spaces information within the network via manifold learning facilitates the comprehensive utilization of information from the data, ultimately enhancing representation learning and yielding superior performance for downstream tasks. Extensive experimental results demonstrate that the proposed method exhibits promising performance and is feasible in the downstream tasks.
Multi-view learning has demonstrated strong potential in processing data from different sources or viewpoints. Despite the significant progress made by Multi-view Graph Neural Networks (MvGNNs) in exploiting graph structures, features, and representations, existing research generally lacks architectures specifically designed for the intrinsic properties of multi-view data. This leads to models that still have deficiencies in fully utilizing consistent and complementary information in multi-view data. Most of current research tends to simply extend the single-view GNN framework to multi-view data, lacking in-depth strategies to handle and leverage the unique properties of these data. To address this issue, we propose a simple yet effective MvGNN framework called Multi-view Representation Learning with Decoupled private and shared Propagation (MvRL-DP). This framework enables multi-view data to be effectively processed as a whole by alternating private and shared operations to integrate cross-view information. In addition, to address possible inconsistencies between views, we present a discriminative loss that promotes class separability and prevents the model from being misled by noise hidden in multi-view data. Experiments demonstrate that the proposed framework is superior to current state-of-the-art methods in the multi-view semi-supervised classification task.
Existing representation learning approaches lie predominantly in designing models empirically without rigorous mathematical guidelines, neglecting interpretation in terms of modeling. In this work, we propose an optimization-derived representation learning network that embraces both interpretation and extensibility. To ensure interpretability at the design level, we adopt a transparent approach in customizing the representation learning network from an optimization perspective. This involves modularly stitching together components to meet specific requirements, enhancing flexibility and generality. Then, we convert the iterative solution of the convex optimization objective into the corresponding feed-forward network layers by embedding learnable modules. These above optimization-derived layers are seamlessly integrated into a deep neural network architecture, allowing for training in an end-to-end fashion. Furthermore, extra view-wise weights are introduced for multi-view learning to discriminate the contributions of representations from different views. The proposed method outperforms several advanced approaches on semi-supervised classification tasks, demonstrating its feasibility and effectiveness.
Due to the heterogeneity gap in multi-view data, researchers have been attempting to apply these data to learn a co-latent representation to bridge this gap. However, multi-view representation learning still confronts two challenges: (1) it is hard to simultaneously consider the performance of downstream tasks and the interpretability and transparency of the network; (2) it fails to learn representations that accurately describe the class boundaries of downstream tasks. To overcome these limitations, we propose an interpretable representation learning framework, named interpretable multi-view proximity representation learning network. On the one hand, the proposed network is customized by an explicitly designed optimization objective that enables it to learn semantic co-latent representations while maintaining the interpretability and transparency of the network from the design level. On the other hand, the designed multi-view proximity representation learning objective function encourages its learned co-latent representations to form intuitive class boundaries by increasing the inter-class distance and decreasing the intra-class distance. Driven by a flexible downstream task loss, the learned co-latent representation can adapt to various multi-view scenarios and has been shown to be effective in experiments. As a result, this work provides a feasible solution to a generalized multi-view representation learning framework and is expected to accelerate the research and exploration in this field.
As researchers strive to narrow the gap between machine intelligence and human through the development of artificial intelligence multimedia technologies, it is imperative that we recognize the critical importance of trustworthiness in open-world, which has become ubiquitous in all aspects of daily life for everyone. However, several challenges may create a crisis of trust in current open-world artificial multimedia systems that need to be bridged: 1) Insufficient explanation of predictive results; 2) Inadequate generalization for learning models; 3) Poor adaptability to uncertain environments. Consequently, we explore a neural program to bridge trustworthiness and open-world learning, extending from single-modal to multi-modal scenarios for readers.1) To enhance design-level interpretability, we first customize trustworthy networks with specific physical meanings; 2) We then design environmental well-being task-interfaces via flexible learning regularizers for improving the generalization of trustworthy learning; 3) We propose to increase the robustness of trustworthy learning by integrating open-world recognition losses with agent mechanisms. Eventually, we enhance various trustworthy properties through the establishment of design-level explainability, environmental well-being task-interfaces and open-world recognition programs. As a result, these designed open-world protocols are applicable across a wide range of surroundings, under open-world multimedia recognition scenarios with significant performance improvements observed.