In this paper, we develop a novel local graph pooling method, namely the Separated Subgraph-based Hierarchical Pooling (SSHPool), for graph classification. We commence by assigning the nodes of a sample graph into different clusters, resulting in a family of separated subgraphs. We individually employ the local graph convolution units as the local structure to further compress each subgraph into a coarsened node, transforming the original graph into a coarsened graph. Since these subgraphs are separated by different clusters and the structural information cannot be propagated between them, the local convolution operation can significantly avoid the over-smoothing problem caused by message passing through edges in most existing Graph Neural Networks (GNNs). By hierarchically performing the proposed procedures on the resulting coarsened graph, the proposed SSHPool can effectively extract the hierarchical global features of the original graph structure, encapsulating rich intrinsic structural characteristics. Furthermore, we develop an end-to-end GNN framework associated with the SSHPool module for graph classification. Experimental results demonstrate the superior performance of the proposed model on real-world datasets.
In this work, we develop a family of Aligned Entropic Graph Kernels (AEGK) for graph classification. We commence by performing the Continuous-time Quantum Walk (CTQW) on each graph structure, and compute the Averaged Mixing Matrix (AMM) to describe how the CTQW visits all vertices from a starting vertex. More specifically, we show how this AMM matrix allows us to compute a quantum Shannon entropy of each vertex for either un-attributed or attributed graphs. For pairwise graphs, the proposed AEGK kernels are defined by computing the kernel-based similarity between the quantum Shannon entropies of their pairwise aligned vertices. The analysis of theoretical properties reveals that the proposed AEGK kernels cannot only address the shortcoming of neglecting the structural correspondence information between graphs arising in most existing R-convolution graph kernels, but also overcome the problems of neglecting the structural differences and vertex-attributed information arising in existing vertex-based matching kernels. Moreover, unlike most existing classical graph kernels that only focus on the global or local structural information of graphs, the proposed AEGK kernels can simultaneously capture both global and local structural characteristics through the quantum Shannon entropies, reflecting more precise kernel-based similarity measures between pairwise graphs. The above theoretical properties explain the effectiveness of the proposed AEGK kernels. Experimental evaluations demonstrate that the proposed kernels can outperform state-of-the-art graph kernels and deep learning models for graph classification.
The over-smoothing has emerged as a major challenge in the development of Graph Neural Networks (GNNs). While existing state-of-the-art methods effectively mitigate the diminishing distance between nodes and improve the performance of node classification, they tend to be elusive for graph-level tasks. This paper introduces a novel entropy-based perspective to explore the over-smoothing problem, simultaneously enhancing the distinguishability of non-isomorphic graphs. We provide a theoretical analysis of the relationship between the smoothness and the entropy for graphs, highlighting how the over-smoothing in high-entropic regions negatively impact the graph classification performance. To tackle this issue, we propose a simple yet effective method to Sample and Discretize node features in high-Entropic regions (SDE), aiming to preserve the critical and complicated structural information. Moreover, we introduce a new evaluation metric to assess the over-smoothing for graph-level tasks, focusing on node distributions. Experimental results demonstrate that the proposed SDE method significantly outperforms existing state-of-the-art methods, establishing a new benchmark in the field of GNNs.
Graph Neural Networks (GNNs) have emerged as powerful tools for graph learning, and one key challenge arising in GNNs is the development of effective pooling operations for learning meaningful graph representations. In this paper, we propose a novel Edge-Node Attention-based Hierarchical Pooling (ENAHPool) operation for GNNs. Unlike existing cluster-based pooling methods that suffer from ambiguous node assignments and uniform edge-node information aggregation, ENAHPool assigns each node exclusively to a cluster and employs attention mechanisms to perform weighted aggregation of both node features within clusters and edge connectivity strengths between clusters, resulting in more informative hierarchical representations. To further enhance the model performance, we introduce a Multi-Distance Message Passing Neural Network (MD-MPNN) that utilizes edge connectivity strength information to enable direct and selective message propagation across multiple distances, effectively mitigating the over-squashing problem in classical MPNNs. Experimental results demonstrate the effectiveness of the proposed method.
Graphs have a superior ability to represent relational data, like chemical compounds, proteins, and social networks. Hence, graph-level learning, which takes a set of graphs as input, has been applied to many tasks including comparison, regression, classification, and more. Traditional approaches to learning a set of graphs heavily rely on hand-crafted features, such as substructures. But while these methods benefit from good interpretability, they often suffer from computational bottlenecks as they cannot skirt the graph isomorphism problem. Conversely, deep learning has helped graph-level learning adapt to the growing scale of graphs by extracting features automatically and encoding graphs into low-dimensional representations. As a result, these deep graph learning methods have been responsible for many successes. Yet, there is no comprehensive survey that reviews graph-level learning starting with traditional learning and moving through to the deep learning approaches. This article fills this gap and frames the representative algorithms into a systematic taxonomy covering traditional learning, graph-level deep neural networks, graph-level graph neural networks, and graph pooling. To ensure a thoroughly comprehensive survey, the evolutions, interactions, and communications between methods from four different branches of development are also examined. This is followed by a brief review of the benchmark data sets, evaluation metrics, and common downstream applications. The survey concludes with a broad overview of 12 current and future directions in this booming field.
Complex networks usually demonstrate intrinsic uncertainty, imprecision, and ambiguity in their structural and dynamical characteristics, which positions them as suitable candidates for analysis through fuzzy systems. The statistical characterization of these networks has been associated with both the Ihara Zeta function, which evaluates prime cycles, and partition functions from thermodynamics. However, traditionally, these functions have been viewed as separate entities, thereby neglecting their potential interconnection. This paper establishes a link between the Ihara Zeta function in algebraic graph theory and the partition function in statistical mechanics to elucidate network structure within the fuzzy systems framework. This linkage offers a fresh perspective on the relationship between microscopic structure and macroscopic properties of networks, taking into account the inherent uncertainty and ambiguity frequently encountered in real-world applications. We derive thermodynamic quantities, such as entropy, that correlate with the configurations of prime cycles of varying lengths and extend these ideas to fuzzy entropy measures to accommodate imprecise or incomplete information. Employing the n-th derivative of the Ihara Zeta function facilitates the computation of prime cycle quantities within a network. This method is conceptually similar to using the partition function within the Bose-Einstein statistical model. The derived entropy measures enable us to investigate phase transitions in network structure, especially in fuzzy conditions, where transitions may not be distinctly defined. Numerical experiments and empirical data validate the efficacy of our approach in characterizing network structure, thus offering a holistic understanding of both microscopic and macroscopic network attributes in the context of fuzzy systems.
In this paper, we propose a Hierarchical Aligned Subtree Convolutional Network (HA-SCN) for graph classification. Our idea is to transform graphs of arbitrary sizes into fixed-sized aligned graphs and construct a normalized K-layer m-ary subtree for each node in the aligned graphs. By sliding convolutional filters over the entire subtree at each node, we define a novel subtree convolution and pooling operation that hierarchically abstracts node-level information. We demonstrate that the proposed HA-SCN model not only realizes the convolution mechanism similar to the Convolutional Neural Networks (CNNs), which have the characteristics of weight sharing and fixed-sized receptive fields, but also effectively mitigates the oversquashing problem. Meanwhile, it establishes the correspondence information between nodes, alleviating the information loss issue. Experimental results on various benchmark graph datasets show that our approach achieves state-of-the-art performance in graph classification tasks.
Infrared and visible image fusion (IVF) plays an important role in intelligent transportation system (ITS). The early works predominantly focus on boosting the visual appeal of the fused result, and only several recent approaches have tried to combine the high-level vision task with IVF. However, they prioritize the design of cascaded structure to seek unified suitable features and fit different tasks. Thus, they tend to typically bias toward to reconstructing raw pixels without considering the significance of semantic features. Therefore, we propose a novel prior semantic guided image fusion method based on the dual-modality strategy, improving the performance of IVF in ITS. Specifically, to explore the independent significant semantic of each modality, we first design two parallel semantic segmentation branches with a refined feature adaptive-modulation (RFaM) mechanism. RFaM can perceive the features that are semantically distinct enough in each semantic segmentation branch. Then, two pilot experiments based on the two branches are conducted to capture the significant prior semantic of two images, which then is applied to guide the fusion task in the integration of semantic segmentation branches and fusion branches. In addition, to aggregate both high-level semantics and impressive visual effects, we further investigate the frequency response of the prior semantics, and propose a multi-level representation-adaptive fusion (MRaF) module to explicitly integrate the low-frequent prior semantic with the high-frequent details. Extensive experiments on two public datasets demonstrate the superiority of our method over the state-of-the-art image fusion approaches, in terms of either the visual appeal or the high-level semantics.
Graph Neural Networks (GNNs) are powerful tools for graph learning, but one of the important challenges is how to effectively extract representations for graph-level tasks. In this paper, we propose an end-to-end Simple Clustering Hierarchical Pooling (SCHPool) operation, which is based on Top-K node selection for learning expressive graph representations. Specifically, SCHPool considers each node and its local neighborhood as a cluster, and introduces a novel multi-view scoring function to evaluate node importance. Based on these scores, clusters centered around the Top-K nodes are retained. This design eliminates the need for complex clustering operations, significantly reducing computational overhead. Furthermore, during the coarsening process, SCHPool employs a lightweight yet comprehensive attention mechanism to adaptively aggregate both the node features within clusters and the edge connectivity strengths between clusters. This facilitates the construction of more informative coarsened graphs, enhancing model performance. Experimental results demonstrate the effectiveness of the proposed model.
The problem of over-smoothing has emerged as a fundamental issue for Graph Convolutional Networks (GCNs). While existing efforts primarily focus on enhancing the discriminability of node representations for node classification, they tend to overlook the over-smoothing at the graph level, significantly influencing the performance of graph classification. In this paper, we provide an explanation of the graph-level over-smoothing phenomenon, and propose a novel Adaptive Multi-Viewed Subgraph Convolutional Network (MultiNet) to address this challenge. Specifically, the MultiNet introduces a local subgraph convolution module that adaptively divides each input graph into multiple subgraph views. Then a number of subgraph-based view-specific convolution operations are applied to constrain the extent of node information propagation over the original global graph structure, not only mitigating the over-smoothing issue but also generating more discriminative local node representations. Moreover, we develop an alignment-based readout that establishes correspondences between nodes over different graphs, thereby effectively preserving the local node-level structure information and improving the discriminative ability of the resulting graph-level representations. Theoretical analysis and empirical studies show that the MultiNet mitigates the graph-level over-smoothing and achieves excellent performance for graph classification.
With advancements in deep stereo matching, recent networks have achieved impressive accuracy in estimating depth information from image pairs. However, stereo matching networks require sufficient disparity labels, which always come at high annotation costs. In this paper, we propose the ALStereo framework for training stereo matching networks under limited labeling budgets, which selects informative samples for manual labeling and conducts semi-supervised learning to propagate the knowledge to unlabeled samples. Specifically, we embed image pairs as nodes in a graph representation, where edges denote the similarity in terms of stereo matching challenges. Based on the graph representation, we divide the labeling budget into two parts for conducting representativeness-based and uncertainty-based strategies, balancing the selection of the most representative and challenging samples. To fully exploit the labeled samples to train networks, we propose a two-stage semi-supervised training pipeline, where the first stage mitigates the domain shifts and the second stage propagates the knowledge of manually annotated samples to unlabeled samples. We set the first benchmark for evaluating training stereo matching networks under limited labeling budgets and demonstrate our method significantly outperforms the compared methods. We also provide analysis to demonstrate our graph representation effectively models the similarity between samples in terms of stereo matching challenges.
Graph-based representations are powerful tools for analyzing structured data. In this paper, we propose a novel model to learn Deep Hierarchical Attention-based Kernelized Representations (DHAKR) for graph classification. To this end, we commence by learning an assignment matrix to hierarchically map the substructure invariants into a set of composite invariants, resulting in hierarchical kernelized representations for graphs. Moreover, we introduce the feature-channel attention mechanism to capture the interdependencies between different substructure invariants that will be converged into the composite invariants, addressing the shortcoming of discarding the importance of different substructures arising in most existing R-convolution graph kernels. We show that the proposed DHAKR model can adaptively compute the kernel-based similarity between graphs, identifying the common structural patterns over all graphs. Experiments demonstrate the effectiveness of the proposed DHAKR model.
In this paper, we propose a new model to learn Adaptive Kernel-based Representations (AKBR) for graph classification. Unlike state-of-the-art R-convolution graph kernels that are defined by merely counting any pair of isomorphic substructures between graphs and cannot provide an end-to-end learning mechanism for the classifier, the proposed AKBR approach aims to define an end-to-end representation learning model to construct an adaptive kernel matrix for graphs. To this end, we commence by leveraging a novel feature-channel attention mechanism to capture the interdependencies between different substructure invariants of original graphs. The proposed AKBR model can thus effectively identify the structural importance of different substructures, and compute the R-convolution kernel between pairwise graphs associated with the more significant substructures specified by their structural attentions. Furthermore, the proposed AKBR model employs all sample graphs as the prototype graphs, naturally providing an end-to-end learning architecture between the kernel computation as well as the classifier. Experimental results show that the proposed AKBR model outperforms existing state-of-the-art graph kernels and deep learning methods on standard graph benchmarks.
Autoencoders, a type of unsupervised model, are capable of learning effective latent representations of data without supervision, only requiring the decoder to be able to reconstruct the original data-point from its latent representation obtained through the encoder. When dealing with structured data, graph-autoencoders are one of the few effective ways to obtain such latent representations of graphs. However, when dealing with large, diverse graph datasets, autoencoders struggle to adapt to varying structures, leading to suboptimal encoding. In our paper, we introduce a novel approach called Mixture of Variational Graph Autoencoders, which addresses this limitation by introducing a mixture of encoder/decoder models which provide multiple local and class-specific models that better adapt to different patches of the data-space. An exhaustive experimental evaluation shows that our approach greatly outperforms the state of the art in reconstruction precision (Code: https://github.com/gdl-unive/MVGAE ).
In this paper, we propose a family of novel Deep Hierarchical Transitive-Aligned Graph Kernels (DHTAGK) for graph classification. To this end, we commence by developing a new Hierarchical Aligned Graph Auto-Encoder (HA-GAE) to construct transitive-aligned embedding graphs that encapsulate the structural correspondence information between graphs. The DHTAGK kernels then measure either the Jensen-Shannon Divergence between the adjacency matrices or the Gaussian kernel between the node feature matrices of the embedding graphs. Unlike the classical Rconvolution kernels and node-based alignment kernels, the DHTAGK kernels can capture the transitive structural correspondence information and thus ensure the positive definiteness. Furthermore, the HA-GAE enables the DHTAGK kernels to simultaneously reflect both local and global graph structures and identify common structural patterns. Experimental results show that the DHTAGK kernels outperform state-of-the-art graph kernels and deep learning methods on benchmark datasets.
Comprehensive dialogue understanding and effective cross-modal interaction remain challenging in multimodal Emotion Recognition in Conversations (ERC). Existing methods often struggle to efficiently process lengthy conversations, with cross-modal interactions biased towards text modality, resulting in some degree of misrepresentation. Graph Neural Networks (GNNs) offer promise by structuring dialogue sequences into graphs, yet they tend to overlook semantic details in high-frequency signals due to their reliance on low-pass filters. To address these limitations, we propose FrameERC, a novel framework that utilizes graph framelet transforms to analyze dialogues. FrameERC first transforms multimodal features into efficient graph signals using two distinct decoupling strategies. By decomposing the graph signals across a broad frequency spectrum with low- and high-pass filters, FrameERC captures critical emotional subtleties that traditional spatial message-passing GNN models may ignore, thereby enriching detailed emotion recognition. Moreover, we introduce a dual-reminder fusion mechanism to enhance the meaningful contribution of non-textual modalities in ERC, ensuring a comprehensive semantic understanding. Extensive experimental results demonstrate that FrameERC achieves superior performance compared to current state-of-the-art methods on two widely-used multimodal ERC datasets.
For many vision tasks, utilizing pre-trained features results in improved performance and consistently benefits from the rapid advancement of pre-training technologies. However, in the field of stereo matching, the use of pre-trained features has not been extensively researched. In this paper, we present the first systematical exploration into the utilization of pre-trained features for stereo matching. To provide flexible employment for any combination of pre-trained backbones and stereo matching networks, we develop the deformable neck (DN) that decouples the network architectures of these two components. The core idea of DN is to utilize the deformable attention mechanism to iteratively fuse pre-trained features from shallow to deep layers. Empirically, our exploration reveals the crucial factors that influence using pre-trained features for stereo matching. We further investigate the role of instance-level information of pre-trained features, demonstrating it benefits stereo matching while can be suppressed during convolution-based feature fusion. Built on the attention mechanism, the proposed DN module effectively utilizes the instance-level information in pre-trained features. Besides, we provide an understanding of the efficiency-accuracy tradeoff, concluding that using pre-trained features can also be a good alternative with efficiency consideration.
In this paper, we propose a novel framework of computing the Quantum-based Entropic Representations (QBER) for un-attributed graphs, through the Continuous-time Quantum Walk (CTQW). To achieve this, we commence by transforming each original graph into a family of k-level neighborhood graphs, where each k-level neighborhood graph encapsulates the connected information between each vertex and its k-hop neighbor vertices, providing a fine representation to reflect the multi-level topological information for the original global graph structure. To further capture the complicated structural characteristics of the original graph through its neighborhood graphs, we propose to characterize the structure of each neighborhood graph with the Average Mixing Matrix (AMM) of the CTQW, that encapsulates the time-averaged behavior of the CTQW evolved on the neighborhood graph. More specifically, we show how the AMM matrix allows us to compute a Quantum Shannon Entropy for each vertex, and thus compute an entropic signature for each neighborhood graph by measuring the averaged value or the Jensen–Shannon Divergence between the entropies of its vertices. For each original graph, the resulting QBER is defined by gauging how the entropic signat ures vary on its k-level neighborhood graphs with increasing k, reflecting the multi-dimensional entropy-based structure information of the original graph. Experiments on standard graph datasets demonstrate the effectiveness of the proposed QBER approach in terms of the classification accuracies. The proposed approach can significantly outperform state-of-the-art entropic complexity measuring methods, graph kernel methods, as well as graph deep learning methods.
F. Escolano合作论文数Dpto. de Ciencia de la Computaci??n e IA;Universidad de Alicante29