
The proliferation of high-velocity big data streams from contemporary technologies, such as the internet of things, social media platforms, wireless sensor networks, blockchain systems, etc., has intensified the need for scalable methods to model and analyze evolving big data representations. Advanced evolutionary representation learning techniques, employing deep learning, graph neural networks, and transformer architectures, generate time-varying feature vectors for diverse entities, including users, sensor nodes, and cryptocurrency wallets. Unsupervised learning over these evolving representations enables the discovery of latent structure, emerging patterns, and anomalies, which are critical for characterizing concept drift. Existing evolutionary clustering approaches, however, often assume fixed data points and a constant number of clusters across timestamps, or lack the scalability required for large-scale, real-world applications. Furthermore, most methods are unable to account for essential dynamics, such as cluster splitting and merging, which limits their ability to capture complex drift behaviors. To overcome these limitations, this paper presents evolVAT, a fast and scalable evolutionary clustering algorithm that incrementally updates clustering results by leveraging previously inferred structure. The proposed method accommodates multiple data-point transitions between consecutive snapshots and supports the addition and removal of entities over time. Experimental evaluations on a broad set of synthetic and real-world datasets demonstrate the effectiveness, robustness, and adaptability of evolVAT in diverse application domains.
Concept drift and class imbalance are two critical challenges in data stream analysis, often interrelated in their effects. Concept drift can intensify the impact of class imbalance, while class imbalance can impede the detection and management of concept drift. To address these challenges, this paper introduces a novel semi-supervised classification approach, Dynamic Ensemble Active Learning for Drifting Imbalanced Data Streams (DEAL-DI). DEAL-DI employs a hybrid dynamic labeling strategy that combines an uncertainty threshold with a random selection mechanism. By dynamically adjusting the threshold based on the current imbalance ratio, this strategy effectively identifies the most informative instances, thereby reducing labeling costs. Additionally, a comprehensive imbalance-handling strategy is proposed, focusing on minority class samples not only during the creation of new classifiers but also by reusing minority instances at the sample level. A novel dynamic subensemble method is further developed to select the top-performing classifiers under current conditions, ensuring their inclusion in subsequent predictions and enhancing model performance. Extensive experiments on both real-world and synthetic data streams demonstrate that DEAL-DI significantly improves recall and G-mean compared to other semi-supervised methods, while simultaneously reducing labeling costs.
With the development of artificial intelligence, multi-modal and multi-domain fake news detection have attracted increasing attention due to their ability to integrate information from different modalities and adapt to diverse scenarios in different domains. Existing methods face a core challenge when handling multi-domain fake news detection, since they struggle to efficiently extract and clearly distinguish between domain-specific and domain-shared features. This blurred boundary not only disrupts the balance between the model's specialization and generalization, but also directly leads to feature redundancy, limiting performance improvement. To address this challenge, we propose a Disentangled Representation learning based Multi-modal Multi-domain Fake News Detection framework (DRMMFND), which effectively analyzes and separates features by different modalities and domains, thereby improving the accuracy of detection. This framework utilizes three mixture of adapters networks to fuse multi-modal information, integrating both domain-specific and domain-shared expert adapters. It uses a router equipped with learnable prompts and a Top-K strategy to generate various types of domain outputs. Furthermore, an auxiliary loss and an orthogonal loss are introduced, achieving effective decoupling of domain information and reducing information redundancy between domains. We conduct experiments on the Weibo, Weibo21 and FineFake datasets, and experimental results show that this approach improves performance in most multi-domain scenarios.
Error-controlled lossy compression can effectively reduce data storage/transfer costs while preserving reconstructed data fidelity based on user-defined error bounds, making it an invaluable tool for scientific data management. However, many use cases require knowledge of compression ratios and domain-specific measures of reconstruction error a priori. Unfortunately, modern state-of-the-art error-controlled lossy compressors primarily focus on point-wise error control rather than providing guarantees regarding compression size or global data fidelity. In this paper, we propose a novel and efficient framework designed to meet diverse compression requirements for lossy compression - namely Surrogate-based Lossy Compression Quality Estimation Framework (SLyCE). SLyCE supports efficient, accurate empirical estimation of lossy compression quality (such as compression ratio, PSNR, or SSIM) for various lossy compressors. We make three key contributions: (1) we develop lightweight compressor surrogates for six state-of-the-art error-controlled lossy compressors-SZx, SZ3, ZFP, SPERR, SZp, and cuSZp-that can be applied to efficiently estimate compression quality metrics across different error bound settings; (2) we propose a simple method to tune our surrogates to optimize both error rate and latency over datasets grouped by data regularity and analyze how dataset characteristics and compressor dynamics affect the accuracy of our surrogate models' predictions; and (3) we comprehensively evaluate each surrogate's performance and accuracy using seven real-world scientific datasets. Experiments demonstrate that our solutions obtain highly accurate metric estimates (e.g., ∼1% estimation errors for SZx) with low execution overhead (e.g.,∼2% estimation cost for SZx).
The Industrial Internet of Things (IIoT) improves productivity and enables real-time monitoring across industrial systems. Nevertheless, this increased connectivity broadens the attack surface, exposing critical infrastructure controlled by IIoT to a wider range of cyber threats. Vulnerabilities in IIoT systems have serious consequences, as a flaw can disrupt production, compromise safety, or affect the stability of critical infrastructures. Consequently, there is a pressing need for efficient vulnerability assessment methods to identify and mitigate vulnerabilities in IIoT systems. A comprehensive dataset capturing the vulnerabilities of IIoT protocols is therefore indispensable, as it establishes the basis for developing and advancing robust vulnerability assessment methods. Nonetheless, existing publicly available datasets exhibit narrow protocol diversity and limited representation of vulnerabilities. This deficiency significantly constrains the development of rigorous vulnerability assessment methods for IIoT systems. To bridge this gap, we present IIoT-VulnSet, a comprehensive dataset encompassing vulnerabilities in multiple IIoT communication protocols, including Modbus, DNP3, MQTT, OPC UA, and S7 Comm. IIoT-VulnSet offers heterogeneity by incorporating diverse industrial protocols and enhances vulnerability representation by explicitly linking attack vectors to the corresponding exploited vulnerabilities, thereby enabling exhaustive and in-depth vulnerability assessment across diverse IIoT protocols. We further assessed the dataset’s utility by applying deep learning and traditional machine learning models, demonstrating its effectiveness in supporting the design and evaluation of advanced IIoT vulnerability assessment methods.
The increasing incidence of eating disorders (EDs) underscores the critical need for methods that can accurately monitor and assess these complex conditions. In recent years, social media has emerged as a significant source of real-world data, with platforms like Twitter (now X) offering valuable insights into user experiences and behaviours related to EDs. This study investigates whether the severity of EDs can be accurately estimated from patterns in social media posts, using these data as a novel lens to understand the scope and progression of EDs. To this end, we conduct an extensive examination of Twitter, which includes the collection of a large ED-related dataset, manual annotation, and the development of an advanced deep learning model, SHED. Our novel model leverages a semantic heterogeneous graph representation to estimate ED severity from Twitter user data. Specifically, it learns semantic representations from user tweets and biographies, while integrating a heterogeneous ED graph representation that combines Twitter users, named entities mapped to knowledge graphs, and an ED lexicon. This approach enables a comprehensive understanding of users’ tweets, biographies, behaviours, interactions, and contextual factors, thereby improving the accuracy of ED severity estimation at the user level. Extensive experiments demonstrate that SHED outperforms all other models on both balanced and imbalanced datasets, achieving the lowest MAE (0.040 and 0.108) and RMSE (0.145 and 0.258). It also achieves the highest F1 score (85.04% and 81.92%) and accuracy (86.21% and 84.14%) for severity classification, demonstrating superior performance across all evaluation metrics.
Suicide remains a major global health concern, with psychological support hotlines serving as a crucial means of early intervention. However, current artificial intelligence approaches commonly employ shallow fusion strategies, namely direct con catenation of multimodal and multiscale features, without explicitly modeling the relationships among them. Such approaches tend to introduce redundancy and weaken the representation of crucial information, limiting accurate suicide risk prediction in real-time psychological support hotlines. Moreover, hotline conversations typically last around 30 minutes, during which callers' emotions fluctuate dynamically. Current graph learning methods, either based on static frameworks or designed for short duration inputs, struggle to model long-term and nonlinear emotional fluctuations. As a result, they are inadequate for handling the complexity of real-world hotline conversations, which limits their practical effectiveness and predictive accuracy. This study proposes a Multimodal Multiscale Graph Learning (MMGL) framework that models both text and speech across clip, sentence, and segment levels to construct a unified semantic space. MMGL introduces a shared-specific feature decomposition mechanism to distinguish between shared information and specific information. Subsequently, the multimodal multiscale features are deeply fused to enhance the construction of the unified semantic space. This mechanism avoids information redundancy while effectively pre serving context-related features, thus extracting key information relevant to suicide risk identification as comprehensively as possible. Additionally, MMGL adaptively constructs graph structures with dynamic edge weights to model both linear and nonlinear emotional relationships. While linear links retain contextual coherence, nonlinear links connect emotionally similar but distant points. Furthermore, a gating mechanism further highlights key emotional cues across long-duration audio, enhancing the model's ability to detect suicidal tendencies. Experiments on a large-scale real-world dataset of psychological support hotline recordings demonstrate that MMGL significantly outperforms unimodal and shallow fusion baselines, achieving an F1-score of 78.18%. These results validate the efficacy of the proposed MMGL framework in accurately predicting suicide risk, highlighting its promise for future integration into suicide prevention strategies.
Multimodal recommendation systems utilize various types of information, including images, text, and audio, to enhance the effectiveness of recommendations. However, existing multimodal recommendation models primarily emphasize enhancing the representation of users and items using multimodal data while insufficiently investigating individual-level user preferences for specific modal information. Additionally, the latent semantic structure of user-item relationships within these multimodal data sets holds the potential to enhance recommendation model performance. Addressing these concerns, we present AdaMRec, a precise multimodal recommendation framework that forecasts user preferences by considering the influence of item semantics from different modalities on users. Recommendations are generated by aggregating attention-weighted user preference scores for each semantic feature across different modalities. Extensive experiments on three public datasets demonstrate that AdaMRec outperforms previous baseline models, showcasing an average performance improvement of 4.83%, with a maximum enhancement of 9.29%. Our code is available at https://github.com/HubuKG/AdaMRec.
Real-world traffic time series exhibit intricate temporal dynamics characterized by entangled high-frequency and low-frequency components, posing significant challenges for accurate forecasting. Decomposing high and low-frequency information (in both time and frequency domains) is a natural approach. However, time- and frequency-domain analyses emphasize different characteristics of time series, while high-frequency information is often progressively smoothed as network depth increases. Consequently, applying a homogeneous modeling strategy to all frequency components may produce inaccurate representations. While Kolmogorov-Arnold Networks (KANs) offer promising function approximation capabilities, their inherent smoothing effect in deep layers often leads to critical loss of high-frequency details. To address this limitation, we propose Momentum-Kan, a novel architecture integrating frequency-disentangled modulation and energy-driven mechanisms. Our framework features three key components: (1) A Frequency-Disentangled Modulator (FDM) dynamically balances high-frequency (local) and low-frequency (global) components through adaptive convolutional kernels modulated by traffic fluctuation intensity; (2) An Energy-driven Mechanism (EDM) leverages the energy ratio between Fourier and wavelet basis functions to establish heterogeneous temporal pattern modeling; (3) A momentum-based gating system iteratively refines representations through residual frequency recombination blocks. Theoretical analysis demonstrates our approach's superiority in preserving fine-grained temporal structures while maintaining global trend consistency. Extensive experiments on six real-world traffic datasets show Momentum-Kan achieves state-of-the-art performance with 12.4% MAPE reduction on PEMS-BAY compared to existing methods.
Graph Contrastive Learning (GCL) has become an effective paradigm in graph representation learning by establishing contrastive relationships between nodes. Existing methods typically generate positive samples through manually designed graph data augmentation and obtain negative samples via random sampling. Suboptimal augmentation can corrupt the original graph topology, while false-negative samples introduce noise that degrades node representation quality. To address these issues, this paper proposes a novel graph contrastive representation learning framework called Enhancing Contrastive Learning with Augmentation (ECLA). First, we adopt a learnable augmentation module that generates adaptive augmented views through learnable edge weight adjustments. Second, we propose a joint metric method that integrates graph structural similarity and node feature similarity to model the deep semantic relationships between nodes, and based on this, we construct the distribution of positive and negative samples. Furthermore, we introduce a dual-weight contrastive loss function that enhances the contrastive learning intensity of semantically similar nodes through positive sample weights, while suppressing the negative impact of potential false-negative nodes through negative sample weights. ECLA jointly optimizes graph augmentation and representation learning in an end-to-end manner. Evaluations across diverse datasets show that ECLA consistently surpasses both leading graph contrastive methods and semi-supervised GNN baselines on node classification under severely limited label conditions.
Fraudulent reviewer detection is crucial for the stakeholders in e-commerce to protect their interests from being undermined by dishonest practices. While both textual contents of reviews and behavioral patterns of reviewers have been identified as indispensable for detecting fraudulent reviewers, existing studies mostly fail to account for its inherent interdependencies. To fill this research gap, this paper proposes a novel approach, called NTAM-TransH (Neural Topic-Attention Modeling with TransH), to make use of the interdependence between review texts and reviewer behavior for fraudulent reviewer detection. NTAM-TransH leverages neural topic model (NTM) to capture the textual characteristics of reviews and translation on hyperplanes (TransH) to capture the behavioral patterns of reviewers, respectively. Attention mechanism is adopted to learn the implicit interdependency by assigning different weights to the topics in the review. Kronecker product plus matrix vectorization is adopted to model the explicit interdependency between reviewer behavior and review textual contents to derive the behavioral-textual representation of the reviewer. Extensive experiments on the YelpZip and YelpNYC datasets demonstrate the superior performance of the proposed NTAM-TransH approach over the baseline methods in detecting fraudulent reviewers.
Most Temporal Knowledge Graph (TKG) extrapolation research focuses on predicting events at the next timestamp. However, in the TKG long-term extrapolation task, the performance of existing models tends to decline due to their reliance on recurrent models, which are less effective for long-term extrapolation tasks. To address this challenge, we propose the Aggregative Hawkes Autoformer (AHA), which revises the Autoformer approach in the context of long time series forecasting. Specifically, we introduce two key enhancements: First, we design a relation-focused Neighbor Aggregator that employs an attention mechanism for more accurate relational modeling. Second, we adapt the long-term forecasting paradigm for TKG long-term extrapolation. Extensive experimental results across four benchmark datasets demonstrate AHA's superior performance in the TKG long-term extrapolation tasks, showcasing its potential to advance the field.
Fuzzy c-means clustering for incomplete data has gained special concern, but many results cannot make a delicate imputation and perfect clustering. To this aim, a Nonlinear Representative Non-negative Latent Factor with Knowledge transfer (NRNLF_KT) and Adaptive Fuzzy Double c-means with Sparse Self-representation (AFD_SS) algorithm, namely NA clustering, is developed. The proposed NA algorithm has three remarkable merits: 1) the NRNLF_KT model employs a nonlinear mapping to intentionally relax the non-negativity constraint and constitutes an unconstrained optimization problem; 2) the NRNLF_KT model establishes an auxiliary matrix related to the target matrix but in a different pattern, which contributes greatly to extracting useful knowledge; and 3) the SS technique is introduced to facilitate the simultaneous consideration of global and local information, meanwhile, the particle swarm optimization algorithm is applied to find an appropriate parameter concerning the specific dataset. Finally, experiment results on real-world datasets verify the superiority of the proposed method.
Accurate identification of distracted driving behaviour is essential for improving driving safety. At present, deep learning models such as convolutional neural networks (CNNs) and the vision transformer (ViT) are widely used in driving action recognition. However, balancing the number of network parameters, detection accuracy, and real-time performance is difficult. To overcome this problem, a dual-stream driver distraction driving detection model (DS-Model) that is based on a guiding learning framework is proposed in this paper. The guiding network stream is a pretrained model with prior knowledge that can quickly extract effective features. The learning network stream is a multiscale weight calculation module that optimizes feature selection through weight calculation. The output features of the network layer guide the training of the learning network, which effectively improves the training efficiency and prediction accuracy of the model. Moreover, the DS-Model can be naturally extended to multiview settings and shows promising performance for multiperspective driver distraction detection on controlled datasets. We conducted experimental evaluations on three public datasets: AUCDD-V1, SFD, and 100-Driver. The results show that the accuracy of the DS-Model on the three datasets reached 95.80%, 99.91%, and 99.91%, respectively. With only 6.30M parameters, the proposed model outperforms most previous approaches.
Hypergraphs have gained widespread exploration in various research domains due to their ability to model higher-order correlations among entities. Although existing studies on hypergraph learning have extended graph convolution to hypergraphs, they struggle to effectively extract features from unlabeled data. Recently, generative self-supervised learning, particularly masked autoencoder, has shown significant potential in handling unlabeled graph data, but applying it to hypergraphs still presents two key challenges: (1) How to accurately capture complex higher-order relationships of hypergraphs with generative learning; (2) How to maintain robustness when designing reconstruction strategies. To address these challenges, we propose the HyperGraph Masked AutoEncoder (HGMAE), which leverages generative self-supervised learning techniques to learn from unlabeled data. Firstly, we employ attribute masking with dynamic linear masking rates and hyperedge masking to enable effective learning of hypergraphs. Subsequently, we adopt two distinct training strategies: node attribute reconstruction to capture various node attribute information and hypergraph-based equivalent bipartite graph reconstruction to capture higher-order semantic information in hypergraph structures. We conduct experiments on citation network classification and visual object recognition tasks, and the results demonstrate that the proposed HGMAE outperforms existing methods and baseline approaches on multiple benchmark datasets.