
In the era of large language models (LLMs), the Internet is flooded with LLM-generated content, including both aligned LLM-generated content that contains equivalent facts to human-written content and distorted LLM-generated content with factual discrepancies. However, there is a significant gap in understanding and addressing the impact of such content on Retrieval-Augmented Generation (RAG) systems, which rely on accurate retrieved knowledge to generate response. This paper comprehensively evaluate the impact of aligned and distorted LLM-generated content on the retrieval and generation stages of RAG systems. Notably, we reveal that distorted LLM-generated content is not only more likely to be retrieved but also significantly degrades generation quality, whereas aligned content can, in some cases, enhance it. To mitigate this, we propose a factual consistency-aware adaptive filtering approach to selectively filter out distorted LLM-generated content from retrieved documents. Experimental results demonstrate that our method significantly improves generative performance in scenarios that mix with LLM-generated content and is broadly applicable across various RAG systems.
Local Differential Privacy (LDP) enables users to perturb data locally from the untrusted data collectors. Recent studies validate the vulnerability of LDP to data poisoning attacks where an attacker injects the deliberately crafted data into the LDP protocols to manipulate the results of data analytic tasks. In this work, we advance the knowledge by proposing the disguised data poisoning attack against the three state-of-the-art LDP protocols, i.e., Optimized Unary Encoding, Prefix Extending Method, and Piecewise Mechanism, that manipulates the results of three popular data analytic tasks, Frequency Estimation, Heavy Hitter Identification, and Mean-Variance Estimation, and moreover can disguise the attack behavior by deeply considering the inherent properties of LDP protocols. The main idea is to bundle the target item with the neighboring items, and leverage the enhancement of the target item to drive sophisticated changes within its neighboring items to maximize the attack gain and disguise the attack behavior. Both the theoretical analysis and the experimental results validate the superior performance of our attack compared to the five existing attacks on five datasets. Furthermore, we explore an defense to mitigate the proposed attack, using the Discrete Cosine Transform, \(\alpha\) sampling, and clustering in frequency domain. The extensive results validate the effectiveness of the proposed defense on five datasets in some scenarios, compared to two latest existing defenses. But, in other cases, the proposed defense is not very effective, and thus new defenses in the further are needed.
Reliably inferring causal effects from observational data is a cornerstone of scientific inquiry and data-driven decision-making, yet its validity is frequently undermined by the inappropriate selection of covariates. Standard adjustment methods often operate on the fragile assumption that all observed covariates are beneficial confounders, a premise that is rarely tenable in high-dimensional settings. The uncritical inclusion of non-confounding variables can introduce significant estimation bias, leading to flawed scientific conclusions and ineffective policies. To confront this foundational challenge, we first introduce the Adjustment by Proxy framework, a unified analytical lens to systematically diagnose the bias introduced by distinct classes of non-confounding covariates. This framework enables a rigorous theoretical derivation of the harm caused by including variables such as instruments, colliders, or mediators. Building upon this diagnostic foundation, we then propose a novel graph-based criterion demonstrating that the Proximal Confounder Set, defined as the minimal set of confounders that d-separates the treatment from all other confounders, constitutes a minimal and theoretically optimal adjustment set for causal effect inference. Extensive experiments on both synthetic and real-world benchmark datasets validate our theoretical claims, demonstrating that adjusting on the Proximal Confounder Set yields consistently more accurate causal estimates than conventional strategies. Together, these contributions provide a rigorous and practical guide for covariate selection, paving the way for more robust and automated causal inference.
Multi-Label Learning (MLL) refers to inducing multi-label prediction models from the precisely labeled training dataset. However, in many real-world scenarios, e.g ., crowdsourcing annotations, the training datasets are often only partially valid, where each training instance is associated with a candidate label set, covering ground-truth labels but also with irrelevant ones. Naturally, learning with such datasets, formally referred to as Partial Multi-label Learning (PML), involves many noisy supervised signals, hence imposing a significant challenge to the prediction model induction. To meet this challenge, we purify the noisy supervised signals by formulating the latent label distribution, i.e ., the probability of a candidate label being a ground-truth one, and then jointly learn it with the prediction model by minimizing their regularized Wasserstein distance, i.e ., a robust distance for distributions as well as involving label correlations. Therefore, we propose a novel PML method, namely Wasserstein Partial Multi-Label Learning with dual Label Correlation Perspectives ( Wpml 3 cp ), solved by the gradient descent with an augmented Lagrange multiplier technique. To further enhance the robustness of Wpml 3 cp against exceptionally high ratios of irrelevant labels, we extend it with a Dual-branch Competitive Cleansing mechanism, leading to Wpml 3 cp -D. Besides, we also analyze the generalization error bound and time complexity of Wpml 3 cp and Wpml 3 cp -D. The extensive experiments are constructed by comparing Wpml 3 cp and Wpml 3 cp -D with existing PML baselines across synthetic and real-world datasets, and empirical results demonstrate that Wpml 3 cp and Wpml 3 cp -D can outperform the PML baselines in various noisy levels.
Non-Euclidean spaces inherently enable high-fidelity embeddings for hierarchical and cyclical data due to their geometric properties. Existing approaches unify hyperbolic and spherical embeddings within the framework of constant curvature spaces. However, current methods for Lipschitz regularization remain limited to non-positive curvature geometries, such as hyperbolic and Euclidean spaces, and cannot be naturally extended to the general constant curvature setting. In this article, we present a rigorous Lipschitz analysis for constant curvature graph convolutional networks ( \(\kappa\) -GCNs) and enhance their robustness through Lipschitz regularization. We derive upper bounds for the Lipschitz constants across constant curvature spaces, thereby standardizing the Lipschitz limits of the \(\kappa\) -stereographic model. Furthermore, we incorporate these bounds into a regularization framework for \(\kappa\) -GCNs to improve stability and robustness. Experimental results demonstrate that the proposed regularization method often strengthens the robustness of \(\kappa\) -GCNs across various curvature regimes, particularly under Gaussian feature noise.
Multiple complementary-label learning (MCLL) is a machine learning task that involves learning a classifier from instances with multiple complementary labels (MCLs). MCLs are labels that indicate the incorrect labels of an instance. Previous methods for learning with ambiguous supervised information may not be effective because MCLs only make up a small proportion of all labels. In this article, we propose MulCo, a simple yet effective framework that uses contrastive learning to enhance the representation capability in MCLL. Contrastive learning involves contrasting semantically similar and dissimilar pairs of instances, with the goal of benefiting from negatives whose ground-truth labels differ from those of anchors. However, it is possible for dissimilar pairs to have the same label due to the random sampling of negatives from inaccurately labeled data. To solve this problem, we design a sifted contrastive loss for MulCo to correct the sampling of same-label negative pairs. We also provide theoretical evidence for the feasibility of the sifted contrastive loss by establishing an upper bound on the ideal contrastive loss. Correspondingly, we develop two progressive solutions using the properties of complementary labels to approximate the ideal contrastive loss through weighting. Our empirical study demonstrates the effectiveness of the proposed method. The code of this article is available at https://github.com/gaoyi439/MulCo .
Bitcoin, as the most valuable cryptocurrency, has become a significant target for criminal activities, leading to substantial financial losses. To timely detect criminal activities and reduce losses, the accurate detection of illicit addresses is crucial. Current research typically involves extracting features from addresses to classify them as either licit or illicit. However, they neglect the transaction topology information associated with the addresses, thereby compromising the efficiency. In this article, we focus on improving the utilization of address information and propose DIBA detector, an effective framework designed for the automatic D etection of I llicit B itcoin A ddresses. DIBA first incorporates a transaction graph construction module that constructs per-address transaction graph based on UTXO model, thereby mapping to the transaction topology for each address. Subsequently, a novel hybrid spatiotemporal network is designed to learn graph representation for each per-address transaction graph, generating comprehensive graph embeddings that serve as critical inputs for the final classification model. Experimental results demonstrate that our proposed framework outperforms existing state-of-the-art methods for detecting illicit Bitcoin addresses, achieving precision and F1-score values of 92.77% and 93.58%, respectively. As a side contribution, we construct a dataset comprising over 180,000 addresses, with half labeled as licit and half as illicit. Our dataset is released at https://github.com/dlvlb123/DIBA .
Deep reinforcement learning (DRL), has shown promise in solving intractable challenges in interactive recommendation systems (IRS). In DRL-based interactive recommendation, state modeling is vital for well-capturing users’ continuous interaction behaviors with recommendation systems. To effectively capture the behavior of users, existing works for state modeling have evolved from sequential-based modeling to session-based modeling. However, existing session-based state modeling works in IRS are still not fully explored with premature session models and insufficient fusion for different session features. As a result, they cannot capture complicated session patterns during interaction, leading to significant information loss. In this article, we propose a Knowledge-enhanced Multi-Level Session Graph (KMSG) model for interactive recommendation to address the above challenge. KMSG models the user’s interactive data into multi-level session graphs and effectively encodes the states via graph neural networks. Specifically, a novel 3-level item transition graph is designed to capture the common session patterns and intra-session item transitions. We further utilize the information from the knowledge graph to enhance the item relations in KMSG. We then design an attention-based graph neural network to propagate the information in KMSG. Extensive experiments on four real-world benchmark datasets demonstrate the superiority of KMSG over state-of-the-art baselines and the rationality of our design in KMSG.
Causality has been integrated with machine learning in uncovering and understanding the causal relationship between variables and observed outcomes. However, the centralized training setting of causal machine learning is not adaptable to most practical scenarios, where datasets are distributed, stored, and unsharable due to privacy concerns. Federated learning (FL), a distributed learning framework that allows collaborative training across multiple devices without raw data sharing, emerges as a potential solution to this problem. By integrating FL into causal problems, the discovery and inference of causal relationships across dispersed datasets can be achieved. On the other hand, causality can also enhance FL models in various dimensions, including model interpretability and explainability, generalizability, adversarial robustness, and fairness and bias mitigation. In this article, we provide a comprehensive review of the above two directions and summarize the interplays between causality and FL (short for Causal-FL ) by organizing our discussion around two key questions: (1) how FL enable decentralized causal analysis; and (2) how causality tackles FL challenges. The potential applications of these methods are also introduced, including healthcare, recommendation, economics, social equity, and so on. Moreover, we discuss promising future directions and future challenges to be explored.
We investigate a practical yet underexplored variant of the influence maximization problem, in which the goal is to activate a given sequence of nested subgraphs as far outward as possible. A subgraph is considered activated if either it reaches a prescribed influence threshold within itself, or a larger subgraph containing it in the nested sequence is activated. We formally define this problem as Structure-Aware Influence Maximization (SAIM) , and show that it is NP-hard and its objective function is non-submodular. To tackle SAIM, we propose a decomposition framework that reduces it to a sequence of subproblems, each called Target Subgraph Activation Probability Maximization (TSAPM) . For TSAPM, we propose two oracles: a Sandwich Approximation (SA) framework with data-dependent guarantees, and a heuristic algorithm tailored to the TSAPM objective. Combined with the SA-based oracle, the overall decomposition framework provides a solution-dependent approximation guarantee for the SAIM problem. Extensive experiments on six real-world datasets validate the effectiveness and efficiency of our proposed methods.
A Digital Twin (DT) is an innovative area of computer science and technology research that generates an imitation of a realistic system, procedure, or object. The usage of DT technology in innovation across several industries has made it a vital and revolutionary tool. The CiteSpace tool can be utilised for carrying out a scientometric evaluation of the effectiveness of DT technologies in engineering. The current analysis used data acquired through the Scopus repository from 2014 to 2025. To highlight emerging areas of research across various engineering areas, it offers link–walkthrough, keyword co-occurrence and author co-citation analysis of networks. The report highlights the present state of DTs in engineering, their impact on numerous fields and the potential for collaborative research.
Most evaluation metrics for binary classification are derived from the confusion matrix, which is inherently non-differentiable because it relies on discrete predictions. This limits their direct use as loss functions in gradient-based learning, creating a mismatch between training objectives and evaluation criteria. To that end, we offer a general-purpose approach, AnyLoss , that transforms any confusion-matrix-based metric into a differentiable loss function. AnyLoss employs a distinct approximation strategy to estimate a specific, targeted metric score for the prediction model. This is followed by a theoretical and practical analysis of the method, which involves conducting extensive experiments with neural network architectures ranging from simple to advanced across diverse data modalities, including tabular, image, and text. The experimental results demonstrate the generality of our new method, which can target any evaluation metrics derived from a confusion matrix, and highlight that it excels at handling imbalanced datasets.
Graph Neural Networks (GNNs) have been widely used for learning representations of graph-structured data, achieving remarkable success in various graph-related Web applications, such as fraud detection. To generate node representations, GNN-based models operate message-passing mechanisms that aim to smooth the learned representations in a local neighborhood. However, fraudsters increasingly employ sophisticated “camouflage” tactics, exhibiting normal behaviors by strategically forming numerous connections with legitimate entities. As a result, existing GNN-based methods struggle to effectively tackle such fraudulent activities due to their reliance on homophily-based message-passing architectures. These methods fail to generate discriminative representations, which is crucial for distinguishing fraudsters from benign entities. To address this problem, we propose a novel Discriminative Enhanced Aggregation Graph Neural Network-based FraudDEtectioNMoDel (DEFEND) . DEFEND incorporates tailored discriminative mechanisms that strengthen representation learning at two complementary levels: (i) intra-relation and (ii) inter-relation. While prior approaches primarily focus on intra-relation patterns and overlook inter-relation information, DEFEND integrates both to capture subtle inconsistencies in fraudster behavior. Specifically, an edge discriminating mechanism classifies neighborhoods into homophily or heterophily-based views by leveraging node attributes and structural characteristics, and a camouflage-aware dual-channel aggregation module captures different frequencies of information tailored to these views to generate rich intra-relation node representations. While prior approaches typically rely on intra-relation information within each relation type, they overlook the discriminative signals that arise from correlations across different relations. In DEFEND, we observe that fraudsters often avoid forming consistent cross-relation interactions, whereas benign entities tend to establish them more frequently. This discrepancy creates a distinctive behavioral pattern. To capture this, we introduce an inter-relation correlation mechanism that correlates a node’s intra-relation representations across multiple relation types using an attention-based weighting scheme. By adaptively weighing the importance of each relation and integrating their contributions, DEFEND enhances the discriminative power of node representations. This mechanism enables the model to leverage both intra-relation and inter-relation levels of information, leading to richer and more robust representations for fraud detection. Finally, a multi-relation combination module aggregates information across different relation types, emphasizing the importance of node–relation pairs in the embedding. We conducted extensive experiments on two real-world fraud datasets to demonstrate the effectiveness of our proposed model, and our results show that DEFEND outperforms the state-of-the-art baselines. The source codes and datasets of our work are available at https://github.com/VenusHaghighi/DEFEND .
Social influence plays a crucial role in shaping user preferences and behaviors, making social recommendation an effective approach for alleviating the cold-start problem. However, most existing social recommendation methods either model social influence at the individual level or assume non-overlapping community structures, which fails to reflect the fact that users typically belong to multiple communities simultaneously. As a result, the influence from different communities and their cross-community interactions are not explicitly modeled. In this paper, we formally study the problem of overlapping community-aware social recommendation, where a user's preference is jointly influenced by personal behavior, social neighbors, and multiple overlapping communities, each contributing differently depending on the target item. To address this problem, we propose GANOC, a unified framework that decomposes user preference into three complementary domains: personal, social, and community. We employ graph attention networks to model influence propagation in both social and community graphs, and design an item-aware attention mechanism to selectively aggregate cross-community influences. Furthermore, a domain attention network is introduced to adaptively integrate preference representations from different domains for rating prediction. Extensive experiments on three real-world benchmark datasets demonstrate that GANOC outperforms state-of-the-art social recommendation methods, particularly for cold-start users, validating the effectiveness of explicitly modeling overlapping community influence.
Session-based recommendation systems focus on capturing users’ evolving intents from short interaction sequences, yet they persistently face three key challenges: the difficulty in dynamically discriminating between short-term and long-term interests, the inherent tradeoff between sequential modeling and relational dependency learning, and the pervasive noise and sparsity in real-world session data. To tackle these issues, we propose Multivariate Relationship Graph Embedding (MRGE), a novel framework that synergizes enhanced recurrent modeling with graph-structured representations. Specifically, MRGE leverages a self-attention–enhanced RNN to concurrently model short-term intents and long-term preferences within sessions, while constructing a heterogeneous session graph that captures multi-relational item dependencies without compromising temporal fidelity. In addition, we introduce an auxiliary edge augmentation mechanism based on neighbor similarity to mitigate data sparsity and noise, thereby facilitating more robust information propagation. Extensive experiments on three public benchmarks— Delicious , Gowalla , and Foursquare —show that MRGE consistently surpasses state-of-the-art baselines and achieves significant improvements in top- \(K\) recommendation accuracy. Our implementation is available at: https://github.com/July-jz/MRGEcode .
Legal Judgment Prediction (LJP) focuses on predicting judgment results based on the facts of cases. While State-of-the-Art (SOTA) methods have shown impressive performance in law article prediction and charge prediction, they still exhibit weaknesses in prison term prediction. One major reason is that existing models fail to mimic human legal quantitative reasoning to understand monetary features in case facts. Consequently, they do not rigorously quantify the severity of the crime, which is essential for prison term prediction. In this article, we explore and explain how to leverage monetary features to improve LJP via quantitative reasoning. Specifically, we propose QR-LJP, a quantitative reasoning-based LJP model, to integrate legal reasoning knowledge into the prediction process. QR-LJP first employs a curated LLM to extract monetary values from case facts and uses legal quantitative reasoning logic to determine the total crime amount, serving as the quantitative measure of the crime’s severity. This measure is subsequently used to make judgment predictions. We evaluate our model on the real-world dataset CAIL-2018. Experimental results demonstrate that our model outperforms current SOTAs, highlighting the effectiveness of legal quantitative reasoning. Moreover, applying our quantitative reasoning strategy to existing SOTA methods yields significant improvements, especially in macro-F1 scores.
Time series forecasting is essential in many real-world applications, yet developing models that generalize well to unseen and related domains – such as forecasting web traffic on new websites/platforms or predicting e-commerce demand in new regions – remains a big challenge. Prior work addressing this problem, known as domain generalization, focuses on identifying common patterns but often overlooks the complex characteristics of time series data and fails to use the available information from test samples of unseen domains. We propose a novel approach, adaptive latent decomposition (ALD) for domain generalization in time series forecasting, which consists of a decomposed variational autoencoder (VAE) and an adaptive inference mechanism to improve predictive performance in unseen domains. The decomposed VAE involves a learnable kernel-selection mechanism that learns latent variables of decomposed components of time series, i.e., trend-cyclical and seasonal components. These latent variables capture hidden temporal dependencies in time series data, allowing forecasting models to learn general patterns from various seen domains. The adaptive inference mechanism bridges the gap between seen and unseen domains with a sample-wise optimization strategy specifically designed for time series forecasting. ALD builds a latent variables-aware pre-trained model and tailors it for each test sample, improving the generalization on unseen test domains. We validate ALD across six real-world datasets, from online behavior to complex temporal systems, demonstrating its superior generalization performance compared to state-of-the-art methods.
The task of dynamic graph link prediction is to forecast the evolution of complex systems. Empirical observations reveal that interactions within these systems exhibit an Entangled Spatio-Temporal Pattern, which manifests through three interrelated phenomena, namely Latent High-Order Bridges, Multi-Frequency Temporal Dynamics, and Spatio-Temporal Entanglement, with stronger structural ties facilitating tolerance for longer temporal gaps. However, limited by computationally prohibitive multi-hop sampling or inefficient long-sequence modeling, existing methods struggle to capture this complex pattern. Inspired by State-Space Models (SSMs) like Mamba for efficient long-range modeling yet aiming to address their native agnosticism to structural and multi-frequency dynamics, we propose a framework named DyGHydra, which couples a tailored Continuous-Time Hierarchical Mamba (CT-HMamba) backbone with a multi-hop structural encoder. The framework first employs the multi-hop structural encoder to reveal latent high-order interactions, extracting interaction-level cross-hop features. Subsequently, the CT-HMamba backbone utilizes these features to address multi-frequency dynamics through a hierarchical architecture, decomposing interaction history to simultaneously model high-frequency bursts and long-term trends. To capture the spatio-temporal entanglement, CT-HMamba further tailors its core state-space mechanism to be co-driven by physical time and structural context. Specifically, physical time governs the state transition decay to reflect temporal forgetting, while structural context modulates the input-output projections to prioritize topologically significant events. Extensive experiments on eleven real-world datasets show that DyGHydra achieves state-of-the-art performance across most settings for both transductive and inductive link prediction, validating its effectiveness in modeling complex temporal dynamics with superior efficiency.
Quasi-clique is one of the most fundamental models for characterizing cohesive subgraphs in network analysis. However, existing quasi-clique definitions and identification algorithms are designed for unsigned graphs, while many real-world networks are modeled as signed graphs with positive and negative edges representing cooperative and adversarial interactions between entities. Therefore, it remains an open problem to define a quasi-clique model tailored for signed graphs. Motivated by this, we propose the maximal balanced \( (\gamma_{1},\gamma_{2}) \) -quasi-clique (MBQC) model, which not only preserves the essence of quasi-completeness but also aligns with the foremost structural balance theory for signed graphs. Specifically, we formulate the problem of MBQCs enumeration in a given signed graph and prove its NP-hardness. To address this problem, we devise a novel branch-and-bound algorithm to efficiently enumerate all MBQCs in a signed graph, which is further optimized with several carefully-crafted techniques to prune unpromising search spaces and enhance enumeration efficiency. Extensive experiments on real-world datasets demonstrate the efficiency, scalability, and effectiveness of our MBQC model and algorithms.
Sequential recommender systems, as a core technology for personalized services, have been widely applied across various domains including e-commerce and social networks. Transformer-based sequential recommendation models face challenges such as high resource consumption and inefficient sequential processing, limiting their deployment and real-time recommendation performance. Inspired by the efficiency of the RWKV architecture in handling long sequences in natural language processing, this article proposes RWKV4Rec (RWKV for Sequential Recommendation), the first application of RWKV to sequential recommendation tasks. Our RWKV4Rec model leverages RWKV’s efficient long-sequence processing capability and linear computational complexity, reducing resource demands and eliminating additional sequential processing operations. We design an item-RWKV block module and propose a Low-Rank Time Mix based on LoRA technology, which adaptively assigns weights to historical items at each timestep to generate expressive behavior sequence features. Combined with explicit user embeddings and personalized local item matrices, it enhances sequence modeling for efficient and accurate interest prediction. Experiments on four benchmark datasets demonstrate that RWKV4Rec outperforms state-of-the-art methods, achieving 1.80–3.76% improvements in NDCG@10.