Multimedia information floods the Internet, subtly influencing human society. Combining multimedia information to alleviate the data sparsity problem is a popular way within the rapid development of recommender systems. However, many studies reveal that multimodal information can introduce cross-modality noise in some cases. A feasible solution to alleviate cross-modality noises is to enhance the common information among modalities. Recent advanced works enhance modality common information between users (via user-user graphs) or items (via item-item graphs) using extra homogeneous graphs. However, these additional homogeneous graph structures will inevitably bring huge computational costs. To better extract common information among modalities while reducing computational costs, we propose a biLateral glOBal SemanTic Enhancement for multimedia Recommendation, which is called LOBSTER. Specifically, LOBSTER constructs two global semantic spaces for user and item representations, enhances global/common semantic features on both the user and item sides through additional learnable representations shared across multiple modalities. LOBSTER further incorporates a layer-refined Graph Convolutional Network (GCN) and a dynamic optimization to alleviate the over-smoothing problem and adjust attention levels for different modalities. Extensive experiments on three real-world datasets demonstrate that LOBSTER achieves competitive or superior performance compared to models incorporating homogeneous graphs, while providing an average 2.45x speedup and a 60.26% reduction in memory usage. Our code is available at https://github.com/Jinfeng-Xu/LOBSTER.
The "pre-training, prompt-tuning" has emerged as a pivotal paradigm in advancing the performance of graph representation learning models across a wide range of downstream tasks. This paradigm leverages the power of pre-trained models and task-specific prompts to bridge the gap between general graph representations and task-specific requirements. Early graph prompt tuning approaches relied on task-specific designs for Graph Neural Networks (GNNs), limiting their adaptability across diverse pre-training strategies. In contrast, another promising line of research has investigated universal graph prompt tuning, which operates directly in the input graph's feature space and builds a theoretical foundation that universal graph prompt tuning can theoretically achieve an equivalent effect of any prompting function, eliminating dependence on specific pre-training strategies. Recent works propose selective node-based graph prompt tuning to pursue more ideal prompts. However, we argue that selective node-based graph prompt tuning inevitably compromises the theoretical foundation of universal graph prompt tuning. In this paper, we strengthen the theoretical foundation of universal graph prompt tuning by introducing stricter constraints, demonstrating that adding prompts to all nodes is a necessary condition for achieving the universality of graph prompts. To this end, we propose a novel model and paradigm, Learning and Editing Universal GrAph Prompt Tuning (LEAP), which preserves the theoretical foundation of universal graph prompt tuning while pursuing more ideal prompts. Specifically, we first build the basic universal graph prompts to preserve the theoretical foundation and then employ actor-critic reinforcement learning to select nodes and edit prompts. Extensive experiments on graph- and node-level tasks across various pre-training strategies in both full-shot and few-shot scenarios show that LEAP consistently outperforms fine-tuning and other prompt-based approaches.
Constructing radio maps traditionally requires extensive site surveys with precise location labels, resulting in costly and time-consuming calibration. Conventional approaches derive labels from inertial measurement units (IMUs) but are constrained by device-level access permissions and the need for pre-installed IMU hardware. In this article, we present a calibration-free radio map construction method that relies solely on channel state information (CSI) measurements, thereby obviating the need for location labels. Our key insight is to embed CSI data into a 2-D geographic space using a neural network, without any label information. The primary challenge is to preserve the real-world topology in the embedded space. To address this issue, we propose a novel topology-guided manifold learning (TGML) approach that learns a low-dimensional embedding through self-supervision based on its mapping to a topological map in the geographic space. Specifically, we introduce an embedding network that learns locally smoothed proximity, thus creating a compact low-dimensional representation in the latent space. We further employ a regularized transport method with a differentiable Sinkhorn distance to establish an optimal mapping between the latent and geographic spaces. We use the resulting mapping deviation to supervise the training of the embedding network, enabling continuous refinement through self-supervision. Experiments in an office environment show that TGML achieves an average localization error of 2.37 m and a relative beam estimation error of 6.68%, outperforming state-of-the-art methods.
Recent advancements in multimodal recommendations, which leverage diverse modality information to mitigate data sparsity and improve recommendation accuracy, have gained significant attention. However, existing multimodal recommendations overlook the critical role of user representation initialization. Unlike items, which are naturally associated with rich modality information, users lack such inherent information. Consequently, item representations initialized based on meaningful modality information and user representations initialized randomly exhibit a significant semantic gap. To this end, we propose a Semantically Guaranteed User Representation Initialization (SG-URInit). SG-URInit constructs the initial representation for each user by integrating both the modality features of the items they have interacted with and the global features of their corresponding clusters. SG-URInit enables the initialization of semantically enriched user representations that effectively capture both local (item-level) and global (cluster-level) semantics. Our SG-URInit is training-free and model-agnostic, meaning it can be seamlessly integrated into existing multimodal recommendation models without incurring any additional computational overhead during training. Extensive experiments on multiple real-world datasets demonstrate that incorporating SG-URInit into advanced multimodal recommendation models significantly enhances recommendation performance. Furthermore, the results show that SG-URInit can further alleviate the item cold-start problem and also accelerate model convergence, making it an efficient and practical solution for multimodal recommendations.
Multimodal large language models (MLLMs) are increasingly integrated into video conferencing, but mostly for speech-centric tasks such as transcription and summarization. Conferencing adaptation remains dominated by signal-driven congestion control that reacts to bandwidth changes without understanding why certain moments or streams will soon become quality-critical. This separation wastes predictive structure in conversation: text often reveals impending role shifts and interaction-mode transitions-for example, taking the floor or initiating screen sharing-seconds before the corresponding media and bandwidth demands materialize. We present Intent-aware Video conferencing Adaptation (IVA), a cross-modal control framework that converts text-derived intent cues into proactive bitrate orchestration. IVA extracts activity, control, and intent semantics from live streams, maps them to per-stream priority weights via a gated-bias model, and solves constrained bitrate allocation to maximize semantic QoE under uplink/downlink and latency limits. IVA learns this semantics-to-priority mapping with reinforcement learning over trace-driven simulations. Experiments with real-world network traces and emulated conferencing scenarios show that IVA improves QoE over signal-only ABR baselines and a semantic-aware conferencing baseline, while preserving low end-to-end latency and stabilizing quality around intent-triggered transitions.
Urban water distribution networks face significant challenges from pipeline leakage, which leads to water loss and operational inefficiencies. Existing data-driven detection methods often neglect inherent hydraulic principles, resulting in poor model generalizability and a lack of quantitative leakage severity assessment. To address these issues, this paper proposes a physics-informed graph transformer fusion (PI-GTF) framework that integrates hydraulic mechanisms with deep learning for leakage detection and grading. The model embeds hydraulic governing equations and signal propagation rules into a graph convolutional network (GCN) and a transformer to capture spatial pipeline topology and long-term temporal dependencies of leakage signals. A novel physics-aware hierarchical adversarial gating attention (PHAGA) module is designed to align and fuse these heterogeneous features effectively. Furthermore, a five-level leakage grading system is established by combining hydraulic model outputs with sensor-based features such as pressure fluctuations and abnormal flow durations. The experimental results of a high-fidelity simulation model of Shenyang’s water network show that PI-GTF outperforms existing methods in terms of accuracy, precision, and F1 score, with zero cross-level misclassification. Migration tests on real residential networks demonstrate strong generalizability, with performance degradation within 2%. This study provides a reliable dual-driven framework for end-to-end leakage management and supports intelligent decision-making in water network maintenance.
Irregular Medical Time Series play a critical role in the clinical domain to better understand the patient's condition. However, inherent irregularity arising from heterogeneous sampling rates, asynchronous observations, and variable gaps poses key challenges for reliable modeling. Existing methods often distort temporal sampling irregularity and missingness patterns while failing to capture variable decay irregularity, resulting in suboptimal representations. To address these limitations, we introduce DBGL, Decay-Aware Bipartite Graph Learning for Irregular Medical Time Series. DBGL first introduces a patient-variable bipartite graph that simultaneously captures irregular sampling patterns without artificial alignment and adaptively models variable relationships for temporal sampling irregularity modeling, enhancing representation learning. To model variable decay irregularity, DBGL designs a novel node-specific temporal decay encoding mechanism that captures each variable's decay rates based on sampling interval, yielding a more accurate and faithful representation of irregular temporal dynamics. We evaluate the performance of DBGL on four publicly available datasets, and the results show that DBGL outperforms all baselines.
Federated Learning (FL) emerged as a promising distributed machine learning paradigm. However, extending FL to the class incremental learning scenarios introduces unique challenges: 1) Capacity conflict and catastrophic forgetting from the shared model overloading, 2) Heterogeneity from Non-Independent and Identically Distributed (Non-IID) data, and 3) Synchronized class misalignment. In this paper, we propose Fisher-Routed MiXture of Experts for Federated Class-Incremental Learning (FedFMX), a novel framework to address these challenges via adaptive expert specialization across clients. The crucial insight is to route each sample to an expert subset that jointly optimizes knowledge acquisition and retention. Specifically, we introduce a Fisher-Routed Expert Scoring (FRES) module to estimate expert importance via Fisher-based stability cost and gradient-based plasticity gain. Then, we design an Adaptive Expert Selection (AES) module by quantifying marginal contributions for adaptive expert subset determination. Finally, by the routing-aware regularization (RAR), we achieve load balance and efficient FL training. We theoretically prove the 𝒪(T^-1) convergence rate. Extensive experiments on multiple benchmarks compared with state-of-the-art methods demonstrate the superiority of FedFMX.
Multimodal recommendation leverages item multimodal features alongside collaborative signals to capture user preferences. While item-item graphs have become a key building block in advanced models, existing methods typically construct them with noisy similarity edges and limit their role to a single function of item-item representation propagation, leaving substantial potential untapped. In this paper, we propose IIMRec, a framework that constructs a single high-quality item-item graph during preprocessing and systematically reuses it across three stages of the recommendation pipeline: representation enhancement, interaction graph enhancement, and optimization enhancement. The graph is built by fusing semantic and co-occurrence signals, then refined via Neighborhood Consistency Edge Reweighting (NCER), which applies the triadic closure principle to amplify structurally reliable edges and suppress spurious ones. Once constructed, the graph is leveraged in three complementary ways: (1) Item-item propagation with a Residual II Gate (RIG) that adaptively controls per-item absorption of semantic neighborhood signals for representation enhancement; (2) A content-guided UI graph expansion that introduces virtual user-item edges through high-confidence semantic neighbors for interaction graph enhancement; (3) II-Neighbor BPR Augmentation (INA) that treats top neighbors of positive items as discounted soft positives for optimization enhancement. We provide theoretical analysis showing that NCER reduces the spectral noise-to-signal ratio, RIG converges to a non-degenerate gating regime, and INA yields a tighter generalization bound. Extensive experiments on four datasets demonstrate that IIMRec consistently outperforms state-of-the-art baselines while running faster and consuming less GPU memory, with particularly strong gains under cold-start and sparse-interaction conditions.
Multiple clustering aims to discover diverse latent structures from different perspectives, yet existing methods generate exhaustive clusterings without discerning user interest, necessitating laborious manual screening. Current multi-modal solutions suffer from static semantic rigidity: predefined candidate words fail to adapt to dataset-specific concepts, and fixed fusion strategies ignore evolving feature interactions. To overcome these limitations, we propose Multi-DProxy, a novel multi-modal dynamic proxy learning framework that leverages cross-modal alignment through learnable textual proxies. Multi-DProxy introduces 1) gated cross-modal fusion that synthesizes discriminative joint representations by adaptively modeling feature interactions. 2) dual-constraint proxy optimization where user interest constraints enforce semantic consistency with domain concepts while concept constraints employ hard example mining to enhance cluster discrimination. 3) dynamic candidate management that refines textual proxies through iterative clustering feedback. Therefore, Multi-DProxy not only effectively captures a user's interest through proxies but also enables the identification of relevant clusterings with greater precision. Extensive experiments demonstrate state-of-the-art performance with significant improvements over existing methods across a broad set of multi-clustering benchmarks.
Although existing multimodal recommendation models have shown promising performance, their effectiveness continues to be limited by the pervasive data sparsity problem. This problem arises because users typically interact with only a small subset of available items, leading existing models to arbitrarily treat unobserved items as negative samples. To this end, we propose VI-MMRec, a model-agnostic and training cost-free framework that enriches sparse user-item interactions via similarity-aware virtual user-item interactions. These virtual interactions are constructed based on modality-specific feature similarities of user-interacted items. Specifically, VI-MMRec introduces two different strategies: (1) Overlay, which independently aggregates modality-specific similarities to preserve modality-specific user preferences, and (2) Synergistic, which holistically fuses cross-modal similarities to capture complementary user preferences. To ensure high-quality augmentation, we design a statistically informed weight allocation mechanism that adaptively assigns weights to virtual user-item interactions based on dataset-specific modality relevance. As a plug-and-play framework, VI-MMRec seamlessly integrates with existing models to enhance their performance without modifying their core architecture. Its flexibility allows it to be easily incorporated into various existing models, maximizing performance with minimal implementation effort. Moreover, VI-MMRec introduces no additional overhead during training, making it significantly advantageous for practical deployment. Comprehensive experiments conducted on six real-world datasets using seven state-of-the-art multimodal recommendation models validate the effectiveness of our VI-MMRec.
Causality has been integrated with machine learning in uncovering and understanding the causal relationship between variables and observed outcomes. However, the centralized training setting of causal machine learning is not adaptable to most practical scenarios, where datasets are distributed, stored, and unsharable due to privacy concerns. Federated learning (FL), a distributed learning framework that allows collaborative training across multiple devices without raw data sharing, emerges as a potential solution to this problem. By integrating FL into causal problems, the discovery and inference of causal relationships across dispersed datasets can be achieved. On the other hand, causality can also enhance FL models in various dimensions, including model interpretability and explainability, generalizability, adversarial robustness, and fairness and bias mitigation. In this article, we provide a comprehensive review of the above two directions and summarize the interplays between causality and FL (short for Causal-FL ) by organizing our discussion around two key questions: (1) how FL enable decentralized causal analysis; and (2) how causality tackles FL challenges. The potential applications of these methods are also introduced, including healthcare, recommendation, economics, social equity, and so on. Moreover, we discuss promising future directions and future challenges to be explored.
Federated crowdsourcing has emerged as a promising paradigm for collaboratively solving learning tasks over mobile edge devices, enabling decentralized model training without direct data uploading. However, conventional federated crowdsourcing systems largely overlook two fundamental challenges: (i) the lack of effective incentive mechanisms under unknown participant quality, (ii) the risk of privacy leakage from model uploading. In this paper, we propose an incentive mechanism for federated crowdsourcing based on the Stackelberg game considering rho-zero-concentrated differential privacy and Combinatorial Multi-Armed Bandit mechanism, called FedCMAB, to tackle the client selection problem with unknown quality and incentive design while preserving confidential information. We model the interaction between the server and clients as a Stackelberg game, where the server dynamically selects a subset of clients and designs rewards to minimize the cost while each client strategically determines its privacy budget to maximize its individual utility. We theoretically establish the existence of the Stackelberg Nash equilibrium. Next, we leverage the combinatorial multi-armed bandit (CMAB) method to learn optimal client selection strategies with provable regret guarantees and derive an upper bound on the cumulative regret of the proposed mechanism. Moreover, we conduct a rigorous Price of Anarchy (PoA) analysis to quantify the efficiency gap between the decentralized equilibrium induced by self-interested clients and the socially optimal solution. Our analyses demonstrate that conventional incentive-agnostic strategy can lead to an unbounded PoA, resulting in severe efficiency loss. In contrast, our FedCMAB framework provably bounds the PoA by a finite constant. Extensive experimental results on multiple datasets demonstrate that Fed-CMAB consistently outperforms state-of-the-art baseline methods on both independent and identically distributed (IID) and non-IID data.
Graph Collaborative Filtering (GCF) has become the dominant paradigm in modern recommender systems by modeling user-item interactions as a bipartite graph and propagating embeddings through a fixed number of message-passing layers. However, applying a uniform propagation depth to every node ignores a fundamental property of real interaction graphs: nodes differ substantially in their local connectivity, so peripheral nodes quickly suffer from over-smoothing while hub-like nodes remain under-explored beyond their immediate neighborhood. In this paper, we revisit GCF from a tree-structured perspective and propose Neural Tree Collaborative Filtering (NTCF), a framework that re-interprets each node's local neighborhood as a rooted tree and assigns a node-specific propagation depth based on a closed-form local-degree-imbalance score that serves as a discrete Ricci-curvature proxy. We provide a theoretical analysis showing that (i) NTCF strictly generalizes NGCF, degenerating to NGCF when all curvature-induced depth adjustments vanish (a lower bound on its representation power), and (ii) the curvature-aware schedule retains strictly more discriminative information at deep layers on positively-curved (peripheral) nodes than uniform-depth propagation. NTCF can achieve higher performance than most widely used GCF backbone models and can be integrated into existing advanced self-supervised models as a backbone, replacing their original backbone to achieve enhanced performance. Extensive experiments on three public datasets demonstrate the superiority of NTCF.
The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices. Cloud-hosted LLMs are powerful but expensive and difficult to scale due to vendor lock-in and high resource needs, resulting in high expenses and unstable performance under load. Recent efforts focus on deploying small language models (SLMs), distilled or pruned from LLMs, on resource-constrained edge devices to reduce costs and improve scalability. However, edge-based SLMs face limited knowledge coverage and notable accuracy gap compared to cloud-based LLMs. To address this, we present DEFRAG, a decentralized edge collaboration system for retrieval-augmented generation (RAG) that optimizes both retrieval and generation across heterogeneous edge devices. For retrieval, DEFRAG compresses and shares knowledge graphs, using hybrid retrieval to expand knowledge coverage. For generation, DEFRAG introduces an optimizer that adaptively selects SLMs and RAG parameters per query, balancing accuracy and cost. We implement DEFRAG on a heterogeneous edge testbed and evaluate it on benchmark QA datasets. We also test it under mobile route stress, non-uniform data placement, and a domain-specific QA workload. The results show that DEFRAG maintains stable service quality and cost efficiency under these broader settings. Results show that DEFRAG narrows the SLM-LLM accuracy gap, while reducing cost by up to 98.4
The proliferation of IoT devices has driven the demand for machine-centric video analytics in vision-critical applications. Unlike human-centric video streaming, these tasks focus on extracting sparse and high-value insights from continuous data streams. To support resource-intensive visual reasoning on constrained edge devices, edge-cloud collaboration is widely adopted. However, existing frameworks face a two-fold challenge. On the edge side, offloading strategies often “blindly” transmit large volumes of redundant background data, overloading network bandwidth and cloud resources. Meanwhile, on the cloud side, resource provisioning faces a fundamental trade-off between minimizing resource costs and ensuring real-time performance under dynamic workloads. To address these challenges, we propose Semantic-aware Edge-cloud Collaboration for cost-efficient video Understanding (SECU), a framework for efficient real-time video analytics. SECU integrates two core components: (1) an edge-side Lightweight Cross-modal Understanding (LCU) module, which distills visual features into semantic keywords to enable selective offloading; and (2) a cloud-side Content Probing-based Optimization (CPO) module, which dynamically requests video content and optimizes cloud resource leasing. This design avoids deploying heavyweight pretrained models on edge devices while enabling scalable and cost-efficient semantic video analytics. Extensive experiments on public datasets demonstrate that SECU achieves up to 78.6% cost savings with high semantic accuracy, outperforming existing methods in resource-constrained environments.
The prevalence of voice-related interaction and communication has raised concerns about privacy leakage and security. For example, millimeter-wave (mmWave) radio signals have been exploited as a potential attacker for acoustic eavesdropping. However, speaker variability and low-quality input pose significant challenges for the practical deployment of mmWave-based eavesdropping. In this paper, we propose SPACE, an acoustic eavesdropping system to recover intelligible speech from low-quality mmWave signals, which can adapt to numerous different speakers and unseen ones. SPACE is a two-stage system that first reconstructs the spectrogram using a novel Radio TransUNet and then synthesizes the waveform through a neural vocoder. Specifically, to alleviate the negative effect of speaker variability, we introduce a speaker encoder to capture speaker features and a fusion network to condition the spectrogram reconstruction based on the extracted speaker characteristics. Further, to facilitate intelligible speech recovery from low-quality input, we design a Frequency Transformation Layer to exploit the correlation among all frequency harmonics and incorporate the neural vocoder to synthesize the speech waveform from the reconstructed spectrogram without using the contaminated phase. The experimental results show that SPACE outperforms existing mmWave-based approaches in scenarios with numerous different speakers and unseen speakers.
Vision-Language Models (VLMs) demonstrate remarkable general-purpose capabilities but often fall short in specialized domains such as medical imaging or geometric problem-solving. Supervised Fine-Tuning (SFT) can enhance performance within a target domain, but it typically causes catastrophic forgetting, limiting its generalization. The central challenge, therefore, is to adapt VLMs to new domains while preserving their general-purpose capabilities. Continual pretraining is effective for expanding knowledge in Large Language Models (LLMs), but it is less feasible for VLMs due to prohibitive computational costs and the unavailability of pretraining data for most open-source models. This necessitates efficient post-training adaptation methods. Reinforcement learning (RL)-based approaches such as Group Relative Policy Optimization (GRPO) have shown promise in preserving general abilities, yet they often fail in domain adaptation scenarios where the model initially lacks sufficient domain knowledge, leading to optimization collapse. To bridge this gap, we propose Reinforced Curriculum Pre-Alignment (RCPA), a novel post-training paradigm that introduces a curriculum-aware progressive modulation mechanism. In the early phase, RCPA applies partial output constraints to safely expose the model to new domain concepts. As the model's domain familiarity increases, training gradually transitions to full generation optimization, refining responses and aligning them with domain-specific preferences. This staged adaptation balances domain knowledge acquisition with the preservation of general multimodal capabilities. Extensive experiments across specialized domains and general benchmarks validate the effectiveness of RCPA, establishing a practical pathway toward building high-performing and domain-adaptive VLMs.