We present MobiFuse, a high-precision depth perception system on mobile devices that combines dual RGB and Time-of-Flight (ToF) cameras. To achieve this, we leverage physical principles from various environmental factors to propose the Depth Error Indication (DEI) modality, characterizing the depth error of ToF and stereo-matching. Furthermore, we employ a progressive fusion strategy, merging geometric features from ToF and stereo depth maps with depth error features from the DEI modality to create precise depth maps. Additionally, we create a new ToF-Stereo depth dataset, RealToF, to train and validate our model. Our experiments demonstrate that MobiFuse excels over baselines by significantly reducing depth measurement errors by up to 77.7%. It also showcases strong generalization across diverse datasets and proves effectiveness in two downstream tasks: 3D reconstruction and 3D segmentation.
Environment sensing and fusion via onboard sensors are envisioned to be widely applied in future autonomous driving networks. This paper considers a vehicular system with multiple self-driving vehicles that is assisted by multi-access edge computing (MEC), where image data collected by the sensors is offloaded from cellular vehicles to the MEC server using vehicle-to-infrastructure (V2I) links. Sensory data can also be shared among surrounding vehicles via vehicle-to-vehicle (V2V) communication links. To improve spectrum utilization, the V2V links may reuse the same frequency spectrum as the V2I links, which may cause severe interference. To tackle this issue, we leverage reconfigurable intelligent computational surfaces (RICSs) to jointly enable V2I reflective links and mitigate interference appearing at the V2V links. Considering the limitations of traditional algorithms in addressing this problem, such as the assumption of quasi-static channel state information, which restricts their ability to adapt to dynamic environmental changes and leads to poor performance under frequently varying channel conditions, in this paper, we formulate the problem at hand as a Markov game. Our novel formulation is applied to time-varying channels subject to multi-user interference and introduces a collaborative learning mechanism among users. The considered optimization problem is solved via a driving safety-enabled multi-agent deep reinforcement learning (DS-MADRL) approach that capitalizes on the RICS presence. Our extensive numerical investigations showcase that the proposed reinforcement learning approach achieves faster convergence and significant enhancements in both data rate and driving safety, as compared to various state-of-the-art benchmarks.
In underwater covert cooperative missions, autonomous underwater vehicles (AUVs) often cannot rely on active sonar to continuously obtain complete information, since active sensing and frequent communications increase the risk of exposure. As a result, AUVs primarily rely on passive observation, an approach that yields incomplete local perception and limited task efficiency. Although underwater acoustic communications can mitigate this limitation through information sharing, they are simultaneously constrained by long delays, severe interference, low reliability, and the risk of covert exposure. Existing communications-oriented multi-agent reinforcement learning (MARL) studies often model communication as an ideal information flow, whereas traditional communication optimization primarily focuses on link-level performance. However, both are insufficient to characterize the actual contribution of perceptual information to cooperative tasks under realistic conditions of covert physical communications. This paper proposes a Sensed Information Value Realization Multi-Agent Reinforcement Learning (SVR-MARL) framework that leverages practical information to characterize the utility of information for cooperative tasks and learns distributed cooperative policies under realistic communication and covert constraints. Through a case study of covert multi-AUV cooperative localization and tracking, the potential of the proposed framework to improve collaborative task efficiency while reducing unnecessary communication and exposure risks is demonstrated.
Federated Learning (FL) faces significant challenges arising from both data and system heterogeneity. While Clustered Federated Learning (CFL) mitigates data heterogeneity by grouping clients with similar data distributions, it remains vulnerable to system heterogeneity, which can slow convergence due to performance disparities among clients. Moreover, data drift may degrade clustering accuracy and training efficiency over time. In this work, we propose a Model Structure-aware Clustered Federated Learning (MSCFL) framework that simultaneously addresses the issues of data heterogeneity, system heterogeneity, and data drift. MSCFL incorporates model pruning (MP) into the CFL framework to enhance training efficiency under system heterogeneity. To enable this integration, we address the key challenge of performing effective clustering based on heterogeneous, pruned local models with varying structures. To this end, we design a model structure-based similarity computation algorithm to integrate CFL with MP. To effectively address data drift, we propose a dynamic cluster migration strategy that efficiently monitors model structures via Hamming Distance and triggers re-clustering only when necessary. Extensive experimental results show that MSCFL improves the accuracy and convergence speed of cluster models, outperforming traditional CFL in various settings.
Offline Federated Reinforcement Learning (FRL), a marriage of federated learning and offline reinforcement learning, has attracted increasing interest recently. Albeit with some advancement, we find that the performance of most existing offline FRL methods drops dramatically when provided with mixed-quality data, that is, the logging behaviors (offline data) are collected by policies with varying qualities across clients. To overcome this limitation, this paper introduces a new vote-based offline FRL framework, named FOVA. It exploits a vote mechanism to identify high-return actions during local policy evaluation, alleviating the negative effect of low-quality behaviors from diverse local learning policies. Besides, building on advantage-weighted regression (AWR), we construct consistent local and global training objectives, significantly enhancing the efficiency and stability of FOVA. Further, we conduct an extensive theoretical analysis and rigorously show that the policy learned by FOVA enjoys strict policy improvement over the behavioral policy. Extensive experiments corroborate the significant performance gains of our proposed algorithm over existing baselines on widely used benchmarks.
Mobile advertising dominates app monetization but introduces risks ranging from intrusive user experience to malware delivery. Existing detection methods rely either on static analysis, which misses runtime behaviors, or on heuristic UI exploration, which struggles with sparse and obfuscated ads. In this paper, we present MANA, the first agentic multimodal reasoning framework for mobile ad detection. MANA integrates static, visual, temporal, and experiential signals into a reasoning-guided navigation strategy that determines not only how to traverse interfaces but also where to focus, enabling efficient and robust exploration. We implement and evaluate MANA on commercial smartphones over 200 apps, achieving state-of-the-art accuracy and efficiency. Compared to baselines, it improves detection accuracy by 30.5
Wireless charging is a cornerstone technology for next-generation mobile and ubiquitous computing. However, its practical deployment has long been constrained by short range, poor flexibility, and lack of support for dynamic multi-device scenarios. In this paper, we propose ChargeX—a system that enables long-range and mobility-resilient wireless charging for multiple small devices. ChargeX pioneers the integration of metasurface-assisted magnetic beamforming, a high-frequency compact transceiver design, and a real-time closed-loop feedback-control mechanism. It further advances the field by introducing a joint optimization framework for dynamically allocating energy across mobile receivers with heterogeneous priorities and spatial-temporal demands. Experimental results demonstrate that it achieves meter-level charging distance, real-time response to device movement, and efficient coordination among multiple receivers, significantly outperforming state-of-the-art prototypes.
Reconfigurable intelligent surfaces (RISs) have gained significant attention in recent years due to their ability to control the reflection of radio-frequency signals and reshape the wireless propagation environment. Unlike traditional studies that primarily focus on the advantages of RISs, this paper examines the negative impacts of RISs by investigating interference propagation caused by user mobility in downlink wireless systems. We employ a stochastic geometric model to simulate the locations of base stations and RISs using the Matérn hard core point process, while user locations are modeled with the homogeneous Poisson point process. We derive novel closed-form expressions for the power distributions of the received signal at the users and the interfering signal. Additionally, we present a novel expression for coverage probability and introduce the concept of interference propagation intensity. To characterize the dynamics of interference caused by user mobility, we adopt an epidemiological approach using the susceptible-infected-susceptible model. Finally, crucial factors influencing the propagation of interference are analyzed. Numerical results validate our theoretical analysis and provide suggestions for managing interference propagation in large-scale multi-RIS wireless communication networks.
Mobile network operators (e.g., China Mobile, Verizon) are significant for providing communication services and collecting massive amounts of data. However, operators are increasingly concerned about customer data breaches involving third-party application providers (e.g., Tencent, Apple, Netflix). This concern is particularly aggravated when anonymous datasets shared with third-party providers or publicly released can be linked to user data compromised in breaches, leading to severe re-identification attacks and privacy threats. However, comprehensive methods for identifying such privacy risks on a large scale are lacking due to limited networking behavioral data. To address this, we aim to measure the re-identification privacy risk associated with sharing or releasing cellular traces amidst data breaches. Based on the analysis of key privacyimpacting features in traffic usage and base station association data, we propose a novel re-identification method, SURE, which learns similarities between cellular traces to classify if traces belong to the same user. Extensive experiments on a largescale dataset of 10,000 users over four months demonstrate SURE's superior performance, with AUC scores exceeding 0.9. Our findings reveal significant re-identification risks in data sharing/release, influenced by data scale and user attributes, corroborated by a public dataset.
In modern Video-on-Demand (VOD) services, non linear viewing behaviors such as seeking and skipping are increasingly common. However, current streaming systems often rely on aggressive prefetching and large fixed playback buffers to ensure smooth playback, which can result in substantial traffic waste when prefetched segments are discarded due to frequent user interaction. To mitigate this inefficiency, we propose BufTune, a seek-aware client-side buffer tuning strategy. BufTune dynamically adjusts the playback buffer size in response to user seek behavior—shrinking the buffer during frequent interactions to avoid downloading segments unlikely to be watched, and gradually restoring it during stable playback to maintain robustness, thereby mitigating traffic waste while preserving, or even improving, QoE. BufTune comprises two core modules: a User Behavior Detector that detects seek events and coordinates playback adaptation, and a Buffer Tuner that adjusts the buffer size—shrinking it during frequent seeks to reduce unnecessary prefetching, and progressively restoring it during stable playback to ensure smooth viewing experience. Experimental results across multiple ABR algorithms and network conditions demonstrate that, compared with the fixed buffer approach, BufTune consistently reduces traffic waste by up to 38.5%, while delivering QoE improvements of up to 4.9%. These results highlight BufTune's effectiveness and generality in addressing the emerging challenges of interactive VOD streaming.
The Federated Multi-Armed Bandit (FMAB) framework is proposed to facilitate collaborative model training in cloud-edge environments. Most existing FMAB-based systems assume that participants have personal datasets. This assumption becomes biased in scenarios with limited local data or when the discrepancy between historical data and online data cannot be quantified. Consequently, participants must passively collect data by offering services and gathering feedback from users within their charging areas. In addition, due to the mobility of the users, the number of requests is not stable. In such scenarios, effective model training requires addressing two key challenges. First, the allocation of resources for data collection at the edge is often mismatched with the actual number of service requests, resulting in limited training data and wasted resources. Second, in areas with sparse service requests, the lack of data further delays model adaptation. In this article, networks with these challenges are summarized as the training while collecting data federated bandit (TCF-bandit), and the over-area over-period upper confidence bound ($\mathcal {O}^{2}$-UCB) algorithm is proposed to address two challenges. In the $\mathcal {O}^{2}$-UCB algorithm, the cloud records request-generation patterns to mitigate resource waste caused by allocation mismatches resulting from unpredictable user mobility. Additionally, a weight-based global confidence radius is computed to assist areas with limited data in quickly identifying their optimal arms. Finally, we prove that the resource allocation waste and regret of the $\mathcal {O}^{2}$-UCB algorithm exhibit sub-linear growth. We conduct experiments in different scenarios on the MovieLens, CIFAR-10, and CIFAR-100 datasets to illustrate its superiority over SOTA methods by around 16.2%.
The exponential growth of smartphones and mobile applications (apps) has generated vast app usage traces, which are critical for stakeholders but accompanied by high costs during data collection, as well as significant privacy risks when data is released or shared. Traditional privacy-preserving methods, such as anonymization and differential privacy, often fail to reconcile data utility with robust privacy guarantees, suffering from re-identification vulnerabilities or excessive noise. In this paper, we propose AttrAppGen, a novel framework leveraging large language models (LLMs) to synthesize app usage traces. AttrAppGen comprises four key stages: i) attribute-aware app text dataset construction via attribute desensitization; ii) cluster-based conditional prompt design to discover behavioral patterns of user and app attributes, and then generate prompt-output app data training pairs; iii) data generation LLM fine-tuning based on app data training pairs; and iv) synthetic app usage data sampling for final data generation. Extensive experiments on a real-world dataset comprising 260,000 records from 6,102 users and 1,129 apps demonstrate the superior performance of AttrAppGen in preserving overall data distribution, ensuring high utility across three downstream tasks, and maintaining strong privacy protection. By effectively balancing the trade-off between utility and privacy, AttrAppGen provides a scalable and secure solution for data sharing in mobile data ecosystems.
How to preserve the data privacy during the training of deep neural network (DNN) is a key security concern in the artificial intelligence era. However, most existing solutions based on homomorphic encryption and Trusted Execution Environment (TEE) are incompatible with heterogeneous Neural Network Accelerators (NNAs), leading to significant performance loss. We propose a novel method based on vector decomposition to allocate operators across different NNAs, ensuring both throughput and privacy simultaneously. Furthermore, based on this approach, we have designed a compiler that automatically converts front-end model descriptions into backend encrypted computation graphs, which is running securely over trusted and untrusted hardware. This compiler heuristically determines the allocation scheme based on hardware affinity and cross-hardware communication costs, significantly reducing additional overhead. Experimental results demonstrate that our method does not incur extra accuracy costs and achieves a throughput significantly higher than existing methods. Deploying our approach at scale on a platform with a billion users, we have verified its negligible impact on real-world operations while ensuring the privacy protection capability for cross-domain data.
Real-time video matting is essential for applications like online video conferencing but faces challenges in human-object interaction (HOI) scenarios, known as the HOI-matting problem. This problem is challenging due to its open-recognition nature, where no dataset can cover the wide range of potential HOI cases, making it difficult for feature-learning-based methods to generalize effectively. To address this issue, we present an HOI-matting dataset and introduce a Model-Agnostic Meta-Learning-based rule-aware learning approach (MAML-RAL). MAML-RAL combines transfer learning and meta-learning to capture domain-invariant HOI rules, complemented by a fast local adaptation strategy to counter domain shifts and background interference. Our method achieves a mean intersection-over-union (mIoU) of 92.3%, outperforming current algorithms, with local adaptation further boosting performance to a remarkable mIoU of 95.84%.
The escalating demand for long-context applications has intensified the necessity of extending the LLM context windows. Despite recent fine-tuning approaches successfully expanding context lengths, their high memory footprints, especially for activations, present a critical practical limitation. Current parameter-efficient fine-tuning methods prioritize reducing parameter update overhead over addressing activation memory constraints. Similarly, existing sparsity mechanisms improve computational efficiency but overlook activation memory optimization due to the phenomenon of Shadowy Activation. In this paper, we propose LeMo, the first LLM fine-tuning system that explores and exploits a new token-level sparsity mechanism inherent in long-context scenarios, termed Contextual Token Sparsity. LeMo minimizes redundant token involvement by assessing the informativeness of token embeddings while preserving model accuracy. Specifically, LeMo introduces three key techniques: (1) Token Elimination, dynamically identifying and excluding redundant tokens across varying inputs and layers. (2) Pattern Prediction, utilizing well-trained predictors to approximate token sparsity patterns with minimal overhead. (3) Kernel Optimization, employing permutation-free and segment-based strategies to boost system performance. We implement LeMo as an end-to-end fine-tuning system compatible with various LLM architectures and other optimization techniques. Comprehensive evaluations demonstrate that LeMo reduces memory consumption by up to 1.93x and achieves up to 1.36x speedups, outperforming state-of-the-art fine-tuning systems.
Mobile contact exhibits user co-traveling events within the same transportation tool, which is crucial for resident profiling, face-to-face interaction detection, etc. In this paper, we investigate urban user mobile contact detection with cellular signaling traces, which is cost-efficient to enable large-scale detection. Specifically, we develop a data collection platform to collect substantial user signaling traces, covering different types of road scenarios within a city. With the collected traces, we perform systematic data analysis to reveal several technical challenges, which are sparsity of signaling trajectory, remote base station noise, and fuzzy matching difficulties. To address challenges, we propose a mobile contact detection method named MoCo. In MoCo framework, we first conduct data denoising to remove the noise from remote base stations. Then, we devise a spatio-temporal filter to eliminate unlikely mobile contact traces in both spatial and temporal domains, reducing the computational overhead. Finally, we design a detection network that integrates the submodules of data alignment, feature encoder, spatio-temporal representation learner, and user mobile contact detector. Extensive evaluation results demonstrate the superiority of MoCo in comparison with state-of-the-art baselines. Robust experiments show that MoCo can work efficiently in different transportation modes and urban densities.
This paper addresses the challenges of Online Action Recognition (OAR), a framework that involves instantaneous analysis and classification of behaviors in video streams. OAR must operate under stringent latency constraints, making it an indispensable component for real-time feedback for edge computing. Existing methods, which typically rely on the processing of entire video clips, fall short in scenarios requiring immediate recognition. To address this, we designed EdgeOAR, a novel framework specifically designed for OAR on edge devices. EdgeOAR includes the Early Exit-oriented Task-specific Feature Enhancement Module (TFEM), which comprises lightweight submodules to optimize features in both temporal and spatial dimensions. We design an iterative training method to enable TFEM learning features from the beginning of the video. Additionally, EdgeOAR includes an Inverse Information Entropy (IIE) and Modality Consistency (MC)-driven fusion module to fuse features and make better exit decisions. This design overcomes the two main challenges: robust modeling of spatio-temporal action representations with limited initial frames in online video streams and balancing accuracy and efficiency on resource-constrained edge devices. Experiments show that on the UCF-101 dataset, our method EdgeOAR reduces latency by 99.23% and energy consumption by 99.28% compared to state-of-the-art (SOTA) method. And achieves an adequate accuracy on edge devices.
Federated multimodal learning enables decentralized devices with diverse modalities to collaboratively train multimodal models without sharing their raw data. In most existing federated multimodal learning approaches, multimodal data is indispensable. However, in reality, a considerable number of devices can only collect unimodal data, and the data labels are often incomplete. Therefore, this paper proposes FedCMD, a federated multimodal learning approach that enables unimodal devices with missing labels to collaboratively train a multimodal model. To effectively leverage the label-missing samples, FedCMD performs unimodal federated learning first to learn unimodal encoders and make pseudo-labels for the unlabeled samples. Then it calculates and shares the prototypes of various modalities among devices for cross-modal feature alignment. The prototypes are finally served as the complementary of missing modalities to learn a multimodal fusion and classification network. With FedCMD, the learned multimodal network can support both unimodal inputs and multimodal inputs with any missing modalities. Extensive experiments demonstrate the efficacy of FedCMD compared to state-of-the-art baselines.
Metasurfaces are a transformative class of artificial electromagnetic materials with significant potential in communication, sensing, and security. However, existing design methods require detailed physical properties as input and lack flexibility under complex constraints, limiting their applicability. In this paper, we propose MetaGen, a general, efficient, and user-friendly generation framework for intelligent metasurface elements. MetaGen employs a fine-tuned large language model to translate natural language instructions into formatted physical properties and integrates a diffusion-based model to generate metasurface elements. Furthermore, we develop a metasurface element dataset with granular frequency sampling and extended geometric parameters to enable MetaGen to learn the complex relationships between metasurface element geometries and electromagnetic responses. Experimental results demonstrate that MetaGen effectively satisfies complex constraints, achieving electromagnetic responses closely aligned with target specifications.