Robots usually appear in the form of robot swarms when they execute complex tasks. Given cost constraints that lead to insufficient computational capabilities in robots, they can offload computation-intensive tasks to mobile edge computing (MEC) servers via satellite-ground network, substantially improving their efficiency in real-world task execution. However, robots' mobility causes their positions relative to MEC servers to change continuously, resulting in dynamically evolving computing offloading environments. To address this issue, this paper aims to ensure the efficiency of multiple robots jointly completing a complex real-world task and at the same time minimize the average completion time of the robot's computing task. We consider a scenario in which multiple robots move to multiple target locations, where the destination of each robot is not assigned in advance. The robots need to move to one of the target locations individually, and at the same time, their own computing tasks can be offloaded to base stations or satellites around them. We conceptualize this problem as a multi-objective optimization framework, which is decoupled into two components: movement control and computation offloading, and propose the DAOMAN algorithm. Simulation results show that the proposed method minimizes the time for robots to reach their respective targets while also reducing the completion time of their computation tasks.
Cooperative multi-agent reinforcement learning (CMARL) policies are vulnerable to action hijacking even when only a few timesteps are compromised. Recent adversarial attacks and adversarial training methods have been explored, but under an explicit attack budget, existing attacks often fail to accurately expose critical coordination weaknesses and incur substantial training cost. We propose Budgeted Hierarchical Efficient Attack (BHEA), a budgeted hierarchical adversarial attack that separates decisions on when and which agents to hijack from action replacement, enabling more precise vulnerability discovery under limited attack opportunities. We further show that training cooperative policies against BHEA substantially improves robustness to limited-step action hijacking while reducing training overhead. Experiments on the StarCraft Multi-Agent Challenge (SMAC) demonstrate stronger attacks under the same attack budget and improved robustness. Code is available at https://anonymous.4open.science/r/BHEA-068D.
Coverless image steganography (CIS) aims to map secret images into container images without modifying the original images for concealment purposes. However, most current CIS methods rely heavily on deep learning models, which require extensive training datasets and demonstrate limited robustness to variations in image styles. Particularly when applied to images with substantial stylistic variations, these methods often produce unsatisfactory steganographic results, leading to significant degradation in the quality of steganographic image (stego-image). Furthermore, existing diffusion model-based CIS approaches can only achieve effective concealment between images with similar styles, thereby limiting the diversity of application scenarios. To address these limitations, we propose a training-free CIS method based on the diffusion model (DStyleStego), which does not rely on the traditional training, and can effectively handle different styles of images, guaranteeing the image quality and the security of steganographic information. Specifically, we design a two-stage latent transformation method to improve the security and flexibility of image steganography. In addition, we introduce a detail compensation function to recover detail information lost during the diffusion process to improve the quality and fidelity of the generated images. Extensive experimental results demonstrate that DStyleStego achieves efficient and stable steganography across diverse image datasets (Stego260 and UniStega) while exhibiting significant advantages in terms of image quality preservation.
This paper investigates a dynamic beam hopping and power allocation problem in rate-splitting multiple access (RSMA) assisted multi-beam satellite system, in which beam services are allocated for terrestrial cells with non-uniform traffic demands and spectrum resource is full-frequency reused among beams for serving multiple ground terminals through RSMA technique. The objective is to jointly optimize beam hopping and power allocation strategy for maximizing the long-term sum rate while simultaneously satisfying delay fairness demand. However, the interference caused by full frequency reuse among beams limits the efficacy of RSMA, and the frequent beam patterns changes make instantaneous channel state information (CSI) detection a challenge. To address these challenges, this paper designs a novel deep reinforcement learning (DRL) based resource allocation framework for discrete-continuous action spaces. To alleviate explosion of action space, each beam is treated as a DRL agent, and a multi-agent deep reinforcement learning method is proposed to learn the beam scheduling policy. In addition, an action mapping mechanism is employed to further manage the DRL power allocation for satisfying the power input constraint arising from limited satellite power budget. Simulation results demonstrate that the proposed method can achieve the highest sum rate of the system while ensuring fairness among non-uniform traffic load cells compared with several baseline methods.
Image steganography aims to conceal secret content within another image while avoiding detection by steganalysis techniques. However, existing learning-based methods typically rely on fixed cover datasets or artificially designed prompts, which often result in insufficient diversity in the generated cover textures. This mismatch between cover distribution and embedding strategies significantly increases detectability under steganalysis attacks. To address this challenge, we propose E2E-Steg, an end-to-end steganography framework that integrates Large Language Model (LLM)-driven prompt optimization with the specialized image hiding and recovery networks. Specifically, we design a prompt optimization module where the LLM acts as an expert quality discriminator. By analyzing visual features such as entropy and texture complexity, the LLM can iteratively refine prompts to guide Stable Diffusion in generating cover images with rich textures. Furthermore, we introduce an image hiding and recovery network based on a multi-scale feature fusion mechanism. This network adaptively embeds secret images into the complex texture regions of the optimized covers, ensuring high-fidelity information recovery while maintaining visual imperceptibility. Extensive experiments demonstrate that E2E-Steg achieves state-of-the-art performance. On average, E2E-Steg achieves a significant PSNR improvement of over 6.08 dB for steganographic images (stego-images) across the DIV2K, COCO, and ImageNet datasets, outperforming existing steganography methods. The source code is available at: https://github.com/421893969/E2E-Steg.
Sparse adversarial attacks perturb only a few pixels to achieve an attack, making them harder to detect and more dangerous. Recently, generative sparse attacks decouple the generation of sparse adversarial examples (AEs) into dense perturbations and sparse masks. By modeling the data distribution from clean examples to sparse AEs, generative sparse attacks mitigate the poor transferability that arises from over-reliance on gradients. These methods put effort into deriving optimal sparse masks on the generated perturbation. However, the quality of perturbation generation has always been overlooked, which limits the transferability of sparse AEs. To explore the influence of perturbation quality, we conduct empirical analyses of sparse gradient-based perturbations. The results show that directly applying sparsity to gradient-based perturbations disrupts their holistic adversarial information, leading to degraded attack performance. Therefore, it is critical to extract key adversarial knowledge from gradient-based perturbations while preserving their overall integrity to guide sparse adversarial attacks. Motivated by this observation, we propose to extract essential adversarial information from gradient-based AEs to guide the generator to produce higher-quality dense perturbations and stronger transferable sparse AEs. Specifically, we introduce the Gradient Perturbation Guidance (GPG) sparse adversarial attack, which integrates gradient adversarial feature guidance and gradient perturbation guidance regularization. The former guides the generator to capture gradient-based adversarial features during encoding, while the latter refines adversarial knowledge from gradient-based perturbations during decoding. Extensive experiments on ImageNet-1K show that our GPG significantly boosts transferability compared to state-of-the-art methods under consistent sparsity constraints. Our code is available at https://github.com/bookman233/GPG
Traffic crashes potentially lead to significant fatalities and injuries, in addition to causing major disruptions to traffic efficiency and substantial economic burdens. Accurate post-crash traffic prediction provides essential information for evaluating traffic perturbations and developing effective solutions for mitigating crash impacts. Previous studies have established a series of deep learning-based models to predict post-crash traffic conditions. Most of them were developed by learning statistical correlations among variables, however, time-varying confounder biases and the heterogeneous effects of crashes cannot be modeled by the correlation learning strategy. Especially for the post-crash prediction with hypothetical crash occurrence, the misunderstanding of the causal relationships among influential factors greatly hurts the effectiveness of the counterfactual prediction functions. By considering the aforementioned issues, the current study proposes the Marginal Structural Causal Transformer (MSCT), a novel causal deep learning model designed for counterfactual post-crash traffic prediction. To accommodate time-varying confounding biases, MSCT incorporates a structure inspired by Marginal Structural Models and introduces a balanced loss function to facilitate learning of invariant causal features. The proposed model is treatment-aware, which enables comprehending and predicting traffic speed under hypothetical crash intervention strategies. To validate the proposed MSCT, a synthetic data generation procedure is introduced to emulate the causal mechanisms between crashes and traffic dynamics. Extensive experiments are conducted on both synthetic and real-world datasets for validating the proposed MSCT, including the prediction accuracy comparison under different degrees of time-varying confounding, varying crash ratios, consecutive crash scenarios, and ablation studies. The experimental results indicate that the proposed MSCT outperforms, for both factual and counterfactual predictions, not only non-causal deep learning models, e.g., Transformer and LLM, but also the cutting-edge causal deep learning methods. Notably, MSCT maintains robust performance even under complex scenarios such as multi-crash conditions. The research findings demonstrate the effectiveness of MSCT for assessing hypothetical crash scenarios and supporting proactive traffic management. Future research directions include spatial causal model development and exploring broader applications across various transportation systems.
Cooperative adaptive cruise control (CACC) leverages vehicle-to-vehicle communication to achieve tighter distance control and better formation maintenance, improving efficiency and safety. However, cross-task robustness and multi-objective decision-making remain challenging. This paper introduces a Multi-Agent Reinforcement Learning (MARL) framework tailored for multi-objective CACC in cross-task environments. The proposed approach employs a synergistic cognitive fusion and dynamic weight adaptation strategy to optimize the allocation of multiple driving objectives. By dynamically adjusting the relative importance of objectives such as safety, efficiency, and comfort, the framework adapts to varying driving scenarios. Simulation experiments demonstrate the method’s effectiveness in enhancing overall system performance and driving safety. Furthermore, comparisons with real-world driving data underscore the approach’s potential for practical application.
Federated learning (FL) collaboratively trains a global model across multiple clients without sharing local data, effectively utilizing data while preserving privacy. However, real-world data are often non-independently and identically distributed ( non-IID) and heterogeneous across clients, making standard FL less suitable for practical deployment. To tackle these challenges, we propose a novel framework named GBG which includes grouping, block training and global distillation. First, it introduces a novel grouping mechanism that clusters heterogeneous clients into different groups. Clients in these groups perform different training tasks. Second, it employs Symmetric Balanced Incomplete Block Design (SBIBD) to construct intra-group blocks and establishes a new training paradigm called block training. Finally, by incorporating mutual learning, the GBG framework enables the effective development of both personalized block-level models and a global model. Moreover, we demonstrate the convergence of block training in combination with existing works. Additional experiments further demonstrate that the GBG framework achieves favorable results in both testing accuracy and training error.
While deep learning has advanced underwater image enhancement (UIE), progress is fundamentally hampered by the scarcity of paired training data. Semi-supervised methods, which leverage abundant unlabeled images, offer a promising solution but often employ naive, static weighting for pseudo-labels. This overlooks a key principle: the most informative training samples are those that represent a successful restoration from severe degradation, thus possessing a greater transformation difficulty. To address this, we propose a novel adaptive training paradigm, the Self-Distilled Degradation-Aware Network (SDDA-Net), governed by an incentive-driven weighting mechanism. We introduce a scoring function that evaluates each pseudo-label by considering both its perceptual quality and a term that is inversely proportional to its structural similarity to the input. This score, which quantifies the difficulty and success of the enhancement, is then used to dynamically modulate the sample's influence on the unsupervised loss. By allocating greater learning weights to successfully restored, complex examples, our framework forces the model to focus on mastering challenging degradations. Extensive experiments on multiple benchmark datasets demonstrate that this adaptive approach significantly outperforms current state-of-the-art methods, validating that the strategic, sample-adaptive modulation of the learning process is critical for success in data-scarce restoration tasks.
Dynamic pricing for electric vehicle (EV) charging stations is critical in power market management. However, existing pricing schemes fail to consider the interaction process among charging stations and the diverse charging preferences of users. Consequently, these schemes struggle to capture users’ sensitivity to pricing, resulting in a failure to guarantee users’ welfare. To fill this gap, a dynamic pricing scheme for multiple charging stations with a diversion mechanism (MCSDP-DM) is proposed in this work. The proposed scheme aims to optimize pricing strategies for EV charging stations, considering user preferences to maximize their welfare, while maximizing the long-term revenue of the charging network. Specifically, we consider charging stations as agents and establish a collaborative multi-agent reinforcement learning (MARL) framework by comprehensively modeling the inter-station interaction process and user-station interaction process. Furthermore, we design a long short-term memory (LSTM)-based congestion level prediction model and a user diversion mechanism, guiding unserved users to charging stations with the lowest predicted congestion levels to ensure efficient utilization of resources among charging stations. Simulation experiments conducted on real-world arrival data validate the effectiveness of the proposed scheme. MCSDP-DM achieves an average daily team revenue of $1664.40, outperforming the TOU scheme by 34.24%, the optimization-based scheme by 23.30%, and the two RL-based schemes by 1.18% and 11.66%, respectively. It also achieves the highest average user utility of 0.8804 and service rate of 0.9940, indicating that the proposed scheme improves both charging network revenue and user welfare.
This study aims to solve the energy management problem of multi-microgrids (MMGs) through a personalized federated reinforcement learning (PFRL) approach to minimizing the operational cost of microgrid while ensuring a good user experience. In this study, we introduce the Partially Observable Markov Decision Process (POMDP) framework to describe the energy management problem of MMGs systems, instead of using sensitive information directly, non-sensitive historical data including photovoltaic power output, load power, outdoor temperature, etc. In order to solve the data silos problem among the microgrids and to promote knowledge sharing, PFRL combines locally trained and personalized models. PFRL allows each microgrid to develop personalized reinforcement learning models based on its own environmental characteristics, and global parameter aggregation through a federated learning (FL) mechanism to overcome the problem of locally optimal solutions. Simulation results obtained under heterogeneous microgrid environments indicate the advantage of PFRL and the comparisons with other algorithms are carried out to verify the effectiveness of PFRL.
With the rapidly evolving landscape of the Internet of Things (IoT) and the ubiquitous utilization of mobile devices, unmanned aerial vehicles (UAVs) play a crucial role in data collection. To address this challenging problem of limited onboard energy and flight time, we propose a novel approach leveraging reconfigurable intelligent surfaces (RIS) to enhance energy efficiency and optimize uplink transmission rates. Our system aims to determine optimal UAV paths in complex 3D urban environments, enabling rapid information collection while maximizing energy efficiency. We introduce the deep reinforcement learning (DRL) called RIS-prioritized twin delayed deep deterministic policy gradient (RIS-PTD3) algorithm, which maximizes communication rates while effectively exploring surroundings. Experimental validation across various user distributions and environmental conditions demonstrates the robustness and stability of our approach.
Dear Editor, This letter is concerned with the problem of stable high-quality sig-nal transmission of unmanned aerial vehicle(UAV)-assisted multi-ple-input multiple-output(MIMO)communication system.The parti-cle swarm optimization(PSO)algorithm is used to achieve optimal beamforming and power allocation for this system.
The security of wireless cyber-physical systems (CPSs) faces significant challenges from channel congestion attacks, which disrupt data transmission and threaten system stability. Existing research on attack energy allocation primarily focuses on single-sensor, single-channel scenarios or assumes discrete energy distribution, limiting the development of optimal strategies. To address this, we model the remote state estimation problem in wireless CPSs as a Markov decision process (MDP) and propose an action mapping mechanism (AM). Based on these, we propose two reinforcement learning algorithms to address the continuous attack energy allocation problem in multisensor, multichannel environments. The first algorithm, genetic algorithm-driven model-based action mapping algorithm (GA-MB-AM), is a model-based reinforcement learning approach that uses a genetic algorithm for network initialization. The second algorithm is soft actor-critic-based action mapping algorithm (SAC-AM), which improves performance in dynamic environments, such as systems with energy harvesting. Simulations demonstrate the superiority of our approach over traditional methods.
In this paper, we propose a quality of experience (QoE) driven multi-UAV networks deployment strategy based on deep reinforcement learning (DRL) and secure consensus protocols. UAVs can adjust their moving direction and distance in three-dimensional (3D) space to serve users who move randomly in the target area, which ensures that all the users are able to get the communication service while making the total communication service quality maximized. Additionally, the proposed method accounts for UAV speed variations and endogenous safey. Through the implementation of multivariate constraints and secure consensus protocols, all normal nodes maintain a robust topology, effectively resisting interference from malicious nodes, ensuring speed remains within a safe range, and ultimately achieving state consensus. Simulation results show that the method can quickly converge to the optimal state with fewer iterations, outperform the traditional method in terms of QoE performance, and ensures the flight safety of UAV formations in the case of attacks by malicious nodes.
Considering the satellites' wide coverage and independence from geographical constraints, satellite edge computing (SEC) has demonstrated broad application prospects. In this paper, a joint offloading and caching framework is proposed, addressing issues, such as redundant data transmission, heterogeneous resources, and the high-speed movement of satellites in SEC. In the framework, latency and system energy consumption are reduced by dynamically caching reusable data and formulating an appropriate offloading strategy. Considering the complexity of the problem, we propose a gated recurrent unit (GRU)-soft actor-critic (SAC) algorithm that trains multiple distinct deep neural networks (DNNs) to output offloading and caching actions. It combines maximum entropy with policy gradient to enhance strategy stability and avoid local optima. Furthermore, the algorithm predicts the future request probabilities of tasks as part of the state space, aiding in strategy training and adapting to dynamically changing environments. Simulation results substantiate the effectiveness of our proposed method in reducing latency and system energy consumption. In the simulation environment comprising five users and two satellites, the latency decreased from nearly 132 to around 129, and the energy consumption decreased from nearly 22 to around 18.
In this article, a distributed Takagi--Sugeno (T-S) fuzzy security control strategy based on memory events is proposed for a class of multi-input-multi-output cyber-physical systems that subjected to replay attacks. This strategy can also address common uncertainties in practical systems, including communication latency, unmodeled dynamics, and external unknown disturbances. It is worth noting that the replay attack model in this article is highly generalized, and the attack location, target, and frequency are all uncertain. Therefore, a suitable distributed memory event-based strategy is designed. It can dynamically adjust the usage of historical data to optimize the release of sampled data, thereby determining when to update the control laws of each subsystem, ensuring system performance while greatly saving communication resources. In addition, the stability of the system is demonstrated by establishing a suitable Lyapunov-Krasovskii functional, ensuring the elimination of the Zeno phenomenon. Finally, the effectiveness of the proposed method is validated through simulations conducted on two commonly encountered practical systems.
Accurate vehicle trajectory prediction process depends on seamless data sharing within the Internet of Vehicles. However, such interconnected data exchange introduces significant security risks. Specifically, network attacks can compromise data integrity, thereby degrading prediction accuracy. Concurrently, the need to protect sensitive vehicle data, such as driving trajectories and user account information, results in data silos that hinder the free flow of information essential for effective prediction. Existing studies have largely addressed either privacy preservation or attack mitigation in isolation, lacking a unified solution that simultaneously tackles both challenges. To address this gap, we propose Fed-SecTP, an integrated dual-module secure federated learning framework. The first module employs a Temporal Convolutional Network (TCN) with multi-head attention to detect and filter network attacks in real-time. The second module combines TCN with a Bidirectional Long Short-Term Memory (Bi-LSTM) network for trajectory prediction and leverages FedProx for federated learning, thereby enabling privacy-preserving model training without sharing raw data. Experimental results demonstrate that Fed-SecTP achieves high prediction accuracy and robustness even when up to 50% of the data is compromised by attacks, while ensuring secure data processing. This framework offers a reliable and comprehensive solution for autonomous vehicle trajectory prediction.