Intelligent reflecting surfaces (IRSs) mounted on maneuverable aerial platforms to form aerial IRS (AIRS) relays represent a novel paradigm for large-scale downlink transmission in smart cities. However, the challenge of multivariate dynamic coupling hinders most existing studies due to high computational complexity and limited scalability. To address these issues, this paper proposes a soft-attention-actor-critic (SAAC) optimization framework that efficiently decomposes the joint optimization of multi-AIRS deployment, passive beamforming, and active beamforming at the base station into two sequential subproblems. The objective is to maximize average downlink spectral efficiency and service fairness, while minimizing deployment energy consumption. In the first stage, a conservative lower bound of spectral efficiency is formulated to guide multiple AIRSs toward near-optimal deployment positions. In the second stage, refined optimization is performed for both passive and active beamforming matrices. Furthermore, multi-head attention modules are incorporated into the critic and actor networks in each phase, enabling AIRS to adaptively attend to the observations and actions of other agents, and enhancing the ability to handle high-dimensional observation-action spaces. Extensive simulation results validate that the proposed SAAC framework consistently outperforms mainstream deep reinforcement learning baselines across diverse network conditions, highlighting its superior performance and scalability.
This paper investigates spectrum and service management in space-air-ground integrated networks (SAGIN) for smart construction site scenarios. In such environments, the deployment of temporary infrastructure and the dynamic evolution of work areas cause rapid changes in access relationships and co-channel interference, making traditional offline optimization methods unable to maintain stable performance over time. To address this, we formulate a joint optimization model that aims to enhance both system spectral efficiency and user service fairness while satisfying the minimum rate requirements of all terminals, explicitly considering the differing timescales of cross-layer resource allocation. Building on this model, we propose a multi-timescale hierarchical deep reinforcement learning (HDRL) framework in which upper-layer policies configure spectrum boundaries and lower-layer policies perform user association and power control. All three layers share a unified global reward and employ proximal policy optimization (PPO)-based policy updates to enable stable training and practical online decision-making. Simulation results demonstrate that the proposed HDRL framework achieves stable convergence and delivers superior overall performance across key metrics.
In recent years, unmanned aerial vehicles (UAVs) have become a crucial component in the development of wireless sensor networks (WSNs) due to their significant advantages, such as high flexibility, low cost, and broad applicability. In certain scenarios, the target area can be extremely large and often divided into multiple sub-regions due to geographic features and other factors. In large-scale target areas divided into multiple sub-regions, existing multi-UAV data collection methods suffer from three core limitations: 1. high path redundancy due to suboptimal inter-UAV coordination; 2. excessive energy consumption from inefficient trajectory planning; 3. unstable connectivity caused by insufficient hierarchical collaboration mechanisms. To address this challenge, this paper proposes a trajectory path plan algorithm for collaborative UAV swarm data collection across multiple sub-regions in a large-scale target area. Specifically, we introduce the enhanced dueling double DQN - genetic particle swarm optimization (ED3QN-GPSO) algorithm, which is based on the hierarchical leader-follower formation strategy. This algorithm stratifies the UAV swarms into upper and lower layers, assigning distinct tasks to each layer for efficient deployment. Simulation results demonstrate that the proposed algorithm effectively reduces both overall path redundancy and energy consumption of the UAV swarm compared to other benchmark algorithms, while simultaneously maintaining high connectivity and coverage rate.
This paper proposes a collaborative offloading paradigm in Unmanned Aerial Vehicle (UAV)-assisted Roadside Unit (RSU)-enabled Vehicular Edge Computing (VEC) scenarios. To address the challenge of long-term task queue stability, we propose a Lyapunov-guided Deep Reinforcement Learning (DRL) framework. The original delay minimization problem is transformed into per-slot decision-making subproblems. Then, the Multi-Head Attention mechanism is integrated into the Soft Actor-Critic (SAC) algorithm to learn optimal offloading strategies between RSUs and UAVs. Simulation results demonstrate that under the queue stability constraints, our proposed Lyapunov-guided Attention-based SAC (LA-SAC) algorithm achieves superior performance in key metrics compared to other DRL baseline methods.
The combination of unmanned aerial vehicles (UAVs) and reconfigurable intelligent surfaces (RISs) forms a cooperative framework that exploits the mutual advantages of both technologies. In this work, multiple RISs mounted on fixed-wing UAVs are adopted to adaptively reconfigure the transmission environments of communications. Given the practical significance of the proposed research, this study incorporates the phase-dependent amplitude response under realistic discrete phase shifts in RISs and accounts for statistical channel state information (CSI). In view of the practical designs of fixed-wing UAVs’ mobilities, the flight turning and pitch angles are considered in it to enhance the application value. Specifically, an optimization framework is designed to maximize the average data rate under statistical CSI over the UAVs’ flight duration by simultaneously optimizing BS’s transmit precoding, practical phase shifts of RISs, UAVs’ 3D trajectories, velocities, accelerations, and flight timeslot durations. To address this highly coupled and non-convex problem, an algorithm based on penalty dual decomposition (PDD) is introduced, incorporating weighted minimum mean square error (WMMSE) and successive convex approximation (SCA). The non-convex objective function is initially converted into a more manageable form by using WMMSE and analysis under statistical CSI. Then, variables are decoupled via auxiliary variables and alternately optimized in the inner iteration process, while operating the update of penalty and dual factors in the outer iteration process. Simulation results demonstrate the effectiveness of the proposed method.
The unmanned aerial vehicle (UAV)-assisted space-air-ground integrated networks (SAGIN) can provide communication and computing services for Internet of Remote Things (IoRT) in the absence of ground cellular network coverage. In this paper, we propose an edge computing architecture based on SAGIN, comprising three parts: a satellite, UAVs, and ground-based IoRT devices. The satellite is responsible for providing access to cloud computing resources. UAVs are equipped with mobile edge computing (MEC) servers. And IoRT devices generate latency-sensitive tasks but possess limited computing capabilities. Our objective is to minimize task processing delays by jointly optimizing UAV deployment, computation offloading, and time slot resource allocation to meet the increasing demands of IoRT devices. Specifically, the proposed problem is decomposed into two components. First, for the deployment of multiple UAVs, we propose a multi-agent softmax deep double deterministic policy gradient (MASD3) approach, which enables UAVs to adjust their flight trajectories based solely on observed information for adaptive deployment. Second, for the computation offloading and time slot resource allocation problems in SAGIN, we employ a numerical computation-based iterative optimization method to minimize the occupation of time slots by computation offloading. Simulation results demonstrate that our proposed solution significantly reduces overall delay compared to alternative benchmark schemes.
This paper investigates the data dissemination for ground user equipments (GUEs) in a multi-unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) network. To address the challenges posed by the high mobility of GUEs and the curse of dimensionality and the credit assignment issue in multi-agent systems, we propose a multi-actor-attention-critic (MAAC) framework that jointly optimizes UAV trajectory design and transmission power control, with the objective of maximizing the minimum achievable rate among GUEs while minimizing UAV flight energy consumption. By incorporating a multi-head attention mechanism and a counterfactual baseline, the proposed approach effectively alleviates performance degradation caused by increasing dimensionality. Simulation results demonstrate that MAAC outperforms representative baseline methods in terms of achievable rate and energy efficiency, and maintains robust convergence performance when scaled to networks with varying numbers of UAVs.
This paper investigates a cell-free massive multiple-input multiple-output system over Weichselberger joint-correlated Rician channels. To enhance uplink distributed processing for multi-antenna user equipment, a joint bilinear equalizer (BE) framework is developed. Based on the minimum-mean-squared-error (MMSE) criterion, we formulate the optimal bilinear equalizer (OBE) matrices using local (L-OBE) and global (G-OBE) channel statistics. These formulations apply to arbitrary channel estimators. Furthermore, closed-form expressions for L-OBE and G-OBE matrices are derived under MMSE channel estimation. A novel closed-form expression of the achievable spectral efficiency (SE) is also derived under the joint BE scheme. Numerical results show that the joint OBE scheme achieves significant SE gains and accommodates more spatial data streams than maximum ratio combining.
Stacked intelligent metasurfaces (SIMs) have emerged as a promising paradigm for wave-domain signal processing in next-generation wireless networks. In this paper, we investigate the joint optimization of antenna selection, SIM phase-shift design, and power allocation in SIM-assisted multiuser multiple-input single-output (MU-MISO) systems to maximize the achievable downlink sum rate. To handle the resulting mixed discrete-continuous constraints, we introduce differentiable reparameterizations and derive analytical gradients for the three coupled variable blocks under the cascaded multilayer propagation model. Building on these physics-aware gradient directions, we propose a model-driven unrolled full-gradient optimization (UFGO) network that unfolds the analytical joint-gradient update process into a trainable architecture with a fixed number of unfolded layers. Each unfolded layer performs learnable block-specific updates guided by the physics-aware analytical gradients and employs differentiable reparameterizations to preserve variable feasibility, thereby enabling efficient online inference while retaining the interpretability of model-based optimization. Simulation results show that UFGO achieves sum-rate performance comparable to the high-iteration FGJO benchmark while requiring substantially lower runtime than the high-iteration AO and FGJO methods. Comparisons with the corresponding variants without antenna selection further demonstrate the performance benefit of adaptive antenna selection. Additional evaluations under electromagnetic propagation model mismatch, imperfect channel state information, and frequency-selective fading demonstrate the performance resilience of UFGO under practical model and channel uncertainties, supporting its applicability to latency-sensitive SIM configuration.
Vehicular edge computing (VEC) enhances computational efficiency by strategically offloading tasks from vehicles to edge servers. Integrating semantic communication into vehicular networks introduces further benefits by leveraging semantic information to reduce task transmission delays. However, semantic-aware computation offloading encounters dual challenges: adaptively selecting semantic features to preserve task-critical meaning and dynamically allocating communication and semantic resources under varying network conditions. To cope with these challenges, we propose an importance-based hybrid-action multi-agent proximal policy optimization (I-HAMAPPO) algorithm for the semantic-aware vehicular computation offloading system in this paper. By assessing the importance scores of semantic features, an importance evaluation module (IEM) is designed to selectively transmit task-relevant information. A utility function, integrating task delay, energy consumption, and semantic similarity, is developed to provide a multi-dimensional performance evaluation of the system. Subsequently, the optimization problem is formulated with the objective of maximizing system utility by optimizing communication resources and semantic compression ratios. Considering the presence of mixed decision variables in the formulated problem, we employ the proposed I-HAMAPPO algorithm to optimize the continuous and discrete actions jointly. Based on real-world vehicle trajectories from the highD dataset, extensive experimental results demonstrate the convergence of I-HAMAPPO and its efficacy in maximizing system utility.
For better flexibility and greater coverage areas, Unmanned Aerial Vehicles (UAVs) have been applied in Flying Mobile Edge Computing (F-MEC) systems to offer offloading services for the User Equipment (UEs). This paper considers a disaster-affected scenario where UAVs undertake the role of MEC servers to provide computing resources for Disaster Relief Devices (DRDs). Considering the fairness of DRDs, a max-min problem is formulated to optimize the saved time by jointly designing the trajectory of the UAVs, the offloading policy and serving time under the constraint of the UAVs' energy capacity. To solve the above non-convex problem, we first model the service process as a Markov Decision Process (MDP) with the Reward Shaping (RS) technique, and then propose a Deep Reinforcement Learning (DRL) based algorithm to find the optimal solution for the MDP. Simulations show that the proposed RS-DRL algorithm is valid and effective, and has better performance than the baseline algorithms.
This paper investigates the joint design of the transmit beamforming matrix and the intelligent reflecting surface (IRS) phase shift matrix in a multi-user, multiple-input singleoutput (MU-MISO) system integrated with an aerial IRS. The proposed approach leverages the high mobility and flexibility of unmanned aerial vehicles (UAVs) to enhance the probability of a line-of-sight (LoS) link, thereby maximizing the sum spectral efficiency of the user equipments. Specifically, to adapt to the time-varying nature of real-world communication environments, we propose a twin delayed deep deterministic policy gradient (TD3) algorithm incorporating self-attention mechanisms for the joint optimization of beamforming and phase shift strategies. Furthermore, the batch normalization technique is employed to improve the algorithm's capability to process extensive state and action spaces, thereby accelerating convergence. Simulation results demonstrate that the proposed algorithm outperforms other mainstream deep reinforcement learning (DRL)-based baseline methods.
In urban environments, direct communication links between a base station (BS) and user equipment (UEs) are often obstructed by buildings. To mitigate these blockages, we integrate uncrewed aerial vehicles (UAVs) and reconfigurable intelligent surfaces (RISs) to enhance system flexibility and improve transmission efficiency. This paper investigates an RIS-assisted multi-user multiple-input single-output (MU-MISO) downlink system, where the RIS is mounted on a UAV. To maximize the system rate while minimizing the UAV's energy consumption and flight duration, we formulate a multi-objective optimization problem. To address this problem, we propose a hybrid algorithm that integrates the soft deep deterministic policy gradient (SD3) algorithm with a graph neural network (GNN) architecture, named SD3-GNN-RIS. The original problem is decomposed into two subproblems: joint active beamforming at the BS and passive beamforming at the RIS, optimized via a GNN-based approach, and three-dimensional (3D) UAV trajectory optimization, formulated as a Markov decision process and solved using the SD3 algorithm. Simulation results demonstrate the superior performance of the proposed algorithm compared to baseline methods in terms of system rate, energy efficiency, and UAV trajectory optimization.
Integrated satellite-terrestrial networks (ISTNs) enable global connectivity but face challenges in efficient resource allocation due to increasing service demands. To address Quality of Service (QoS) degradation caused by inefficient resource allocation in ISTN's heterogeneous network, we propose a network slicing (NS) resource allocation algorithm based on deep reinforcement learning (DRL). First, an ISTN system model is constructed using NS, along with an evaluation approach for slicing services. Next, a satisfaction utility function is defined to quantify the QoS of slicing services, and an optimization problem is formulated. Then, based on the Markov decision process (MDP) and dueling double deep Q-learning (D3QN) theory, an NS resource allocation algorithm is designed, comprising both training and execution phases. Simulation results demonstrate that the proposed algorithm outperforms baseline approaches in system satisfaction, bandwidth allocation, and satellite network utilization.
In the evolution of the Internet of vehicles (IoV), the increasing demand for vehicular computation tasks presents significant challenges, particularly in the context of constrained local computation resources and high processing delays. To mitigate these challenges, multi-access edge computing (MEC) offers a potential solution by leveraging edge servers for lowlatency processing. However, it also encounters issues such as sub-channel competition and workload imbalance owing to the uneven distribution of vehicle densities. This paper introduces a novel IoV architecture that incorporates multi-task and multi-roadside unit (RSU) capabilities, enabling edge-toedge collaboration for efficient task offloading among RSUs. The optimization problem is formulated with the objective of minimizing the overall task delay, which is further divided into two sub-problems: communication resource allocation and load balancing. Considering the non-deterministic polynomial (NP)- hard nature of these sub-problems, we propose a two-stage deep reinforcement learning-based communication resource allocation and load balancing (DRLCL) algorithm to address them sequentially. Based on realistic vehicle trajectories, comprehensive evaluation results demonstrate the superiority of the proposed algorithm in reducing system delay compared to existing stateof-the-art baselines, offering an effective approach for optimizing the performance of vehicular edge computing (VEC) networks.
With the rapid development of Intelligent Transportation Systems (ITS), many new applications for Intelligent Connected Vehicles (ICVs) have sprung up. In order to tackle the conflict between delay-sensitive applications and resource-constrained vehicles, computation offloading paradigm that transfers computation tasks from ICVs to edge computing nodes has received extensive attention. However, the dynamic network conditions caused by the mobility of vehicles and the unbalanced computing load of edge nodes make ITS face challenges. In this paper, we propose a heterogeneous Vehicular Edge Computing (VEC) architecture with Task Vehicles (TaVs), Service Vehicles (SeVs) and Roadside Units (RSUs), and propose a distributed algorithm, namely PG-MRL, which jointly optimizes offloading decision and resource allocation. In the first stage, the offloading decisions of TaVs are obtained through a potential game. In the second stage, a multi-agent Deep Deterministic Policy Gradient (DDPG), one of deep reinforcement learning algorithms, with centralized training and distributed execution is proposed to optimize the real-time transmission power and subchannel selection. The simulation results show that the proposed PG-MRL algorithm has significant improvements over baseline algorithms in terms of system delay.
Due to the dynamic nature of service requests and the uneven distribution of services in the Internet of Vehicles (IoV), Multi-access Edge Computing (MEC) networks with pre-installed servers are often susceptible to insufficient computing power at certain times or in certain areas. In addition, Vehicular Users (VUs) need to share their observations for centralized neural network training, resulting in additional communication overhead. In this paper, we present a hybrid MEC server architecture, where fixed RoadSide Units (RSUs) and Mobile Edge Servers (MESs) cooperate to provide computation offloading services to VUs. We propose a distributed federated learning and Deep Reinforcement Learning (DRL) based algorithm, namely Federated Dueling Double Deep Q-Network (FD3QN), with the objective of minimizing the weighted sum of service latency and energy consumption. Horizontal federated learning is incorporated into the Dueling Double Deep Q-Network (D3QN) to allocate cross-domain resources after the offload decision process. A client-server framework with federated aggregation is used to maintain the global model. The proposed FD3QN algorithm can jointly optimize power, sub-band, and computational resources. Simulation results show that the proposed algorithm outperforms baselines in terms of system cost and exhibits better robustness in uncertain IoV environments.
In recent years, semantic communication has garnered significant attention for its potential to address challenges in traditional communication systems. However, in complex communication environments, semantic communication still faces challenges, such as semantic information loss, low transmission efficiency, and poor adaptability. This article proposes a novel semantic communication with adaptive semantic reconstruction (SCASR) scheme to enhance transmission efficiency and adaptability in complex communication environments. First, a compression mechanism based on semantic importance is designed to achieve flexible and efficient semantic compression. Then, we develop an adaptive semantic reconstruction network to predict and reconstruct lost semantic information. Finally, we integrate an attention mechanism into the reconstruction network, dynamically adjusting parameter weights based on signal-to-noise ratio (SNR), semantic compression rate (SCR), and packet loss rate (PLR) to improve reconstruction quality and adaptability. To evaluate the efficiency of SCASR, we conduct extensive simulation experiments on semantic segmentation tasks using the Cityscapes dataset. Results demonstrate that SCASR outperforms existing semantic communication and traditional schemes, offering higher mean Intersection over Union (mIoU), and enhanced semantic transmission benefit (STB).
Vehicular networks face significant challenges in achieving high energy efficiency (EE) while guaranteeing diverse quality of service (QoS) requirements of users, especially under limited bandwidth and power budgets in highly dynamic and dense topologies. To address these challenges, this study formulates a joint resource optimization problem to maximize the average EE of cellular users (CUs) and vehicle-to-vehicle (V2V) users by jointly optimizing subchannel assignment, frequency reuse patterns, and power allocation while ensuring the required QoS of both users. To solve the non-convex optimization problem, we propose a semi-persistent scheduling (SPS)-based energy-efficient resource allocation scheme that integrates non-orthogonal multiple access (NOMA) with network slicing (NS). Specifically, during the frequency reservation phase of SPS periods, CUs are assigned to network slices using the proposed NS grouping strategy, and V2V users are clustered into V2V NOMA clusters using the proposed clustering optimization algorithm. Frequency reuse patterns are then determined for network slices and V2V NOMA clusters. In the subsequent data transmission phase, a centralized energy-efficient iterative power control algorithm is introduced to enhance the average CU EE, and a distributed heuristic power control method is leveraged to improve the average V2V EE. Simulation results demonstrate that the proposed scheme outperforms the baseline methods in improving EE and satisfying the required QoS of both CUs and V2V users while avoiding over-allocation of frequency resources.
Mobile crowdsensing (MCS) enables data collection by leveraging the sensing capabilities of distributed mobile devices (MDs). However, its performance is often constrained by limited coverage and connectivity of ground-based sensors. To overcome these limitations, this paper presents an efficient data collection framework in unmanned aerial vehicle (UAV)-assisted MCS systems. The proposed framework leverages the aerial mobility of UAVs to enhance the spatial coverage of MCS. By optimizing flight trajectories of UAVs, we aim to balance the age of information (AoI) and energy consumption. Specifically, this work considers the mobility of MDs and introduces a communication-enhanced trajectory optimization (CETO) algorithm to improve UAV coordination and adaptability in dynamic environments. Simulation results demonstrate that the proposed algorithm significantly outperforms other mainstream baseline methods employing deep reinforcement learning (DRL).