Achieving high-precision billet length control in continuous casting and rolling (CCR) systems remains a challenging task. This is mainly due to strong inter-stage coupling and cumulative disturbance propagation throughout the production process. Multi-source disturbances continuously accumulate and interact under dynamic operating conditions, while production rhythms exhibit long-term temporal dependency and non-stationary behavior. Thermal expansion, rolling force fluctuations, oxidation loss, steel accumulation, and equipment interventions propagate across stages and amplify over time, resulting in hysteresis, nonlinear error evolution, and cross-stage deviation amplification. Conventional rule-based or static optimization strategies fail to characterize global process coupling and temporal correlations. Meanwhile, standard deep reinforcement learning (DRL) methods often experience policy degradation, poor robustness, and unstable convergence when confronted with rhythm mutations, process switching, or truncated sequences. To address these challenges, this paper proposes an intelligent billet length control framework based on Long Short-Term Memory enhanced Proximal Policy Optimization (LSTM-PPO). A full-process physical model is first established under mass propagation constraints, integrating deformation dynamics, thermal effects, and comprehensive loss propagation mechanisms. The control problem is formulated as a Markov Decision Process with structured state representation and industrially feasible discrete actions. By embedding LSTM into the PPO architecture, long-term dependencies and irregular production rhythms are effectively captured. Bayesian optimization (BO) is further adopted for global hyperparameter tuning to enhance convergence efficiency and policy robustness. Extensive experiments on a high-fidelity CCR simulation platform demonstrate stable convergence under rhythm mutations and strong noise disturbances, significantly outperforming traditional controllers and standard DRL algorithms in billet accuracy, residual control performance, and overall stability.
The rapid proliferation of large-scale artificial intelligence (AI) applications has led to unprecedented demand for distributed, efficient, and highly scalable computing infrastructures. Traditional communication networks have evolved to support cloud-edge-end collaborative computing. However, they remain fundamentally communication-centric and lack native mechanisms for coordinating widespread computing resources. Computing power networks (CPNs) address these limitations by adopting a computation-first design that treats computing power as a unified, schedulable network resource. CPNs enable on-demand computing resource scheduling. This allows heterogeneous tasks to be dynamically matched with the most suitable cloud, edge, or terminal nodes based on real-time workload conditions, device capabilities, and network states. Moreover, CPNs achieve deep computing-network collaboration through joint optimization of routing, task placement, data exchange, and resource orchestration, providing performance gains beyond traditional decoupled architectures. Their unified and programmable control framework further supports seamless integration across heterogeneous devices, enabling flexible resource virtualization and efficient management of diverse network environments. Motivated by these unique characteristics, this survey provides a comprehensive overview of CPN technologies, their architectural foundations, resource allocation mechanisms, and their emerging role in supporting large-scale distributed AI training and inference. We highlight key challenges, discuss open research problems, and provide insights into future directions for advancing computation-centric networked intelligence.
Many real-world problems involve hierarchical multi-objective optimization over coupled sequential decisions. A common strategy is to combine heuristic planning with deep reinforcement learning (DRL). While existing hybrid methods often show strong empirical performance, the coupled closed-loop learning dynamics between upper-tier planning and lower-tier control are typically analyzed only empirically, and explicit convergence guarantees for the integrated scheme remain limited. To address this gap, we propose a heuristic-supervised-DRL (HSD) framework that tightly couples (i) a heuristic planner for upper-tier decision-making, (ii) a DRL agent for lower-tier execution, and (iii) an online supervised predictor that serves as an adaptive bridge between planning and execution. The key novelty of HSD lies in this closed-loop architecture and its accompanying theoretical treatment. By formulating the coupled updates as a two-timescale stochastic approximation process, we show that, under standard conditions, the supervised predictor tracks a quasi-stationary regression target and the overall joint process converges almost surely to an asymptotically stable equilibrium. We further analyze robustness under approximate planning errors. As a case study, we instantiate HSD in a multi-UAV-assisted mobile edge computing system. Experimental results show that the proposed framework consistently outperforms representative baselines in key performance metrics, demonstrating both its practical effectiveness and its value as a principled framework for hierarchical optimization in dynamic environments.
Unmanned aerial vehicles (UAVs) are playing an increasingly pivotal role in modern communication networks,offering flexibility and enhanced coverage for a variety of applica-tions. However, UAV networks pose significant challenges due to their dynamic and distributed nature, particularly when dealing with tasks such as power allocation, channel assignment, caching,and task offloading. Traditional optimization techniques often struggle to handle the complexity and unpredictability of these environments, leading to suboptimal performance. This survey provides a comprehensive examination of how deep reinforcement learning (DRL) can be applied to solve these mathematical optimization problems in UAV communications and networking.Rather than simply introducing DRL methods, the focus is on demonstrating how these methods can be utilized to solve complex mathematical models of the underlying problems. We begin by reviewing the fundamental concepts of DRL, including value-based, policy-based, and actor-critic approaches. Then,we illustrate how DRL algorithms are applied to specific UAV network tasks by discussing from problem formulations to DRL implementation. By framing UAV communication challenges as optimization problems, this survey emphasizes the practical value of DRL in dynamic and uncertain environments. We also explore the strengths of DRL in handling large-scale network scenarios and the ability to continuously adapt to changes in the environment. In addition, future research directions are outlined, highlighting the potential for DRL to further enhance UAV communications and expand its applicability to more complex,multi-agent settings.
Low Earth Orbit (LEO) satellite communication can complement to terrestrial networks by providing extended Internet access to remote regions, especially those in harsh environments. However, some operational satellites may work as Eavesdroppers (Eves), referred to as ESats, intercepting confidential signals during communication between the Ground Station (GS) and the Serving Satellite (SSat). Besides, Eavesdropping Unmanned Aerial Vehicle (EUAV) also poses a potential threat to satellite-ground communication along the extended propagation paths. In view of this, this paper deploys an Intelligent Reflecting Surface (IRS) near the GS within the satellite system, where several ESats and EUAV coexist and the satellite follows a predictable trajectory. We discuss the optimal transmit beamforming vector and IRS configuration to maximize the Ergodic Secrecy Rate (ESR) and minimize the Secrecy Outage Probability (SOP). Besides, to mitigate the multipath effect induced by the satellite mobility, the Doppler spread is taken into consideration, which is to be minimized. A Genetic Algorithm (GA) is introduced to solve this intractable non-convex multi-objective problem. To evaluate the efficiency of IRS deployment and IRS phase shift optimization, we compare our proposed scheme against benchmark strategies, including systems without IRS, with random IRS, and with Artificial Noise (AN). Simulation results demonstrate that the proposed system achieves enhanced security and communication performance and outperforms baselines.
Satellite edge computing offers promising solutions for extending the coverage of terrestrial networks, particularly in remote and harsh environments. By offloading ground data to satellites for processing, this paradigm enables real-time data handling in non-terrestrial networks (NTNs). However, due to limited onboard resources and dynamic task requirements from Internet of Things (IoT) devices, online energy management becomes a critical challenge that hinders the scalability and deployment of satellite edge computing systems. In this paper, we propose an online energy management framework based on satellite-edge collaboration. A queue-based collaboration scheme is designed to coordinate satellites in handling tasks with random arrival patterns. Building upon this scheme, we formulate a joint optimization problem that integrates transmission resource allocation, computational resource assignment, power control, routing strategies, and collaboration policies, aiming to minimize the energy consumption of the satellite network. Given the dynamic, complex, and distributed nature of the problem, we present a Lyapunov-based multi-agent deep reinforcement learning (MADRL) algorithm. Specifically, we first transform the long-term stochastic optimization problem into a sequence of deterministic subproblems using the Lyapunov theory. Sub-sequently, each subproblem is decomposed into a resource allocation subproblem and an edge collaboration subproblem. Correspondingly, we devise a MADRL-based algorithm for edge collaboration and a Lagrangian multiplier iteration (LMI)-based algorithm for resource allocation. Finally, the overall problem is iteratively optimized using the block coordinate descent (BCD) framework. Simulation results demonstrate that the proposed approach achieves significant performance improvements over both baseline methods and several state-of-the-art pure deep reinforcement learning (DRL) approaches.
In distribution networks, incipient faults often manifest as faint and transient electrical disturbances before fully developing. Incipient fault detection is challenging due to the weak and nonstationary characteristics of fault signatures. Moreover, fault feeder identification is more difficult, as residuals across feeders tend to appear highly similar. To address these challenges, we present FD-Mamba, a Mamba-based neural state-space model that integrates control-theoretic principles with signal-processing techniques. Specifically, we propose a Kalman-inspired neural correction (KNC) mechanism that performs residual-driven state updates with learnable gain factors. In addition, we introduce a frequency-momentum updating (FMU) mechanism that stabilizes frequency tracking under nonstationary perturbations. Experimental results on two datasets show that FD-Mamba outperforms existing methods. It achieves a root-mean-square error (RMSE) of 0.427 and fault feeder detection accuracy of 98.1% on real-world field dataset.
Unmanned aerial vehicles serving as aerial base stations can rapidly restore connectivity after disasters, yet abrupt changes in user mobility and traffic demands shift the quality of service trade-offs and induce strong non-stationarity. Deep reinforcement learning policies suffer from plasticity loss under such shifts, as representation collapse and neuron dormancy impair adaptation. We propose plasticity enhanced multi-agent mixture of experts (PE-MAMoE), a centralized training with decentralized execution framework built on multi-agent proximal policy optimization. PE-MAMoE equips each UAV with a sparsely gated mixture of experts actor whose router selects a single specialist per step. A non-parametric Phase Controller injects brief, expert-only stochastic perturbations after phase switches, resets the action log-standard-deviation, anneals entropy and learning rate, and schedules the router temperature, all to re-plasticize the policy without destabilizing safe behaviors. We derive a dynamic regret bound showing the tracking error scales with both environment variation and cumulative noise energy. In a phase-driven simulator with mobile users and 3GPP-style channels, PE-MAMoE improves normalized interquartile mean return by 26.3% over the best baseline, increases served-user capacity by 12.8%, and reduces collisions by approximately 75%. Diagnostics confirm persistently higher expert feature rank and periodic dormant-neuron recovery at regime switches.
Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state unexamined. We show that sustained non-stationarity damages this internal state directly: as objectives shift, neurons progressively fall dormant and the shared policy loses the capacity to learn. The obvious remedy, resetting dormant neurons, is unsafe under shared-parameter multi-agent training: many neurons that appear inactive are still receiving strong training gradients, and whether a neuron appears dormant depends on which agent's observations it processes. PRIME (Plasticity Recovery In Multi-agent Environments) therefore verifies both directions before intervening. Extending the bidirectional Silent Neuron framework to cooperative multi-agent reinforcement learning, it aggregates activation and gradient statistics over the full team batch, reads the backward signal from the gradient the training loss has already deposited , not from a hand-crafted proxy, and reinitializes only neurons that are simultaneously activation-dormant and gradient-silent. Useful representations are preserved while learning capacity is restored. On a phase-switching UAV emergency communication simulator, PRIME improves interquartile mean return by 24.9% over MAPPO and holds dormant neuron fractions at 10–20% versus 40–45%; ablations attribute the gains to the gradient signal and team-level aggregation rather than to the specific reset operator. A dynamic regret bound shows that the perturbation cost scales with the small silent-subspace dimension rather than the full parameter count.
In RSU-enhanced Internet of Vehicle (IoV) networks, tasks can be offloaded from roadside units (RSUs) to vehicles in order to improve network resource utilization. To achieve this, multi-modal data fusion, such as LiDAR, radar and cameras, is introduced to obtain both environmental and vehicular information. In practical scenarios, multi-modal data differ significantly in spatial resolution, temporal density, and data volume, resulting in alignment challenges and inconsistent fusion quality. On the one hand, multi-modal data typically enhance environmental perception precision, which benefits task offloading. On the other hand, although increasing the volume of multi-modal data improves perception precision, the transmission of such data incurs substantial latency. To address these challenges, this paper proposes a perception-transmission-fusion-offloading (PTFO) collaborative optimization framework that jointly considers multi-modal feature alignment, size-dependent transmission scheduling, and fusion-aware offloading decisions. Unlike traditional approaches that treat the four stages independently, PTFO explicitly models their mutual influences: fusion quality depends on transmission decisions, offloading delay is affected by the size of fused features, and perception precision feeds back into scheduling mechanisms. For perception and transmission, a deep reinforcement learning (DRL) strategy is employed to approximate optimal policies that balance precision, latency, and energy efficiency under dynamic environment.
Link scheduling in wireless communication aims to minimize interference and maximize channel utilization. When accurate channel state information (CSI) is available, optimal scheduling can be achieved by modeling link interference as a conflict graph and solving the weighted maximum independent set problem (WMIS). However, the NP-hard nature of WMIS presents significant computational challenges in obtaining high-quality solutions efficiently. To tackle this challenge, we first propose a divided QUBO model, which formulates the original problem into sub-models within non-overlapping subspaces. Building upon this, we then develop a corresponding algorithm called DQAOA, which further narrows the search space within solution bounds established by classical greedy algorithms and semi-definite programming. Using a specialized mixing operator (the XY-mixer), the algorithm confines the evolution to each reduced subspace and runs in parallel. Additionally, an auxiliary Hamiltonian inspired by counter-diabatic driving is introduced into the core evolution equation to accelerate convergence. The experimental results on the IBM Qiskit platform demonstrate that DQAOA outperforms classical methods and other QAOA variants, providing an efficient and reliable solution for optimal link scheduling in wireless networks.
Federated learning enables privacy-preserving collaborative model training, but Non-independent and identically distributed (Non-IID) client data severely constraints its performance. Personalized federated learning balances global knowledge sharing and local adaptation through layer decoupling. However, static or heuristic model partitioning strategies ignore the differences in update directions across network layers among clients, leading to a significant decline in client performance fairness. To address this issue, this paper proposes the FedLUD method. This approach dynamically partitions layers by analyzing the consistency of their update directions: layers with consistent directions are designated as shared layers to aggregate global knowledge, while layers exhibiting significant divergence are designated as personalized layers to preserve local adaptation. Experiments on multiple benchmark datasets demonstrate that FedLUD effectively mitigates the performance fairness issue in personalized federated learning. Compared to state-of-the-art methods, it achieves a 15.7% improvement in worst-client accuracy and a 3.49% increase in overall accuracy, providing a novel solution for enhancing system fairness.
In continuous casting and rolling (CCR) systems, the precise billet cutting is critical for ensuring product dimensional accuracy and minimizing material waste. However, conventional rule-based cutting strategies rely on static parameters and heuristic rules, which often fail to adapt to real-time disturbances such as thermal expansion, cross-sectional fluctuations, and trimming losses. These limitations frequently lead to suboptimal control performance and excessive residual lengths. Deep reinforcement learning (DRL), with its dynamic decision-making and self-adaptive capabilities, offers a promising alternative. This study proposes an intelligent billet cutting strategy based on the proximal policy optimization (PPO) algorithm, integrating process disturbances, thermodynamic characteristics, and physical constraints into a unified control framework. Extensive evaluations conducted in a high-fidelity CCR simulation environment demonstrate that the proposed method outperforms both traditional approaches and existing RL-based methods in terms of control accuracy, generalization ability, and robustness to disturbances, highlighting its practical potential in intelligent steel manufacturing systems.
To address the challenges of localization and coverage for randomly moving users in unmanned aerial vehicle (UAV)-assisted communication networks, this paper proposes a collaborative framework that integrates particle filter-based localization with an interacting multiple model (IMM) Kalman filter for motion prediction. First, initial user positions are estimated using a particle filter algorithm. The received signal strength indicator (RSSI) is converted into distance estimates via a path loss model. User positions are then derived through a weighted average of particle states, and the associated state uncertainty is quantified using the particle covariance matrix. Second, an IMM Kalman filter is employed to model heterogeneous user mobility patterns. This filter dynamically predicts user motion by probabilistically weighting different motion models based on their likelihoods. Building upon these estimations and predictions, an Markov decision process (MDP) is formulated to optimize UAV trajectories. A composite reward function balances immediate and future coverage performance. Mulit-agent deep reinforcement learning (MADRL) is then applied to learn policies that maximize long-term coverage efficiency. Experimental results validate the effectiveness of the proposed framework, showing significant improvements in localization accuracy and UAV trajectory optimization when compared with conventional approaches.
With the rapid advancement of intelligent transportation systems, Road Side Units (RSUs) in the Internet of Vehicles (IoV) play a crucial role in sensing environments information for vehicles to make driving decisions and transportation management. By offloading computation tasks from RSUs to other nodes such as cloud servers, other RSUs, and vehicles, the network computation resources can be fully utilized, however, at the expense of increasing task delay. In multi-hop task offloading, tasks can be forwarded to vehicles outside the coverage area of the source RSU where tasks are generated. However, due to the movement of vehicles, stable transmission among nodes is not guaranteed, posing a challenge in determining the next node for task forwarding. In this paper, we ensure the connectivity among nodes by establishing a mobility model and design a mechanism for selecting forwarding vehicles. The goal is to find those vehicles that can communicate stably with the source RSU and determine the optimal communication path. We formulate task offloading as a 0-1 mathematical model with the objective of minimizing the task delay. Subsequently, we propose a solution based on asynchronous deep reinforcement learning A3C. Through extensive simulations, we validate the effectiveness of our approach.
Recent improvements in communication and artificial intelligence (AI) are accelerating the realization of roadside unit (RSU) assisted automated vehicle networks. It is expected to greatly boost the level of automated driving by introducing distributed RSUs to Everything (R2X) networks, in which RSUs not only behave as data forwarders but also actively and collectively perceive and analyze the environment to assist automated vehicles in making driving decisions. However, there are still many challenges for offloading computation-intensive and delay-sensitive tasks from RSUs to other network nodes in a distributed way, including tradeoffs among multiple objectives, substantial computation complexity for the problem, stringent autonomy requirements, and so forth. In this article, we present the characteristics and enabling technologies of a distributed Deep Reinforcement Learning (DRL) framework for R2X in the AI landscape. As case studies, we evaluate the performance of task offloading in the distributed learning framework from three perspectives of QoS, caching, and battery lifetime of RSUs, which demonstrates the effectiveness of our framework. Finally, we discuss the open challenges for future researches.
Roadside units (RSUs) with strong sensing abilities enhance the viability of the RSU-to-Everything (R2X) paradigm, offering crucial infrastructure support for Mobile Edge Computing (MEC) that enables real-time data processing and reduced delay. Since RSUs collect a large volume of data but have limited computing capability, data analysis tasks are usually offloaded to other network nodes, such as the cloud, other RSUs, or even vehicles. The multi-hop distributed collaborative task offloading scheme is expected to achieve high resource utilization efficiency and low task delay in this scenario, despite increasing energy consumption in data transmission. However, the highly dynamic nature of the R2X network topology makes it challenging for a node to independently select the next hop and collaboratively allocate tasks to neighbors in a multi-hop transmission path. Specifically, offloading decisions made by an individual node are influenced not only by its immediate neighbors but also by other nodes along the multi-hop path, referred to in this paper as the effect value. Additionally, the heterogeneity in computing resources and link delays among network nodes further increases the difficulty. To address these challenges, we first apply a Long Short-Term Memory (LSTM) model to predict and update the neighbors for each node while considering effect values, allowing them to independently adapt to environmental changes. Then, we design a two-layer Deep Reinforcement Learning (DRL) algorithm for network nodes to make decisions. The first-layer DRL algorithm is implemented by RSUs to determine task offloading modes. When an RSU decides to offload tasks to multiple vehicles for collaborative computing, the second-layer DRL algorithm is used by a vehicle to select its next hop vehicle and allocate tasks. Simulation results show that our proposed approach effectively adapts to topology changes in complex and highly dynamic network environments. Compared with existing methods, this approach significantly improves task completion ratio and reduces processing delay.
This article addresses the resource allocation problem in a multi-unmanned aerial vehicle (UAV)-assisted full-duplex wireless network, where the goal is to optimize UAV trajectory, user association, and power allocation in order to maximize the total uplink and downlink transmission rates for user devices. The proposed optimization framework accounts for the challenges introduced by imperfect self-interference cancellation and incorporates a generalized probabilistic air-ground channel model to increase its practical applicability. However, due to the mobility of UAVs and the dynamic nature of the environment, solving this problem using conventional optimization methods becomes both computationally expensive and impractical. To overcome these challenges, we model the problem as a partially observable stochastic game, where the macro base station (MBS) and UAVs act as agents, each interacting with the environment and receiving distinct observations. Given the complexity of this setting, including large state and action spaces, we adopt the Proximal Policy Optimization (PPO) algorithm. PPO is well-suited to this problem due to its efficiency in handling high-dimensional state-action spaces, its ability to ensure stable learning, and its effectiveness in managing the partial observability of the environment. These features make PPO an ideal approach for solving the resource allocation problem in UAV-assisted full-duplex networks.
Cloud-edge collaboration and edge intelligence have greatly driven the growth of the Industrial Internet of Things (IIoT). However, the jittery network delay and limited computational resources of edge servers make it difficult to meet the stringent latency requirements in IIoT, and so far there is no good solution to solve this problem. To this end, we introduce scalable neurocomputing, which provides neural networks with different utilities and computation resource requirements, to be deployed on edge servers of cloud-edge IIoT systems. We then optimize such systems by formulating data scheduling and system computational resource allocation as an infinite horizon optimization problem, considering that the data collection from end devices is an infinite long-term process. To solve this hard problem, we design a rolling prediction-optimization framework that transforms the infinite horizon problem into a truncated finite horizon optimization that maximizes the average system utility while satisfying the stringent delay constraints. We have conducted extensive simulations and built a prototype system, which verify the feasibility and performance of our proposed scheme.
With the rapid increase in cloud services and the increasing shift toward them, balancing the cloud load has become a critical research issue. The increasing demand from customers for technology services worldwide is largely due to its direct impact on performance quality. Therefore, to provide better service quality, it is essential to consider reducing response time and cost in load balancing. This paper addresses the load-balancing problem by optimizing response time and computing cost in cloud computing systems. Specifically, we first formulated the load balancing problem in cloud computing mathematically by designing an objective function that optimizes response time and cost, including constraints to ensure task assignment to a single virtual machine (VM) and resource limits are not exceeded. To solve this optimization problem, we proposed a hybrid algorithm, ACOCSA, that combines Ant Colony Optimization (ACO) and Crow Search Algorithm (CSA). We implemented the ACOCSA algorithm, evaluated the response time, cost, and load balancing metrics, and found that our algorithm performed better in terms of response time and computational cost, and showed good fairness among VMs.