Emergency response missions employing heterogeneous multi-UAVs are challenged by both the tight coupling between task assignment and motion control and the inherent difficulty in obtaining full probability distributions of stochastic disturbances. In practice, only partial moment information—typically the mean, covariance, and support set—can be estimated. To address this, the paper proposes a distributionally robust integrated “decision–control” framework. A bi-level optimization model is established: the upper level minimizes the maximum mission completion time across all platforms, while the lower level solves minimum-time optimal control problems under kinematic constraints, with the two levels coupled through task execution times. Given only the mean, covariance, and support set of disturbances, an ambiguity set is constructed, and by leveraging duality theory and semidefinite programming, the distributionally robust chance constraints on site reachability are equivalently transformed into deterministic safety margins. A two-stage trajectory planning method is further designed to decouple accumulated time estimation from robust constraint enforcement, ensuring computational tractability. Simulation results across multiple disturbance configurations show that, whereas deterministic planning yields an overall mission success rate of only about 0.7%, the proposed framework consistently achieves success rates above 99.9% and effectively balances workload among multiple UAVs. These results validate the practical benefit of the framework in providing reliable emergency response plans under limited distributional information.
Highlights What are the main findings? A hierarchical target tracking architecture is proposed to reduce task complexity while enhancing formation adaptability and scalability. The proposed method exhibits superior convergence, flexibility, and scalability for target tracking in unmanned aerial vehicle swarms, as validated through simulations. What are the implications of the main finding? This study presents a hierarchical architecture integrating distributed optimization with affine control, enabling UAV swarms to track dynamic target swarms while simultaneously evading threat zones. The proposed architecture supports large-scale UAV swarms in both the leader and follower layers, and it enables scalable addition of new nodes based on existing ones.Highlights What are the main findings? A hierarchical target tracking architecture is proposed to reduce task complexity while enhancing formation adaptability and scalability. The proposed method exhibits superior convergence, flexibility, and scalability for target tracking in unmanned aerial vehicle swarms, as validated through simulations. What are the implications of the main finding? This study presents a hierarchical architecture integrating distributed optimization with affine control, enabling UAV swarms to track dynamic target swarms while simultaneously evading threat zones. The proposed architecture supports large-scale UAV swarms in both the leader and follower layers, and it enables scalable addition of new nodes based on existing ones.Abstract Target tracking of unmanned aerial vehicle (UAV) swarms remains a significant challenge due to highly maneuverable target swarms and complex environments. To address these challenges, a hierarchical target tracking architecture is proposed, comprising a leader layer and a follower layer. This design reduces task complexity while improving formation adaptability and system scalability. In the leader layer, a distributed time-varying optimization model and a distributed protocol are developed to enable the UAV swarm to track highly maneuverable target swarms in real time. In the follower layer, a control protocol based on an affine transformation is employed to enable adaptive formation control under complex environmental constraints (e.g., threat avoidance). Moreover, the convergence performance of the proposed method is rigorously demonstrated through theoretical analysis. Finally, simulation results validate the convergence, feasibility, and scalability of the proposed method. Comparative simulations further demonstrate the superiority of the proposed method.
This paper proposes a multiple unmanned aerial vehicle (multi-UAV) formation obstacle avoidance strategy based on the dynamic vector velocity obstacle (DVVO) method, aiming to address the design challenges of the VO method in three-dimensional (3D) dynamic space applications. By considering the relative position and velocity between the UAV and obstacle, a dynamic avoidance plane (DA-Plane) is constructed, which can accurately reflect the collision risk of UAVs in 3D space. On this basis, a dynamic vector velocity obstacle space is designed, and the optimal constant critical escape velocity is selected outside this space to accomplish the obstacle avoidance task, thereby ensuring the smoothness of the flight process. Subsequently, this paper extends the DVVO method to multi-UAV formation based on fuzzy control and the virtual spring algorithm, and designs a differentiated obstacle avoidance strategy to ensure the safe operation of multi-UAV formation in dynamic environments. Simulation and real-world experiments demonstrate that the proposed method enables multi-UAV formation to effectively avoid dynamic obstacles in 3D space and reach the target in a desired formation shape. Compared to the citation algorithm, the DVVO method shows superior performance in terms of smooth motion and computational efficiency. This study effectively simplifies the complex obstacle avoidance problems of multi-UAV formation in 3D dynamic environments, demonstrating significant potential for applications.
To reveal the autonomous emergence mechanism of cooperative behavior in complex systems, this study integrates reinforcement learning with evolutionary game theory to construct a Q-learning-based multi-agent system for the spatial Prisoner’s Dilemma. The concept of spatiotemporal cell is proposed as a microscopic analysis unit to systematically explore the laws governing the emergence, growth, and extinction of cooperation, while quantifying the regulatory effects of parameters through an α -γ phase diagram. Our results show that the emergence of cooperation in the spatial Prisoner’s Dilemma system does not require external interventions as in traditional models. Instead, it relies on the autonomous exploration-belief update-strategy coordination process of agents. Specifically, only when agents within a spatiotemporal cell undergo long-term interactive learning with their neighbors and are triggered by synchronous cooperative-state pulses will they transition from the ground-state belief mode to the excited states. Macroscopically, cooperative behavior exhibits a stable two-periodic oscillation mode. The oscillation amplitude increases with the learning rate α and decreases with the discount factor γ , and a high γ (valuing future rewards) conversely inhibits cooperation. We develop a spatiotemporal cell theoretical framework. Based on this theory, four phases of the system’s collective decision-making behaviors are identified in the α -γ phase diagram: Ground state phase, Cooperative phase, Cooperation growth phase, and Cooperation extinction phase, which are highly consistent with the results of simulation experiments. This study provides a new paradigm and theoretical support for understanding the behavior emergence in ecological and social systems, as well as for the design of engineering systems such as unmanned swarms.
The emergence and stable evolution of cooperation among self-interested individuals is a central issue spanning evolutionary biology, social dynamics, and artificial intelligence. Conventional imitation-based evolutionary game models, lacking mechanisms for active exploration and experiential accumulation, often trap populations in suboptimal steady states and fail to explain the persistence of complex cooperative patterns. In this study, we construct a multi-agent reinforcement learning framework for the spatial snowdrift game and propose a spatiotemporal cell theory that systematically elucidates the mechanisms underlying cooperation driven by autonomous learning. Our results show that agents accumulating interaction experience via Q-learning achieve cooperation levels significantly surpassing classical replicator dynamics across a broad parameter range, with the system self-organizing into robust collective decision-making structures. From an experiential learning perspective, we reveal an endogenous mechanism of cooperative emergence, demonstrating that efficient cooperation can arise solely from individual exploration and local feedback, without external punishment, reputation mechanisms, or centralized control. The spatiotemporal cell theory provides a unified analytical framework that quantifies the coupling between microscopic learning trajectories and macroscopic pattern evolution. Based on this theory, we derive a contour plot of the fraction of cooperators in the α-γ parameter plane that delineates cooperative stability, breakdown, and frozen defect-line phases, and uncover two distinct evolution pathways: cooperative amplification induced by synchronous exploration and noise accumulation driven by asynchronous exploration. This work deepens the understanding of cooperative evolution and provides theoretical support for designing decentralized adaptive multi-agent systems.
The survivability of the infrastructure network, especially the combat system, is a core index to reflect the merits of the system. The cascade effect caused by the failure of a single or a few nodes is the focus of the research on the resistance. Many scholars have studied this problem and put forward some useful cascading failure models. Considering node resilience, a cascade failure model is proposed in which nodes have three states: normal, overload and failure. In addition, the conditions for the failure of connected edges are given. The model fully considers the redundancy design of the real system, and can objectively reflect the performance of the system to deal with cascade failure. We study the effect of node resilience on network cascade failure in both typical network and real network. Experimental results show that the proposed model can reduce the scale of cascade failures with higher cost utilization compared with the classical model that is widely used, especially when the network capacity is small.
Addressing prevalent challenges in current cooperative task assignment methods for cross-domain unmanned swarm, such as the disconnection between decision-making and execution processes, and the inadequate incorporation of platform kinematic constraints, this study introduces an integrated decision-control cooperative task assignment approach based on a bi-level optimization framework. The proposed framework formulates a bi-level programming model that tightly couples upper-level task assignment with lower-level optimal control. The upper-level model aims to minimize the maximum task completion time by optimizing the assignment and visitation sequences of diverse target types across heterogeneous unmanned platforms. The lower-level model, given the task sequences from the upper level, addresses a minimum-time optimal control problem based on a comprehensive nonlinear kinematic model. This approach enables precise computation of task execution times, which are subsequently fed back to the decision-making layer, thereby establishing a closed-loop optimization mechanism. To solve this complex model efficiently, the lower-level employs differential flatness transformation to eliminate trigonometric functions in the kinematic equations and discretizes the continuous-time optimal control problem into a nonlinear programming problem via the Radau pseudospectral method. For the upper-level combinatorial optimization, an improved genetic algorithm is developed, integrating hybrid encoding, dual-archive elitism preservation, adaptive crossover and mutation strategies, and periodic local search. Simulation results demonstrate that, compared with traditional Euclidean-distance-based assignment methods, the proposed approach generates kinematically feasible and smooth trajectories while thoroughly accounting for the kinematic constraints of heterogeneous platforms, thereby demonstrating its effectiveness and superiority in improving the comprehensive mission performance of cross-domain unmanned swarms.
The tracking of multiple maritime moving targets by cross-domain unmanned swarms under constrained communication conditions poses significant challenges, including target separation, incomplete and delayed information, heterogeneous system dynamics, and limited energy. To address these issues, a distributed self-organizing control approach is proposed for full mission-cycle management. The approach utilizes the high mobility and sensing accuracy of Unmanned Aerial Vehicles (UAVs) and the long endurance of Unmanned Surface Vehicles (USVs) to achieve real-time subgroup reorganization in response to target behavior.A timestamp-based update protocol and a delay-compensated prediction method are designed to mitigate the use of outdated information, ensuring consistent decision-making and accurate tracking even under communication delays. Furthermore, a cooperative control framework is developed, incorporating distributed task assignment, conflict resolution, optimal control model for target tracking, and a State-Event-Condition-Action (SECA) model for autonomous state management. This framework enables autonomous state switching of UAVs among Tracking, Returning, Energy Replenishment, and Standby states, supporting continuous and flexible operations. Simulation results demonstrate that the proposed method achieves rapid self-organization, stable high-precision tracking, and efficient energy utilization during target separation and communication delays. It is shown to outperforms traditional centralized and fixed-formation strategies in both mission continuity and overall adaptability.
Facing the increasing complexity of multi-dimensional maritime operations, cross-domain heterogeneous unmanned swarms provide an appealing paradigm for comprehensive ocean perception. However, platform disparities and strict spatiotemporal synchronization requirements challenge mission planning where task assignment and path planning are tightly coupled. To address this, we design a single-stage cooperative planning framework tailored for air-sea-underwater swarms. First, a recursive temporal deduction mechanism is introduced to accurately quantify the cooperative waiting costs induced by platform heterogeneity. Subsequently, a Q-learning enhanced Genetic Algorithm integrating chaotic initialization and Opposition-Based Learning, named COGAQ, is proposed for rapid solution. This algorithm adopts a feasible coalition set index encoding strategy to limit the search space and integrates variable neighborhood search to rectify temporal asynchrony. Comparative experiments against mainstream algorithms validate the effectiveness of the designed framework, highlighting the optimum-seeking capability of COGAQ in handling tightly coupled constraints. Furthermore, a large-scale scenario containing 55 heterogeneous platforms and 102 targets verifies the algorithm's scalability. This paper provides a precise and efficient solution for air-sea-underwater heterogeneous temporal coordination problems, presenting an appealing approach for cooperative mission planning in complex scenarios like maritime search and rescue, ocean resource monitoring, and joint anti-submarine operations.
Highlights What are the main findings? A DMPC-SQP-based cooperative tracking framework for UAV swarm is proposed. In dense threat environments, it achieves a 93.4% QP feasibility rate and reduces the mean tracking error by 25.4% compared to fixed-altitude DMPC and by 48.7% compared to LQR. An altitude-cooperative target recapture strategy is designed. After target loss due to threat occlusion, it reduces the total loss duration from 16.80 s to 9.10 s, increasing the target coverage rate to 90.90%. A decentralized formation reconfiguration strategy is proposed. Following member failure, the swarm autonomously completes leader election and formation reorganization within 10 s, while maintaining a safe inter-UAV separation. What are the implications of the main findings? This study provides an autonomous fault-tolerant cooperative tracking and obstacle avoidance method for UAV swarm to stably track moving targets in complex maritime environments featuring threat zones and platform failure risks, thereby enhancing system robustness and mission continuity. The proposed decentralized formation reconfiguration strategy and DMPC-SQP hierarchical optimization architecture are independent of a central node and offer good scalability, making them applicable to other multi-UAV cooperative tasks, such as cooperative search and encirclement.Highlights What are the main findings? A DMPC-SQP-based cooperative tracking framework for UAV swarm is proposed. In dense threat environments, it achieves a 93.4% QP feasibility rate and reduces the mean tracking error by 25.4% compared to fixed-altitude DMPC and by 48.7% compared to LQR. An altitude-cooperative target recapture strategy is designed. After target loss due to threat occlusion, it reduces the total loss duration from 16.80 s to 9.10 s, increasing the target coverage rate to 90.90%. A decentralized formation reconfiguration strategy is proposed. Following member failure, the swarm autonomously completes leader election and formation reorganization within 10 s, while maintaining a safe inter-UAV separation. What are the implications of the main findings? This study provides an autonomous fault-tolerant cooperative tracking and obstacle avoidance method for UAV swarm to stably track moving targets in complex maritime environments featuring threat zones and platform failure risks, thereby enhancing system robustness and mission continuity. The proposed decentralized formation reconfiguration strategy and DMPC-SQP hierarchical optimization architecture are independent of a central node and offer good scalability, making them applicable to other multi-UAV cooperative tasks, such as cooperative search and encirclement.Abstract To address the challenge of stable tracking of moving maritime targets by unmanned aerial vehicle(UAV) swarm in environments with threat zones and platform failure risks, this paper proposes a cooperative tracking and guidance strategy integrating Distributed Model Predictive Control (DMPC) with Sequential Quadratic Programming (SQP). A cooperative tracking model is developed incorporating UAV kinematics, environmental threats, stereo-vision positioning, and field-of-view constraints. Two original strategies are introduced within the DMPC framework: an altitude-cooperative target recapture strategy reduces target total loss duration by approximately 7 s compared to fixed-altitude baselines, while a distributed formation reconfiguration strategy restores stable tracking within 10 s after member failure and ensures safe inter-UAV separation. A multi-constraint trajectory tracking controller based on DMPC-SQP achieves real-time co-optimization of threat avoidance, formation maintenance, and tracking accuracy. Simulation results in dense threat environments demonstrate a 93.4% Quadratic Programming feasibility rate, with mean tracking error reduced by 25.4% over fixed-altitude DMPC and 48.7% over methods based on the Linear Quadratic Regulator (LQR), while maintaining robust performance under 300 ms communication delay, sensor noise, and moderate wind disturbance.
Compared with single-domain unmanned swarms, cross-domain unmanned swarms continue to face new challenges in terms of platform performance and constraints. In this paper, a joint unmanned swarm target assignment and mission trajectory planning method is proposed to meet the requirements of cross-domain unmanned swarm mission planning. Firstly, the different performances of cross-domain heterogeneous platforms and mission requirements of targets are characterised by using a collection of operational resources. Secondly, an algorithmic framework for joint target assignment and mission trajectory planning is proposed, in which the initial planning of the trajectory is performed in the target assignment phase, while the trajectory is further optimised afterwards. Next, the estimation of the distribution algorithms is combined with the genetic algorithm to solve the objective function. Finally, the algorithm is numerically simulated by specific cases. Simulation results indicate that the proposed algorithm can perform effective task assignment and trajectory planning for cross-domain unmanned swarms. Furthermore, the solution performance of the hybrid estimation of distribution algorithm (EDA)-genetic algorithm (GA) algorithm is better than that of GA and EDA.
To address the insufficiency of timeliness and accuracy in multi-UAV task allocation under dynamic environments, this paper proposes an intelligent decisionmaking optimization method based on Long Short-Term Memory (LSTM) networks. By analyzing task types and allocation constraints, the task allocation problem is transformed into an optimization problem. A four-layer model (input layer, LSTM layer, fully connected layer, output layer) is constructed, leveraging LSTM's gating mechanism to capture temporal dependencies between tasks and UAV states, thereby achieving precise matching of task priorities and UAV capabilities. Experiments show that compared to traditional RNN, the proposed method improves task allocation accuracy by 12.3% (reaching 94.0%), reduces average decision time to 0.89 seconds, and achieves a task completion rate of 96.8%, providing an effective technical approach for multi-UAV collaboration in complex dynamic scenarios.
To address high dynamics, strong uncertainty, and decision-dimensional explosion in air combat, this paper constructs a PPO-based hierarchical tactical decision-making algorithm (PHT-PPO) based on AFSIM simulation for UAV single-aircraft intelligent tactical decision-making. Firstly, a three-degree-of-freedom model enables dynamic modeling of UAVs and missiles, with a state space integrating global situation, local observation and identity coding, and a multi-dimensional discrete action space realizing joint decision-making of high-level maneuvering commands and firepower control. Secondly, a "high-level PPO-tactical decision + low-level rule-PID control" architecture is proposed: the upper layer outputs discrete maneuvering commands via an improved PPO network (including GRU temporal modeling and invalid action masking); the lower layer uses expert knowledge rules for maneuver execution, compressing action space and improving sample efficiency. The reward function combines win-loss rewards with radar illumination, attack, and boundary event rewards. Training shows rapid convergence with decreasing variance; in 1000 tests, the hierarchical strategy achieves 98.8% win rate (vs 51.1% win/29.9% draw for non-hierarchical). Ablation experiments identify attack and radar illumination rewards as key tactical drivers (boundary rewards have limited impact). Trajectory visualization confirms spontaneous formation of "concealed approach-radar evasion-flank positioning-one-shot escape" tactics. This research provides an interpretable, transferable, and scalable technical paradigm for air combat agents, laying foundations for multi-aircraft coordination, strong electromagnetic countermeasures, and online migration.
Policy training against diverse opponents remains a challenge when using Multi-Agent Reinforcement Learning (MARL) in multiple Unmanned Combat Aerial Vehicle (UCAV) air combat scenarios. In view of this, this paper proposes a novel Dominant and Non-dominant strategy sample selection (DoNot) mechanism and a Local Observation Enhanced Multi-Agent Proximal Policy Optimization (LOE-MAPPO) algorithm to train the multi-UCAV air combat policy and improve its generalization. Specifically, the LOE-MAPPO algorithm adopts a mixed state that concatenates the global state and individual agent’s local observation to enable efficient value function learning in multi-UCAV air combat. The DoNot mechanism classifies opponents into dominant or non-dominant strategy opponents, and samples from easier to more challenging opponents to form an adaptive training curriculum. Empirical results demonstrate that the proposed LOE-MAPPO algorithm outperforms baseline MARL algorithms in multi-UCAV air combat scenarios, and the DoNot mechanism leads to stronger policy generalization when facing diverse opponents. The results pave the way for the fast generation of cooperative strategies for air combat agents with MARL algorithms.
In response to the need for a coordinated attack on multiple ground-moving targets within a three-dimensional setting by unmanned aerial vehicle (UAV) swarm, this research has crafted a method for target allocation within the UAV swarm. This method is predicated on the dynamic adjustment of distance matrices. Furthermore, a time-coordinated control law for UAV swarm is proposed contingent upon the convergence of the error variable of the attack position under the terminal time conditions. First, the UAV swarm collaborative attack mission profile was analyzed according to the time sequence relationship, and the expected configuration and motion model of the UAV swarm were established. Secondly, in order to avoid conflicts, attack positions are allocated to the UAV swarm, and on this basis, a time-coordinated control law for the UAV swarm is proposed. The law passed the theoretical verification of Lyapunov function. Finally, simulation experiments are used for verification. The results show that the proposed UAV swarm collaborative attack positions allocation method and time-coordinated control law can ensure accurate arrival at relatively preset attack positions within a specified terminal time and complete multi-angle coordinated attacks on multiple ground-moving targets.
To tackle the multiconstraint unbalanced target allocation challenge for UAV swarms in distributed mission scenarios, we introduce the prescribed-time distributed consensus-based target allocation (PDC-TA) algorithm. Initially, drawing on the "capability-complexity" decomposition principle from Mosaic Warfare's system effectiveness evaluation, we reformulate the unbalanced allocation issue into a balanced allocation framework. This step lays the groundwork for a prescribed-time distributed multiconstraint target allocation model. Building on this foundation, we devise a PDC-TA protocol. This protocol integrates consensus theory with the Hungarian algorithm, effectively decoupling the distributed information negotiation from the allocation computation. This separation capitalizes on the swarm's strengths in local information interaction and parallel computation, streamlining the allocation process. Rigorous theoretical proofs affirm both the consistency and global optimality of the allocation scheme under constrained conditions. Empirical validation through balanced and unbalanced target allocation scenarios in UAV swarm operations, such as return-to-base and interception missions, underscores the algorithm's efficacy. Comparative analyses with classical distributed task allocation algorithms further highlight its superiority in runtime efficiency and solution optimality. The PDC-TA algorithm consistently delivers globally optimal and constraint-compliant target allocations within the prescribed time frame, setting it apart from traditional methods.
This paper investigates multiple unmanned aerial vehicles dynamic target allocation problem with capability-requirement matching constraints. The objective is to achieve conflict-free allocation through distributed collaboration under a dynamic environment. To address this problem, we propose a decentralized method named Dynamic Consensus-Based Group Algorithm (DCBGA). This algorithm utilizes a rule-based method known as State-Event-Condition-Action (SECA) as the decision-making framework. In addition, this algorithm consists of iterations between two phases: target selection and conflict resolution. In the target selection phase, each unmanned aerial vehicle (UAV) selects targets in a market-based method and follows the principle of "capability-requirement matching". In the conflict resolution phase, a consensus-based mechanism is designed so that each UAV communicates with its neighbors to generate a conflict-free allocation result. In the simulation section, this paper performs convergence analysis, sensitivity analysis, and optimization comparisons. The simulation results confirm both the convergence and robustness of the DCBGA. Furthermore, the comparison results demonstrate that the proposed algorithm exhibits better optimization performance than other advanced algorithms.
Cascading failures in infrastructure networks have serious impacts on network function. The limited capacity of network nodes provides a necessary condition for cascade failure. However, the network capacity cannot be infinite in the real network system. Therefore, how to reasonably allocate the limited capacity resources is of great significance. In this article, we put forward a capacity allocation strategy based on community structure against cascading failure. Experimental results indicate that the proposed method can reduce the scale of cascade failures with higher capacity utilization compared with Motter-Lai (ML) model. The advantage of our method is more obvious in scale-free network. Furthermore, the experiment shows that the cascade effect is more obvious when the vertex load is randomly varying. It is known to all that the growth of network capacity can make the network more resistant to destruction, but in this paper it is found that the contribution rate of unit capacity rises first and then decreases with the growth of network capacity cost.
This paper addresses the global optimal consensus problem with time-varying objective functions and state-coupled inequality constraints. By integrating the predefined-time consensus and prediction-correction schemes, a distributed optimization protocol is proposed for the cooperative interception scenarios involving the unmanned aerial vehicle (UAV) swarm. Specifically, the δ-exact penalty method is utilized to eliminate the coupled constraints. A prediction-correction term with the penalty function is then established to track the trajectory of the time-varying optimal solution. Additionally, an edge-based predefined-time consensus term is introduced to facilitate the state consensus among UAVs. Under the undirected and connected topology, the proposed protocol ensures the predefined-time state consensus and asymptotically minimizes the sum of individual convex objective functions. Furthermore, the protocol is extended to other scenarios involving switching topology, directed detail-balanced topology, and external disturbances. Numerical simulations illustrate that the UAV swarm can form the predesigned interception formation with a predefined-time frame and effectively track the target swarm.