To address the challenges of limited autonomy, low decision-making efficiency, and poor generalization in UAV task planning for tracking mobile target under uncertain situations, this paper proposes a transfer-fusion algorithm based on the integration of three-way decision-making and self-attention mechanism into an optimized Soft Actor-Critic framework (TW-AM-SAC). Unlike research that mostly turns to deterministic reinforcement learning strategy, this one introduces a non-deterministic SAC algorithm to integrate the exploration and improvement into a single strategy to help realize the UAV's autonomous decision-making. Subsequently, to mitigate the issues of singular reward functions with fixed weights in task planning, three-way decision-making theory is incorporated to design autonomous reward functions tailored to different situations, while a self-attention mechanism is fused to assign dynamic weight distributions to the reward components. Furthermore, to enhance the adaptability of the intelligent algorithm across varying situations, a transfer learning model incorporating self- game is constructed to improve generalization performance. The simulation verification can be known that the TW-AM-SAC transfer-algorithm proposed in this paper has more effective tracking frequency and greater advantages in autonomous tracking when applied to UAV tracking of moving targets, and meanwhile converges faster with better generalization, compared with the single SAC algorithm.
In Multi-UAV path planning tasks, conventional multi-agent reinforcement learning algorithms often suffer from low sample efficiency and poor training stability. To address these issues, this paper proposes 3E-MADDPG, a novel reinforcement learning algorithm that integrates Triple Experience Modules into the MADDPG framework. Specifically, 3E-MADDPG incorporates enhanced experiences generated by Hindsight Experience Replay (HER), expert experiences derived from the Artificial Potential Field (APF) method, and counter-example experiences constructed via an adversarial mechanism. These experiences are stored in separate replay buffers and fused through a stratified sampling strategy to guide stable and efficient policy learning. Experiments in a Gym-based 2D simulation show that 3E-MADDPG outperforms baseline methods (MADDPG and 2E-MADDPG) in terms of average cumulative reward and success rate. Furthermore, the proposed method maintains robust performance in increasingly complex scenarios with a larger number of UAVs, validating its scalability and potential for real-world cooperative applications.
Under GPS-denied adversarial conditions, unmanned aerial vehicle (UAV) formations experience positioning failures caused by adversarial interference. This paper proposes a collaborative strategy enabling UAVs to maintain real-time positioning via intra-formation relative spatial measurements after GPS loss. A Partially Observable Markov Decision Process (POMDP) model is formulated for collaborative positioning and scheduling. Belief states are updated using an Extended Kalman Filter (EKF), while Q-values are estimated via a Deep Q-Network (DQN) to achieve accurate real-time collaborative positioning. Through specific-scenario simulations, the effectiveness of the proposed model is demonstrated in achieving efficient UAV formation management/scheduling and facilitating operational UAVs to locate GPS-denied UAV peers.
Traffic flow forecasting is crucial for smart city development. Existing methods primarily focus on spatial-temporal correlation learning but often overlook the distinct temporal characteristics of traffic flow. From the temporal perspective, traffic flow can be decomposed into low frequency components-representing periodic patterns inherent in the transportation system and high frequency components-reflecting short-term variations caused by external factors. However, current approaches tend to capture high frequency components while neglecting low frequency ones, which has resulted in a performance bottleneck. To address this problem, we propose a novel framework, named Low-High Frequency Network(LHFNet), for traffic flow forecasting. Our framework comprises three key components: a low frequency encoder, a high frequency encoder, and a frequency feature fusion block. The low frequency encoder employs bi-level routing attention as its core module. To enhance the stability of low frequency representations, patch embedding and patch merging operations are integrated, and the benefit of this integration is that the model complexity can be reduced while enabling longer input series. For the high frequency encoder, multilayer perceptrons are used with residual connections as the primary structure. Finally, the frequency feature fusion block dynamically integrates both frequency features through a gated selection mechanism. Extensive comparative experiments have been conducted on four real-world datasets and the results have demonstrated that LHFNet outperforms state-of-the-art models.
In order to improve the overall performance of the reinforcement learning algorithm and enable multi-agent to complete collaborative navigation in a highly dynamic and multi-constrained complex environment, this paper proposes multi-agent collaborative navigation based on the theory of deep reinforcement learning algorithm, adopting the deep deterministic strategy gradient (MADDPG) technology and introducing the experience priority extraction mechanism and the information calculation optimization method to carry out simulation verification on the multi-agent autonomous navigation problem. The paper proposes PD-MADDPG, an enhanced multi-agent navigation algorithm. Through methods such as prioritized experience replay and adaptive message dropout, the multi-agent collaborative navigation algorithm significantly enhances navigation efficiency, robustness, and scalability in dynamic environments with obstacle, advancing autonomous like drone swarms and logistics robots.
To address the curse of dimensionality inherent in applying dynamic programming to air combat maneuver decision-making tasks, this paper proposes an unmanned combat aerial vehicle (UCAV) maneuver strategy generation method based on approximate dynamic programming (ADP). Through scenario analysis of one-versus-one (1v1) within-visual-range (WVR) air combat, a Markov Decision Process (MDP) based air combat maneuver model is established. A Neural Network-based Approximate Policy Iteration (NN-API) algorithm is designed, utilizing neural network statistical approximation to replace the true value function. Furthermore, a specialized sampling scheme is devised to mitigate the sparse reward problem, increasing the proportion of high-quality samples to facilitate effective value function approximation. Simulation experiments demonstrate that the maneuver strategies generated by this method exhibit superior effectiveness and robustness compared to baseline strategies.
This paper addresses the NP-hard nature of the Job Shop Scheduling Problem (JSSP) by proposing a deep reinforcement learning approach based on the Schlably framework. The method integrates Deep Q-Network (DQN) with a normalized processing time mapping rule, employing an $\epsilon$ -greedy strategy to balance exploration and exploitation while incorporating an action masking mechanism to ensure scheduling feasibility. With the objective of minimizing total tardiness, experimental results on LA06-LA10 benchmark instances demonstrate that the proposed method significantly outperforms classical rules like SPT in reducing total tardiness, validating the superiority of deep reinforcement learning in complex scheduling scenarios.
The remarkable feature extraction capability of deep learning has garnered significant attention. However, with increasing data dimensionality, clustering, as a common data preprocessing method, can transform the high-dimensional feature space into multiple low-dimensional subspaces. A multi-granularity deep network, after being partitioned through clustering, often faces the issue of imbalanced features among different clusters. Therefore, we propose a feature balance strategy in this paper. Through three basic assumptions, minimum feature number, and compression ratio constraints, the final model can learn and represent information in a balanced manner at different levels, thus improving the overall model performance. The experiments show that the proposed strategy can balance the features effectively, enabling the model to achieve higher performance and stability.
As an essential component of intelligent transportation, traffic flow prediction is crucial in decision-making and traffic system optimization. However, traffic flow prediction is a challenging task. Traffic flow is essentially a type of time series with distinct temporal characteristics. Furthermore, due to the mutual influence between traffic nodes, there are complex dependencies in the flow sequences of each node, which contribute to the complexity of traffic flow prediction. This paper primarily analyzes the temporal dimension of traffic flow, which can be decomposed into different frequencies. Existing models often fail to handle specific frequency bands, leading to bottlenecks in prediction performance. In recent years, Transformers have proven to possess powerful sequence modelling capabilities. However, the self-attention mechanism is insensitive to high-frequency information, and the loss of such information results in decreased prediction performance. To address the aforementioned issues, we proposed decomposing the multi-head attention mechanism, with different attention heads handling information of different frequencies in traffic flow, and named it VHL-Attention. Based on VHL-Attention, we established VHLformer. Experiments on real-world datasets demonstrate that VHLformer achieves the state-of-the-art performance.
To evaluate radar performance in complex electromagnetic environments, a compact and efficient causal model is required to model such a complex, nonlinear high-stakes problem. Hence, in this paper, we propose a feature reduction causal network (FRCN). Firstly, to determine the number of hidden layer features in the FRCN, a feature extraction strategy is designed using the intrinsic dimension (ID) of raw data as key prior knowledge, thereby reducing modeling complexity and improving computational efficiency. Then, to further reveal the causal relationships between features and the final objective, a Bayesian network (BN) is constructed in the task layer, intuitively showing the coupling relationships through a directed graph and providing interpretability for decisions on high-stakes problems. Moreover, we extend the layer-wise relevance propagation to the BN in the FRCN, enabling bidirectional reasoning throughout the entire process, which is beneficial to understand the model and its behavior in a human-understandable way. In experiments, it is proved that ID plays a significance role in feature number selection. Next, we design a new interpretable evaluation indicator, called decisionspecific average edge relevance, to quantify interpretability. Compared to eight representative models, FRCN not only achieves higher accuracy but also provides stronger interpretability in terms of relevance, informativeness, and trustworthiness. A detailed analysis of a radar system enhances the understanding of coupling relationships among various factors, thereby validating the effectiveness of FRCN in feature reduction, interpretability, and trustworthiness for high-dimensional, complex, and nonlinear data.
Deep reinforcement learning (DRL) is extensively applied in autonomous unmanned aerial vehicle (UAV) control yet faces critical challenges regarding adaptability and generalization in dynamic environments. To address these limitations, this paper proposes the Meta Gated Transformer-XL Soft Actor-Critic (Meta-GSAC) algorithm. This framework integrates a Gated Transformer-XL module to capture long-term temporal dependencies from multimodal inputs and incorporates the Reptile algorithm to facilitate multi-task meta-learning. Experimental results demonstrate that Meta-GSAC significantly outperforms standard baselines. Notably, it achieves optimal policy convergence with approximately 50% fewer training epochs while effectively eliminating the high-frequency control oscillations observed in the GSAC baseline. Moreover, the proposed method exhibits superior few-shot adaptation capabilities, enabling the UAV to rapidly adapt to novel task scenarios with minimal gradient updates.
Due to the complexity of the multi-UAV rounding up maneuvering target task in continuous and complex environments, it is difficult for the UAVs to quickly and accurately capture maneuvering targets. Therefore, this paper proposes CEL-MADDPG algorithm based on Curriculum Experience Learning. It improves the efficiency of multi-UAVs rounding up maneuvering target, and has certain generalization. which is better applied to the multi-UAV roundup task in complex dynamic environments. The main contributions are the following two: By introducing the Curriculum Experience Learning, the multi-UAV rounding up task is divided into target tracking, encircling transition, and shrinking capture to learn, and designed corresponding reward function according to the task characteristics of each subtask. Which improves the learning efficiency of the model. Additionally, the CEL-MADDPG adopts the Preferential Experience Replay strategy to select experiences that are conducive to accelerating network convergence, and the experience most similar to the current state is further selected as a learning sample by using Relative Experience Learning (REL). This improves the sampling efficiency of samples and the training and optimization efficiency of the model. Simulation experiments show that the CEL-MADDPG algorithm can effectively improve the training efficiency of the model and has higher task completion efficiency.
The manned/unmanned aircraft collaborative task allocation is developed on the basis of the multi-UAV task planning method, which can give full play to the pilot's human intelligence at critical moments, and is a complement to the lower intelligence of UAVs to improve the system's environmental adaptability and effectiveness. In this paper, we propose a task allocation method based on hierarchical decision-making mechanism and contract net auction algorithm improvement. From the perspective of confrontation between two sides of the air war, the types of tasks that UAVs need to perform in the battlefield are analysed. In the process of task bundle construction, each attribute value of the tasks in the battlefield is quantified, and then handed over to the manned aircraft decision-making mechanism to construct a reasonable auction sequence. In the auction process, target coverage and penalty terms are introduced to improve the objective function, which is used to balance the task allocation load among UAVs. Finally, the effectiveness of the proposed method is verified by experimental simulation.
With the widespread application of Deep Learning (DL), the black-box characteristics of DL raise questions, especially in high-stake decision-making fields like autonomous driving. Consequently, there is a growing demand for research on the interpretability of DL, leading to the emergence of eXplainable Artificial Intelligence as a current research hotspot. Current research on DL interpretability primarily focuses on transparency and post-hoc interpretability. Enhancing interpretability in transparency often requires targeted modifications to the model structure, potentially compromising the model's accuracy. Conversely, improving the interpretability of DL models based on post-hoc interpretability usually does not necessitate adjustments to the model itself. To provide a fast and accurate counterfactual explanation of DL without compromising its performance, this paper proposes a post-hoc interpretation method called relevance inference based on direct contribution to employ counterfactual reasoning in DL. In this method, direct contribution is first designed by improving Layer-wise Relevance Propagation to measure the relevance between the outputs and the inputs. Subsequently, we produce counterfactual examples based on direct contribution. Ultimately, counterfactual results for the DL model are obtained with these counterfactual examples. These counterfactual results effectively describe the behavioral boundaries of the model, facilitating a better understanding of its behavior. Additionally, direct contribution offers an easily implementable interpretable analysis method for studying model behavior. Experiments conducted on various datasets demonstrate that relevance inference can be more efficiently and accurately generate counterfactual examples compared to the state-of-the-art methods, aiding in the analysis of behavioral boundaries in intelligent decision-making models for vehicles.
The advancement of defense capabilities relies heavily on improving air combat proficiency. Effective pilot training plays a pivotal role in achieving this goal. Simulated flight training is a critical method for training pilots, and leveraging intelligent scoring algorithms can significantly enhance pilot proficiency. In this study, we propose an enhanced SVM algorithm that incorporates PCA for dimensionality reduction. By combining pilot training-related data from flight simulators with advanced machine learning techniques, we aim to develop an intelligent digital instructor system. This system provides real-time, objective, and quantitative assessments, along with detailed diagnostic feedback to pilot trainees. Furthermore, the algorithm’s potential extends beyond civilian pilot training to autonomous air combat scenarios involving UCAVs in the future.
Deep Learning (DL) stands out as a leading model for processing high-dimensional data, where the nonlinear transformation of hidden layers effectively extracts features. However, these unexplainable features make DL a low interpretability model. Conversely, Bayesian network (BN) is transparent and highly interpretable, and it can be helpful for interpreting DL. To improve the interpretability of DL from the perspective of feature cognition, we propose the feature analysis network (FAN), a DL structure fused with BN. FAN retains the DL feature extraction capability and applies BN as the output layer to learn the relationships between the features and the outputs. These relationships can be probabilistically represented by the structure and parameters of the BN, intuitively. In a further study, a correlation clustering-based feature analysis network (cc-FAN) is proposed to detect the correlations among inputs and to preserve this information to explain the features’ physical meaning to a certain extent. To quantitatively evaluate the interpretability of the model, we design the network simplification and interpretability indicators separately. Experiments on eight datasets show that FAN has better interpretability than that of the other models with basically unchanged model accuracy and similar model complexities. On the radar effect mechanism dataset, from the feature structure-based relevance interpretability indicator, FAN is up to 4.8 times better than that of the other models, and cc-FAN is up to 21.5 times better than that of the other models. FAN and cc-FAN enhance the interpretability of the DL model structure from the aspects of features; moreover, based on the input correlations, cc-FAN can help us to better understand the physical meaning of features.
When constructing a Bayesian network for high-dimensional data, due to the complex relationships among distinct nodes, the difficulty in detecting the community structure will directly restrict the feasibility of the divide-and-conquer learning algorithm. This study attempts to solve this problem from the shortest path perspective and proposes the heuristic K-standard deviation algorithm. Firstly, we design a novel heuristic function and drive the A* algorithm to find the shortest paths between nodes in the Bayesian network. In addition, the new heuristic function is theoretically proved to be admissible and consistent. Experiments on different synthetic datasets and benchmark Bayesian networks verify that the proposed heuristic K-standard deviation algorithm generally gets better clustering performance than other representative algorithms and improves the efficiency and accuracy of the conventional Bayesian network structure learning algorithms, especially for high-dimensional data.
Compared to existing traffic alert and collision avoidance systems (TCAS), the development of the new Airborne Collision Avoidance System X (ACAS X) adopts a model-based optimization approach to enhance airspace safety and operational efficiency. However, limitations such as the generation of massive numerical tables during development and the separation of development and evaluation processes hinder the system's maintenance and further application in avionics systems. Therefore, in this study, we tackle the aircraft collision avoidance problem using deep reinforcement learning methods, which substantially reduce storage requirements and enable self-updating during interaction with the environment, thus streamlining the development process. Our contributions include constructing a simulation environment for aircraft collision avoidance and establishing a reward system. Through three different reinforcement learning methods, we address collision avoidance while considering aircraft scheduling issues. Simulation results demonstrate the effectiveness of reinforcement learning in tackling aircraft collision avoidance and airspace scheduling problems.
Aiming at the problem that the counterfactual inference methods will cause relevance drift when interpreting the deep networks with multiple feature extraction structures, which leads to inaccurate interpretation, this paper proposes a benchmark conservation relevance inference method based on direct contribution (BCRI). By ensuring the consistency of the relevance propagation of multiple feature extraction structures, BCRI can overcome the relevance drift, and can accurately and quickly analyze the relevance of each input variable. This method can carry out the reasonable analysis of the model and understand the model behavior pattern. Experimental results show that the proposed method can generate more trustworthy counterfactual interpretations efficiently than other methods.
In this paper, an intelligent algorithm integrating model predictive control and Standoff algorithm is proposed to solve trajectory planning that UAVs may face while tracking a moving target cooperatively in a complex three-dimensional environment. A fusion model using model predictive control and Standoff algorithm is thus constructed to ensure trajectory planning and formation maintenance, maximizing UAV sensors’ detection range while minimizing target loss probability. Meanwhile, with this model, a fully connected communication topology is used to complete the UAV communication, multi-UAV formation can be reconfigured and planned at the minimum cost, keeping off deficiency in avoiding real-time obstacles facing the Standoff algorithm. Simulation validation suggests that the fusion algorithm proves to be more capable of maintaining UAVs in stable formation and detecting the target, compared with the model predictive control algorithm alone, in the process of tracking the moving target in a complex 3D environment.