This paper deals with the problem of attitude tracking control for a flight vehicle in the terminal guidance phase under multi-source disturbances. First, a nonlinear dynamic model for the flight vehicle is formulated in strict-feedback form. Second, external disturbances, aerodynamic parameter variations, and unmodeled dynamics during the terminal guidance phase are treated as a lumped disturbance. An extended state observer (ESO) is then employed to estimate this lumped disturbance, which is further compensated for within the controller. Subsequently, a tracking controller is designed by combining the fully actuated system (FAS) approach with backstepping control, which transforms the original nonlinear system into a desired linear closed-loop system. Based on Lyapunov stability theory, it is proven that the attitude tracking errors converge into the neighborhood of the origin. Finally, the effectiveness of the proposed controller is verified through numerical simulations.
Aiming at the requirements of time consistency and impact angle for loitering munition swarms in cooperative strike missions, this paper first proposes a time-varying desired formation, transforming the guidance problem with multiple constraints into a cooperative tracking problem of the desired formation. To enable the swarm to reach the constrained state and strike the target under the guidance of the desired formation, a virtual leader is set for the swarm, and an intelligent guidance policy based on the Trust Region Policy Optimization (TRPO) algorithm is designed for the virtual leader, which ensures that the virtual leader can hit the target accurately under the impact angle constraint. The time-varying desired formation is composed of the state vectors of each member in the swarm relative to the virtual leader, where the direction of each vector is determined by the expected impact angle, and the modulus of all vectors converges to 0 at the moment when the virtual leader hits the target. The cooperative guidance command proposed for the swarm to track the time-varying desired formation will drive the linear system constituted by the swarm tracking errors to be stable. The results of numerical simulation show that the swarm tracking error always converges during the guidance process, which ensures that the swarm can hit the target at the expected impact angle simultaneously.
Deterministic policy gradient methods with neural network actors, such as Twin Delayed Deep Deterministic Policy Gradient (TD3), offer strong performance but remain difficult to interpret. Although existing fuzzy reinforcement learning methods can improve interpretability to some extent, embedding a fully interpretable fuzzy actor into the TD3 pipeline with rule-level prior knowledge remains challenging. To address these challenges, we propose Fuzzy-TD3, a deterministic actor-critic algorithm that replaces the neural actor with a Takagi-Sugeno (T-S) fuzzy policy and integrates prior knowledge at the rule level. In this framework, the consequents of fuzzy rules associated with locally linearized operating regions are fixed to an optimal feedback law, while the remaining rule consequents are learned directly from data. Antecedents are reparameterized to maintain valid fuzzy partitions and are trained via gradient descent, ensuring interpretability. Simulation results on two inverted-pendulum benchmarks show that Fuzzy-TD3 improves sample efficiency and steady-state regulation while maintaining competitive transient and input-usage performance compared with neural actor-critic baselines. This work provides an interpretable and practical reinforcement learning framework that unites fuzzy theory with classical control, offering a robust solution for data-driven control in industrial applications.
In the terminal guidance stage, when the target's maneuverability approaches that of the interceptor, traditional guidance laws are prone to overloading saturation, resulting in a decrease in interception accuracy. To address this problem of intercepting highly-maneuvering targets, a softmax-based reinforcement learning (SMRL) guidance law is proposed. To overcome the susceptibility of conventional reinforcement learning-based guidance laws to value estimation bias, this guidance law combines the softmax deep double deterministic policy gradients (SD3) algorithm, which significantly reduces inaccuracies in optimal policy approximation due to value estimation bias. Furthermore, the interception problem is formulated as a Markov decision process (MDP), and a dimensionless state normalization method is introduced to enhance the generalization ability of guidance law. A heuristic reward function tailored for highly-maneuvering targets is developed, considering the characteristics of aircraft autopilots and actuator saturation. Simulation results demonstrate that SMRL outperforms conventional methods in terms of generalization capability, guidance accuracy, and maximum tolerable overload ratio, exhibiting great potential for practical engineering applications in high-maneuverability interception missions.
The scarcity of space orbit resources has intensified the issue of orbital games between spacecraft. This paper presents a control strategy for the pursuing spacecraft in a multiple-to-one orbital pursuit–evasion game problem. Initially, a mathematical model is established for the multiple-to-one orbital pursuit–evasion game problem, incorporating impulse maneuvers for both the pursuing and evasive spacecraft. This model also considers the maximum single velocity increment and maximum total velocity increment restrictions imposed on the spacecraft. Subsequently, a method for constructing a pursuing policy, which integrates evasive spacecraft action prediction and the deep deterministic policy gradient algorithm, is proposed. The pursuing policy network and the evasive spacecraft action prediction network are trained under a centralized training and decentralized execution framework. The evasive spacecraft action prediction network can predict the actions of the evasive spacecraft based on the states observed by the pursuing spacecraft. It then predicts the subsequent state of the evasive spacecraft through the control model, providing supplementary information for the decision-making process of the pursuing policy. The simulation results indicate that the pursuing policy, enhanced by the prediction of the evasive spacecraft’s actions, outperforms the traditional deep deterministic policy gradient algorithm in both the training stage and statistical performance. Furthermore, the prediction of the evasive spacecraft’s actions can compensate for the numerical and maneuverability deficiencies of the pursuing spacecraft. However, a key limitation of this approach is its reliance on prior action data of the evasive spacecraft.
The variational Bayesian (VB) adaptive Kalman filter (VBAKF) has emerged as a prominent solution for state estimation under unknown and time-varying noise in navigation and control systems. However, excessive computational demand precludes adoption in resource-limited or latency-critical applications. To address this limitation, we present a VB-unscented transform (UT) adaptive Kalman filter (VU-AKF) framework that seamlessly integrates variational inference with unscented sampling. First, the noise statistics are propagated via a minimal set of deterministically selected sigma points. These sigma points are then analytically updated within the VB formalism and fused using statistically consistent weights, thereby obviating both fixed-point iterations and sliding window storage. A single Kalman update subsequently yields the state estimate. Extensive Monte Carlo simulations demonstrate that the proposed VB-UT scheme preserves the estimation accuracy of the VBAKF while reducing the computational time by 86% and the memory footprint by 37.7% relative to the fixed-point approach and by 82.3% and 29.1% relative to the sliding window strategy. This work provides a practical and efficient estimation-theoretic foundation for high-performance, real-time measurement of the instrumentation under complex and dynamic noise environments.
This paper presents an information-driven bias compensated pseudolinear Kalman filter framework for three-dimensional (3D) bearings-only target tracking under colored measurement noise. To overcome the limitations of traditional centralized methods, we introduce a consensus-based distributed filtering architecture that enables peer-to-peer sensor networks to achieve fundamentally unbiased and consistent state estimation in the presence of colored noise. A novel generalized Send-on-Delta (GSoD) event-based mechanism is incorporated to adaptively control communication by quantifying the incremental information contribution of the local posterior information pair relative to its last broadcast instant. The triggering decision jointly evaluates the information-weighted state deviation and the variation of the local information matrix, reflecting both state novelty and estimation confidence evolution. Furthermore, by considering the causes of bias formation, we propose a new event-based distributed bias compensated pseudolinear Kalman filter (EB-PLKF-CBC) specifically under colored noise environments. This approach reduces redundant data transmission while maintaining high estimation accuracy. Finally, the proposed algorithm was validated using a representative tracking example, and the results confirm its effectiveness.
This paper proposes an H infinity control strategy for rational energy management in fuel cell electric vehicles (FCEVs). The energy management system (EMS) is formulated as a Takagi-Sugeno (T-S) fuzzy model with parameter variations, referred to as a fuzzy parameter-varying (FPV) model. Based on this formulation, a general FPV control approach combining parameter sampling and linear regression is developed to stabilize the battery state of charge (SoC) while ensuring DC bus voltage safety. To further reduce fuel cost, the method is optimized by dynamically constraining the increase in hydrogen consumption. Simulation results on an FCEV case study demonstrate that the proposed strategy effectively coordinates power sources and reduces hydrogen consumption, while significantly lowering the online computational burden compared with existing methods.
Fuzzy parameter-varying (FPV) systems constitute an emerging framework for analyzing nonlinear systems. However, due to the coexistence of fuzzy logic and time-varying parameters, the stabilization methods for FPV systems often suffer from severe conservatism. To address this issue, this study proposes a dimension-reduced design method to construct an FPV sliding mode control (SMC) law for general FPV systems. Furthermore, the global asymptotic stability of the resulting control system is rigorously proved by the stability theory of switching systems. Finally, two simulation cases of different dimensions demonstrate the superiority of the proposed method in terms of stability and computational complexity.
In the era of data-driven decision-making, selecting appropriate nonlinear modeling techniques is critical for building robust and interpretable predictive systems. While both Takagi-Sugeno (T-S) fuzzy models and Back-Propagation (BP) neural networks are well-established universal approximators in the data science domain, their fundamentally different structural characteristics lead to varied performance across application scenarios. This paper presents a systematic comparative study of these two modeling approaches through a series of simulation experiments designed to reflect key tasks in predictive analytics, including static function approximation, dynamic system modeling and forecasting, and real-time state tracking of time-varying systems. By evaluating performance across multiple dimensions modeling accuracy, robustness to noise, and adaptability to temporal dynamics, this work provides actionable insights into the practical strengths and limitations of each model type. The results show that T-S fuzzy models offer superior accuracy in clean, stable environments, while BP neural networks demonstrate strong resilience to noise and generalization ability in uncertain conditions. Additionally, T-S fuzzy models exhibit higher adaptability in real-time, dynamic contexts, making them valuable for time-sensitive predictive applications. This study contributes to the broader data science community by offering a structured framework for model selection based on application-specific demands, helping practitioners and researchers alike to optimize predictive modeling strategies for complex, nonlinear systems.
In scenarios involving unknown data associations, Gaussian Mixture Probability Hypothesis Density (GM-PHD) filtering, rooted in Random Finite Set (RFS) theory, presents a significant advantage over traditional data association methods. However, the performance of the GM-PHD filter can degrade substantially when targets are in close proximity, as multiple measurements may be associated with a single target. To address this limitation, we propose an improved Gaussian component merging and extraction method for GM-PHD filters based on measurement marks. This approach assigns a unique mark to each instantaneous measurement and subsequently merges Gaussian components with the same mark through a filtering process. Employing this approach ensures comprehensive collaboration of the weights, means, and covariances of the corresponding target components, thereby facilitating the efficient merging and extraction of similar target intensities. Importantly, this method effectively prevents incorrect fusion of genuine target components. Simulation results indicate that the proposed algorithm achieves a lower localization error and a more accurate estimation of target counts compared to the standard GM-PHD filter, particularly in scenarios involving closely spaced targets with varying clutter means and detection probabilities.
This research introduces an innovative approach to optimal control for a class of linear systems with input saturation. It leverages the synergy of Takagi-Sugeno (T-S) fuzzy models and reinforcement learning (RL) techniques. To enhance interpretability and analytical accessibility, our approach applies T-S models to approximate the value function and generate optimal control laws while incorporating prior knowledge. By addressing the challenge of limited interpretability associated with conventional neural network utilization in RL, our approach utilizes segmented functions for saturation derivative characteristics approximation, effectively handling non-differentiability issues at saturation boundaries. Furthermore, our research presents a novel gradient identification method to overcome the impractical reliance on next-time-step State variables in RL for current-time-step policy improvements. This enables the derivation of optimal control laws corresponding to each fuzzy rule, ensuring practical applicability in the control field. The proposed methodology is rigorously evaluated through computer simulations, confirming its effectiveness, optimality, and convergence properties. This research contributes valuable insights and practical solutions to input-saturation control systems, offering a versatile and robust framework for real-world applications.
In the context of three-dimensional angle-of-arrival (3D-AOA) target tracking, this study proposes a state-constrained and noise-separated pseudo-linear Kalman filtering (SC-NS-PLKF) algorithm to address nonlinear filtering challenges. Whereas existing bias-compensated (BC), instrumental-variable (IV) and unbiased (UB) PLKF methods only correct the pseudo-linear bias, SC-NS-PLKF achieves high-precision unbiased estimation with enhanced algorithmic stability. Specifically, this method (i) derives a new pseudo-linear measurement model through nonlinear equivalent transformation and noise separation; (ii) employs auxiliary filtering to supply the target position required by NS-PLKF and constructs an ellipsoidal constraint domain that guarantees divergence prevention; (iii) provides a rigorous proof of bounded estimation error and includes a complexity analysis to validate computational efficiency. Extensive Monte-Carlo simulations demonstrate significant gains in accuracy and stability compared to state-of-the-art methods.
Intelligence advancements have significantly escalated the challenges in flight vehicle engagement, particularly in penetration-interception games. To address the issue of unknown interceptor models, we propose a Gaussian process-based interceptor trajectory prediction algorithm. By incorporating the evader's maneuver trajectories, this algorithm promises to greatly improve the evaluation of penetration effectiveness. We investigate the application of various Gaussian process models and identify advantageous scenarios for exact and sparse variational models. The simulation results demonstrate that these models achieve accurate predictions within their advantageous scenarios, which will provide valuable references for penetration strategy evaluation. Copyright (c) 2025 The Authors. This is an open access article under the CC BY-NC-ND license (https://creativecommons.org/licenses/by-nc-nd/4.0/)
To address the challenge of engaging highly protected and maneuverable targets, this paper proposes a policy that employs cruise missile cluster to strike mobile targets. This method is based on a fully cooperative game framework, employing targeted optimization of the Trust Region Policy Optimization (TRPO) algorithm to train guidance policies for cluster members adhering to fully cooperative game rules. To enhance the damage effectiveness of cruise missiles, this paper considers the terminal impact angle constraint, transforming the convergence of the terminal impact angle into the convergence of the line-ofsight angle. This ensures that the velocity direction at the terminal guidance phase aligns with the target-munition line, thereby improving the damage effectiveness of the cruise missiles.
Near space hypersonic vehicles present severe challenges for real-time trajectory prediction due to their high dynamics, expansive operating domains, and unpredictable maneuvering patterns. This article proposes an online trajectory prediction method based on Gaussian Process (GP) regression. Unlike existing approaches, the proposed method does not require complex pretraining or extensive prior knowledge, yet it achieves an effective balance between high accuracy and real-time performance. Specifically, the method first employs an extended Kalman filter (EKF) to fuse the vehicle's fast- and slow-time-varying kinematic models, yielding an efficient online estimate of acceleration. This acceleration estimate sequence is then used to train the GP model, enabling rolling prediction of future accelerations within a sequential iterative framework. Simulation results in two representative maneuvering scenarios demonstrate that the proposed method accurately captures the vehicle's motion trends from the online estimates and delivers superior real-time prediction accuracy.
This paper proposes a novel missile guidance law optimization method based on deep reinforcement learning, specifically targeting terminal guidance for missiles engaging highly maneuverable targets in near-space environments. In scenarios where both the missile and target have comparable overload capabilities, effective interception becomes a significant challenge. Existing methods, such as the Saturated Super-Twisting Algorithms, demonstrate strong performance in maneuvering target interception but face difficulties in parameter tuning and control input saturation. To overcome these limitations, this study introduces the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm to optimize the parameters of missile guidance laws, offering an innovative solution to these complex challenges. The TD3 algorithm, known for its ability to handle noisy environments and mitigate Q-value overestimation, enhances the guidance system's capability to intercept highly maneuverable targets with greater precision. Simulation results validate the proposed approach, demonstrating a substantial performance improvement over traditional methods, thus providing both theoretical and practical contributions to missile guidance system optimization for next-generation missile defense applications.
This paper focuses on the control problem for a two-degree-of-freedom flight simulator experimental setup, proposing a reinforcement learning-based flight attitude controller. The flight simulator aims to simulate the aircraft attitude control system, requiring consideration of its nonlinearity, model uncertainty, and the impact of external disturbances when designing the controller. Proximal Policy Optimization (PPO), as a policy gradient-based deep reinforcement learning algorithm, autonomously learns an approximately optimal controller based on a given objective function without the need for a mathematical model of the controlled object. Thanks to the application of the Actor-Critic framework and neural networks, the training of the two-degree-of-freedom flight simulator controller can rapidly converge within a short period. Simulations validate the generalization capability of the trained PPO controller and its robustness to external disturbances.
Parameter optimization is a crucial area within the field of control theory. This study introduces a novel framework based on reinforcement learning (RL) for controlling quadrotors. Initially, fast nonsingular terminal sliding mode control (FNTSMC) serves as the fundamental trajectory tracking controller for the quadrotor. Subsequently, fixed-time disturbance observers (FTDO) are employed to mitigate disturbances. Ultimately, an RL training framework is introduced to optimize the hyperparameters within the FNTSMCs. Extensive simulation and physical experiments are conducted to validate the efficacy and superiority of the proposed control framework.
To address the challenge of target tracking for non-maneuvering,single-station setups in long-range scenarios,we propose a target tracking algorithm leveraging three-dimensional angle of arrival data,characterized by its asymptotically unbiased nature.Initially,we construct a motion and observation model centered on a non-maneuvering single station,assuming a known rate prior,and examine the system's ob-servability.To tackle the bias inherent in the pseudo linear least squares algorithm,we introduce a con-strained total least squares method that demonstrates asymptotically unbiased properties,with its effective-ness validated through simulations.In tests involving three-dimensional angle tracking over distances in the hundred-kilometer range,with angle measurement standard deviations at 0.1°,0.2°,and 0.3°,the con-strained total least squares method achieves a time-average relative distance error of 6%,12%,and 21%within 50-100 seconds,respectively,and an absolute position error of 9 km,19 km,and 35 km;at initial distances of 70,140,and 280 km,the errors are 1%,6%,and 30%for the same duration,with absolute errors of 0.7 km,9 km,and 30 km.Notably,the relative distance error can be reduced to below 10%within 100 seconds,marking a significant precision enhancement,while maintaining operational speed comparable to the pseudo linear least squares method.The constrained total least squares approach exhibits rapid convergence,high accuracy,and swift processing,showing resilience against angle measurement er-rors and initial distance variations.It offers a robust solution for 3D angle of arrival tracking of non-maneu-vering single-station targets in distant settings.
Changhua Hu (胡昌华)合作论文数中国人民解放军火箭军工程大学导弹工程学院4