This article addresses the optimal control problem of nonlinear multiagent pursuit-evasion (PE) games and proposes a novel adaptive dynamic programming (ADP) framework that simultaneously accounts for intrateam cooperation and interteam antagonism. Subsequently, for the continuous-time, nonaffine, and nonlinear structure induced by the PE error dynamics, a $Q$ -function is defined, and a continuous-time $Q$ -learning recursion is derived via integral reinforcement learning. An actor-critic neural network architecture is then employed to approximate the $Q$ -function and the optimal policy, respectively. To address a key difficulty in Lyapunov stability analysis for the nonaffine case, where the actor-network weight-update rate is hard to construct as a negative-definite quadratic form in the weight error and thus does not readily yield uniformly ultimately bounded (UUB), a structured weight-update law is proposed along with a closed-loop stability proof. Under the stated compact-domain, persistent-excitation, and small-gain conditions, the closed-loop analysis establishes uniform ultimate boundedness of the system errors and the weight-estimation errors. Simulation results for a scenario with two pursuers against three evaders demonstrate that the pursuers can achieve capture.
更多
查看译文
关键词
Adaptive dynamic programming (ADP),general nonlinear systems,pursuit–evasion (PE) games