This paper investigates the reinforcement-learning-based structured optimal control for linear stochastic systems with an unknown state matrix under communication topology constraints. A data-driven structured policy iteration algorithm is developed within the framework of adaptive dynamic programming (ADP) to learn a structured feedback policy directly from system trajectory data. By reconstructing the policy evaluation equation from data and updating the feedback gain matrix through analytical orthogonal projection, the proposed method avoids explicit dependence on the exact system dynamics while preserving the desired sparsity pattern. The convergence of the proposed algorithm is established theoretically. Sufficient conditions for the mean-square stability of the closed-loop system are derived, and suboptimal performance bounds induced by topology constraints are characterized. Simulation results show that, even when specific communication links are removed, the learned structured policy yields a steady-state control cost close to that of the unconstrained optimal controller. These results verify the effectiveness of the proposed reinforcement-learning-based structured control method.
更多
查看译文
关键词
Linear stochastic systems,Structured optimal control,Reinforcement learning,Adaptive dynamic programming,Policy iteration,Communication topology