2026 International Conference on Unmanned Aircraft Systems (ICUAS)(2026)
被引用0|浏览0
摘要
This work presents a lightweight, high-level planner for UAV navigation based on a Soft Actor-Critic (SAC) reinforcement learning policy. Trained in a simplified 3D simulation, the policy generates velocity commands to reach static goals and generalizes zero-shot to dynamic tasks, including trajectory tracking and pursuit of moving targets. A parallelized PyTorch implementation accelerates training, enabling convergence in under five minutes on accessible computing hardware. The policy was validated in SITL and controlled indoor flight experiments using a real UAV with Vicon-based localization. Results demonstrate that a policy trained under simplified assumptions can generalize to multiple navigation-related tasks while requiring modest onboard computational resources. A video demonstration of the main experiments can be found at https://youtu.be/R2PkxlgmO74.