Unmanned aerial vehicle (UAV) swarm agents operating in radio-frequency (RF)-degraded environments require coordination mechanisms that connect directly observable signals to executable responses. This work presents an embodied visual-communication approach in which bio-inspired motion–LED glyphs are represented by a reduced six-parameter semantic chart embedded within a full 24-dimensional hybrid execution manifold. A learned translator large language model (LLM) maps a perceived glyph to a response that is instantiated as an executable trajectory and propagated through closed-loop quadrotor dynamics. The system is evaluated in 200 three-UAV search-and-rescue trials with receiver-specific degradation from sensing range, field of view, occlusion, and relative motion. Under clean observations, the quantized small model and rule-based translator both achieve 100% semantic correctness, but under single- and two-parameter corruption, the quantized small model achieves 64.7% and 64.1% , compared with 44.8% and 35.2% for rule-based, while reducing mean multi-agent trajectory error from 2.902 m to 0.993 m . All methods maintain 100% rotor-allocation and finite-horizon feasibility. The quantized small model also has comparable overall semantic correctness with the large model ( 83.3% vs. 86.3% ) while dropping latency from 5.115 s to 2.789 s and is validated onboard a Jetson–Pixhawk UAV platform, where airborne inference and bounded command execution produce measurable physical motion. These results establish a unified pathway from degraded visual observation to semantically meaningful, dynamically grounded UAV coordination.
更多
查看译文
关键词
Multi-UAV systems,Autonomous aerial robotics,Embodied intelligence,Bio-inspired communication,Learning-enabled control,Large language models