2026 IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE International Conference on Robotics, Automation and Mechatronics (RAM)(2026)
Jilin University College of Communication Engineering Jilin University
被引用0|浏览0
摘要
Vision-language models (VLMs) have shown strong performance in autonomous driving (AD) tasks, supporting scene understanding and safety-related multimodal reasoning. However, robustness under adversarial perturbations remains critical, and the alignment vulnerability between visual evidence and task semantics under sequential observations is insufficiently explored. This paper proposes TCMA, a Temporal Cross-Modal Alignment Attack combining a task-oriented objective, an alignment disruption loss, and a lightweight temporal propagation mechanism to attack perception-oriented VLMs in AD. Specifically, TCMA constructs a semantic anchor from the task prompt to suppress correct visual-text alignment, and warm-starts each frame’s attack from the previous perturbation while enforcing temporal consistency. On BDD100K with Dolphins, TCMA achieves 50.0\% overall center-frame targeted ASR across traffic-light, pedestrian, and rider tasks. Transfer evaluation on Qwen2.5-VL further reaches 73.3\% center-frame and 76.7\% vote-level overall targeted ASR, demonstrating strong cross-model generalizability.