2025 IEEE 31TH INTERNATIONAL CONFERENCE ON PARALLEL AND DISTRIBUTED SYSTEMS, ICPADS(2025)
Univ Elect Sci & Technol China
被引用0|浏览0
摘要
The security of vision-language pre-trained models (VLMs) has become an increasingly critical concern, particularly due to their vulnerability to adversarial attacks in open environments. Most existing methods focus on image perturbations, ignoring image-text structural relationships and the need for dynamic, semantics-aware attack strategies. To address these limitations, we propose a dynamically adaptive framework for generating multi-behavior adversarial patches. A topological neighborhood graph models cross-modal semantic structures, while a lightweight text classifier detects sensitive or jailbreak instructions to switch attack strategies. A two-stage optimization-semantic stripping and target binding-precisely controls perturbations for target-specific outputs. Experimental results demonstrate that the proposed method achieves high attack success rates across a variety of tasks, including image-text retrieval, image classification, and multimodal question answering. Moreover, the approach exhibits strong robustness and adaptability, exposing critical security vulnerabilities in current VLM systems and offering valuable insights for future research on the defense of multimodal models.