ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)(2026)
Zhejiang University
被引用0|浏览11
摘要
Large language models (LLMs) excel in natural language generation but remain susceptible to malicious prompts that can elicit harmful outputs. To address this, alignment techniques such as reinforcement learning from human feedback (RLHF), instruction tuning, and adversarial training are employed to enforce safety constraints. However, recent studies reveal vulnerabilities to white-box and black-box attacks that undermine these safeguards. This paper introduces a novel, interpretable jailbreak framework for LLMs, comprising three key steps: (i) activating safety-related parameters by inputting pre-processed harmful and harmless prompts; (ii) identifying critical safety-constraint parameters using gradient-driven conditional localization combined with similarity-based neuron recognition; and (iii) relearning selected parameters with a weighted cross-entropy loss to restore the model’s ability to generate previously restricted responses. Experiments on mainstream models, including Llama2, Qwen2.5, DeepSeek, and Mistral, demonstrate a 100% Attack Success Rate (ASR), surpassing existing fine-tuning-based white-box attacks while minimally impacting overall model performance. Our approach achieves jailbreaks by updating only a small fraction of parameters, offering high interpretability and transferability across diverse LLM architectures. While aimed at advancing safety research, we acknowledge the potential misuse of open-source LLMs and advocate for strengthened safety strategies to foster a secure AI ecosystem.