🤖 AI Summary
This work addresses the security vulnerability of large language models (LLMs) to jailbreaking attacks. Methodologically, it proposes a dynamic, multi-round adversarial prompt generation framework that integrates role-playing, context manipulation, and semantic obfuscation into a parameterized attack model. Guided by failure analysis, the framework iteratively refines prompts across dialogue turns and enhances attack efficacy via strategic system prompt engineering and hyperparameter optimization. Its key contribution is the first automated, interpretable, and high-success-rate jailbreaking method tailored to complex, multi-turn conversational scenarios. Experiments on mainstream LLMs—including ChatGPT, Llama-3, and DeepSeek-V2—achieve an 86% jailbreaking success rate, revealing critical structural weaknesses in current safety mechanisms under multi-turn interaction. The approach establishes a novel paradigm for LLM red-teaming evaluation.
📝 Abstract
Large Language Models (LLMs) continue to exhibit vulnerabilities to jailbreaking attacks: carefully crafted malicious inputs intended to circumvent safety guardrails and elicit harmful responses. As such, we present AutoAdv, a novel framework that automates adversarial prompt generation to systematically evaluate and expose vulnerabilities in LLM safety mechanisms. Our approach leverages a parametric attacker LLM to produce semantically disguised malicious prompts through strategic rewriting techniques, specialized system prompts, and optimized hyperparameter configurations. The primary contribution of our work is a dynamic, multi-turn attack methodology that analyzes failed jailbreak attempts and iteratively generates refined follow-up prompts, leveraging techniques such as roleplaying, misdirection, and contextual manipulation. We quantitatively evaluate attack success rate (ASR) using the StrongREJECT (arXiv:2402.10260 [cs.CL]) framework across sequential interaction turns. Through extensive empirical evaluation of state-of-the-art models--including ChatGPT, Llama, and DeepSeek--we reveal significant vulnerabilities, with our automated attacks achieving jailbreak success rates of up to 86% for harmful content generation. Our findings reveal that current safety mechanisms remain susceptible to sophisticated multi-turn attacks, emphasizing the urgent need for more robust defense strategies.