AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models

📅 2025-04-18
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the security vulnerability of large language models (LLMs) to jailbreaking attacks. Methodologically, it proposes a dynamic, multi-round adversarial prompt generation framework that integrates role-playing, context manipulation, and semantic obfuscation into a parameterized attack model. Guided by failure analysis, the framework iteratively refines prompts across dialogue turns and enhances attack efficacy via strategic system prompt engineering and hyperparameter optimization. Its key contribution is the first automated, interpretable, and high-success-rate jailbreaking method tailored to complex, multi-turn conversational scenarios. Experiments on mainstream LLMs—including ChatGPT, Llama-3, and DeepSeek-V2—achieve an 86% jailbreaking success rate, revealing critical structural weaknesses in current safety mechanisms under multi-turn interaction. The approach establishes a novel paradigm for LLM red-teaming evaluation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessMultiagent Systems: Adversarial Agents

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large Language Models (LLMs) continue to exhibit vulnerabilities to jailbreaking attacks: carefully crafted malicious inputs intended to circumvent safety guardrails and elicit harmful responses. As such, we present AutoAdv, a novel framework that automates adversarial prompt generation to systematically evaluate and expose vulnerabilities in LLM safety mechanisms. Our approach leverages a parametric attacker LLM to produce semantically disguised malicious prompts through strategic rewriting techniques, specialized system prompts, and optimized hyperparameter configurations. The primary contribution of our work is a dynamic, multi-turn attack methodology that analyzes failed jailbreak attempts and iteratively generates refined follow-up prompts, leveraging techniques such as roleplaying, misdirection, and contextual manipulation. We quantitatively evaluate attack success rate (ASR) using the StrongREJECT (arXiv:2402.10260 [cs.CL]) framework across sequential interaction turns. Through extensive empirical evaluation of state-of-the-art models--including ChatGPT, Llama, and DeepSeek--we reveal significant vulnerabilities, with our automated attacks achieving jailbreak success rates of up to 86% for harmful content generation. Our findings reveal that current safety mechanisms remain susceptible to sophisticated multi-turn attacks, emphasizing the urgent need for more robust defense strategies.
Problem

Research questions and friction points this paper is trying to address.

Automates adversarial prompt generation to test LLM safety
Exposes vulnerabilities in LLMs through multi-turn jailbreaking attacks
Evaluates attack success on models like ChatGPT and Llama
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated adversarial prompt generation for systematic LLM vulnerability evaluation
Dynamic multi-turn attack methodology with iterative prompt refinement
Leveraging parametric attacker LLM with strategic rewriting and roleplaying techniques
💼 Related Jobs
No related jobs found.
A
Aashray Reddy
Del Norte High School
A
Andrew Zagula
Bridgewater-Raritan HS
N
Nicholas Saban
West Valley College