Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing gradient-based jailbreak detection methods, such as GradSafe, in multi-turn dialogues, where significant distributional discrepancies exist between synthetic and real-world data. Through gradient alignment analysis and ROC-AUC evaluation, this work reveals this gap and proposes a sliding-window context scanner, establishing the necessity of short windows and length-aware thresholds. Empirical evaluations on the WildChat dataset demonstrate that the detection AUC degrades to 0.76, underscoring the urgent need for calibration against authentic benign data. Ultimately, this research provides critical empirical evidence and methodological guidance for enhancing the reliability of jailbreak detection in multi-turn conversational scenarios.
📝 Abstract
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an input by the alignment between its induced gradient and a fixed unsafe reference direction, and their effectiveness in multi-turn dialogue remains unclear. We conduct a controlled evaluation of gradient-based jailbreak detection in multi-turn settings. We extend GradSafe with a Context Window Scanner that applies the detector to fixed-size windows of user turns and uses the maximum window score as the conversation-level score. We evaluate different window sizes, attack families, benign conversation distributions, and target models. The results differ sharply between synthetic and realistic benign settings. Against synthetic benign conversations, the detector achieves an ROC-AUC of 0.98 on human-authored multi-turn jailbreaks. On WildChat benign conversations, ROC-AUC drops to 0.76, and a threshold calibrated on synthetic data flags more than 90% of benign conversations as unsafe. Under realistic benign distributions, single-turn windows give the highest separability, whereas longer windows and accumulated contexts reduce performance. The detector is also sensitive to the attack-generation method and target model: successful Crescendo attacks receive scores comparable to or lower than benign conversations, and Qwen2.5-7B-Instruct yields near-random separability with a different optimal window size. These findings show that gradient-based signals can support multi-turn jailbreak detection, but reliable deployment requires calibration on realistic benign conversations, short-window scoring, length-aware thresholds, and evaluation across attack types and model architectures.
Problem

Research questions and friction points this paper is trying to address.

multi-turn dialogue
jailbreak detection
gradient-based detection
safety alignment
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-turn Jailbreak Detection
Gradient-based Detection
Context Window Scanner
Safety Alignment
WildChat Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Omar Sheta
University of Louisville
R
Rinku Deuja
University of Louisville
H
Hadi Masoudi
University of Louisville
Minghong Fang
Minghong Fang
University of Louisville
SecurityPrivacyAI SafetyMachine Learning