Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inconsistent safety feedback behaviors and cross-task prediction challenges in tool-use agents caused by residual jailbreak contexts. To this end, it proposes a paired continuation framework and analyzes over 12,000 samples to reveal the differential routing mechanisms induced by safety feedback. By employing inter-layer activation patching and causal intervention techniques, the authors identify a shared "late commitment" pattern, demonstrating a causal relationship between critical layer representations and macro-level routing outcomes. Based on these findings, a predictive baseline is established using leave-one-parent-task-out cross-validation. Experimental results show that for response-based agents, the model achieves ROC AUC scores of 0.675, 0.777, and 0.702 for rescue, collateral damage, and persistent unsafe behavior, respectively.
📝 Abstract
As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we introduce a paired continuation framework across 192 parent tasks spanning 42 domains, evaluating 12,148 analyzed continuation pairs (curated from a 12,288-pair initially design) across eight diverse agents. We find that identical safety feedback induces sharply model-dependent behavioral routing rather than uniform protection: redirecting unsafe trajectories toward legitimate completion (\emph{rescue}), sustaining unauthorized execution (\emph{persistent unsafe}), or triggering over-refusal on benign tasks (\emph{collateral loss}). Through layer-wise activation patching, we discover a shared \emph{late-commit pattern} where causal intervention effects surge sharply near the final layers (relative depths of 0.958--0.984) despite an over 30-fold variation in peak magnitude across architectures. Crucially, critical-layer representations correlate with macroscopic routing outcomes, and intervening at these layers causally alters concrete next-step tool actions. Building on this causal foundation, we test whether localized intervention-derived features can serve as predictive proxies for full-trajectory routing outcomes on unseen parent tasks under leave-one-parent-task-out evaluation, finding that they provide viable predictive signals in responsive agents with peak ROC AUCs reaching 0.675 for \emph{rescue}, 0.777 for \emph{collateral loss}, and 0.702 for \emph{persistent unsafe}. These findings establish a mechanistic lens and a predictive baseline for anticipating the safety and utility trade-offs of post-jailbreak feedback in autonomous agents.
Problem

Research questions and friction points this paper is trying to address.

jailbreak context
tool agents
safety routing
post-jailbreak feedback
behavioral divergence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Jailbreak context
Tool agents
Activation patching
Late-commit pattern
Safety routing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xi Wang
Xi Wang
University of Defense Technology
LLM SafetyJailbreakSafety Alignment
Songlei Jian
Songlei Jian
NUDT
representation learningmachine learningdata science
Yiming Zhang
Yiming Zhang
University of Science and Technology of China
Computer Vision
B
Bin Ji
National University of Defense Technology
Z
Zhaoye Li
National University of Defense Technology
M
Ma Jun
National University of Defense Technology
B
Baosheng Wang
National University of Defense Technology
J
Jie Yu
National University of Defense Technology