How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions

📅 2026-07-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of jailbreaking attacks predominantly focus on attack success rates, which poorly reflect their actual value in improving model safety and alignment. This work proposes the first defender-centric evaluation paradigm, treating jailbreak samples as red-teaming resources and measuring attack utility through their downstream effectiveness in enhancing model robustness. To this end, we introduce the A-MESS framework, which employs an AttackSHAP score—based on Shapley values—to quantify the marginal contribution of individual attacks. By integrating black-box utility estimation with surrogate model optimization, our approach efficiently selects high-value attack subsets under query budget constraints. Experiments demonstrate that attack success rate exhibits weak correlation with defensive utility, that AttackSHAP can be accurately estimated with few queries, and that the selected subsets significantly improve model safety.
📝 Abstract
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting-agnostic framework for attributing and selecting jailbreak attacks from black-box subset utility observations. A-MESS estimates AttackSHAP, a Shapley-based score that attributes marginal utility to individual attacks and selects compact attack subsets under user-specified budgets via greedy or surrogate-based optimization. Across controlled utility landscapes and real LLM safety settings, we find that ASR rankings are weakly aligned with defender-centric utility, that AttackSHAP can be estimated accurately with limited utility queries, and that directly optimizing subsets yields stronger safety utility than attacker-centric or attribution-only selection. These results suggest evaluating jailbreak attacks as resources for improving safety, not only as tools for breaking models.
Problem

Research questions and friction points this paper is trying to address.

jailbreak attacks
safety alignment
defender-centric evaluation
red-teaming
attack utility
Innovation

Methods, ideas, or system contributions that make the work stand out.

jailbreak attacks
defender-centric evaluation
Shapley value
safety alignment
attack subset selection
Y
Yukai Zhou
ShanghaiTech University
F
Feiyang Lu
ShanghaiTech University
X
Xiaokai Mao
Zhejiang University
J
Jinfei Liu
Zhejiang University
W
Wenjie Wang
ShanghaiTech University