TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of large language models to multi-turn jailbreak attacks, wherein distributed requests circumvent single-turn safety alignment. Grounded in multi-turn trajectory risk theory, we propose the TRACE framework. Its core innovations include reformulating safety principles as token-level optimization objectives and designing a trajectory reward attribution with contrastive erasure mechanism, which leverages refusal-attributed advantage weighting and gradient norm penalties for precise training. Furthermore, advantages are computed via a frozen reference model and its ablated counterpart to perform high-gap position contrastive erasure. Evaluated across five open-source models under seven attack strategies, TRACE achieves the lowest attack success rate while limiting performance degradation on general benchmarks such as MMLU to no more than 1.23 points, effectively balancing safety and utility.
📝 Abstract
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1\.23 points. Source code can be found in the supplemental material.
Problem

Research questions and friction points this paper is trying to address.

multi-turn safety
large language models
safety alignment
jailbreak attacks
preference optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn safety
trajectory return attribution
contrastive erasure
token-level objective
refusal-ablated advantage
🔎 Similar Papers
No similar papers found.