How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of language models to multilingual narrative wrapping attacks that bypass safety refusal mechanisms. To this end, it constructs GUISE, a cross-lingual attack benchmark, and proposes AXIS, a defense framework that enhances model safety through representation alignment and commitment objectives. Furthermore, this work introduces a pioneering "warn-before-answer" evaluation criterion, revealing the significant shifting effect of narrative wrapping on refusal behavior. By integrating preference optimization with rotation objective functions, the proposed approach achieves robust training. Extensive experiments demonstrate that AXIS attains state-of-the-art comprehensive safety and utility scores across multiple mainstream models, effectively defending against cross-lingual narrative attacks while preserving helpfulness.
📝 Abstract
Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
Problem

Research questions and friction points this paper is trying to address.

narrative wrapping
LLM refusal
jailbreak
cross-language safety
safety alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Narrative Wrapping
Cross-Language Benchmark
Representation Analysis
Preference Optimization
LLM Safety
🔎 Similar Papers