DeShortcut-Align: Decoupling Spurious Shortcuts for Robust Safety Alignment in Large Reasoning Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-refusal and general capability degradation in safety alignment of large reasoning models caused by learned formatting and lexical shortcuts. We propose a decoupled alignment framework featuring a novel three-stage synergistic mechanism. Specifically, refusal sensitivity attribution precisely identifies superficial cues, while attribution-guided contrastive data augmentation disentangles spurious correlations. Furthermore, attention-blinded counterfactual consistency regularization is introduced to eliminate shortcut reliance. Experimental results demonstrate that our approach reduces template attack risks by 72% and decreases over-refusal by over 58%, achieving robust, intent-sensitive safety evaluation while effectively preserving the models' general reasoning capabilities.
📝 Abstract
Safety alignment of large reasoning models (LRMs) via supervised fine-tuning (SFT) and reinforcement learning (RL) often yields near-perfect safety scores, yet this apparent success comes at the cost of severe over-refusal and degraded general capabilities. Through systematic empirical analysis, we find that these failures are closely associated with the learning of spurious shortcuts rather than robust intent-sensitive safety evaluation. Specifically, we identify two dominant shortcuts: formatting shortcuts, where refusal behaviors are overly bound to structural prompt templates that frequently appear in safety alignment corpora; and lexical shortcuts, where sensitive keywords reflexively trigger refusals on benign queries. To mitigate reliance on these shortcuts, we propose DeShortcut-Align, a shortcut-decoupling alignment framework that reduces dependence on superficial cues. DeShortcut-Align operates across three coordinated stages: (1) Refusal Sensitivity Attribution, which masks input tokens to quantify their impact on the final refusal response distribution; (2) Attribution-Guided Contrastive Augmentation, which constructs benign contrastive samples using high-sensitivity tokens to mitigate lexical shortcuts; and (3) Counterfactual Consistency Regularization, which constructs template-ablated states via attention blinding to enforce decision consistency across SFT and RL, mitigating formatting shortcut dependence. Experiments on 7B and 14B models demonstrate that DeShortcut-Align significantly improves robustness against template-stripping bypass attacks (reducing performance drops by up to 72%), substantially reduces over-refusal by over 58%, and better preserves general-purpose reasoning capabilities, thereby mitigating the alignment tax commonly observed in safety training.
Problem

Research questions and friction points this paper is trying to address.

Safety Alignment
Large Reasoning Models
Spurious Shortcuts
Over-refusal
Alignment Tax
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Alignment
Spurious Shortcuts
Contrastive Augmentation
Counterfactual Consistency Regularization
Large Reasoning Models
🔎 Similar Papers
No similar papers found.
Q
Qirui Liu
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Y
Yichen Sun
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Y
Yan Wang
Ant Group
Zhixuan Chu
Zhixuan Chu
Associate Professor, Zhejiang University; Alibaba Group; Ant Group
L
Linbo Jiang
Chongqing Ant Consumer Finance Co., Ltd
J
Jianan Lin
Chongqing Ant Consumer Finance Co., Ltd
Kui Ren
Kui Ren
Professor and Dean of Computer Science, Zhejiang University, ACM/IEEE Fellow
Data Security & PrivacyAI SecurityIoT & Vehicular Security