AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出AUDITPLAN方法,通过先制定安全计划再作答的方式,解决了模型仅凭最终答案难以区分稳健拒绝与不当捷径的问题,提高了安全性对齐的可靠性与可审计性。
📝 Abstract
Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.
Problem

Research questions and friction points this paper is trying to address.

safety alignment
robust refusal
shortcuts
Innovation

Methods, ideas, or system contributions that make the work stand out.

AUDITPLAN
FAITHGATE
safety plan
robustness
auditability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sai Sri Pushpa Jampani
Indian Institute of Technology Patna, Bihar, India
K
Kshitij Mishra
Indian Institute of Technology Patna, Bihar, India
Asif Ekbal
Asif Ekbal
Department of Computer Science and Engineering, IIT Patna
Artificial IntelligenceNatural Language ProcessingMachine Learning Application