EOPSA: Efficient On-Policy Self-Distilled Safety Alignment

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training inefficiency and inference degradation in online self-distillation for safety alignment caused by supervision collapse and style gradient dilution. To overcome these challenges, we propose a critical-token-focused optimization framework. Methodologically, we design an adaptive rollout scheduling strategy and a selective distillation mechanism, introducing a teacher rescue rate to delineate reliable supervision intervals. Furthermore, precise gradient updates are achieved through dynamic generation boundary control combined with safety signal filtering. Experimental results demonstrate that the proposed approach reduces computational overhead by approximately 50% while backpropagating gradients for only 2% of tokens. It significantly outperforms full-distillation baselines in maintaining both safety compliance and reasoning capabilities.
📝 Abstract
On-Policy Self-Distillation (OPSD) has emerged as a promising paradigm for safety alignment, delivering dense, token-level supervision by distilling from a teacher conditioned on refusal-oriented privileged prompts. However, we reveal that this paradigm suffers from critical inefficiencies that degrade both training efficiency and general reasoning capabilities. Specifically, we diagnose two fundamental bottlenecks: (1) supervisory collapse over extended rollouts, where the teacher's corrective efficacy degrades precipitously as the student's generation prefix lengthens, injecting noisy gradients into late-stage tokens; and (2) gradient dilution from stylistic shifts, where the distillation objective is dominated by safety-irrelevant stylistic discrepancies induced by privileged prompting, washing out genuine safety signals and impairing base reasoning. To resolve these issues, we propose Efficient On-Policy Self-Distilled Safety Alignment (EOPSA), which concentrates computational and gradient budgets exclusively on reliably supervised, safety-critical tokens. EOPSA incorporates two coordinated mechanisms: (i) Adaptive Rollout Scheduling, which dynamically bounds the generation horizon guided by a novel Teacher Rescue Rate (TRR) metric to operate strictly within reliable supervision regimes; and (ii) Selective Distillation, which filters out safety-neutral tokens to restrict gradient updates exclusively to safety-pivotal transitions. Extensive evaluations across reasoning models up to 32B parameters demonstrate that EOPSA slashes rollout computation by $\sim$50% and backpropagates through merely $\sim$2% of tokens, consistently outperforming full-token distillation baselines in both safety compliance and reasoning retention.
Problem

Research questions and friction points this paper is trying to address.

safety alignment
on-policy self-distillation
supervisory collapse
gradient dilution
reasoning degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Self-Distillation
Safety Alignment
Adaptive Rollout Scheduling
Selective Distillation
Teacher Rescue Rate
🔎 Similar Papers
No similar papers found.
Q
Qirui Liu
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Y
Yichen Sun
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Y
Yan Wang
Ant Group
Y
Yu Mi
The State Key Laboratory of Blockchain and Data Security, Zhejiang University
Wei Cao
Wei Cao
PhD, Oklahoma State University
Metamaterialsplasmonicsphotonicsterahertz
Yue Shen
Yue Shen
Ant Group
recommend user growth
Zhixuan Chu
Zhixuan Chu
Associate Professor, Zhejiang University; Alibaba Group; Ant Group
Kui Ren
Kui Ren
Professor and Dean of Computer Science, Zhejiang University, ACM/IEEE Fellow
Data Security & PrivacyAI SecurityIoT & Vehicular Security