Does On-Policy Distillation for Safety Pose Backdoor Risks?

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically reveals, for the first time, the risk of backdoor propagation in safe alignment via online policy distillation (OPD), demonstrating that untrusted teacher models can efficiently transfer malicious behaviors to student models. Our findings indicate that extremely low poisoning rates suffice to trigger highly successful attacks, with training duration and KL loss selection significantly amplifying this vulnerability. To mitigate this threat, we propose Lazy Defense, a mechanism based on KL reward clipping that delays backdoor acquisition by constraining aggressive parameter updates. Experimental results show that a 3% poisoning rate yields a 70% attack success rate, while merely ten poisoned samples achieve 67% after sixteen training epochs. Crucially, Lazy Defense effectively postpones backdoor transfer under low-poisoning scenarios, providing essential security guarantees for trustworthy policy distillation.
📝 Abstract
On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Distillation
Backdoor Attack
LLM Safety
Knowledge Distillation
Poisoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy distillation
Backdoor attack
Safety alignment
Knowledge distillation
Lazy Defense