🤖 AI Summary
This study addresses the challenge of balancing safety and helpfulness in vision-language models when confronting implicit cross-modal risks. We propose Intent-Privileged Online Self-Distillation, which leverages evidence-grounded intents as privileged supervision signals. By integrating intent-conditioned preference distillation, online policy self-distillation, and one-shot prompting, our method enables efficient training with a single rollout, eliminating the need for additional safety modules during inference. Experimental results demonstrate that the proposed approach substantially reduces data requirements, accelerates training by 5×, and decreases inference length by 7%. Furthermore, it improves the joint safety-helpfulness success rate from 43.9% to 53.5%, achieving effective and lightweight safety alignment for multimodal systems.
📝 Abstract
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.