Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing safety and helpfulness in vision-language models when confronting implicit cross-modal risks. We propose Intent-Privileged Online Self-Distillation, which leverages evidence-grounded intents as privileged supervision signals. By integrating intent-conditioned preference distillation, online policy self-distillation, and one-shot prompting, our method enables efficient training with a single rollout, eliminating the need for additional safety modules during inference. Experimental results demonstrate that the proposed approach substantially reduces data requirements, accelerates training by 5×, and decreases inference length by 7%. Furthermore, it improves the joint safety-helpfulness success rate from 43.9% to 53.5%, achieving effective and lightweight safety alignment for multimodal systems.
📝 Abstract
Vision-Language Models (VLMs) remain vulnerable to cross-modal implicit risks: visual and textual inputs that appear benign in isolation can jointly elicit unsafe responses. Existing safety methods often require large preference datasets, costly multi-rollout training, or additional safeguards at inference time. They may also sacrifice helpfulness by directly refusing requests that could be answered safely. In this paper, we propose Intent-Privilege On-Policy Self-Distillation (OPSD), which leverages evidence-grounded intent as privileged supervision during training to help VLMs recognize implicit risks and provide safe, useful responses instead of blanket refusals. OPSD distills a teacher's intent-conditioned preferences over responses into a student using a single rollout per prompt; the student then responds without intent annotations or an additional safety module. With only 1,447 safety-specific examples - 95% fewer than standard preference datasets - OPSD reduces training time by 5x relative to multi-rollout GRPO-style training and average inference length by 7%. It attains the highest ratio for joint safety-helpfulness success, which measures the proportion of responses that are both safe and helpful, across all five evaluation groups. Remarkably, on pooled SIUO+HoliSafe, this success ratio rises from 43.9% to 53.5%. These results show that training-time intent supervision can improve both safety and helpfulness while substantially reducing data, training, and inference costs.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Implicit Risks
Safety-Helpfulness Alignment
Cross-modal Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
On-Policy Self-Distillation
Implicit Risks
Safety-Helpfulness Alignment
Intent Supervision
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haotian Deng
Haotian Deng
ByteDance
Computer Networking
W
Wenbin Xing
Sun Yat-sen University
G
Gang Xu
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)
T
Tao He
University of Electronic Science and Technology of China
J
Jinkai Zheng
Hangzhou Dianzi University
Chun Li
Chun Li
MD Anderson Cancer Center
diagnostic imagingdrug deliverynanotechnology
Z
Zheng Zhu
GigaAI
Ming Li
Ming Li
Department of Industrial and Systems Engineering, The Hong Kong Polytechnic University
Cyber-Physical SystemBlockchainLarge Language ModelsESG Technologies