Frame Differential On-Policy Self-Distillation for Video Reasoning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in video reasoning where fixed sparse frame sampling struggles to balance computational efficiency with fine-grained visual information. To this end, we propose FD-OPSD, a reinforcement learning-based method that constructs frame-difference signals from token-level preference discrepancies. It transfers dense-frame evidence into a sparse-frame policy via on-policy self-distillation, maintaining inference consistency without requiring external teacher models or dense autoregressive unrolling. Experiments conducted on Qwen-series multimodal large language models demonstrate that FD-OPSD significantly outperforms baselines such as GRPO across six video reasoning benchmarks, effectively improving average performance under varying frame budgets.
📝 Abstract
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.
Problem

Research questions and friction points this paper is trying to address.

Video Reasoning
Reinforcement Learning
Sparse Frame Budget
Multimodal Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Frame Differential Self-Distillation
On-Policy Reinforcement Learning
Video Reasoning
Token-level Preference
Sparse Frame Rollout
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Haiying He
Haiying He
China Agricultural University
LLMMLLMAgent
X
Xin Zheng
HKUST
S
Shaoli Hu
HKUST
S
Shijun Xiao
NKU
X
Xuanhe Liu
SEU
B
Bing Li
KAUST
Harry Yang
Harry Yang
HKUST
computer visionmachine learning