🤖 AI Summary
This study addresses the lack of fine-grained guidance from sequence-level rewards and conflicts in multi-source evidence aggregation within VideoLLMs by proposing a multi-evidence self-distillation framework. Methodologically, it constructs spatiotemporal and answer teacher models with shared parameters, achieving preference alignment through gated residual fusion and reverse KL divergence correction. Furthermore, this work introduces a novel verification-guided optimization mechanism that categorizes trajectories into successful, failed, and uncertain cases to apply differentiated supervision strategies. Experimental results demonstrate that the proposed method consistently outperforms existing reinforcement learning and self-distillation baselines across multiple video benchmarks.
📝 Abstract
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.