MeSD: Multi-Evidence Self-Distillation for VideoLLM

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of fine-grained guidance from sequence-level rewards and conflicts in multi-source evidence aggregation within VideoLLMs by proposing a multi-evidence self-distillation framework. Methodologically, it constructs spatiotemporal and answer teacher models with shared parameters, achieving preference alignment through gated residual fusion and reverse KL divergence correction. Furthermore, this work introduces a novel verification-guided optimization mechanism that categorizes trajectories into successful, failed, and uncertain cases to apply differentiated supervision strategies. Experimental results demonstrate that the proposed method consistently outperforms existing reinforcement learning and self-distillation baselines across multiple video benchmarks.
📝 Abstract
While reinforcement learning with verifiable rewards provides reliable outcome supervision for VideoLLMs, sequence-level rewards offer limited token-level guidance. On-policy self-distillation addresses this limitation by conditioning a self-teacher on privileged information to provide dense token-level supervision. However, aggregating heterogeneous evidence within a single teacher context obscures cross-evidence agreement and conflict. A further challenge lies in determining whether teacher guidance should refine reward-based updates or provide corrective supervision for failed trajectories. To address these issues, we propose MeSD, a multi-evidence self-distillation framework for VideoLLMs. MeSD constructs three evidence-conditioned teachers with shared parameters, using the ground-truth answer as a common semantic context while separately incorporating temporal and spatial evidence. Given the same student-generated prefixes, MeSD evaluates evidence-specific preferences relative to the Answer Teacher and fuses teacher-common preferences with gated teacher-specific residuals. Furthermore, MeSD introduces Verification-Guided Optimization to classify trajectories as Success, Failure, or Indeterminate. For Success and Indeterminate trajectories, MeSD refines token-level advantage magnitudes while preserving reward-derived signs. For verified failure trajectories that contain the required evidence, MeSD applies failure-conditioned distillation, using reverse-KL correction toward the fused distribution. Experiments on multiple video benchmarks demonstrate consistent gains over reinforcement learning and self-distillation baselines.
Problem

Research questions and friction points this paper is trying to address.

VideoLLMs
token-level supervision
self-distillation
heterogeneous evidence
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Evidence Self-Distillation
VideoLLM
Verification-Guided Optimization
Gated Fusion
Reverse-KL Correction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Weijie Zhu
UCAS
Han Fang
Han Fang
TeleAI, China Telecom (中国电信人工智能研究院 TeleAI)
Text-to-imageMLLMVideo-text retrievalFace recognition
H
Hanyu Fu
UCAS
Y
Yuzhe Zhang
PKU
Xin Wei
Xin Wei
Schmidt AI in Science Postdoc, University of Michigan
Natural HazardsAI for GeohazardsResilienceRiskReliability
Z
Zhaoyan Pan
ZJU
F
Feiran Liu
BJTU
X
Xunjie Jin
BUAA
Hongbo Sun
Hongbo Sun
Institute of Artificial Intelligence, China Telecom(中国电信人工智能研究院,TeleAI), Peking University
Fine-grained visual analysisMulti-modal understandingMachine learning
Zhiyu Lin
Zhiyu Lin
Beijing Jiaotong University
Tianyi Gao
Tianyi Gao
Washington University in St. Louis
T
Tianyi Ding
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
Ye Yuan
Ye Yuan
Chinatelecom
computer vison;machine learning
Z
Zhongjiang He
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
H
Hao Sun
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd
Z
Zhiheng Wu
UCAS