Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing knowledge distillation methods assume a constant supervisory value for each sample, overlooking the dynamic evolution of student capabilities and thereby inducing data redundancy and training inefficiency. To address this limitation, this work proposes a student-curriculum coupled framework that innovatively decouples supervision credibility from necessity. By integrating online policy distillation, anchor-frontier curricula, and closed-loop feedback control, the framework enables the student to dynamically activate or suspend specific supervision subsets on demand according to its current task-specific deficiencies. Experimental results demonstrate that the proposed method achieves an average recall improvement of 5.1% across three benchmarks while reducing the number of training samples by 60% and decreasing computational time by 50.4%.
πŸ“ Abstract
On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.
Problem

Research questions and friction points this paper is trying to address.

Temporal Video Grounding
On-Policy Distillation
Training Curriculum
Supervision Value
Vision-Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Temporal Video Grounding
Student-Curriculum Coupling
Curriculum Learning
Vision-Language Models
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
J
Jiacheng Qiu
State University of New York at Stony Brook
Y
Yunsoo Kim
State University of New York at Stony Brook
Ruichen Xu
Ruichen Xu
Stony Brook University
Machine LearningDeep LearningSimulated Annealing AlgorithmNeural OperatorPDE Solver
Jian Luo
Jian Luo
University of California San Diego
Materials ScienceCeramicsGrain Boundary
P
Petar M. Djurić
State University of New York at Stony Brook
Sima Mofakham
Sima Mofakham
State University of New York at Stony Brook