EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of end-to-end optimization in self-distillation and memory entity diversity collapse for streaming video understanding. To this end, it proposes an event-anchored self-distillation framework that models streaming memory as incremental updates over verifiable events. The method introduces multiplicative weight adaptation for reward signals and leverages events as privileged information to reweight the teacher model. Furthermore, it integrates online policy self-distillation with entity coverage rewards to achieve multi-objective reinforcement learning. Experimental results demonstrate that the proposed approach attains 79.8% and 73.4% accuracy on the StreamingBench and OVO-Bench real-time tracks, respectively. It also improves effective entity recall by 17.4% while incurring only a marginal 6.8% increase in memory overhead.
📝 Abstract
Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.
Problem

Research questions and friction points this paper is trying to address.

streaming video understanding
on-policy self-distillation
memory collapse
end-to-end optimization
entity diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Event-Grounded Self-Distillation
Streaming Video Understanding
On-Policy Self-Distillation
Memory Collapse
Entity Coverage Reward
🔎 Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30
💼 Related Jobs
No related jobs found.