🤖 AI Summary
This study addresses the challenges of end-to-end optimization in self-distillation and memory entity diversity collapse for streaming video understanding. To this end, it proposes an event-anchored self-distillation framework that models streaming memory as incremental updates over verifiable events. The method introduces multiplicative weight adaptation for reward signals and leverages events as privileged information to reweight the teacher model. Furthermore, it integrates online policy self-distillation with entity coverage rewards to achieve multi-objective reinforcement learning. Experimental results demonstrate that the proposed approach attains 79.8% and 73.4% accuracy on the StreamingBench and OVO-Bench real-time tracks, respectively. It also improves effective entity recall by 17.4% while incurring only a marginal 6.8% increase in memory overhead.
📝 Abstract
Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.