🤖 AI Summary
This study addresses the inherent tension between real-time perception and long-term memory in streaming vision-language models under limited computational budgets, along with the issue of historical interference. To this end, we propose FlashBack, a training-free framework that introduces a selective multi-level memory mechanism to retrieve historical context on demand via semantic understanding. Furthermore, it designs a Side-KV pathway to fuse recent contextual information with retrieved memories while preserving the model's native states losslessly. Built upon a dual-trajectory architecture, FlashBack enables dynamic switching to balance real-time responsiveness with long-range dependency modeling. Experimental results demonstrate that our approach significantly improves performance on long-horizon tasks across benchmarks such as OVO-Bench, achieving results comparable to strongly supervised methods while fully retaining real-time perceptual capabilities.
📝 Abstract
Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.