FlashBack: Knowing When to Remember in Streaming Vision-Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent tension between real-time perception and long-term memory in streaming vision-language models under limited computational budgets, along with the issue of historical interference. To this end, we propose FlashBack, a training-free framework that introduces a selective multi-level memory mechanism to retrieve historical context on demand via semantic understanding. Furthermore, it designs a Side-KV pathway to fuse recent contextual information with retrieved memories while preserving the model's native states losslessly. Built upon a dual-trajectory architecture, FlashBack enables dynamic switching to balance real-time responsiveness with long-range dependency modeling. Experimental results demonstrate that our approach significantly improves performance on long-horizon tasks across benchmarks such as OVO-Bench, achieving results comparable to strongly supervised methods while fully retaining real-time perceptual capabilities.
📝 Abstract
Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between real-time perception and long-term memory. Retrieving historical information provides a natural remedy, yet historical recall is not uniformly beneficial: unnecessary history may introduce irrelevant context into current reasoning and interfere with native real-time perception. Effective streaming memory should therefore address not only what to remember, but also when and how to access it. To this end, we introduce FlashBack, a training-free framework for selective, multi-level memory in streaming vision-language models. Before retrieving history, FlashBack draws on the semantic understanding of the frozen streaming VLM to infer whether a query calls for historical evidence. This assessment determines whether inference remains on the Native trajectory or invokes an isolated Recall trajectory. The Recall trajectory combines recent context with retrieved long-term memory through a query-local Side-KV pathway, preserving local temporal continuity without modifying the persistent Native state. We instantiate FlashBack on StreamingVLM and Mage-VL-4B and evaluate it on OVO-Bench and StreamingBench. The results show improvements on several long-horizon and memory-dependent tasks while largely preserving real-time perception, with performance competitive with strong training-based streaming methods despite requiring no additional training. Our code will be announced later.
Problem

Research questions and friction points this paper is trying to address.

Streaming Vision-Language Models
Long-term Memory
Real-time Perception
Historical Recall
Memory Retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Vision-Language Models
Training-free Framework
Selective Memory Retrieval
Side-KV Pathway
Long-term Memory
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yi Chen
Rightly Robotics, A4X
M
MingMing Yu
Rightly Robotics, A4X
R
Rui-Qi Wang
Rightly Robotics, A4X
B
Boran Wang
Rightly Robotics, A4X
X
Xiaohang Cao
Rightly Robotics, A4X
C
Chu Tang
Rightly Robotics, A4X
Jingmin Chen
Jingmin Chen
Alibaba Group
Choice BehaviorMachine LearningDeep LearningTransportation
J
Jie Gu
Rightly Robotics, A4X