KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational costs and loss of coherent event evidence caused by visual token accumulation in long video reasoning. To this end, we propose a training-free bounded visual memory framework that decouples recent fine-grained observations from a historical structured event repository. We introduce a query-agnostic write mechanism alongside online CRUD (create, read, update, delete) strategies. Furthermore, leveraging a novelty-driven event proposal algorithm, the framework adaptively allocates a fixed reading budget via a text-only router. Compatible with modular vision-language models and encoder-free architectures, our approach achieves state-of-the-art compression performance across most benchmarks using only 10% of the token budget. It significantly improves real-time question-answering accuracy, comprehensively surpassing existing strongest baselines.
📝 Abstract
Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observations in a visual cache and organizes historical evidence into a structured event bank. Candidate events are proposed according to their novelty relative to previously stored events and maintained through an online add--merge--evict update. When a question arrives, a text-only router adaptively allocates a fixed readout budget between recent and event memory, without reprocessing historical frames. KeyRec operates on model-facing visual embeddings and supports both modular encoder--projector VLMs and the encoder- and projector-free NEO-ov architecture. Across four streaming and long-video benchmarks and three VLM backbones, KeyRec achieves the best compressed performance in 13 of 15 settings using only 10\% of the dense decoder-facing visual-token budget. It outperforms the strongest compressed baseline by 2.21--18.37 points on real-time questions, achieves the best compressed result in five of six long-video settings, and performs best in every NEO-ov 2B setting.
Problem

Research questions and friction points this paper is trying to address.

long-video understanding
streaming video
visual memory
visual token compression
vision-language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bounded Visual Memory
Training-free Framework
Event Bank
Text-only Router
Long-Video Understanding
🔎 Similar Papers
No similar papers found.