🤖 AI Summary
This study addresses the loss of fine-grained visual information caused by coarse-grained representations in streaming video understanding. To this end, we propose a multi-level entity-aware structured memory framework. This method innovatively introduces a hierarchical semantic chunking strategy spanning global, entity, and spatial levels, coupled with a structured indexing mechanism that enables on-demand access to high-resolution content for efficient retrieval and reasoning. Notably, the framework functions as a plug-and-play module requiring no additional training. Experimental results demonstrate that our approach significantly enhances the performance of multiple base models on benchmarks such as StreamingBench, achieving state-of-the-art results.
📝 Abstract
Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.