MEMO: Multi-Level Entity-Aware Memory for Streaming Video Understanding

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of fine-grained visual information caused by coarse-grained representations in streaming video understanding. To this end, we propose a multi-level entity-aware structured memory framework. This method innovatively introduces a hierarchical semantic chunking strategy spanning global, entity, and spatial levels, coupled with a structured indexing mechanism that enables on-demand access to high-resolution content for efficient retrieval and reasoning. Notably, the framework functions as a plug-and-play module requiring no additional training. Experimental results demonstrate that our approach significantly enhances the performance of multiple base models on benchmarks such as StreamingBench, achieving state-of-the-art results.
📝 Abstract
Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons, posing a fundamental challenge for memory modeling. Existing approaches primarily focus on increasing memory capacity, either by compressing historical information into fixed-size representations or by extending storage beyond GPU memory. However, these methods largely rely on global or coarse-grained representations, inevitably losing fine-grained visual information. In this work, we argue that streaming video memory should explicitly encode structured and semantically meaningful representations, particularly at the entity level. To this end, we propose MEMO, a novel framework that models streaming video through multi-level, entity-aware structured memory. MEMO performs multi-level perception to jointly capture global semantics, entity dynamics, and spatial structures, partitioning streaming video into semantically coherent chunks. Each chunk is organized into a structured memory, where lightweight global and entity-level representations serve as retrieval indices, while the corresponding high-resolution visual content is retained separately for on-demand access. At inference time, MEMO performs query-specific retrieval over the structured memory and selectively recalls relevant visual evidence for downstream reasoning. Notably, MEMO is training-free and plug-and-play with existing multimodal large language models. Extensive experiments on StreamingBench and OVO-Bench demonstrate that MEMO consistently improves multiple base models and achieves state-of-the-art performance.
Problem

Research questions and friction points this paper is trying to address.

streaming video understanding
memory modeling
entity-aware representation
fine-grained visual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Streaming Video Understanding
Entity-Aware Memory
Structured Memory
Training-Free
Multi-Level Perception
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yinying Li
East China Normal University
Y
Yuqian Fu
King Abdullah University of Science and Technology
Y
Yulin Dai
East China Normal University
Jingyu Gong
Jingyu Gong
Shanghai Jiao Tong University
3D Computer Vision
Tianwen Qian
Tianwen Qian
East China Normal University
MultimediaVision and LanguageEmbodied AI
X
Xiaoling Wang
East China Normal University