Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of high construction costs and the inability to dynamically update offline static memory in long video understanding. To overcome this, we propose an agent-based online dynamic memory framework that departs from the conventional "construct-then-infer" paradigm by synchronously building a dynamic temporal tree memory during inference. Specifically, the method achieves on-demand refinement through low-frame-rate preliminary screening followed by high-frame-rate detailed reading, while textualizing question-answering records to replace raw video inputs, thereby enabling persistent and dynamically updatable memory. Experimental results demonstrate that the proposed framework significantly reduces context overhead and outperforms representative offline methods while maintaining accuracy.
📝 Abstract
Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and this pipeline is costly at both ends: with few questions, building memory for the whole video costs far more than answering them; with many questions, the memory is never updated, so what is learned while answering questions is lost to the next question. To alleviate these, we introduce Sprout, an agentic framework that builds memory while reasoning: a temporal tree that sprouts detailed nodes as questions are answered. The agent watches the video segment by segment at a low frame rate, stopping when the current question can be answered, remembers each segment as a coarse node of the tree, and revisits key intervals at a higher frame rate to refine the tree with the recovered details. Once a segment is recorded as text, its video input is removed from the context history, while the original video remains reachable through the video tools. The memory tree and prior question--answer records persist across questions, so the memory is online and dynamic: built from the first question onward and updated by every question thereafter. We find that replacing accumulated video inputs with textual memory substantially reduces context usage while maintaining accuracy, with slight improvements in some settings. Across benchmarks on three models, Sprout achieves competitive or improved accuracy relative to representative offline memory methods, with no upfront construction stage and lower context cost per question.
Problem

Research questions and friction points this paper is trying to address.

long video understanding
video memory
multimodal large language models
context limits
dynamic memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Video Understanding
Dynamic Memory
Temporal Tree
Reasoning while Building Memory
Context Efficiency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Wei Chen
Wei Chen
HKUST
Computer VisionVision-Language
X
Xuanyu Zheng
Kling AI
Y
Yancheng Long
Kling AI
Haoyang Xu
Haoyang Xu
TianJin University
Optical Fiber Sensor
Kaiyu Jiang
Kaiyu Jiang
Kuaishou
MLLM
Bin Wen
Bin Wen
快手
MLLM
T
Tingting Gao
Kling AI
H
Han Li
Kling AI
L
Long Chen
HKUST