🤖 AI Summary
This study addresses the limitations of conventional fixed-inference paradigms in video understanding when handling long temporal sequences, multimodal inputs, and sparse evidence by systematically reviewing agent-driven video understanding frameworks. These frameworks enable active information acquisition and verification through adaptive state construction and action selection. The paper introduces a novel "challenge-design" taxonomy that precisely maps core challenges—including context bottlenecks, evidence sparsity, and temporal causality—to corresponding mechanisms such as hierarchical evidence memory, active retrieval, state tracking, and multi-agent collaboration. Furthermore, it surveys prevailing state-space paradigms, learning strategies, and evaluation benchmarks, offering actionable insights to guide future research toward native temporal modeling in video understanding.
📝 Abstract
As large language models (LLMs) become capable of processing increasingly diverse modalities and longer temporal contexts, an emerging line of work is moving beyond fixed video-language inference toward agentic systems that actively decide what information to inspect, retain, verify, and act upon. This survey reviews video understanding agents: systems that use video as the primary information source and solve understanding tasks through adaptive state construction and action selection. We first formalize an agent loop for video understanding, then address a central question: why do agents matter for video understanding? To answer this, we organize the literature through a challenge-to-design taxonomy, linking context bottlenecks to hierarchical evidence memory, evidence sparsity to active evidence acquisition, temporal causality to state and process tracking, and multimodal ambiguity to role-specialized coordination. We further review state space paradigms, learning paradigms, supervision signals, benchmarks, and evaluation protocols. Finally, we identify open directions toward agentic-native temporal modeling and video-native agents. Project page: https://github.com/DXY0711/Awesome-Agentic-Video-Understanding