AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决多跳视频问答问题,提出AgentVidBench基准,评估MLLM代理的空间、时间及因果推理能力。
📝 Abstract
Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic capabilities. To bridge this gap, we introduce AgentVidBench, a multi-hop video question answering benchmark focused on evaluating the spatial, temporal, and causal reasoning capabilities of MLLM agents. Beyond standard question-answer pairs, AgentVidBench provides step-by-step solution traces to support trajectory evaluation that assesses whether agents explicitly acquire the evidence needed to justify their answers. Experiments with 12 proprietary and open-source MLLMs show that single-turn performance remains limited on AgentVidBench, while integrating these models into state-of-the-art agentic workflows generally improves performance with respect to both accuracy and trajectory scores. We further present a simple yet effective agentic strategy that serves as a competitive baseline on AgentVidBench, establishing our benchmark as a holistic testbed for future research on agentic video understanding. Code and datasets are available at https://github.com/krafton-ai/agentvidbench and https://huggingface.co/datasets/agentvidbench/agentvidbench.
Problem

Research questions and friction points this paper is trying to address.

Multi-Hop Reasoning
Video Understanding
Multimodal Large Language Models
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Hop Video Question Answering
Spatial and Temporal Reasoning
Causal Reasoning
Step-by-Step Solution Traces
Agentic Workflows
🔎 Similar Papers
No similar papers found.
S
Seoyeon An
KRAFTON
H
Hyeonseo Jang
KRAFTON
M
Minsu Kim
KRAFTON
Chanho Lee
Chanho Lee
KAIST
Y
Younghan Park
KRAFTON
Kangwook Lee
Kangwook Lee
University of Wisconsin-Madison, KRAFTON AI
Machine LearningInformation Theory