VAMR: Multi-Question Agentic Reasoning for Efficient Long-Form Video Understanding

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency and redundant exploration caused by isolated agent reasoning in multi-question answering over long videos. To overcome this, we propose a collaborative multi-question framework based on shared tool trajectories. Methodologically, the approach constructs unified video understanding through hierarchical memory and condition awareness, and introduces the Question Horizon Policy Optimization (QHPO) algorithm. By integrating supervised fine-tuning initialization with question-level critics, QHPO enables joint reinforcement learning across shared trajectories for questions at varying stages of progress. Experimental results demonstrate that our method achieves state-of-the-art accuracy on benchmarks such as LVBench, while reducing inference rounds by 85.9% and frame processing volume by 61.4%, thereby substantially improving efficiency in long video understanding.
📝 Abstract
Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.
Problem

Research questions and friction points this paper is trying to address.

Long-form video understanding
Multi-question reasoning
Video agents
Shared tool-use trajectory
Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Question Reasoning
Shared Tool-Use Trajectory
Layered Multi-Question Memory
Question-Horizon Policy Optimization
Long-Form Video Understanding
🔎 Similar Papers
No similar papers found.
R
Runquan Gui
University of Science and Technology of China, Tencent
H
Hanzhu Chen
Tencent
Z
Zehao Wang
University of Science and Technology of China, Tencent
H
Hanzxin Zhu
University of Science and Technology of China, Tencent
Xin Li
Xin Li
University of Science and Technology of China
Data MiningArtificial IntelligenceNeuroscienceAI for Science
Zhibo Chen
Zhibo Chen
Professor@University of Science and Technology of China
Generative AIvisual signal representationvideo codingvideo analysis and processing