🤖 AI Summary
This study addresses the inefficiency and redundant exploration caused by isolated agent reasoning in multi-question answering over long videos. To overcome this, we propose a collaborative multi-question framework based on shared tool trajectories. Methodologically, the approach constructs unified video understanding through hierarchical memory and condition awareness, and introduces the Question Horizon Policy Optimization (QHPO) algorithm. By integrating supervised fine-tuning initialization with question-level critics, QHPO enables joint reinforcement learning across shared trajectories for questions at varying stages of progress. Experimental results demonstrate that our method achieves state-of-the-art accuracy on benchmarks such as LVBench, while reducing inference rounds by 85.9% and frame processing volume by 61.4%, thereby substantially improving efficiency in long video understanding.
📝 Abstract
Long-form video understanding often involves multiple questions about different aspects of the same recording. Yet existing video agents typically process each question through an isolated tool-use trajectory. This repeatedly restarts video exploration and memory construction, missing opportunities to acquire evidence jointly and progressively build a shared understanding that supports the complete question set. We introduce \textbf{VAMR} (\textbf{V}ideo \textbf{A}gent for \textbf{M}ulti-Question \textbf{R}easoning), which coordinates all questions about a video through one shared tool-use trajectory. At each round, a persistent policy model can invoke tools for one or more unresolved questions and submit answers for questions with sufficient evidence. Question-conditioned visual perception retrieves fine-grained clues for several questions in one call, while layered multi-question memory integrates reusable context into a shared video story and preserves separate evidence for individual questions. After supervised fine-tuning initializes this interaction protocol, we propose question-horizon policy optimization (\qhpo) to optimize shared trajectories in which questions progress and finish at different rounds. Specifically, a question-level critic estimates the value of each active question, while round alignment maps each question advantage to the rounds that directly serve it before the aligned advantages are aggregated to optimize the shared actor. Across LVBench, Video-Holmes, and LongVideoBench, VAMR achieves the highest accuracy overall and the fewest reasoning rounds among iterative methods. On LVBench, it reaches 62.1\% accuracy, exceeding VideoARM by \textbf{4.3} points while reducing reasoning rounds and processed frames by \textbf{85.9\%} and \textbf{61.4\%}.