Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses two critical limitations in existing video-stream agents: modality bias—manifested as an overreliance on text-based search—and parametric knowledge leakage due to dependence on internal memory. To overcome these challenges, the authors propose a multimodal deep research agent tailored for continuous video streams, which decouples perception from exploration and incorporates a staged tool-unlocking mechanism that compels the model to invoke external tools grounded in cross-frame visual understanding. A two-stage training paradigm—combining supervised fine-tuning with Group Relative Policy Optimization—is introduced to surpass the performance ceiling of conventional imitation learning. The study also establishes Video-DR-Bench, the first human-AI collaborative benchmark for video deep research. Evaluated on this benchmark, the proposed Video-DeepResearch-35B-A3B achieves an average accuracy of 64.0%, substantially outperforming Claude-4.5-Sonnet (59.0%), GPT-5 (52.5%), and Gemini 2.5 Pro (57.5%).
📝 Abstract
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.
Problem

Research questions and friction points this paper is trying to address.

modality bias
parametric knowledge leakage
multimodal agents
video streams
visual grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

video grounding
multimodal agent
tool-augmented reasoning
GRPO
perception-exploration decoupling
🔎 Similar Papers