🤖 AI Summary
This study addresses the challenges of inefficient real-time interaction and poor simulation-to-reality generalization in autonomous agent navigation and visual evidence grounding within open-web video research. We propose a scalable training framework that enables effective sim-to-real transfer through three key components: a multi-hop task synthesis pipeline that filters textual shortcuts, a high-speed local video simulator preserving search interaction dynamics, and a Retrieval Domain Randomization GRPO (RDR-GRPO) algorithm. Experiments based on the Qwen3.5-4B model demonstrate that our approach achieves 40.48% accuracy on the Video-BrowseComp benchmark, performing comparably to Gemini-3-Flash-Preview while reducing API token consumption by 74.9%.
📝 Abstract
Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.