VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of inefficient real-time interaction and poor simulation-to-reality generalization in autonomous agent navigation and visual evidence grounding within open-web video research. We propose a scalable training framework that enables effective sim-to-real transfer through three key components: a multi-hop task synthesis pipeline that filters textual shortcuts, a high-speed local video simulator preserving search interaction dynamics, and a Retrieval Domain Randomization GRPO (RDR-GRPO) algorithm. Experiments based on the Qwen3.5-4B model demonstrate that our approach achieves 40.48% accuracy on the Video-BrowseComp benchmark, performing comparably to Gemini-3-Flash-Preview while reducing API token consumption by 74.9%.
📝 Abstract
Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and ground answers in visual evidence. Training such agents at scale is challenging as live video interaction is slow and unreliable, whereas fixed local simulation can induce retrieval-specific shortcuts that fail to transfer to the open web. We introduce VideoResearchAgent, a scalable training framework to address these challenges. First, we introduce controllable task synthesis pipeline to synthesize multi-hop research tasks from timestamped visual evidence while filtering text-only shortcuts. Second, we build a field-aligned local video simulator that preserves deployment-facing search and watch interactions while accelerating video search by a factor of 34.5-64.6. Third, we introduce Retrieval-Domain-Randomized GRPO (RDR-GRPO), which diversifies candidate rankings, distractors, metadata, and result structure during training to reduce overfitting to simulated retrieval. On Video-BrowseComp, the VideoResearchAgent trained using Qwen3.5-4B achieves 40.48% accuracy, comparable to Gemini-3-Flash-Preview, while reducing cumulative API-token consumption by 74.9% relative to the untrained model. Together, these results establish an accurate and efficient training recipe for open-web video research.
Problem

Research questions and friction points this paper is trying to address.

open-web video research
video reasoning agent
sim-to-real transfer
task synthesis
retrieval overfitting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Web Video Research
Task Synthesis
Sim-to-Real RL
RDR-GRPO
Video Simulator
🔎 Similar Papers
Yuhang Zhou
Yuhang Zhou
Fudan University
Natural Language ProcessingMultimodal LearningAgent
F
Fei Li
Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI
Y
Yuxi Wu
Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI
Bin Zhu
Bin Zhu
Assistant Professor, Singapore Management University
MultimediaComputer Vision
Jingjing Chen
Jingjing Chen
Fudan University
MultimediaComputer VisionMachine LearningPattern recognition