OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck in multimodal deep research where visual anchors, entity relationships, and factual dependencies are difficult to model jointly. To overcome this challenge, we propose an agent framework centered on the Visual Grounded Evidence Graph (VGEG). Methodologically, we design a VGEG-based evidence-aware reward mechanism that unifies data construction and evaluation pipelines, while jointly optimizing supervised fine-tuning and reinforcement learning strategies with a dedicated VGEG data engine and fine-grained benchmarking techniques. Experimental results demonstrate that our approach surpasses Qwen3-VL by over 17 percentage points on both custom-built and general-purpose benchmarks, significantly enhancing deep retrieval and factual synthesis capabilities across images and videos.
📝 Abstract
Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition. A key challenge is to preserve the dependencies linking localized visual anchors, entity relations, source-supported facts, and answer-producing operations. We introduce OneSearch-VL, a unified agent centered on the Visually Grounded Evidence Graph (VGEG), which encodes these dependencies as a shared task-level reference for data construction, process supervision, and operation-level evaluation. Our VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. Using these data, we assemble OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K for SFT and RL, respectively. We further derive the Evidence-aware Visual-Grounded Rubric reward (EVGR) from VGEG annotations to supervise evidence traceability and visual grounding during RL. For fine-grained evaluation, we construct OneSearch-MI-Bench and OneSearch-Video-Bench, organizing questions by the research operations encoded in their VGEGs. Experiments show that OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, while also achieving substantial gains across 7 image benchmarks and VideoDR. Project repository: https://github.com/appletea233/OneSearch-VL
Problem

Research questions and friction points this paper is trying to address.

multimodal deep research
visual grounding
evidence graph
image and video understanding
research agent
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visually Grounded Evidence Graph
Multimodal Deep Research Agent
Evidence-aware Reward
Process Supervision
Unified Visual Reasoning
🔎 Similar Papers