Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of acquiring evidence from local small objects in high-resolution visual question answering by proposing the VPS framework. This method introduces a novel search mechanism that combines parallel tile inspection with adaptive zooming, wherein a primary agent orchestrates sub-agents to retrieve image tiles in parallel while adaptively zooming to precisely localize critical regions. Furthermore, a prompt-free verification module and a role-specific GRPO supervision pipeline are designed to enable decoupled training of the controller and reader. Experimental results demonstrate that the proposed framework outperforms pure zoom-based search in 14 out of 15 comparisons, achieving improvements of up to 8.0 points. It also yields a 4.17-point gain on HR-Bench 4K while significantly reducing tool invocation overhead.
📝 Abstract
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.
Problem

Research questions and friction points this paper is trying to address.

Visual Question Answering
High-Resolution Images
Local Evidence Acquisition
Sequential Zooming
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Parallel Search
High-Resolution VQA
Adaptive Zoom
Role-specific GRPO
Multi-agent Framework
🔎 Similar Papers
No similar papers found.
X
Xijia Tao
The University of Hong Kong
Y
Yihua Teng
Huawei Research
Xinyu Fu
Xinyu Fu
Hong Kong Research Center, Huawei
Large Language ModelsMLLMAgentsHeterogeneous Graphs
C
Cheng Gong
City University of Hong Kong
Z
Ziru Liu
Huawei Research
X
Xudong Xie
Huazhong University of Science and Technology
R
Rui Liu
Huawei Research
Lingpeng Kong
Lingpeng Kong
Google DeepMind, The University of Hong Kong
Natural Language ProcessingMachine Learning