HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of erroneous reasoning paths in multi-turn visual search agents caused by reliance on outcome-only rewards, proposing a pioneering process reinforcement learning framework anchored in human search behavior to achieve a paradigm shift from outcome supervision to process supervision. Methodologically, this work integrates a manual annotation platform, a task-adaptive weight scorer, and a trajectory distillation-based process reward model to supervise search trajectories with fine-grained data, thereby optimizing high-resolution image question-answering performance. Experimental results demonstrate that the proposed framework significantly outperforms conventional outcome-oriented reinforcement learning approaches. Furthermore, applying process supervision during early training stages improves subsequent scaling efficiency by 6.7 times.
📝 Abstract
Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K+ human-annotated data with fine-grained behavioral signals. During training, a carefully designed judge scores each rollout with task-adaptive weights, anchored on the distilled trace of how a human annotator actually searched the same image. Extensive experiments show that HaPRL consistently outperforms outcome-based RL, and early-stage process supervision yields 6.7x more improvement in subsequent outcome-based scaling. Our results also demonstrate the importance of aligning model behavior with human process annotation signals, which offer new insight into the training of foundation models.
Problem

Research questions and friction points this paper is trying to address.

Visual Search Agent
Process Reinforcement Learning
Faulty Reasoning Paths
Outcome-based RL
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process Reinforcement Learning
Visual Search Agent
Human-Anchored
Process Supervision
Behavior Alignment
🔎 Similar Papers
No similar papers found.