From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the shallow shortcut bias in web process reward models (PRMs) caused by the scarcity of contrastive samples. To mitigate this, we propose SURFPRM, a framework that constructs an interactive element graph to systematically synthesize negative actions across spatial, temporal, and spatiotemporal dimensions. By integrating multi-strategy sampling, SURFPRM generates high-quality preference data through environment-grounded negative action proposals. Experimental results demonstrate that the framework increases the proportion of grounded minimal contrastive pairs from 24.19% to 74.60%. Furthermore, it surpasses existing baselines across multiple benchmarks, achieving performance comparable to proprietary large language models, and improves the success rate of the GPT-4o series on complex web tasks by over 12%.
📝 Abstract
Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded Minimal Contrastive Pairs (GMCPs)-where competing candidates target genuine on-page elements with identical action types. In representative baselines preference data (namely, WebArbiter), GMCPs account for merely 24.19%, biasing PRMs during training to rely on shallow shortcuts (such as element hallucinations and action type mismatches) rather than acquiring genuine contextual decision semantics. To address these challenges, we propose SURFPRM, a graph-guided process preference synthesis framework for comparative Web PRMs. SURFPRM structures web demonstrations into a persistent Interaction Element Graph that acts as an environment-grounded negative action proposal mechanism, systematically synthesizing contrastive negative actions across spatial, temporal, and spatiotemporal confusion axes. This elevates the GMCP proportion from 24.19% to 74.60%, producing the curated SURFPRM-DATA dataset. Across six open-source backbones (3B to 9B parameters), PRMs trained on SURFPRM-DATA outperform baseline-trained models on average on WEBPRMBENCH and rival leading proprietary LLMs. In downstream reward-guided trajectory search on WEBARENA-LITE, SURFPRM provides step-level guidance for both GPT-4o (+14.21%) and GPT-4o-mini (+12.83%) policies, yielding substantial improvements in complex web task success rates.
Problem

Research questions and friction points this paper is trying to address.

Process Reward Models
Preference Data Synthesis
Web Agents
Grounded Minimal Contrastive Pairs
Shortcut Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Process Reward Models
Interaction Element Graph
Preference Synthesis
Grounded Minimal Contrastive Pairs
Web Agents
🔎 Similar Papers
No similar papers found.