Probe to Act: Elevating Browser-Use Agent via Active Visual Probing

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of aligning DOM and visual information, as well as context loss during long interactions in browser agents. We propose P2A, a proactive probing framework that introduces a novel "probing-as-action" paradigm. By shifting static alignment to on-demand dynamic verification, P2A maps DOM elements to pixel-level evidence in real time during decision-making and constructs an evidence-based long-term memory. Furthermore, it integrates prompt engineering, cold-start synthetic data, and bootstrapped supervised fine-tuning, ensuring compatibility with both proprietary and open-source models. Experimental results demonstrate that our approach matches full-observation performance with only 1.2× the context cost while significantly improving task success rates, achieving, for instance, a 7.1% gain with Gemini.
📝 Abstract
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve dense DOM-pixel alignment before every operation. We introduce Probe to Act (P2A), an active probing framework for the browser-agent loop that moves this alignment into decision time. P2A addresses an asymmetric bridge between symbolic DOM hypotheses and screenshot layout by rendering on-demand symbolic DOM structure back into pixels. Before committing a state-changing browser operation, the agent can issue lightweight probes to translate DOM handles into pixel evidence, map screen regions back to DOM candidates, register visual-only targets, and commit verified notes. These interleaved processes naturally produce evidence-based memory: only probed, acted-on, or explicitly committed observations are kept across steps, preserving only decision-critical evidence in long-horizon contexts. P2A can be used as a prompting strategy for proprietary models under the standard DOM+SoM interface, and can be distilled into open-weight models through cold-start synthesis and self-bootstrapped SFT. Across three browser-use benchmarks, P2A shows clear gains on task success rate for both proprietary and fine-tuned models; on VisualWebArena, for example, it improves Gemini-3-Pro from 54.1% to 61.2% and Qwen3-VL-8B from 24.6% to 32.9%, while matching the costly full-observation history ($\sim$3$\times$) at only $\sim$1.2$\times$ the peak retained input context of action-only history.
Problem

Research questions and friction points this paper is trying to address.

Browser-use agent
DOM-pixel alignment
Visual probing
Long-horizon context
Set-of-Marks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Active Visual Probing
Browser-Use Agent
DOM-Pixel Alignment
Evidence-Based Memory
Model Distillation