🤖 AI Summary
This study addresses the challenges of uncontrollable difficulty and high refresh costs in browser agent benchmarks by proposing a method that reframes task difficulty as a programmable property. Through controlled environmental interventions, deterministic, detectable, and recoverable state perturbations are introduced across different layers of the web stack while preserving user instructions and success criteria unchanged. These perturbations are further combined with cognitive primitive annotations to construct challenging tasks. Experiments conducted on self-hosted websites using an automated evaluation framework demonstrate that this approach reduces the average pass rate of agents by 22.9% and reveals that 75% of failures stem from belief errors. The associated code and datasets have been made publicly available.
📝 Abstract
As browser-use agents improve, benchmarks keep pace by collecting new tasks, websites, and applications, often making tasks longer or more novel. This makes difficulty expensive to refresh and difficult to control: when many aspects change at once, it is unclear what actually makes a task challenging. We instead construct challenging instances from tasks agents already solve, turning difficulty into a programmable property of the environment. BreakingWeb pairs every base task with an intervention condition that preserves the user instruction, latent target, and backend success criterion while changing the environment at different web stack layers. Each intervention is deterministic, detectable, and recoverable, and is annotated with the cognitive primitive it primarily loads. The benchmark contains 519 clean/intervention task pairs across seven self-hosted websites and 29 intervention families, all graded against outcomes. We evaluate six strong browser-use agents, three GUI-only agents that see only screenshots, and humans. The construction is effective: interventions cut agent pass rate by 22.9% on average and overturn nearly half of the tasks each agent solves cleanly, whereas humans lose 10.0% on a first attempt and 5.7% after one familiarisation attempt. The dominant failure is belief failure: 75% of the six agents' failures end with a declared success although the required change never happened. Our code, data and environment are publicly available at www.breakingweb.app.