Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing web agent evaluations rely solely on final outcomes, lacking fine-grained scrutiny of failure trajectories. This study addresses this limitation by conducting a manual audit on WebArena-Lite tasks to correct automatic evaluation errors and analyze critical error patterns. Building upon these insights, we propose a Memory and Analysis Support Mechanism (MASM) alongside guiding prompts to enable trajectory-aware, fine-grained evaluation. Experiments utilizing GPT-5.5 and Qwen3.5 demonstrate that human review recovers approximately 8% in success rate, while MASM significantly enhances the performance of untrained models. Furthermore, our analysis systematically reveals common failure modes exhibited by web agents, offering valuable diagnostic insights for future agent development.
📝 Abstract
Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.
Problem

Research questions and friction points this paper is trying to address.

Web Agent Evaluation
Trajectory Analysis
Human Verification
Automatic Evaluator
Innovation

Methods, ideas, or system contributions that make the work stand out.

Web Agent Evaluation
Trajectory Analysis
Memory and Analysis Support Mechanism (MASM)
Human-grounded Verification
Guide Text
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chengguang Gan
Techtouch, Inc.
Z
Zimeng He
Institute of Science Tokyo
Y
Yoshihiro Tsujii
Techtouch, Inc.
K
Ken-ichiro Kobayashi
Techtouch, Inc.
H
Hiroki Itoh
Techtouch, Inc.
Kotaro Funakoshi
Kotaro Funakoshi
Tokyo Institute of Technology
Multimodal Dialogue SystemsHuman-Machine InteractionComputational Linguistics