AdaptArena: Evaluating Test-Time Personalization of Web Agents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge faced by web agents in inferring personalized preferences from implicit signals under scenarios involving ambiguous user intents and heterogeneous preferences. To this end, it systematically defines and constructs the first test-time personalization benchmark comprising 480 tasks, and proposes a retrieval-augmented generation (RAG) framework to evaluate agents’ capacity to infer implicit preferences from historical interaction trajectories. The work reveals a substantial performance gap between reasoning and execution. Experimental results demonstrate that while oracle performance achieves an 82.92% success rate, state-of-the-art large language model agents attain at most 15.62%, underscoring the fundamental challenges inherent in implicit preference inference and action grounding.
📝 Abstract
Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference inference. AdaptArena consists of 480 tasks, featuring both single-preference and double-preference scenarios. Each evaluation task must be solved by retrieving and leveraging the most relevant historical user trajectory that implicitly encodes the target preference. In addition, we introduce AdaptiveAgent, a retrieval-based framework for standardized evaluation of implicit preference inference. Experiments reveal a substantial performance gap: while oracle agents with access to ground-truth preferences achieve an 82.92% success rate, the evaluated LLM agents using our framework reach at most 15.62%. Furthermore, we find that correctly inferring user preferences is necessary but not sufficient for task success, as execution failures in downstream web interactions remain a significant bottleneck even when agents align with the target preference. These findings highlight implicit preference inference and robust action grounding as key challenges for deploying reliable, user-facing web agents. We release our code: https://github.com/McGill-NLP/web-agents-test-time-adaptations
Problem

Research questions and friction points this paper is trying to address.

test-time personalization
web agents
implicit preference inference
LLM agents
robust action grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Personalization
Implicit Preference Inference
Web Agents
Retrieval-based Framework
Action Grounding