Beyond Offline A/B Testing: Context-Aware Agent Simulation for Recommender System Evaluation

📅 2026-01-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the persistent gap between offline evaluation and online performance in recommender systems, which stems largely from existing large language model (LLM)-based user simulators neglecting critical contextual factors such as time, location, and user intent. To bridge this gap, the authors propose ContextSim, a novel framework that integrates real-life contextual dynamics into LLM-powered user agents for the first time. ContextSim employs a life-simulation module to generate daily scenarios enriched with temporal, spatial, and motivational cues, and enforces both internal chain-of-thought reasoning and behavioral–trajectory consistency constraints to produce context-aware agents that more faithfully emulate human interactions. Experimental results demonstrate that interactions synthesized by ContextSim closely mirror real user behavior, yielding offline A/B test outcomes highly correlated with live metrics; recommendation strategies optimized using this framework significantly enhance actual user engagement.
📝 Abstract
Recommender systems are central to online services, enabling users to navigate through massive amounts of content across various domains. However, their evaluation remains challenging due to the disconnect between offline metrics and online performance. The emergence of Large Language Model-powered agents offers a promising solution, yet existing studies model users in isolation, neglecting the contextual factors such as time, location, and needs, which fundamentally shape human decision-making. In this paper, we introduce ContextSim, an LLM agent framework that simulates believable user proxies by anchoring interactions in daily life activities. Namely, a life simulation module generates scenarios specifying when, where, and why users engage with recommendations. To align preferences with genuine humans, we model agents'internal thoughts and enforce consistency at both the action and trajectory levels. Experiments across domains show our method generates interactions more closely aligned with human behavior than prior work. We further validate our approach through offline A/B testing correlation and show that RS parameters optimized using ContextSim yield improved real-world engagement.
Problem

Research questions and friction points this paper is trying to address.

Recommender System Evaluation
Offline A/B Testing
Context-Aware Simulation
User Modeling
LLM Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

ContextSim
LLM agent simulation
context-aware evaluation
recommender system
behavioral consistency
🔎 Similar Papers
No similar papers found.
N
Nicolas Bougie
Woven by Toyota
G
Gian Maria Marconi
Woven by Toyota
X
Xiaotong Ye
Woven by Toyota
Narimasa Watanabe
Narimasa Watanabe
Woven by Toyota