CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing user simulators lack outcome calibration, making it difficult to faithfully reproduce the failure modes and success rates observed in real AI agent interactions. To address this limitation, this work proposes Calibration-free User Embedding (CUE), a training-free framework that encodes conversations and samples continuous embeddings to generate persona instructions, guiding large language models to simulate user behavior. This approach enables user-conditioned replay and alignment with aggregate metrics. Experiments on the τ²-Bench benchmark demonstrate that CUE more accurately reproduces real-world user failure modes and success rates compared to existing methods, significantly enhancing both the ecological validity of the simulation and its cross-domain generalization capability.
📝 Abstract
Recent benchmarks rely on user simulators to evaluate AI agents in multi-turn interaction. While existing simulation techniques demonstrate surface fidelity to human style and behavior, ecologically valid interactive benchmarking also requires alignment in when and how agents fail across simulated and real user populations. We find that existing simulators lack outcome calibration: agreement with observed success rates and failure patterns when real users interact with the same agent. We introduce Calibrated User Embeddings (CUE), a framework that both encodes observed sessions and samples continuous representations, then decodes them into persona commands to steer LLMs to act as user simulators without training. Through this, we evaluate user-conditioned replay of past sessions and aggregate metric agreement when sampling novel personas for the same tasks. On $τ^2$-Bench, CUEd simulators commit fewer simulator-attributed errors and more faithfully reproduce real-user agent failure modes, aggregate success rates, and outcomes for specific task-user pairs than other persona-based simulation methods. These gains coexist with competitive user fidelity as measured using metrics established in prior work. After being fit to mostly customer support interactions, the same CUEd simulators generalize to document creation, math tutoring, and casual conversation, and remain effective across different simulator LLMs without CUE retraining.
Problem

Research questions and friction points this paper is trying to address.

user simulators
multi-turn benchmarking
outcome calibration
failure patterns
AI agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Calibrated User Embeddings
User Simulators
Multi-Turn Benchmarking
Outcome Calibration
Persona Steering
🔎 Similar Papers
No similar papers found.