On Evaluating and Improving Conversational Agents in Production

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of non-replayable logs, systemic stochastic fluctuations, and metrics that fail to localize behavioral changes in the offline evaluation of large-scale multi-agent shopping assistants. To overcome these issues, we propose a systematic evaluation and improvement framework. Methodologically, it employs grounded user simulation to generate dialogues as an alternative to log replay, utilizes repeated baseline runs with paired percentile bootstrapping to mitigate noise, and introduces fixed-scenario assertion-based mechanisms for hypothesis-driven fine-tuning verification. In practice, this framework successfully localized a product carousel malfunction, effectively distinguished genuine improvements from runtime variance, and uncovered missing evidence alongside configuration failures, thereby substantially enhancing debugging efficiency in production environments.
📝 Abstract
We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a modified system, because a different response changes every turn that follows. (ii) The unchanged system itself varies from run to run. Its LLM components are stochastic, and in product search the available products, their prices, and the customer's personalization signals change. (iii) Aggregate quality scores combine distinct behaviors, so they show that quality has changed but not which behavior caused the change. Our framework addresses each obstacle in turn. For a reported behavior, an Evaluation Harness generates targeted assertions and a fixed cohort of customer scenarios. It then reproduces the behavior in a local instance of the assistant through grounded user simulation. Instead of replaying the log, the simulator writes new customer turns conditioned on the recorded messages and context. Repeated runs of the unchanged system form a stored baseline. An Improvement Orchestrator turns the assertion results into hypotheses, implements each as an isolated modification, and compares it with the baseline using paired percentile bootstrap intervals over scenario-level differences. When an investigation ends, the harness may propose revisions to future evaluations, subject to human approval and without altering past decisions. We report production investigations with this framework. Assertion profiles showed which positions of a product carousel a failure affected, and repeated runs distinguished a real improvement from run-to-run fluctuation. Audits of the evaluation itself found a judge that lacked the evidence it needed and a model setting that was configured but not applied.
Problem

Research questions and friction points this paper is trying to address.

conversational agents
offline evaluation
multi-agent system
production environment
LLM stochasticity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conversational Agents Evaluation
Grounded User Simulation
Paired Percentile Bootstrap
Multi-Agent System
Improvement Orchestrator
💼 Related Jobs
No related jobs found.
Kasra Hosseini
Kasra Hosseini
Zalando SE, Berlin, Germany
W
Wen-Sen Cheng
Zalando SE, Berlin, Germany
M
Marco-Andrea Buchmann
Zalando, Switzerland
E
Emir Mulabegovic
Zalando SE, Berlin, Germany
Weiwei Cheng
Weiwei Cheng
Zalando SE, Berlin, Germany