Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the insufficient test coverage for customer service AI agents in regulated industries and the risk that online experimentation poses to user trust. To overcome these challenges, this work proposes a hypothesis-driven simulation evaluation workflow. By leveraging the Snowglobe simulator and synthetic data generation techniques, the method enables multi-turn interaction simulation without invoking production backends. Furthermore, end-to-end binary evaluators are employed to pre-screen candidate agents for safety, facilitating large-scale configuration exploration with zero user risk. Experimental results demonstrate that simulated scores align closely with production environments. In A/B testing, the proposed approach yields a 36.69-point improvement in tNPS and an 8.82% increase in self-service rates, thereby achieving safe, iterative development of AI agents within regulated settings.
πŸ“ Abstract
Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.
Problem

Research questions and friction points this paper is trying to address.

Customer Experience AI Agents
Simulation-based Screening
Regulated Industries
Pre-deployment Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simulation-based Screening
Customer Experience AI Agents
Synthetic Customers
Agentic Workflows
Large Language Models
πŸ”Ž Similar Papers
2024-10-06Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)Citations: 13
πŸ’Ό Related Jobs
No related jobs found.
E
Edesio Alcoba
Nubank
K
Kevin Rossell
Nubank
A
Aman Gupta
Nubank
Shao Tang
Shao Tang
Linkedin
LLM Post-TrainingAgentOptimization
Jiwoo Hong
Jiwoo Hong
KAIST AI
Artificial IntelligenceNatural Language Processing
P
Pabel Carrillo-Mendoza
Nubank
W
Wanderson ConceiΓ§Γ£o Ferreira
Nubank
A
Alvaro Tedeschi
Nubank
Z
Zayd Simjee
Guardrails AI
Shreya Rajpal
Shreya Rajpal
University of Illinois, Urbana-Champaign
Machine LearningCrowdsourcingComputer Vision
B
Bruno Finardi Hime
Nubank
C
Christian Sousa
Nubank
L
Luis Moneda
Nubank
H
Herbert Fei
Nubank
D
Daniel Silva
Nubank
R
Rohan Ramanath
Nubank