Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations

πŸ“… 2026-02-08
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the impact of human-AI collaboration involving generative AI on service performance in e-commerce customer support. Conducted as a large-scale randomized controlled trial in Alibaba’s after-sales service setting, the intervention involved AI automatically generating diagnostic suggestions and solution proposals only at the initial stage of each conversation, with human agents retaining full discretion over whether to adopt them. The results demonstrate that AI assistance significantly improves average response speed and customer-perceived satisfaction while effectively narrowing the performance gap between low- and high-performing agents. However, it concurrently introduces interference for top-tier agents, leading to slower response times and higher rates of immediate customer retries. This work provides the first empirical evidence of the heterogeneous effects of generative AI in real-world service environments, offering critical insights for the design of AI-augmented workflows.
πŸ“ Abstract
In collaboration with Alibaba, this study leverages a large-scale field experiment to assess the impact of a generative AI assistant on worker performance in e-commerce after-sales service. Human agents providing digital chat support were randomly assigned with access to a gen AI assistant that offered two core functions: diagnosis of customer issues and solution proposals, presented as text messages. Agents retained discretion to adopt, modify, or disregard AI-generated messages. To evaluate gen AI's impact, we estimate both the intention-to-treat (ITT) effect of gen AI access and the local average treatment effect (LATE) of gen AI usage. Results show that gen AI significantly improved service speed, measured by issue identification time and chat duration. Gen AI also improved subjective service quality reflected in customer ratings and dissatisfaction rates, but it had no significant effect on objective service quality indicated by customer retrial rates. The performance improvements stemmed not only from automation but also from changes in the dynamics of agent-customer interactions: agent communication became more informative and efficient, while customers experienced reduced communication burdens. Low performers achieved the greatest improvements in both service speed and quality, narrowing the performance gap. In contrast, top-performing agents showed little improvement in service speed but experienced declines in both subjective and objective service quality. Evidence suggests that this decline results from increased multitasking tendency, proxied by longer shift-away times across concurrent chats, which slowed customer responses and raised abandonment and retrial rates. These findings suggest that gen AI reshapes work, demanding tailored deployment strategies.
Problem

Research questions and friction points this paper is trying to address.

generative AI
customer service
field experiment
service quality
performance heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

generative AI
field experiment
human-AI collaboration
performance heterogeneity
service operations
πŸ”Ž Similar Papers
Xiao Ni
Xiao Ni
Head of Biometrics, Sarepta Therapeutics
Bayesian statisticsvariable selectionmixed modelsmachine learningrare diseases
Y
Yiwei Wang
Zhejiang University, International Business School, Zhejiang, China
T
Tianjun Feng
Fudan University, School of Management, Shanghai, China
L
Lauren Xiaoyan Lu
Dartmouth College, Tuck School of Business, Hanover, NH 03755
Yitong Wang
Yitong Wang
ByteDance Inc.
computer vision
C
Congyi Zhou
Alibaba Group Inc.