Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba's Customer Service Operations

📅 2026-05-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how human intervention influences service outcomes when autonomous AI fails in customer service contexts, with a focus on the cognitive and emotional dimensions of intervention efficacy. Drawing on a randomized field experiment conducted on Alibaba’s Taobao platform, the research compares customer service agents’ performance in chat tasks with and without AI assistance. It provides the first empirical evidence that intervention effectiveness depends on the type of AI failure (technical versus emotional), the timing of intervention, and the level of subsequent human effort. Findings reveal that while AI reduces average chat duration, it lowers satisfaction ratings for conversations it handles. Human intervention effectively preserves service quality in technical failures but shows limited efficacy in emotional failures. Early intervention not only enhances agents’ subsequent effort but also significantly improves service outcomes for AI-unmanageable chats, generating positive spillover effects through workflow adaptation.
📝 Abstract
Agentic AI systems that autonomously perform service tasks are entering customer service operations. However, limited evidence exists on how human interventions shape service outcomes when agentic AI failures create both cognitive and emotional consequences. We study this issue through a randomized field experiment on Alibaba's Taobao platform. Workers in the treatment condition supervised an agentic AI system that resolved AI-eligible chats while continuing to handle AI-ineligible chats, whereas control workers resolved all chats without agentic AI. The findings show that AI deployment reduces average chat duration and has limited effects on retrial rates, but substantially lowers ratings for AI-eligible chats. Moreover, human intervention effectiveness in AI-eligible chats depends on the nature of AI failure, post-escalation intervention effort, and intervention timing. Human intervention preserves service quality in algorithm-triggered technical escalations, i.e., unresolved customer issues beyond the AI's capability, but is less effective in algorithm-triggered emotional escalations, i.e., where customers express frustration or dissatisfaction. These differences are partly explained by variation in workers' post-escalation intervention effort across escalation types. In algorithm-triggered emotional escalations, workers showed lower engagement: they sent fewer messages, contributed a smaller share of total chat rounds, and showed less proactivity in information seeking and solution provision. We further find that early intervention is essential for sustaining high post-escalation intervention effort. Finally, we document a positive spillover effect on AI-ineligible chats, as treated workers adapted their multitasking workflow to devote greater attention to these chats. These findings offer implications for human-in-the-loop process design in human-AI collaboration systems.
Problem

Research questions and friction points this paper is trying to address.

Agentic AI
Human-in-the-Loop
Customer Service
AI Failure
Human Intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic AI
Human-in-the-Loop
Field Experiment
AI Failure Types
Intervention Timing