Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency

πŸ“… 2026-09-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of existing user simulators that merely mimic surface-level styles while failing to replicate intent evolution and behavioral consistency in real interactions. To this end, we propose TRACER, a multi-turn user simulator that explicitly models intent dynamics and aligns with authentic trajectories through a two-stage training paradigm combining supervised fine-tuning with multi-turn reinforcement learning. We introduce hierarchical structuring, trajectory-level rewards, and bias-aware advantage modulation to effectively mitigate reward sparsity and credit assignment challenges in extended dialogues. Additionally, a dynamic marketing benchmark is constructed to evaluate the persuasive capabilities of large language models. Experimental results demonstrate that TRACER-7B surpasses baselines by 11.4% in conversion rate F1 score and achieves near-random performance in Turing tests, while revealing that high response quality does not necessarily yield higher conversion rates.
πŸ“ Abstract
Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.
Problem

Research questions and friction points this paper is trying to address.

user simulation
behavioral consistency
multi-turn dialogue
intent evolution
LLM evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-turn user simulation
reinforcement learning
intent evolution
hierarchical rewards
trajectory alignment
πŸ’Ό Related Jobs
No related jobs found.
G
Geng Chen
Peking University
R
Ruotong Pan
Kuaishou Technology
Z
Zhirui Yang
Kuaishou Technology
Q
Qiqi He
Xi’an Jiaotong University
Jiawei Chen
Jiawei Chen
Institute of Software, Chinese Academy of Sciences
Large Language Models
Z
Zhang Yunfei
Kuaishou Technology
C
Chongyuan Chen
Kuaishou Technology
M
Minxuan Lv
Institute of Information Engineering
Z
Zheng Yang
Beijing University of Posts and Telecommunications
W
Win-Bin Huang
Peking University
X
Xiangyu Wu
Kuaishou Technology
W
Wenwu Ou
Kuaishou Technology