π€ AI Summary
This study addresses the limitation of existing user simulators that merely mimic surface-level styles while failing to replicate intent evolution and behavioral consistency in real interactions. To this end, we propose TRACER, a multi-turn user simulator that explicitly models intent dynamics and aligns with authentic trajectories through a two-stage training paradigm combining supervised fine-tuning with multi-turn reinforcement learning. We introduce hierarchical structuring, trajectory-level rewards, and bias-aware advantage modulation to effectively mitigate reward sparsity and credit assignment challenges in extended dialogues. Additionally, a dynamic marketing benchmark is constructed to evaluate the persuasive capabilities of large language models. Experimental results demonstrate that TRACER-7B surpasses baselines by 11.4% in conversion rate F1 score and achieves near-random performance in Turing tests, while revealing that high response quality does not necessarily yield higher conversion rates.
π Abstract
Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.