🤖 AI Summary
This study addresses the limitations of existing multi-turn dialogue evaluation frameworks, which often overlook user heterogeneity and rely on monolithic reward models. To overcome these issues, this work proposes a dual-perspective personalized evaluation paradigm. Methodologically, we construct a benchmark comprising thousands of instances that integrates interpretable FACTORS-based user profiles, personalized reward models, and the Sim4Eval simulator. Furthermore, a meta-evaluation mechanism is introduced to transcend the constraints of conventional binary preference prediction. The research yields seven key findings that substantiate the critical value of user modeling in dialogue assessment. Ultimately, this work establishes a novel framework for human-centric evaluation of conversational systems, offering actionable insights to facilitate downstream model optimization.
📝 Abstract
Evaluating user experience (UX) with automated computational methods has gained increasing attention, supported by empirical evidence from UXBench. However, binary preference prediction provides limited insight, while relying on a single user-agnostic reward model overlooks the inherent heterogeneity of users, whose expectations can differ substantially. In this paper, we present UXBench Pro, comprising 1{,}000 test instances derived from real user interactions across 12 task scenarios and 82 domains. Each instance is paired with a FACTORS user profile that characterizes the user through seven interpretable behavioral facets, differentiating user groups. To provide richer evaluation insights, we introduce a dual-perspective paradigm that combines a personalized User Reward Model (URM) for third-person judgment with Sim4Eval, a user simulator that enables multi-turn interactions and provides first-person evaluation across four cognitive state dimensions. To assess the reliability of these based evaluators, we further introduce two meta-benchmarks, URMBench and USimBench, that evaluate how faithfully they reproduce real human preferences and behaviors. Extensive experiments reveal seven key findings that highlight the importance of user modeling and multi-perspective evaluation, offering a fresh perspective on user-centric benchmarking and motivating personalized model optimization.