π€ AI Summary
This study addresses the limitations of existing e-commerce dialogue benchmarks in evaluating dynamic constraint evolution and multi-objective coordination by proposing RealWorldShop, a million-scale product benchmark, alongside the RealShopAgent framework. This framework integrates large language models with structured dialogue modeling and catalog-grounded retrieval to enable agentic conversational interactions through explicit state management and retrieval control. Furthermore, it introduces a pioneering conversation-level decision support evaluation system encompassing user simulation, role-based assessment, and runtime safeguards. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines across state tracking, constraint updating, and multi-intent scenarios.
π Abstract
Large language models are reshaping ecommerce from static recommenders into interactive shopping assistants, yet real-world shopping requires session-level decision support: users reveal and revise constraints, coordinate multiple goals, and expect product-grounded recommendations over a full conversation. Existing benchmarks are mostly outcome-oriented or execution-oriented, leaving this evolving decision process under-evaluated. We introduce REALWORLDSHOP, a benchmark built on 3.28M grounded products, structured shopping episodes, a profile-grounded and actioncontrolled user simulator, and role-play evaluation. Our analysis shows that current systems produce locally plausible responses but struggle with state tracking, constraint updating, and grounded convergence, especially under ambiguous intent, bundle, and multi-intent scenarios. We further propose REALSHOP_AGENT, an executable session-control framework with explicit state management, shopping-flow control, catalog-grounded retrieval, and runtime guards. Experiments show that REALSHOP_AGENT consistently outperforms strong baselines on REALWORLDSHOP.