Personalization Matters: Long-Horizon Conversation Agent with User-Centric Information in Online Shopping Interactions

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of maintaining preference consistency and the limitations of static user profiles in controlling long-horizon interactions within multi-turn dialogues for online shopping. To this end, we propose a multi-agent multimodal Retrieval-Augmented Generation (RAG) framework. By leveraging a role decomposition mechanism and a user-centric retrieval strategy, the framework integrates product metadata, reviews, and image descriptions to enable dynamic state tracking and personalized reasoning. Furthermore, we construct a trajectory-level evaluation protocol encompassing global preference consistency. Experimental results demonstrate that the proposed method significantly outperforms baseline models on automatic metrics and achieves a score of 4.60 in real-user studies, effectively validating its advantages in enhancing personalized experiences during long-horizon interactions.
📝 Abstract
Personalized conversational shopping requires maintaining preference consistency over multi-turn interactions, where users reveal constraints gradually. Existing approaches often rely on static profiles and do not explicitly control long-horizon interaction behavior. We propose a multi-agent, multimodal Retrieval-Augmented Generation (RAG) framework that decomposes dialogue state tracking, recommendation retrieval, preference-aware reasoning, and response generation, while integrating product metadata, product reviews, image-derived descriptions, and user historical reviews. To evaluate interaction-level quality, we adopt a trajectory-level protocol with four dimensions: Global Preference Consistency, Cumulative Information Synthesis, Interaction Trajectory, and Tone Consistency. On an Amazon Reviews 2023 benchmark, retrieval-enabled variants outperform a no-RAG baseline on automatic trajectory metrics (average 4.82 vs. 3.74). In a small real-user study ($n{=}5$), the Full variant achieves the highest mean overall rating (4.60 vs. 2.20 for Baseline), providing exploratory evidence that role decomposition plus user-centric retrieval improves perceived personalization.\footnote{Code and dataset are available at: https://github.com/RenaGao/Multimodel_RAG_Indexing
Problem

Research questions and friction points this paper is trying to address.

Personalized conversational shopping
Long-horizon conversation
Preference consistency
User-centric information
Multi-turn interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent RAG
Multimodal Retrieval-Augmented Generation
Long-Horizon Conversation
Personalized Recommendation
Trajectory-level Evaluation
🔎 Similar Papers