Learning Better Reasoning for Generative Recommendation with Semantic IDs

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of invalid reasoning misleading generation and degrading performance in generative recommendation. To tackle this, we propose Evo-Rec, a framework that optimizes the reasoning process for semantic ID generation through a three-stage strategy. Specifically, supervised fine-tuning is first employed to filter valid reasoning trajectories for semantic ID alignment. A multi-candidate sampling mechanism is then introduced to enhance exploration capabilities. Finally, constrained reinforcement learning is leveraged to further refine the generation policy. Experiments demonstrate that Evo-Rec comprehensively outperforms existing recommendation models across three mainstream benchmarks, validating the effectiveness of reasoning process optimization in improving generative recommendation performance.
📝 Abstract
Generative recommendation reformulates item retrieval as sequence generation, allowing a unified model to directly generate the next item from a user's interaction history. Semantic IDs further make this paradigm effective and scalable by representing each item as discrete codes, enabling knowledge sharing among semantically related items. Recent studies introduce explicit reasoning before Semantic-ID generation, helping models summarize user interests and infer possible preference transitions. However, reasoning is not inherently beneficial: Inaccurate or uninformative reasoning may mislead subsequent item generation and ultimately degrade recommendation performance. This raises a central challenge: how can a recommender select and learn effective reasoning traces and progressively evolve toward better reasoning from its own generations? In this work, we propose Evo-Rec, a three-stage framework for learning better reasoning and further enhancing it through reinforcement learning. First, we align Semantic IDs with their textual and behavioral contexts, enabling the model to understand and generate item identifiers. Second, we sample multiple candidate reasoning traces and retain those that improve the prediction of the ground-truth item, providing a stronger reasoning initialization through supervised fine-tuning. Third, we further optimize the reasoning policy through reinforcement learning with catalog-constrained item generation and ranking-aware recommendation feedback. Experiments on three Amazon Review benchmarks show that Evo-Rec consistently outperforms discriminative, generative, and reasoning-enhanced recommenders across all evaluation metrics. These results demonstrate the effectiveness of our framework in learning better reasoning for SID-based generative recommendation.
Problem

Research questions and friction points this paper is trying to address.

Generative Recommendation
Semantic IDs
Reasoning Traces
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Recommendation
Semantic IDs
Reasoning Traces
Reinforcement Learning
Supervised Fine-Tuning
🔎 Similar Papers
2024-05-12International Conference on Information and Knowledge ManagementCitations: 60
M
Mengdan Zhu
Emory University
Y
Yufan Zhao
Microsoft
S
Sophie Di
Cornell University
Y
Yao Zhao
Microsoft
T
Tao Di
Microsoft
Y
Yulan Yan
Microsoft
Sridhar Iyer
Sridhar Iyer
Microsoft
Liang Zhao
Liang Zhao
Winship Distinguished Professor&Associate Professor, Emory University
data miningmachine learningspatial data mininggraph neural networksgenerative AI