EvoReason: Self-Evolving Reasoning Primitive-Guided On-Policy Distillation for Latent Reasoning in Generative Recommendation

๐Ÿ“… 2026-07-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing implicit reasoning approaches in generative recommendation suffer from limited generalization due to direct distillation of raw chain-of-thought outputs, which are often plagued by redundant expressions and unstable reasoning paths. To address this, this work proposes EvoReason, a novel framework that introduces, for the first time, a reasoning-primitive-guided mechanism for structured chain-of-thought generation. EvoReason extracts reusable reasoning primitives from high-quality agent trajectories and integrates them with a self-evolutionary policy distillation strategy, enabling dynamic alignment and co-evolution between the teacherโ€™s explicit reasoning space and the studentโ€™s implicit reasoning space. This approach significantly enhances the transferability and consistency of implicit reasoning representations while preserving the low-latency deployment advantage of generative recommender systems, thereby substantially improving their reasoning capabilities.
๐Ÿ“ Abstract
Generative recommendation benefits from reasoning-enhanced inference, and latent reasoning offers an efficient paradigm by encoding intermediate reasoning processes into compact continuous representations for latency-sensitive deployment. Despite its efficiency, existing latent reasoning approaches typically rely on directly distilling raw chain-of-thought (CoT) trajectories into latent representations, assuming that textual reasoning traces provide sufficient supervision. However, recommendation reasoning trajectories contain diverse reasoning processes with redundant expressions and unstable reasoning paths, making raw CoT supervision suboptimal for learning transferable latent reasoning representations. To address this challenge, we propose EvoReason, a self-evolving latent reasoning framework that adaptively aligns explicit reasoning supervision with the student's latent reasoning space through primitive-guided on-policy distillation. First, EvoReason extracts reusable reasoning primitives from high-quality agentic recommendation trajectories, where each primitive captures an essential reasoning behavior and serves as a pseudo-tool for structured teacher reasoning. Then, based on these primitives, we equip the teacher with primitive-aware reasoning capabilities, enabling it to generate structured CoT supervision with reduced redundancy and improved consistency. Finally, during latent reasoning optimization, EvoReason introduces a self-evolving on-policy distillation mechanism, where the primitive-guided reasoning process evolves according to the student's latent reasoning outcomes. Through this closed-loop co-evolution, policy updates continuously improve latent reasoning behaviors is refined according to the resulting latent reasoning outcomes, enabling progressively better-aligned CoT supervision and more effective reasoning transfer.
Problem

Research questions and friction points this paper is trying to address.

latent reasoning
generative recommendation
chain-of-thought
reasoning distillation
redundant reasoning paths
Innovation

Methods, ideas, or system contributions that make the work stand out.

reasoning primitives
on-policy distillation
latent reasoning
self-evolving framework
generative recommendation
๐Ÿ”Ž Similar Papers
2024-05-17Annual Meeting of the Association for Computational LinguisticsCitations: 4
Z
Zhuang Zhuang
Kuaishou Technology, Beijing, China
Zhipeng Wei
Zhipeng Wei
ICSI, UC Berkeley
robustness of deep learning
R
Rongfeng Guo
Shenzhen University, Shenzhen, China
S
Shijie Li
Kuaishou Technology, Beijing, China
P
Peng Zhao
Kuaishou Technology, Beijing, China
Jie Chen
Jie Chen
University of Science and Technology of China
Artificial Intelligence3D Vision
Fei Pan
Fei Pan
Unversity of Michigan
Computer VisionMachine Learning