🤖 AI Summary
This work addresses the high inference cost of generative recommendation models at scale and the limitations of existing knowledge distillation approaches in handling unique challenges such as semantic ID hierarchy-induced difficulty imbalance and beam search prefix pruning errors. To overcome these issues, the paper proposes SmartGR, a novel framework that introduces, for the first time, hierarchical-aware semantic ID distillation and beam-search-aware ranking distillation to effectively transfer knowledge from large teacher models to lightweight student models. Experimental results on four benchmark datasets demonstrate that SmartGR improves recommendation performance by 8.6% on average while achieving a 2.39× speedup in inference, significantly reducing computational overhead without compromising effectiveness.
📝 Abstract
Generative recommendation (GR) has emerged as a promising paradigm for recommender systems. Scaling up GR models can improve recommendation performance, but it also substantially increases inference cost. Knowledge distillation provides a practical solution by transferring knowledge from a large GR model to a lightweight one. However, existing distillation methods do not account for two GR-specific challenges: imbalanced distillation difficulty across the semantic ID (SID) hierarchy and incorrect prefix pruning during beam search. To address these challenges, we propose SmartGR, a novel distillation framework that utilizes Hierarchy-Aware SID Distillation to transfer the teacher's modeling capability across the hierarchy and leverages Beam-Aware Ranking Distillation to distill the teacher's ranking preferences during beam search. Extensive experiments on four benchmark datasets demonstrate the effectiveness and efficiency of SmartGR, improving the performance by 8.6% while achieving a 2.39$\times$ inference speedup on average.