🤖 AI Summary
This study addresses the inefficiency of multi-stage cascaded architectures and the difficulty of jointly replacing components in industrial recommender systems by proposing a unified generative recommendation framework. This framework integrates retrieval and ranking within a single encoder-decoder architecture, employing reinforcement learning post-training with frozen module rewards. Furthermore, it introduces the mGRPO algorithm, which optimizes rewards via reference-anchored margins to enhance performance while preserving log-likelihood stability. By incorporating multimodal semantic IDs within an end-to-end architecture, the proposed method facilitates progressive deployment. Experimental results demonstrate that the approach reduces retrieval latency by 69% and yields online improvements of 0.82% in watch time and 2.56% in share volume.
📝 Abstract
Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.