GRP v0.1 Technical Report

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of multi-stage cascaded architectures and the difficulty of jointly replacing components in industrial recommender systems by proposing a unified generative recommendation framework. This framework integrates retrieval and ranking within a single encoder-decoder architecture, employing reinforcement learning post-training with frozen module rewards. Furthermore, it introduces the mGRPO algorithm, which optimizes rewards via reference-anchored margins to enhance performance while preserving log-likelihood stability. By incorporating multimodal semantic IDs within an end-to-end architecture, the proposed method facilitates progressive deployment. Experimental results demonstrate that the approach reduces retrieval latency by 69% and yields online improvements of 0.82% in watch time and 2.56% in share volume.
📝 Abstract
Industrial recommendation systems rely on multi-stage cascades whose retrieval, ranking, and serving components are difficult to replace jointly. We present GRP, a generative recommendation framework that combines retrieval, ranking, and reward modeling in a single encoder-decoder model, and evaluate a progressive path toward end-to-end recommendation. The model generates multimodal Semantic IDs and scores candidates with a jointly trained ranking module. The frozen ranking module then supplies rewards for reinforcement-learning post-training. We introduce mGRPO, which adds a reference-anchored margin to reward optimization to preserve the likelihood of logged targets. Offline experiments examine history encoding, model capacity allocation, event selection, tokenization, and reward discrimination. Serving optimizations reduce end-to-end retrieval latency by 69%. Online experiments evaluate the model as a retrieval source, with early-ranking bypass, and with replacement of weaker sources. In a retrieval-only comparison, view time increases by 0.46% and shares by 0.77% relative to production. A separate comparison combining bypass and source replacement yields increases of 0.82% in view time and 2.56% in shares, with neutral platform-level guardrails. These results support progressive deployment while identifying remaining gaps in ranking quality and performance across recommendation metrics.
Problem

Research questions and friction points this paper is trying to address.

Recommendation Systems
Multi-stage Cascade
End-to-end Recommendation
Generative Recommendation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Recommendation
Semantic IDs
mGRPO
End-to-End Recommendation
Reinforcement Learning
🔎 Similar Papers
No similar papers found.