🤖 AI Summary
This study addresses the challenges of model isolation, redundant encoding, and engineering overhead inherent in the cascaded architecture of industrial recommender systems by proposing a unified Transformer that consolidates the retrieval, pre-ranking, and ranking stages. Methodologically, cross-stage collaboration and intra-model knowledge distillation are achieved through shared user sequence encoding and joint training. The work introduces Decision-Conditioned Generative Retrieval (DCGR) to guide generation via business objectives, alongside Sequence-Native Training (SNT) to amortize encoding costs. Furthermore, sparse Mixture-of-Experts (MoE) and μP parameterization are incorporated to ensure stable model scaling. Deployed in a large-scale production system, the proposed architecture yields a 9.74% increase in Gross Merchandise Volume (GMV) while achieving 3.2× the throughput of the original system under equivalent hardware constraints.
📝 Abstract
Industrial recommendation systems typically operate as a \emph{cascade} of retrieval, pre-rank, and fine-rank, but these stages are usually trained and served as separate models, causing repeated user-sequence encoding, isolated optimization, and duplicated engineering effort. Building on OneTrans' model-level unification, we present OneTrans-V2, one Transformer that unifies the entire cascade. It encodes the user behavior sequence once as a shared context while preserving stage-specific candidate features and computation. Joint training lets the three stages reinforce one another and enables in-model knowledge distillation from fine-rank to pre-rank. We scale the shared backbone with sparse mixture-of-experts (MoE), which increases capacity with bounded activated computation, and stabilize scaling with $μ$P-style parameterization. To consolidate objective-specific retrieval channels, we introduce Decision-Conditioned Generative Retrieval (DCGR). DCGR predicts a decision prefix describing the upcoming interaction and generates items conditioned on it, allowing business objectives to steer a single generative process. Finally, Sequence-Native Training (SNT) organizes training around each user's lifelong behavior sequence and amortizes its encoding across exposures. Deployed across all three stages of a large-scale industrial recommendation system, OneTrans-V2 improves gross merchandise value (GMV) by 9.74\% and, with a co-designed serving stack, delivers $3.2\times$ the throughput of the cascade it replaces under the same hardware budget.