OneTrans-V2: Unifying Retrieval, Pre-rank, and Fine-rank with One Transformer in Industrial Recommender

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of model isolation, redundant encoding, and engineering overhead inherent in the cascaded architecture of industrial recommender systems by proposing a unified Transformer that consolidates the retrieval, pre-ranking, and ranking stages. Methodologically, cross-stage collaboration and intra-model knowledge distillation are achieved through shared user sequence encoding and joint training. The work introduces Decision-Conditioned Generative Retrieval (DCGR) to guide generation via business objectives, alongside Sequence-Native Training (SNT) to amortize encoding costs. Furthermore, sparse Mixture-of-Experts (MoE) and μP parameterization are incorporated to ensure stable model scaling. Deployed in a large-scale production system, the proposed architecture yields a 9.74% increase in Gross Merchandise Volume (GMV) while achieving 3.2× the throughput of the original system under equivalent hardware constraints.
📝 Abstract
Industrial recommendation systems typically operate as a \emph{cascade} of retrieval, pre-rank, and fine-rank, but these stages are usually trained and served as separate models, causing repeated user-sequence encoding, isolated optimization, and duplicated engineering effort. Building on OneTrans' model-level unification, we present OneTrans-V2, one Transformer that unifies the entire cascade. It encodes the user behavior sequence once as a shared context while preserving stage-specific candidate features and computation. Joint training lets the three stages reinforce one another and enables in-model knowledge distillation from fine-rank to pre-rank. We scale the shared backbone with sparse mixture-of-experts (MoE), which increases capacity with bounded activated computation, and stabilize scaling with $μ$P-style parameterization. To consolidate objective-specific retrieval channels, we introduce Decision-Conditioned Generative Retrieval (DCGR). DCGR predicts a decision prefix describing the upcoming interaction and generates items conditioned on it, allowing business objectives to steer a single generative process. Finally, Sequence-Native Training (SNT) organizes training around each user's lifelong behavior sequence and amortizes its encoding across exposures. Deployed across all three stages of a large-scale industrial recommendation system, OneTrans-V2 improves gross merchandise value (GMV) by 9.74\% and, with a co-designed serving stack, delivers $3.2\times$ the throughput of the cascade it replaces under the same hardware budget.
Problem

Research questions and friction points this paper is trying to address.

Industrial Recommender Systems
Cascade Architecture
Model Unification
User Sequence Encoding
Generative Retrieval
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Transformer
Sparse Mixture-of-Experts
Decision-Conditioned Generative Retrieval
Sequence-Native Training
Knowledge Distillation
🔎 Similar Papers