🤖 AI Summary
This study addresses the coupled challenges of representation collapse and item collision in generative retrieval by proposing GEAR, a framework enabling end-to-end adaptation for large-scale advertising systems. GEAR jointly optimizes the tokenizer, generator, and reranker through three key innovations: BasisVQ, which employs orthogonal basis reparameterization to stabilize gradients; prefix-aware BasisRQ to enhance semantic expressiveness; and a context-conditioned reranking head to resolve ambiguity. Successfully deployed on Douyin serving hundreds of millions of daily active users, online A/B testing demonstrates significant improvements in both retrieval accuracy and core business metrics. This work establishes a differentiable and scalable new paradigm for generative advertising retrieval.
📝 Abstract
Generative retrieval reformulates recommendation as the generation of discrete item tokens. However, scaling this paradigm to real-world recommender systems reveals two critical bottlenecks: 1) Representation collapse, where the item tokenizer converges to degenerate results under continuous distribution shifts, fundamentally hindering stable end-to-end adaptation. 2) Item collisions, where the massive candidate pool causes distinct items to share identical token sequences, compromising the final retrieval precision. Crucially, these bottlenecks are inherently coupled: expanding codebook capacity to mitigate collisions inevitably exacerbates collapse. To address them simultaneously, we propose GEAR, an end-to-end framework that jointly optimizes the tokenizer, generator, and reranker. To mitigate representation collapse, we introduce BasisVQ, which re-parameterizes the codebook via an orthogonal basis to enable global gradient sharing and rigid spatial rotation of the latent space, effectively stabilizing gradient dynamics without ad-hoc heuristics. We further extend it to prefix-aware BasisRQ, substantially enhancing the codebook's expressiveness with the same asymptotic time complexity. To resolve item collisions, GEAR integrates a context-conditioned reranking head into the generative process, efficiently disambiguating colliding items with minimal computational overhead. By unifying stable tokenization and joint reranking within an end-to-end generative framework, GEAR establishes a fully differentiable and scalable paradigm. It currently serves hundreds of millions of daily active users on Douyin Ads, yielding substantial empirical improvements in extensive online A/B tests.