Training with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the probability competition issue in generative recommendation, where appending training objectives to mitigate missed detections degrades ranking performance. We propose a framework that decouples supervision signals from probability competition by constructing a matching loss function to isolate the influence of training objectives on inference candidates, and designing a group-normalized intermediate loss to eliminate competition between inference-only appended objectives and inference candidates. Experiments based on the OneRec model using Amazon datasets demonstrate that our approach improves FT-NDCG by 7.8%–22.2%. These results validate the necessity of evaluating completion strategies according to generator behavior and establish a new paradigm for training optimization in generative recommendation systems.
📝 Abstract
Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2\%; 95\% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7\% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.
Problem

Research questions and friction points this paper is trying to address.

Generative Recommendation
Missed Targets
Probability Competition
Training-Inference Mismatch
Reranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Recommendation
Missed Targets
Probability Competition
Matched Losses
Candidate Completion
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xuesi Wang
Independent Researcher, Shanghai, China
Y
Yangbin Shi
Zhejiang University, Hangzhou, China
Xiaolin Zheng
Xiaolin Zheng
Professor of Mechanical Engineering, and Energy Science & Engineering, Stanford University
EnergyPropulsionCombustionNanomaterialsElectrochemistry