Breaking the Loop: An Empirical Comparison of Strategies for Novelty and Freshness in YouTube Music

πŸ“… 2026-07-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the feedback loop problem in continuously trained music recommendation systems, which undermines the timeliness of newly released content and the novelty of unheard tracks. Conducting a multi-layered intervention experiment on YouTube Music’s homepage, the authors present the first empirical comparison of six debiasing and exploration strategies within a real-world, large-scale continuous training system. Using off-policy online A/B testing, they evaluate approaches including serving-layer heuristics, training data reweighting, architectural debiasing, and uncertainty-driven exploration via Spectral Normalization Gaussian Process (SNGP). Their findings reveal that serving-layer interventions are often nullified by the learning dynamics of the feedback loop, whereas SNGP-based exploration substantially increases exposure of new songs. Although architectural debiasing enhances overall diversity, it does not facilitate the discovery of new content.
πŸ“ Abstract
Continuously trained ranking models in music recommenders fall into feedback loops where previously consumed items dominate recommendations. This suppresses two distinct content classes: new releases (temporal freshness) and unlistened catalog items (novelty). Industry practitioners have a wide menu of interventions available, ranging from serving-time heuristics, training-data reweighting, architectural debiasing, to uncertainty-driven exploration, each of which are well understood in academic settings. But live systems offer challenges with continuously ingested content, interconnected components, and practical limitations that counteract the findings from academic research. We report results from off-policy online A/B tests for six interventions and a combination experiment across four conceptual layers (serving, training, architecture, exploration) on the YouTube Music homepage. All interventions modify the ranking model or the serving layer that consumes its scores; candidate generation and other upstream components are held fixed. We discuss key takeaways from our results: first, serving-time interventions on continuously trained systems are neutralized by the learning loop. Second, architectural debiasing reduces popularity dominance and improves diversity but does not create discovery, while carrying hidden integration costs. Finally, uncertainty-driven exploration interventions with a Spectral-normalized Neural Gaussian Process (SNGP) head produce the largest new-release lift, though they come with a measurable engagement or diversity tradeoff. We close with recommendations on which layer to intervene at, and the hidden costs of each choice.
Problem

Research questions and friction points this paper is trying to address.

feedback loop
novelty
freshness
music recommendation
popularity bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

feedback loop
uncertainty-driven exploration
SNGP
architectural debiasing
temporal freshness
πŸ”Ž Similar Papers
No similar papers found.
S
Srivaths Ranganathan
Google LLC
Z
Zihuan Diao
Google LLC
Bernardo Cunha
Bernardo Cunha
Professor of Computer Architectures, Universidade de Aveiro
Computer architecturesinstrumentationartificial visionrobotics
J
Joshua L. Moore
Google LLC
R
Robin Dumas
Google LLC
M
Murat Goksedef
Google LLC
Y
Yanwei Song
Google LLC
M
Mukai Lu
Google LLC
G
Gergo Varady
Google LLC
T
Tracy Pesin
Google LLC