🤖 AI Summary
This study addresses the information loss caused by treating browsing depth as a latent variable in carousel ranking. We propose a ranking bandwidth model based on observable browsing depth that explicitly distinguishes unclicked from unexposed items, leveraging authentic browsing signals to optimize statistical updates. Furthermore, we derive an instance-dependent UCB regret bound and an asymptotically optimal DMED bound, overcoming the limitations of traditional latent-variable assumptions. Extensive experiments combining UCB, TS, and DMED algorithms with RecGaze-based parametric simulations demonstrate that the proposed OD-TS achieves regret comparable to PBM-TS while yielding lower average regret than PBM-UCB, thereby confirming the effectiveness of our approach.
📝 Abstract
Carousel interfaces allow a recommender system to directly observe how far a user has browsed. This signal distinguishes displayed but unclicked items from items that were never displayed, whereas conventional ranking-bandit models, including cascade and position-based models, generally treat examination as latent. We formulate a ranking-bandit problem in which a learner presents a list of $L$ items, observes the user's maximum browsing depth, and receives click feedback only for positions up to that depth. The objective is to maximize the expected number of clicks under an unknown item-attractiveness vector and a browsing-depth distribution. We propose three algorithms based on UCB, Thompson Sampling, and DMED, all of which update item statistics only from observed exposures. We derive an instance-dependent logarithmic upper bound for our UCB-based algorithm and an asymptotic upper bound for our DMED-based algorithm that coincides with the lower bound as its parameter $\alpha\downarrow 0$, establishing asymptotic optimality in this limit. Simulations in synthetic shallow- and deep-browsing environments, together with experiments parameterized from RecGaze interaction logs, show that OD-TS attains final mean regret similar to PBM-TS, while the proposed methods achieve lower final mean regret than PBM-UCB.