🤖 AI Summary
This study investigates whether popularity calibration genuinely enhances user experience in music recommendation and examines its reliability across varying levels of user listening history and item familiarity. The authors construct three types of playlists—high-popularity, low-popularity, and calibrated—and employ a controlled naive recommender to generate personalized lists. Calibration is quantified using Jensen–Shannon divergence (JSD), and subjective user feedback is collected through controlled experiments. This work presents the first systematic validation of JSD’s stability with respect to real users’ perceived calibration. Results indicate that while users can discern differences in popularity, they do not exhibit a significant preference for calibrated recommendations. Moreover, computed popularity labels show only weak alignment with users’ subjective judgments, and the relationship between JSD and perceived calibration is significantly moderated by item familiarity, playlist composition, and the availability of historical interaction data.
📝 Abstract
Popularity calibration in recommender systems has been studied both as a form of user-centered personalization and as an indicator of popularity bias. Most existing work evaluates calibration through offline metrics, often assuming that users prefer recommendation lists whose popularity distribution matches their historical consumption profile. However, user studies on calibration remain limited, and existing findings suggest that calibrated recommendations do not necessarily have a strong effect on user experience. Moreover, although prior work has shown that calibration metrics can correlate with users' perceptions of recommendation lists, the robustness of this relation remains unclear under different levels of item familiarity and incomplete user-history information. In this work, we study the perceived value and measurement reliability of popularity calibration in music recommendation. We construct personalized track lists from users' recent listening histories and use a controlled naive recommender to create lists with different popularity compositions: highpop-heavy, lowpop-heavy, and calibrated. We investigate whether users perceive differences between these lists, whether calibrated lists are preferred, how robust JSD-based popularity calibration is under different familiarity and history-availability conditions, and how computational popularity labels align with users' own popularity judgments. Our results show that users perceive differences in popularity composition, but do not clearly prefer calibrated lists. We further find that the relation between JSD and perceived popularity depends on item familiarity, list composition, and available user history, while computational and user-judged popularity labels only weakly align. These findings contribute to a more critical understanding of popularity calibration as both an offline metric and a user-facing construct.