π€ AI Summary
This study addresses the limitation of existing offline metrics that evaluate similarity or popularity in isolation, thereby failing to accurately assess the serendipity of recommender systems. We propose SPADE, a metric that maps items into a user-specific two-dimensional popularity-similarity space and constructs a Pareto frontier. Serendipity is rigorously quantified by computing the minimum Euclidean distance from test items to this frontier. By integrating multi-dimensional features with actual relevance, SPADE effectively isolates genuinely serendipitous recommendations. Experiments conducted across five datasets and against five baseline algorithms demonstrate that SPADE prevents models from gaming the evaluation through irrelevant or non-personalized recommendations, achieving reliable and robust serendipity assessment.
π Abstract
Recommender systems engineer serendipity to foster active exploration and break predictable consumption cycles. The problem with existing offline beyond-accuracy metrics is that they often either isolate historical similarity or global popularity. We aim to design an evaluation metric that examines similarity, popularity, and actual user relevance. To achieve this, we introduce SPADE (Serendipitous Pareto Distance Evaluation). SPADE maps all items into a two-dimensional space to directly calculate a user-specific Pareto frontier of maximally popular and historically similar items. The final serendipity score is then computed by averaging the minimum Euclidean distance from this boundary strictly for the correctly recommended test-set items. Evaluating SPADE across five datasets and five baseline algorithms confirms its effectiveness; our results show that the metric successfully prevents algorithms from exploiting beyond-accuracy measures with irrelevant or non-personalized recommendations, reliably isolating serendipitous discoveries.