🤖 AI Summary
This work addresses the lack of a unified evaluation benchmark for diffusion-based recommender systems, which has led to irreproducible experiments and unfair comparisons. To this end, we construct the first open-source, unified evaluation framework encompassing 14 models across five recommendation scenarios. We establish a standardized evaluation protocol that eliminates configuration discrepancies through unified data splitting, training, and inference procedures, thereby ensuring a fair basis for comparison. Building upon this framework, we conduct a systematic empirical study that validates the potential of diffusion models in recommendation tasks while identifying key performance factors and existing challenges. By providing a solid foundation for future research, this project makes all code and datasets fully open-source to facilitate reproducibility and further advancement in the field.
📝 Abstract
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.