Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of a unified evaluation benchmark for diffusion-based recommender systems, which has led to irreproducible experiments and unfair comparisons. To this end, we construct the first open-source, unified evaluation framework encompassing 14 models across five recommendation scenarios. We establish a standardized evaluation protocol that eliminates configuration discrepancies through unified data splitting, training, and inference procedures, thereby ensuring a fair basis for comparison. Building upon this framework, we conduct a systematic empirical study that validates the potential of diffusion models in recommendation tasks while identifying key performance factors and existing challenges. By providing a solid foundation for future research, this project makes all code and datasets fully open-source to facilitate reproducibility and further advancement in the field.
📝 Abstract
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of diffusion-based RSs, a critical issue has emerged: the lack of a unified and systematic quantitative evaluation benchmark, which often results in irreproducible experimental results and unfair comparisons across studies due to inconsistent data processing, training configurations, inference procedures, and evaluation protocols. To address this challenge, we propose Eval4DiRec, the first unified and open-source evaluation framework specifically designed for diffusion-based RSs. Eval4DiRec supports 14 representative diffusion-based RS models across five different recommendation scenarios, providing consistent and reproducible experimental settings to systematically assess their performance. Built upon this framework, we conduct extensive empirical studies to benchmark these models under unified protocols. The results highlight the strong potential of diffusion models for recommendation while also revealing key factors and practical challenges that substantially affect their performance, thereby establishing a solid foundation to facilitate fair evaluation and guide future research in this promising field. Our code and data are available at: https://github.com/wangcong2001/Eval4DiRec.
Problem

Research questions and friction points this paper is trying to address.

Diffusion-based Recommender Systems
Evaluation Benchmark
Reproducibility
Fair Comparison
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion-based Recommender Systems
Evaluation Framework
Benchmark
Reproducibility
Empirical Study
🔎 Similar Papers
2024-09-06arXiv.orgCitations: 3
C
Cong Wang
Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, China
Shoujin Wang
Shoujin Wang
University of Technology Sydney
Data ScienceMachine LearningRecommender SystemMisinformationData Science Application
Y
Yishuo Li
Department of Computer Science and Technology, Tongji University, China
Q
Qi Zhang
Department of Computer Science and Technology, Tongji University, China
Liang Hu
Liang Hu
Tongji University
Artificial IntellegenceMachine LearningData Science
W
Wenpeng Lu
Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science; Shandong Academy of Artificial Intelligence, China