🤖 AI Summary
The CTR prediction field has long suffered from a lack of standardized benchmarks and unified evaluation protocols, leading to irreproducible experiments and incomparable results.
Method: We introduce the first open-source, reproducible CTR benchmark platform, systematically re-evaluating 24 state-of-the-art models across five public datasets under consistent preprocessing and evaluation protocols—conducting over 7,000 experiments (>12,000 GPU hours).
Contribution/Results: We propose a standardized CTR evaluation paradigm that reveals widespread overestimation of deep model performance differences; we demonstrate that rigorous hyperparameter optimization and fair experimental design are critical for valid comparisons. After thorough tuning, performance gaps among most models significantly narrow. We fully open-source all code, configurations, and results—including training scripts, data pipelines, and evaluation metrics—to advance reproducible research and trustworthy SOTA assessment.
📝 Abstract
Click-through rate (CTR) prediction is a critical task for many applications, as its accuracy has a direct impact on user experience and platform revenue. In recent years, CTR prediction has been widely studied in both academia and industry, resulting in a wide variety of CTR prediction models. Unfortunately, there is still a lack of standardized benchmarks and uniform evaluation protocols for CTR prediction research. This leads to non-reproducible or even inconsistent experimental results among existing studies, which largely limit the practical value and potential impact of their research. In this work, we aim to perform open benchmarking for CTR prediction and present a rigorous comparison of different models in a reproducible manner. To this end, we ran over 7,000 experiments for more than 12,000 GPU hours in total to re-evaluate 24 existing models on multiple dataset settings. Surprisingly, our experiments show that with sufficient hyper-parameter search and model tuning, many deep models have smaller differences than expected. The results also reveal that making real progress on the modeling of CTR prediction is indeed a very challenging research task. We believe that our benchmarking work could not only allow researchers to gauge the effectiveness of new models conveniently but also make them fairly compare with the state of the arts. We have publicly released the benchmarking tools, evaluation protocols, and experimental settings of our work to promote reproducible research in this field.