🤖 AI Summary
This study investigates the critical impact of data quality and diversity on the effectiveness and robustness of fake news detection models. Addressing prevalent issues in existing datasets—including inconsistent annotations, significant biases, and missing metadata—we propose a systematic data governance framework. Specifically, we construct the first unified, open-source GitHub repository encompassing 120+ publicly available fake news datasets, enabling standardized integration, bias identification, multi-dimensional metadata annotation, and searchable indexing. Our work provides the first empirical evidence demonstrating that intrinsic data characteristics—rather than model architecture alone—fundamentally determine detection performance, thereby establishing a data-centric research paradigm. The repository has been adopted as a benchmark data entry point by over ten academic institutions and industry teams worldwide, catalyzing a paradigm shift in fake news detection from model-centric to data–model co-evolutionary development.
📝 Abstract
This comprehensive survey serves as an indispensable resource for researchers embarking on the journey of fake news detection. By highlighting the pivotal role of dataset quality and diversity, it underscores the significance of these elements in the effectiveness and robustness of detection models. The survey meticulously outlines the key features of datasets, various labeling systems employed, and prevalent biases that can impact model performance. Additionally, it addresses critical ethical issues and best practices, offering a thorough overview of the current state of available datasets. Our contribution to this field is further enriched by the provision of GitHub repository, which consolidates publicly accessible datasets into a single, user-friendly portal. This repository is designed to facilitate and stimulate further research and development efforts aimed at combating the pervasive issue of fake news.