Fake News Detection: It's All in the Data!

📅 2024-07-02
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the critical impact of data quality and diversity on the effectiveness and robustness of fake news detection models. Addressing prevalent issues in existing datasets—including inconsistent annotations, significant biases, and missing metadata—we propose a systematic data governance framework. Specifically, we construct the first unified, open-source GitHub repository encompassing 120+ publicly available fake news datasets, enabling standardized integration, bias identification, multi-dimensional metadata annotation, and searchable indexing. Our work provides the first empirical evidence demonstrating that intrinsic data characteristics—rather than model architecture alone—fundamentally determine detection performance, thereby establishing a data-centric research paradigm. The repository has been adopted as a benchmark data entry point by over ten academic institutions and industry teams worldwide, catalyzing a paradigm shift in fake news detection from model-centric to data–model co-evolutionary development.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Application Domains: Misinformation & Fake NewsMachine Learning: Ethics, Bias, and Fairness

Application Category

Web Mining and Content Analysis: Web data provenance, reliability, and authenticityEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSocial Networks and Social Media: Fairness and bias in social network and social media analysis
📝 Abstract
This comprehensive survey serves as an indispensable resource for researchers embarking on the journey of fake news detection. By highlighting the pivotal role of dataset quality and diversity, it underscores the significance of these elements in the effectiveness and robustness of detection models. The survey meticulously outlines the key features of datasets, various labeling systems employed, and prevalent biases that can impact model performance. Additionally, it addresses critical ethical issues and best practices, offering a thorough overview of the current state of available datasets. Our contribution to this field is further enriched by the provision of GitHub repository, which consolidates publicly accessible datasets into a single, user-friendly portal. This repository is designed to facilitate and stimulate further research and development efforts aimed at combating the pervasive issue of fake news.
Problem

Research questions and friction points this paper is trying to address.

Analyzes dataset quality's impact on fake news detection models
Explores dataset biases and labeling systems affecting model performance
Provides centralized GitHub repository for accessible fake news datasets
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emphasizes dataset quality and diversity
Provides GitHub repository for datasets
Addresses ethical issues and biases
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Warsaw University of Technology | Polish Academy of Sciences
S
Soveatin Kuntur
Warsaw University of Technology, plac Politechniki 1, 00-661 Warsaw, Poland
Anna Wróblewska
Anna Wróblewska
Warsaw University of Technology
machine learningnatural language processingimage processingmultimodal learningrecommendation
M
M. Paprzycki
Systems Research Institute, Polish Academy of Sciences, Nowelska 6, 01-447 Warsaw, Poland
M
M. Ganzha