Claim Extraction for Fact-Checking: Data, Models, and Automated Metrics

📅 2025-02-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of unified definition and evaluation standards for claim extraction in fact-checking. We systematically consolidate fragmented claim extraction objectives and propose a fine-grained modeling framework tailored to atomic factual claims. To support this, we introduce FEVERFact—the first dataset dedicated to atomic claims (17K instances)—and design a multidimensional automated evaluation suite covering atomicity, fluency, decontextualization, faithfulness, focus, and coverage; each dimension is mapped to canonical NLP tasks (e.g., textual entailment, question answering, summarization). Experiments demonstrate strong agreement with human judgments (F1 = 0.89, RMSE < 0.12) and robust model ranking stability—even on the most challenging F_fact metric. Both the dataset and evaluation code are publicly released.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Knowledge Representation and Reasoning: Reasoning with BeliefsConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
In this paper, we explore the problem of Claim Extraction using one-to-many text generation methods, comparing LLMs, small summarization models finetuned for the task, and a previous NER-centric baseline QACG. As the current publications on Claim Extraction, Fact Extraction, Claim Generation and Check-worthy Claim Detection are quite scattered in their means and terminology, we compile their common objectives, releasing the FEVERFact dataset, with 17K atomic factual claims extracted from 4K contextualised Wikipedia sentences, adapted from the original FEVER. We compile the known objectives into an Evaluation framework of: Atomicity, Fluency, Decontextualization, Faithfulness checked for each generated claim separately, and Focus and Coverage measured against the full set of predicted claims for a single input. For each metric, we implement a scale using a reduction to an already-explored NLP task. We validate our metrics against human grading of generic claims, to see that the model ranking on $F_{fact}$, our hardest metric, did not change and the evaluation framework approximates human grading very closely in terms of $F_1$ and RMSE.
Problem

Research questions and friction points this paper is trying to address.

Claim Extraction using text generation
Evaluation of Atomicity, Fluency, Faithfulness
Validation of metrics against human grading
Innovation

Methods, ideas, or system contributions that make the work stand out.

One-to-many text generation
FEVERFact dataset release
Evaluation framework implementation
🔎 Similar Papers