Examining the Metrics for Document-Level Claim Extraction in Czech and Slovak

📅 2025-11-18
📈 Citations: 0
Influential: 0
📄 PDF

career value

161K/year
🤖 AI Summary
Document-level claim extraction for fact-checking lacks robust evaluation methodologies—particularly for low-resource languages (e.g., Czech and Slovak) and informal text. To address this, we propose the first multilingual, semantics-aware evaluation framework for document-level claim extraction, grounded in three core criteria: atomicity, verifiability, and decontextualization. Our framework enables reliable comparison between model outputs and human annotations via claim-set alignment and fine-grained semantic similarity computation. Experiments on a newly curated Czech/Slovak news commentary dataset demonstrate that conventional metrics (e.g., precision/recall) severely underestimate model performance. In contrast, our approach more accurately reflects model capabilities, quantifies inter-annotator agreement, and establishes a reproducible, interpretable evaluation paradigm for cross-lingual fact-checking.

Technology Category

Application Category

📝 Abstract
Document-level claim extraction remains an open challenge in the field of fact-checking, and subsequently, methods for evaluating extracted claims have received limited attention. In this work, we explore approaches to aligning two sets of claims pertaining to the same source document and computing their similarity through an alignment score. We investigate techniques to identify the best possible alignment and evaluation method between claim sets, with the aim of providing a reliable evaluation framework. Our approach enables comparison between model-extracted and human-annotated claim sets, serving as a metric for assessing the extraction performance of models and also as a possible measure of inter-annotator agreement. We conduct experiments on newly collected dataset-claims extracted from comments under Czech and Slovak news articles-domains that pose additional challenges due to the informal language, strong local context, and subtleties of these closely related languages. The results draw attention to the limitations of current evaluation approaches when applied to document-level claim extraction and highlight the need for more advanced methods-ones able to correctly capture semantic similarity and evaluate essential claim properties such as atomicity, checkworthiness, and decontextualization.
Problem

Research questions and friction points this paper is trying to address.

Developing evaluation metrics for document-level claim extraction in fact-checking
Creating alignment methods to compare model-extracted and human-annotated claim sets
Addressing challenges in informal Czech and Slovak news comment analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aligning claim sets from same source documents
Computing similarity through alignment score metrics
Evaluating semantic properties like atomicity and checkworthiness
🔎 Similar Papers
No similar papers found.
L
Lucia Makaiová
Brno University of Technology, Czech Republic
M
Martin Fajčík
Brno University of Technology, Czech Republic
A
Antonín Jarolím
Brno University of Technology, Czech Republic