Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation

πŸ“… 2024-11-25
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 7
✨ Influential: 1
πŸ“„ PDF
πŸ€– AI Summary
This paper addresses the fundamental question of whether automated factual consistency evaluation metrics genuinely assess alignment between summaries and source documents. To this end, we propose the first stress-testing framework based on difficulty-aware sample categorization, systematically benchmarking state-of-the-art metrics. Our findings reveal three critical limitations: (1) Most metrics exhibit substantial performance degradation in deep-reasoning scenarios and are highly susceptible to spurious correlations induced by irrelevant sentences; (2) Several metrics suffer from systematic β€œgaming” vulnerabilities, relying predominantly on shallow surface-level features rather than semantic entailment; (3) LLM-based prompting approaches (e.g., ChatGPT-DA), while comparatively robust, depend on parametric knowledge rather than source-grounded reasoning, introducing coverage bias. Collectively, these results expose foundational flaws in current factual consistency evaluation paradigms, providing both theoretical insights and methodological foundations for developing more reliable, source-aware assessment frameworks.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Reasoning under Uncertainty: Other Foundations of Reasoning under UncertaintyKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
πŸ“ Abstract
Modern LLMs can now produce highly readable abstractive summaries, to the point that traditional automated metrics for evaluating summary quality, such as ROUGE, have saturated. However, LLMs still sometimes introduce inaccuracies into summaries, i.e., information inconsistent with or unsupported by the corresponding source. Measuring the occurrence of these often subtle factual inconsistencies automatically has proved challenging. This in turn has motivated development of metrics intended to measure the factual consistency of generated summaries against sources. But are these approaches measuring what they purport to? Or are they mostly exploiting artifacts? In this work, we stress test a range of automatic factuality metrics, including specialized models and LLM-based prompting methods, to probe what they actually capture. Using a shallow classifier to separate ``easy''examples for factual evaluation where surface features suffice from ``hard''cases requiring deeper reasoning, we find that all metrics show substantial performance drops on the latter. Furthermore, some metrics are more sensitive to benign, fact-preserving edits than to factual corrections. Building on this observation, we demonstrate that most automatic factuality metrics can be gamed, i.e., their scores can be artificially inflated by appending innocuous, content-free sentences to summaries. Among the metrics tested, the prompt based ChatGPT-DA approach is the most robust and reliable. However, this comes with a notable caveat: Prompting LLMs to assess factuality may overly rely on their parametric knowledge rather than the provided reference when making judgments. Taken together, our findings call into question the reliability of current factuality metrics and prompt a broader reflection on what these metrics are truly measuring.
Problem

Research questions and friction points this paper is trying to address.

Evaluating whether automatic factuality metrics truly measure factual consistency in summaries
Testing if current metrics can distinguish subtle factual errors requiring deep reasoning
Investigating whether factuality scores can be artificially inflated without improving content
Innovation

Methods, ideas, or system contributions that make the work stand out.

Stress testing automatic factuality metrics for robustness
Separating easy and hard cases for factual evaluation
Demonstrating metrics can be gamed with content-free sentences
πŸ’Ό Related Jobs
No related jobs found.
Northeastern University
S
S. Ramprasad
Northeastern University
B
Byron C. Wallace
Northeastern University