consistency-aware summarization

Design and build summarization systems, evaluation methods, and contradiction-detection tools that detect, analyze, and mitigate factual or logical contradictions between generated summaries and their source texts, and that optimize models and outputs for faithfulness and consistency metrics.

consistency-awaresummarization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Verify with Caution: The Pitfalls of Relying on Imperfect Factuality Metrics

Jan 24, 2025
AG
Ameya Godbole
🏛️ University of Southern California

Existing factual consistency evaluation metrics exhibit unstable cross-dataset performance and frequently misestimate model capability—particularly under content rewriting or when source information spans long distances. Method: We systematically benchmark five mainstream factuality metrics across 11 summarization, RAG, and question-answering benchmarks, employing multi-metric横向 comparison, cross-task evaluation, and bias attribution experiments—including rewriting sensitivity and source span analysis. Contribution/Results: Our empirical study uncovers four fundamental flaws: systemic bias, poor domain transferability, literal translation preference, and neglect of distant contextual information. Metrics achieve only 0.32 average Spearman correlation with human judgments; 43% of datasets yield incorrect system rankings; and for highly rewritten outputs, misclassification rates exceed 68%. These findings challenge prevailing automatic evaluation practices and motivate a paradigm shift toward “human verification before deployment.”

Consistency IssuesFactuality EvaluationLanguage Models

Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation

Nov 25, 2024
SR
S. Ramprasad
🏛️ Northeastern University

This paper addresses the fundamental question of whether automated factual consistency evaluation metrics genuinely assess alignment between summaries and source documents. To this end, we propose the first stress-testing framework based on difficulty-aware sample categorization, systematically benchmarking state-of-the-art metrics. Our findings reveal three critical limitations: (1) Most metrics exhibit substantial performance degradation in deep-reasoning scenarios and are highly susceptible to spurious correlations induced by irrelevant sentences; (2) Several metrics suffer from systematic “gaming” vulnerabilities, relying predominantly on shallow surface-level features rather than semantic entailment; (3) LLM-based prompting approaches (e.g., ChatGPT-DA), while comparatively robust, depend on parametric knowledge rather than source-grounded reasoning, introducing coverage bias. Collectively, these results expose foundational flaws in current factual consistency evaluation paradigms, providing both theoretical insights and methodological foundations for developing more reliable, source-aware assessment frameworks.

Evaluating whether automatic factuality metrics truly measure factual consistency in summariesInvestigating whether factuality scores can be artificially inflated without improving contentTesting if current metrics can distinguish subtle factual errors requiring deep reasoning

STORYSUMM: Evaluating Faithfulness in Story Summarization

Jul 09, 2024
MS
Melanie Subbiah
🏛️ Columbia University | Answer.AI

This paper addresses the challenge of evaluating faithfulness in narrative text summarization by introducing StorySumm, the first fine-grained benchmark dedicated to detecting latent inconsistencies and errors in abstractive summaries. Methodologically, it constructs a dataset of short stories paired with LLM-generated summaries, annotated manually with localized error types, precise error spans, and semantic attributions; it further proposes the first explainable, span-level faithfulness evaluation framework tailored to narrative domains. Key contributions are threefold: (1) it exposes systematic blind spots in single-annotator human evaluation protocols and advocates a multi-source ground-truth fusion paradigm; (2) it provides the first faithfulness evaluation resource with explicit error localization and semantic attribution; and (3) empirical results show that state-of-the-art automatic metrics achieve at most 70% balanced accuracy, confirming StorySumm’s rigor as a challenging new benchmark.

Assessing automatic metrics for faithfulness evaluation accuracyDetecting challenging inconsistencies in narrative summariesEvaluating faithfulness in abstractive story summarization

In text summarization, diverse XAI methods frequently yield contradictory attributions for the same model output—a phenomenon termed the “disagreement problem”—which critically undermines explanation credibility and AI accountability. This work presents the first systematic empirical investigation of this issue in summarization. We propose Regionalized eXplainable AI (RXAI), a novel framework that abandons the global attribution assumption by partitioning the source document into semantically coherent segments and performing attribution independently per segment. RXAI integrates Sentence-BERT embeddings, hierarchical clustering for segmentation, and multiple XAI techniques—including Integrated Gradients and LIME—augmented with interactive sentence-level visualization. Evaluated on XSum and CNN/Daily Mail, RXAI significantly reduces cross-method attribution disagreement and improves local attribution consistency. Our approach establishes a new paradigm for fine-grained, reliable, and verifiable summarization explanations.

Addressing inconsistent explanations that reduce trust in AI-generated summariesInvestigating conflicting explanations from different XAI methods in text summarization modelsProposing segmentation-based approach to reduce disagreement between explanation methods

Logical Consistency of Large Language Models in Fact-checking

Dec 20, 2024
BG
Bishwamittra Ghosh
🏛️ Max Planck Institute for Software Systems | Aalborg University | Independent Researcher

This work addresses the critical issue of logical inconsistency in large language models (LLMs) when performing knowledge graph (KG)-augmented propositional logic fact-checking—particularly under complex logical queries involving negation, conjunction, and disjunction. To tackle this challenge, we propose a systematic solution comprising three key components: (1) the first curated benchmark suite explicitly designed to evaluate logical consistency across three categories of propositional logic queries; (2) a novel consistency metric for assessing LLM responses to formalized propositional logic queries; and (3) an integrated methodology combining retrieval-augmented generation (RAG), propositional logic formalization, KG embedding, and supervised fine-tuning to enhance reasoning stability. Empirical evaluation reveals severe logical inconsistency in state-of-the-art LLMs; our fine-tuned models achieve a 32.7% absolute improvement in consistency. All code and benchmarks are publicly released.

Addresses logical inconsistency in large language models (LLMs) under complex logical queries.Focuses on fact-checking tasks involving propositional logic queries from knowledge graphs (KGs).Proposes measures and fine-tuning to improve LLMs' logical consistency on complex queries.

Latest Papers

What's happening recently
View more

This work addresses the challenge that existing automatic evaluation metrics struggle to accurately assess the factual consistency of opinion summaries generated by large language models. To this end, the authors propose FactSim, an end-to-end fully automated evaluation method that extracts factual claims from both the generated summary and the original user reviews, and introduces a robust fact similarity scoring mechanism designed to handle negations, paraphrases, and elaborations. This approach effectively measures both factual consistency and coverage. Experimental results demonstrate that FactSim achieves significantly higher correlation with human judgments than current state-of-the-art automatic metrics, offering a more reliable proxy for human evaluation in assessing factual faithfulness.

evaluation metricsfact-checkingfactual consistency

This work proposes a novel re-ranking approach to enhance the factuality of generated summaries—defined as their consistency with the source document—by integrating, for the first time, a consensus mechanism based on Minimum Bayes Risk (MBR) decoding with a factuality-oriented evaluation of source-document consistency. The method jointly re-ranks multiple candidate summaries, effectively balancing diversity preservation with improved factual accuracy. Experimental results demonstrate that the proposed approach achieves strong performance on automatic metrics and significantly outperforms existing systems in human evaluations, thereby validating its effectiveness and novelty in advancing summary factuality.

consensusconsistencyfactuality

AlignCheck: a Semantic Open-Domain Metric for Factual Consistency Assessment

Dec 03, 2025
AA
Ahmad Aghaebrahimian
🏛️ Zurich University of Applied Sciences | Swiss Institute of Bioinformatics

Large language models (LLMs) frequently generate “hallucinations”—plausible yet factually incorrect statements—posing significant risks in high-stakes domains such as clinical decision-making. Existing evaluation metrics struggle to jointly ensure factual consistency and interpretability, hindering precise error diagnosis and correction. To address this, we propose FACT-DECOMP: a pattern-agnostic, semantic decomposition–based evaluation framework that quantifies factual consistency at fine-grained, interpretable levels. It comprises atomic fact extraction, semantic alignment modeling, weighted consistency scoring, and dynamic complexity control. Evaluated on both general-purpose and clinical benchmarks, FACT-DECOMP consistently outperforms state-of-the-art metrics (e.g., FactScore, FEVERScore) in accuracy and robustness, while supporting unified assessment across open-domain and domain-specific texts. The implementation is publicly available, providing a reproducible, debuggable foundation for developing fact-aware LLMs.

Addresses hallucination in high-stakes domains like clinical applicationsAssesses factual consistency in open-domain textsImproves interpretability and flexibility over existing evaluation metrics

Existing metrics for factuality and faithfulness struggle to evaluate how language models handle documents containing both supporting and contradictory evidence. This work proposes ConflictScore, the first formal and quantitative framework for assessing a model’s ability to recognize and articulate conflicting evidence. It decomposes model responses into atomic claims, fine-grained labels their relationships with all source documents, and introduces two complementary metrics: CS-C (Conflict Sensitivity) and CS-R (Response Reasonableness). Built upon this framework, the ConflictBench benchmark encompasses diverse conflict types. Experiments demonstrate that ConflictScore effectively identifies overconfident claims across domains and serves as a feedback signal that significantly improves model truthfulness on TruthfulQA.

conflicting evidenceevaluation metricsfactuality

Existing metrics for evaluating factual consistency in abstractive summarization lack sufficient reliability to effectively guide model training. To address this limitation, this work proposes an automated preference learning framework that obviates the need for complex reward shaping. The approach generates preference signals by aggregating multiple weak factuality indicators, filters out samples with high disagreement, and constructs a high-quality preference dataset relying solely on source documents. Training leverages pairs of summaries that are lexically similar yet differ in factual accuracy, combined with decoding strategy variations and reinforcement learning to optimize model performance. Experiments demonstrate consistent and significant improvements in factual consistency across models of varying scales, with smaller models achieving performance comparable to much larger counterparts.

evaluation metricsfactual consistencyfactuality

Hot Scholars

JZ

Jeff Z. Pan

Professor of Knowledge Computing, University of Edinburgh
Artificial IntelligenceKnowledge Representation and ReasoningKnowledge Based Learning
DR

Dan Roth

Professor of Computer Science, University of Pennsylvania
Natural Language ProcessingMachine LearningKnowledge Representation and ReasoningArtificial Intelligence
JC

Jiaoyan Chen

Department of Computer Science, University of Manchester
Knowledge GraphOntologyMachine LearningLarge Language Model
RL

Ru Li

Harbin Institute of Technology
SL

Siyi Liu

Hong Kong University of Science and Technology (Guangzhou)
Recommender SystemsInformation Retrieval