Score
Design and build summarization systems, evaluation methods, and contradiction-detection tools that detect, analyze, and mitigate factual or logical contradictions between generated summaries and their source texts, and that optimize models and outputs for faithfulness and consistency metrics.
Large language models (LLMs) are susceptible to factual hallucinations induced by erroneous information in training data, undermining their reliability. Method: This paper systematically surveys factuality evaluation methodologies, addressing three core challenges: hallucination detection, limitations of existing benchmark datasets, and the reliability of evaluation metrics. We formulate five key research questions and propose a domain-customized fact-checking framework integrating instruction tuning, retrieval-augmented generation (RAG), multi-agent reasoning, and external knowledge integration. Enhanced interpretability and output consistency are achieved via advanced prompting strategies and domain-specific fine-tuning. Contribution/Results: Empirical results demonstrate that evidence-aligned evaluation—leveraging external verifiable sources—significantly outperforms purely autoregressive metrics in hallucination mitigation. The proposed framework advances the development of high-fidelity, context-aware, and domain-adapted trustworthy language models.
Existing factual consistency evaluation metrics exhibit unstable cross-dataset performance and frequently misestimate model capability—particularly under content rewriting or when source information spans long distances. Method: We systematically benchmark five mainstream factuality metrics across 11 summarization, RAG, and question-answering benchmarks, employing multi-metric横向 comparison, cross-task evaluation, and bias attribution experiments—including rewriting sensitivity and source span analysis. Contribution/Results: Our empirical study uncovers four fundamental flaws: systemic bias, poor domain transferability, literal translation preference, and neglect of distant contextual information. Metrics achieve only 0.32 average Spearman correlation with human judgments; 43% of datasets yield incorrect system rankings; and for highly rewritten outputs, misclassification rates exceed 68%. These findings challenge prevailing automatic evaluation practices and motivate a paradigm shift toward “human verification before deployment.”
This paper addresses the fundamental question of whether automated factual consistency evaluation metrics genuinely assess alignment between summaries and source documents. To this end, we propose the first stress-testing framework based on difficulty-aware sample categorization, systematically benchmarking state-of-the-art metrics. Our findings reveal three critical limitations: (1) Most metrics exhibit substantial performance degradation in deep-reasoning scenarios and are highly susceptible to spurious correlations induced by irrelevant sentences; (2) Several metrics suffer from systematic “gaming” vulnerabilities, relying predominantly on shallow surface-level features rather than semantic entailment; (3) LLM-based prompting approaches (e.g., ChatGPT-DA), while comparatively robust, depend on parametric knowledge rather than source-grounded reasoning, introducing coverage bias. Collectively, these results expose foundational flaws in current factual consistency evaluation paradigms, providing both theoretical insights and methodological foundations for developing more reliable, source-aware assessment frameworks.
This paper addresses the challenge of evaluating faithfulness in narrative text summarization by introducing StorySumm, the first fine-grained benchmark dedicated to detecting latent inconsistencies and errors in abstractive summaries. Methodologically, it constructs a dataset of short stories paired with LLM-generated summaries, annotated manually with localized error types, precise error spans, and semantic attributions; it further proposes the first explainable, span-level faithfulness evaluation framework tailored to narrative domains. Key contributions are threefold: (1) it exposes systematic blind spots in single-annotator human evaluation protocols and advocates a multi-source ground-truth fusion paradigm; (2) it provides the first faithfulness evaluation resource with explicit error localization and semantic attribution; and (3) empirical results show that state-of-the-art automatic metrics achieve at most 70% balanced accuracy, confirming StorySumm’s rigor as a challenging new benchmark.
In text summarization, diverse XAI methods frequently yield contradictory attributions for the same model output—a phenomenon termed the “disagreement problem”—which critically undermines explanation credibility and AI accountability. This work presents the first systematic empirical investigation of this issue in summarization. We propose Regionalized eXplainable AI (RXAI), a novel framework that abandons the global attribution assumption by partitioning the source document into semantically coherent segments and performing attribution independently per segment. RXAI integrates Sentence-BERT embeddings, hierarchical clustering for segmentation, and multiple XAI techniques—including Integrated Gradients and LIME—augmented with interactive sentence-level visualization. Evaluated on XSum and CNN/Daily Mail, RXAI significantly reduces cross-method attribution disagreement and improves local attribution consistency. Our approach establishes a new paradigm for fine-grained, reliable, and verifiable summarization explanations.
This work addresses the critical issue of logical inconsistency in large language models (LLMs) when performing knowledge graph (KG)-augmented propositional logic fact-checking—particularly under complex logical queries involving negation, conjunction, and disjunction. To tackle this challenge, we propose a systematic solution comprising three key components: (1) the first curated benchmark suite explicitly designed to evaluate logical consistency across three categories of propositional logic queries; (2) a novel consistency metric for assessing LLM responses to formalized propositional logic queries; and (3) an integrated methodology combining retrieval-augmented generation (RAG), propositional logic formalization, KG embedding, and supervised fine-tuning to enhance reasoning stability. Empirical evaluation reveals severe logical inconsistency in state-of-the-art LLMs; our fine-tuned models achieve a 32.7% absolute improvement in consistency. All code and benchmarks are publicly released.
This work addresses the challenge that existing automatic evaluation metrics struggle to accurately assess the factual consistency of opinion summaries generated by large language models. To this end, the authors propose FactSim, an end-to-end fully automated evaluation method that extracts factual claims from both the generated summary and the original user reviews, and introduces a robust fact similarity scoring mechanism designed to handle negations, paraphrases, and elaborations. This approach effectively measures both factual consistency and coverage. Experimental results demonstrate that FactSim achieves significantly higher correlation with human judgments than current state-of-the-art automatic metrics, offering a more reliable proxy for human evaluation in assessing factual faithfulness.
This work proposes a novel re-ranking approach to enhance the factuality of generated summaries—defined as their consistency with the source document—by integrating, for the first time, a consensus mechanism based on Minimum Bayes Risk (MBR) decoding with a factuality-oriented evaluation of source-document consistency. The method jointly re-ranks multiple candidate summaries, effectively balancing diversity preservation with improved factual accuracy. Experimental results demonstrate that the proposed approach achieves strong performance on automatic metrics and significantly outperforms existing systems in human evaluations, thereby validating its effectiveness and novelty in advancing summary factuality.
Large language models (LLMs) frequently generate “hallucinations”—plausible yet factually incorrect statements—posing significant risks in high-stakes domains such as clinical decision-making. Existing evaluation metrics struggle to jointly ensure factual consistency and interpretability, hindering precise error diagnosis and correction. To address this, we propose FACT-DECOMP: a pattern-agnostic, semantic decomposition–based evaluation framework that quantifies factual consistency at fine-grained, interpretable levels. It comprises atomic fact extraction, semantic alignment modeling, weighted consistency scoring, and dynamic complexity control. Evaluated on both general-purpose and clinical benchmarks, FACT-DECOMP consistently outperforms state-of-the-art metrics (e.g., FactScore, FEVERScore) in accuracy and robustness, while supporting unified assessment across open-domain and domain-specific texts. The implementation is publicly available, providing a reproducible, debuggable foundation for developing fact-aware LLMs.
Existing metrics for factuality and faithfulness struggle to evaluate how language models handle documents containing both supporting and contradictory evidence. This work proposes ConflictScore, the first formal and quantitative framework for assessing a model’s ability to recognize and articulate conflicting evidence. It decomposes model responses into atomic claims, fine-grained labels their relationships with all source documents, and introduces two complementary metrics: CS-C (Conflict Sensitivity) and CS-R (Response Reasonableness). Built upon this framework, the ConflictBench benchmark encompasses diverse conflict types. Experiments demonstrate that ConflictScore effectively identifies overconfident claims across domains and serves as a feedback signal that significantly improves model truthfulness on TruthfulQA.
Existing metrics for evaluating factual consistency in abstractive summarization lack sufficient reliability to effectively guide model training. To address this limitation, this work proposes an automated preference learning framework that obviates the need for complex reward shaping. The approach generates preference signals by aggregating multiple weak factuality indicators, filters out samples with high disagreement, and constructs a high-quality preference dataset relying solely on source documents. Training leverages pairs of summaries that are lexically similar yet differ in factual accuracy, combined with decoding strategy variations and reinforcement learning to optimize model performance. Experiments demonstrate consistent and significant improvements in factual consistency across models of varying scales, with smaller models achieving performance comparable to much larger counterparts.