🤖 AI Summary
Current abstractive text summarization systems often generate fluent yet unfaithful content, particularly exhibiting hallucinations in the representation of inter-entity relationships, which compromises reliability in high-stakes applications. This work proposes a fine-grained evaluation framework grounded in relation triple extraction, leveraging dependency-aware relation extraction and multiple linguistically informed enhancements—such as named entity anchoring, passive agent recovery, and negation-aware verb modeling—to achieve more accurate relation identification. The framework further introduces a normalized Relation Hallucination Index (RHI) to enable fair comparisons across models and datasets. Experimental results demonstrate that the proposed approach yields more stable and discriminative faithfulness evaluations, advancing the development of automated, relation-level hallucination detection.
📝 Abstract
Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and
events. Such relation-level hallucinations undermine the reliability of generated summaries, particularly in high-stakes domains. In this work, we present a
refined and grounded framework for evaluating relation hallucination in abstractive summarization. We present the empirical Relation Hallucination Index (RHI) by
introducing a dependency-aware relation extraction algorithm that incorporates lemmatization-based normalization, named entity grounded subject resolution,
passive agent recovery, negation-aware verb modeling, reporting verb filtering, nominal relation fallback, clausal propagation, and systematic deduplication.
These enhancements improve the structural fidelity of extracted relation triples and reduce spurious matches during evaluation. In addition, we introduce a
normalized formulation of RHI to ensure scale-invariant comparison between datasets and models. The revised metric decomposes hallucination into interpretable
components, aggregates relation hallucination metric into a normalized relation faithfulness score. Extensive evaluation across multiple state-of-the-art
summarization models demonstrates that the grounded extraction process yields more stable and discriminative hallucination measurements. The proposed framework
advances automated relation-level faithfulness evaluation and supports coherence-aware, hallucination-sensitive model analysis.