🤖 AI Summary
Existing summarization evaluation metrics, such as ROUGE, struggle to effectively quantify the abstractiveness of generated summaries and fail to capture the fundamental distinction between extractive and abstractive approaches. This work proposes a novel framework for measuring abstractiveness through three heuristic indicators: Reference Abstractiveness (RA), Summary Abstractiveness (SA), and Abstractiveness Ratio (AR), augmented by document-length modulation and a cubic non-overlap factor to quantify the degree of deviation from the source text. Notably, the introduction of AR enables the detection of potential hallucinations, yielding an evaluation system that is dimensionally consistent, bounded, and nonlinearly sensitive to the extractive–abstractive boundary. Experiments on XSUM demonstrate the metric’s efficacy, clearly differentiating extractive models (SA ≈ 0.12–0.26) from abstractive ones (SA ≈ 0.96–1.77), thereby validating its practical utility and theoretical soundness.
📝 Abstract
Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond
surface-level metrics like ROUGE. We introduce Reference Abstraction (RA), Summary Abstraction (SA), and Abstraction
Ratio (AR) -- a set of principled heuristic metrics that measure how much a summary diverges from extractive copying
of the source text. The formulation uses the harmonic mean of document lengths modulated by a cubic non-overlap
factor, yielding dimensionally consistent, bounded output with non-linear sensitivity to the extractive-abstractive
boundary. Evaluation on 100 XSUM documents across four summarization models (BART-large-cnn, Pegasus-xsum, DistilBart,
MT5-small) demonstrates that the metrics successfully discriminate between extractive models (SA ~ 0.12-0.26) and
abstractive models (SA ~ 0.96-1.77), and that the Abstraction Ratio identifies summaries requiring manual evaluation
for potential hallucination. Code and results are available at https://github.com/katweNLP/AbstractionStudy.