🤖 AI Summary
To address named entity hallucination—i.e., the generation of entities absent from or inconsistent with the source text—in abstractive summarization, this paper proposes a reinforcement learning framework that requires no human-annotated factual labels. The core contribution is the Entity Hallucination Index (EHI), a computationally tractable and differentiable reward signal quantifying factual deviation by automatically extracting and aligning named entities between summaries and source documents. Built upon a pre-trained language model, our method jointly optimizes EHI alongside standard language modeling objectives during summary generation. Experiments demonstrate significant reductions in EHI, substantial improvements in entity accuracy and factual consistency, and preservation of summary fluency and informativeness. To ensure reproducibility, we release both source code and a fully executable Colab notebook.
📝 Abstract
Reducing hallucinations in abstractive summarization remains a critical challenge for deploying language models (LMs) in real-world settings. In this work, we introduce a rewarddriven fine-tuning framework that explicitly optimizes for Entity Hallucination Index (EHI), a metric designed to quantify the presence, correctness, and grounding of named entities in generated summaries. Given a corpus of meeting transcripts, we first generate baseline summaries using a pre-trained LM and compute EHI scores via automatic entity extraction and matching. We then apply reinforcement learning to fine-tune the model parameters, using EHI as a reward signal to bias generation toward entity-faithful outputs. Our approach does not rely on human-written factuality annotations, enabling scalable fine-tuning. Experiments demonstrate consistent improvements in EHI across datasets, with qualitative analysis revealing a significant reduction in entity-level hallucinations without degradation in fluency or informativeness. We release a reproducible Colab pipeline, facilitating further research on hallucination-aware model fine-tuning using lightweight, hallucintion metrics like EHI.