LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

📅 2026-06-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究提出LongNovel,一个针对长篇小说摘要中幻觉检测的多尺度基准,通过多种方法确保数据真实性和幻觉类型均衡分布。
📝 Abstract
Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. https://github.com/BDML-lab/LongNovel
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
long-context novel summarization
multi-scale benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Scale Benchmark
Hallucination Detection
Long-Context Summarization
Bilingual Novels
Data Authenticity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ruizhi Zhang
Ruizhi Zhang
East China Normal University
Jinwei Chen
Jinwei Chen
vivo
computer vision
X
Xiangju Lu
iQIYI Inc
H
He Yan
iQIYI Inc
Mo Yu
Mo Yu
WeChat AI, Tencent
NLPQuestion AnsweringInformation ExtractionMachine Reading ComprehensionStory Understanding
J
Junmin Zhu
iQIYI Inc
W
Wei Zhang
East China Normal University, Shanghai Innovation Institute