🤖 AI Summary
This study addresses the challenges of localizing hallucinations and retrieving supporting evidence in large language models by proposing a novel joint modeling framework for hallucination detection and input evidence alignment. Integrating encoder masked prediction, confidence estimation, and token-level alignment mechanisms, the approach enables fine-grained hallucination identification with interpretable traceability. This method overcomes limitations of existing detection techniques by precisely pinpointing hallucinated segments while associating them with high-quality input evidence. Human evaluation confirms the reliability of the alignment results, demonstrating significant improvements in the trustworthiness and transparency of generated content. Collectively, this work establishes a new paradigm for mitigating hallucinations through verifiable evidence grounding.
📝 Abstract
Hallucinations remain a major obstacle to the reliable use of large language models (LLMs) in conditional text generation. Existing methods primarily assess the factuality of an entire generated text, providing limited insight into which output spans are hallucinated or how they relate to the input. We introduce the task of hallucination span detection with input-side evidence alignment, which jointly identifies hallucinated spans and aligns output tokens with the corresponding input evidence. Our approach is based on the observation that faithful output tokens are predictable from the input, whereas hallucinated tokens are not. We therefore train an encoder-based model to predict masked output tokens from the input representation, using prediction confidence for hallucination detection while naturally producing alignments to the input. Experiments show that the proposed method effectively detects hallucinated spans and identifies meaningful input-side evidence. Human evaluation confirms the quality of the predicted alignments.