External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of localizing semantic drift boundaries and the coarse granularity of conventional token-level binary classification in hallucination detection for large language models. We propose a cross-model span-level hallucination detection framework based on hidden state probing. This method introduces a novel external observer mechanism that analyzes inter-layer activation patterns, enabling a small model to monitor the internal representations of a larger model and precisely identify the onset and continuation spans of hallucinations. Experimental results demonstrate that the proposed framework significantly improves Precision-Recall AUC under extreme class imbalance conditions, achieving fine-grained hallucination isolation. These findings validate the effectiveness of cross-model collaboration in surpassing self-detection capabilities.
📝 Abstract
As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary classification task, failing to capture the structured, sequential boundaries of semantic drift. Here, we introduce an internal hidden state framework for fine-grained, span-level hallucination detection. By inspecting layer-wise activation patterns, we attempt to detect the exact hallucination onset and continuation tokens in an LLM generation. Our experiments show that this approach successfully isolates hallucination onsets, achieving substantial improvements in Precision-Recall AUC over random baselines despite extreme class imbalance. Ultimately, we propose a novel cross-model detection framework in which one model observes the internal representations elicited by another model's generation. We find that an external observer can match or exceed a generator's self-detection of its own hallucination onsets, including when the observer is the smaller model, suggesting that self-detection is not the ceiling for onset localisation.
Problem

Research questions and friction points this paper is trying to address.

Hallucination Detection
Large Language Models
Span-Level Detection
Hidden State Probing
Cross-Model Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Span-Level Hallucination Detection
Cross-Model Probing
Hidden State Analysis
Hallucination Onset Localization
Layer-wise Activation Patterns
🔎 Similar Papers
No similar papers found.