The First Token Knows: Single-Decode Confidence for Hallucination Detection

📅 2026-05-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing hallucination detection methods rely on multiple sampling passes, incurring high computational costs and exhibiting sensitivity to lexical perturbations. This work proposes phi_first, a novel approach that estimates model confidence efficiently using only a single token—the first content word in a greedy decoding sequence—by computing the normalized top-K logit entropy. Requiring no repeated sampling, phi_first achieves an average AUROC of 0.820 across three instruction-tuned models (7–8B parameters) and two benchmarks, significantly outperforming semantic self-consistency (0.793) and surface-form self-consistency (0.791). This study demonstrates for the first time that uncertainty in text generation can be effectively captured through the confidence signal of a single token.
📝 Abstract
Self-consistency detects hallucinations by generating multiple sampled answers to a question and measuring agreement, but this requires repeated decoding and can be sensitive to lexical variation. Semantic self-consistency improves this by clustering sampled answers by meaning using natural language inference, but it adds both sampling cost and external inference overhead. We show that first-token confidence, phi_first, computed from the normalized entropy of the top-K logits at the first content-bearing answer token of a single greedy decode, matches or modestly exceeds semantic self-consistency on closed-book short-answer factual question answering. Across three 7-8B instruction-tuned models and two benchmarks, phi_first achieves a mean AUROC of 0.820, compared with 0.793 for semantic agreement and 0.791 for standard surface-form self-consistency. A subsumption test shows that phi_first is moderately to strongly correlated with semantic agreement, and combining the two signals yields only a small AUROC improvement over phi_first alone. These results suggest that much of the uncertainty information captured by multi-sample agreement is already available in the model's initial token distribution. We argue that phi_first should be reported as a default low-cost baseline before invoking sampling-based uncertainty estimation.
Problem

Research questions and friction points this paper is trying to address.

hallucination detection
self-consistency
confidence estimation
language models
uncertainty quantification
Innovation

Methods, ideas, or system contributions that make the work stand out.

first-token confidence
hallucination detection
self-consistency
uncertainty estimation
greedy decoding
💼 Related Jobs
No related jobs found.