U-Space: Uncovering When and Why Uncertainty Arises in Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the difficulty large language models face in accurately determining when to abstain from answering, noting that existing uncertainty estimation methods rely on repeated generation or additional training while failing to reveal the sources and evolution of uncertainty during reasoning. To overcome these limitations, this study proposes U-Space, a low-dimensional subspace approach that constructs orthogonal bases via semantic anchor mapping. By introducing mechanistic interpretability into uncertainty quantification for the first time, it achieves interpretable, token-level uncertainty measurement without requiring labels, training, or repeated generation, precisely localizing and tracking dynamic uncertainty shifts throughout reasoning independently of output length. Experiments demonstrate that U-Space outperforms mainstream baselines in confidence evaluation on reasoning benchmarks, exhibits robustness under length-controlled settings, and achieves significantly higher cross-task transfer reliability than supervised estimators.
📝 Abstract
Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: https://github.com/s2labres/U-Space.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Quantification
Large Language Models
Mechanistic Interpretability
Token-level Uncertainty
Generation Length Confounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Quantification
Mechanistic Interpretability
Token-level Uncertainty
U-Space
Length-controlled Evaluation
🔎 Similar Papers
No similar papers found.