When Confidence Signals Disagree: Local and Global Confidence in Autoregressive Language Models

📅 2026-09-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了自回归语言模型中局部与全局置信度的差异及其对预测正确性和采样稳定性的影响,指出不同置信度读数不可互换。
📝 Abstract
Modern predictive systems expose multiple quantities that are commonly interpreted as measures of confidence. However, these quantities can summarize different aspects of the predictive process. This distinction matters when confidence is used to evaluate reliability or inform downstream oversight and control. We investigate whether different confidence readouts are empirically interchangeable in an autoregressive language model by comparing local confidence, defined from the probability of the greedy-selected answer token, with global confidence, defined from modal-answer frequency under repeated sampling. Across MMLU and ARC Challenge, the two signals are weakly correlated and differ substantially in their association with correctness: global confidence is moderately associated with correctness, whereas local confidence shows little association. We further test whether question-level disagreement between the signals is associated with sampling instability. On ARC, larger local--global confidence gaps are associated with higher answer entropy, more distinct sampled answers, and lower modal-answer concentration. The gap--entropy association persists when disagreement and instability are estimated from disjoint stochastic samples, indicating that it is not explained by shared finite-sample variation. The corresponding relationship is substantially weaker on MMLU, where only 4% of questions exhibit sampling instability. These results show that confidence readouts derived from the same predictive system are not empirically interchangeable and that their disagreement can provide a diagnostic of unstable sampling behavior. Confidence should therefore be treated as an explicitly defined measurement rather than as a single intrinsic scalar property of a model, particularly when it is used to inform downstream evaluation, oversight, or control.
Problem

Research questions and friction points this paper is trying to address.

confidence signals
autoregressive language models
local confidence
global confidence
predictive reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

local confidence
global confidence
sampling instability
autoregressive language models
confidence disagreement
🔎 Similar Papers
2024-08-21BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLPCitations: 1
💼 Related Jobs
No related jobs found.
J
Julio C. Amador Díaz López
SiftyML, LTD, London, United Kingdom