🤖 AI Summary
This study addresses whether the internal probabilities of large language models align with and are interchangeable with their verbalized confidence. By intervening on uncertainty sources within training data, this work systematically investigates the coupling between these two probability readout mechanisms through causal interventions, distributional statistical analysis, and model probing techniques. The findings reveal that both mechanisms are driven by a common source and exhibit strong alignment, demonstrating that verbalized probabilities can serve as an effective proxy for probing internal distributions. This establishes a significant coupling relationship between internal and verbalized probabilities, providing theoretical grounding and methodological support for reliably assessing model uncertainty via textual confidence.
📝 Abstract
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that internal probabilities track relative frequencies in the training data, and that verbalized probabilities track explicit probabilistic assertions in the training data. However, we do not know whether these two readouts are aligned, except when frequencies and probabilistic assertions in the training data happen to align. This limits our understanding of when we can use verbalized uncertainties as a proxy for either training data frequencies, or a model's internal distribution. We resolve this gap by systematically exploring how LLMs probability readouts are impacted by training and in-context data, via intervening on the underlying uncertainty sources in the data. We find that both internal and verbalized probability readouts are impacted by both distributional and asserted uncertainty in the training data. Further, we find that verbalized and internal probabilities are aligned beyond what would be expected by independently tracking the same uncertainty sources, suggesting that verbalized probabilities can be used to probe a model's internal distribution.