🤖 AI Summary
This study investigates whether the uncertainty expressed by large language models (LLMs) in the absence of sufficient context appropriately reflects the degree of missing information. Treating LLMs as implicit imputers, the work adapts uncertainty principles from multiple imputation and introduces a black-box diagnostic metric, ρ_R(α), to quantify how contextual information mitigates uncertainty. Through a five-level context-missing experimental setup on SQuAD, the authors estimate answer-level uncertainty—using response entropy and sampling-based confidence—via repeated sampling. Results show that response entropy reliably increases with greater context缺失 and exhibits a strong correlation with accuracy (R² gap as low as 0.057), substantially outperforming conventional confidence metrics.
📝 Abstract
Large language models (LLMs) are increasingly deployed in settings where the available context is incomplete or degraded. We argue that an LLM generating answers under incomplete context can be viewed as an implicit imputer, and evaluated against a criterion from the multiple imputation (MI) literature: uncertainty should scale with the amount of missing information. We assess this criterion on SQuAD, using a controlled framework in which context availability is varied across five levels. We evaluate two answer-level uncertainty measures that can be estimated from repeated sampling: sampling-based confidence (empirical mode frequency) and response entropy. Confidence fails to reflect increasing missingness: it remains high even as accuracy collapses. Entropy, by contrast, increases with context removal, consistent with the MI analogy, and explains substantially more variance in accuracy than confidence across all evidence levels (quadratic $R^2$ gap up to 0.057). We further introduce a black-box diagnostic $ρ_R(α)$ that estimates the proportion of baseline uncertainty resolved by context level $α$, requiring only repeated sampling with and without context. These results suggest that entropy is a more responsive black-box uncertainty measure than confidence under incomplete context.