Also Small Models Can Reasonably Self-Evaluate Their Confidence

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the reliability of self-assessed confidence in language models of varying scales on question-answering (QA) tasks. By employing multiple self-evaluation methods and conducting systematic comparative experiments across cross-domain QA benchmarks, we comprehensively assess the uncertainty quantification capabilities of different-sized models over diverse knowledge domains. Our results demonstrate that the reliability of self-assessed confidence is independent of both model scale and absolute accuracy; despite lower overall performance, smaller models yield consistently reliable confidence signals. These findings challenge the prevailing assumption that only large-scale models possess robust self-awareness. Consequently, this work provides a critical theoretical foundation for the efficient deployment of lightweight yet trustworthy models in resource-constrained scenarios.
📝 Abstract
This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions. This independence means the most capable model is not necessarily the best at self-assessing prediction reliability. These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.
Problem

Research questions and friction points this paper is trying to address.

uncertainty quantification
self-evaluation
language models
confidence estimation
model scale
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Evaluation
Uncertainty Quantification
Confidence Reliability
Small Language Models
Model Scale
🔎 Similar Papers
No similar papers found.