Calibration as a First-Class Criterion in LLM Evaluation

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
论文针对语言模型校准问题,提出在NLP研究中将校准作为评价模型的重要标准,通过已有数据和方法提升模型可信度。
📝 Abstract
Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.
Problem

Research questions and friction points this paper is trying to address.

Calibration
NLP
LLM Evaluation
Adoption Gap
Trustworthy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Calibration
LLM Evaluation
Confidence Scores
NLP Subfields
🔎 Similar Papers
No similar papers found.