Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models

📅 2026-08-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型在决策中产生错误信息及信心错位问题,提出基于声明级别的置信度校准方法,通过分解响应并使用推理时信号进行校准。
📝 Abstract
Large Language Models (LLMs) increasingly support decision-making in high-stakes domains, but they often hallucinate and express confidence that is misaligned with factual correctness. Response-level confidence is a coarse signal: a single generation can mix correct and incorrect statements, so a single number is not actionable for users that must accept, reject, or verify individual pieces of information. We study claim-level confidence calibration as a decision-relevant uncertainty signal: each response is decomposed into atomic, verifiable claims, and each claim is assigned a calibrated confidence using inference-time signals from consistency across samples and self-verification. Our framework operates in closed-box settings (no logits, no fine-tuning) and applies post-hoc calibration directly at the claim level, enabling selective intervention such as evidence retrieval or human review for low-confidence claims. Across TriviaQA and TruthfulQA we evaluate seven baselines on six recent models (Llama-3.1, Mistral, Qwen2.5, DeepSeek-R1, GPT-4, GPT-4o), and show that claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Confidence Calibration
Decision Making
Innovation

Methods, ideas, or system contributions that make the work stand out.

claim-level confidence calibration
decision-relevant uncertainty signal
post-hoc calibration
selective intervention
💼 Related Jobs
No related jobs found.
T
Toghrul Abbasli
Tsinghua University
K
Kentaroh Toyoda
Vulcan Research, AIFT; Keio Global Research Institute (KGRI)
Y
Yuan Wang
China Mobile Research Institute
L
Li Chen
Zhongguancun Laboratory