TRACE: Single-Pass Decoding-Trace Risk Localization for Generation Calibration

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that local errors in large language model generation are often obscured by global confidence scores, thereby complicating calibration. To overcome this, it proposes a single-pass decoding trajectory risk localization method that pioneers modeling decoding uncertainty as trajectories. By recording token-level surprisal and predictive entropy while employing local risk operators to preserve uncertainty spikes, the approach achieves answer-level risk scoring and probability calibration without requiring labeled data. Evaluated across four tasks, the proposed method significantly outperforms nineteen baselines, reducing the Brier score to 0.137 and improving the AUROC to 0.792, which demonstrates its generalizability and effectiveness.
📝 Abstract
Reliable confidence estimation is essential for large language model deployment. However, answer-level calibration remains challenging because generation errors are often localized: a response may be fluent and high-probability overall while still failing at a critical number, entity, or factual claim. Existing estimators compress token probabilities, sequence likelihoods, entropy, or beam statistics into a global score, which can dilute such local risk signals. We propose TRACE, a single-pass, decoded-answer-preserving confidence estimator that treats decoding-time uncertainty as a trajectory through three steps: (i) recording token-level surprisal and predictive entropy during decoding, (ii) applying local risk operators to preserve uncertainty spikes, and (iii) converting localized trace risk into answer-level confidence. TRACE produces a label-free risk score, while TRACE+ calibrates trace-only features into probabilities using a held-out split, without extra generations or external verifiers. We evaluate four tasks against 19 calibration baselines, and TRACE+ reduces Brier from 0.149 to 0.137 and improves AUROC from 0.758 to 0.792 over the strongest likelihood baseline. Across seven LLMs, TRACE+ improves over the best non-TRACE baseline pool from 0.136 to 0.120 Brier and from 0.764 to 0.817 AUROC. Results show that localizing decoding-time risk provides a general approach to calibration.
Problem

Research questions and friction points this paper is trying to address.

confidence estimation
generation calibration
large language models
local risk localization
decoding uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

confidence estimation
decoding-trace
risk localization
generation calibration
single-pass decoding
🔎 Similar Papers
2023-03-04International Conference on Learning RepresentationsCitations: 13
Y
Yuebin Xu
The Hong Kong University of Science and Technology (Guangzhou)
X
Xuemei Peng
The Hong Kong University of Science and Technology (Guangzhou)
J
Junlan Chen
The Hong Kong University of Science and Technology (Guangzhou)
Z
Zhiyi Chen
The Hong Kong University of Science and Technology (Guangzhou)
Zeyi Wen
Zeyi Wen
Assistant Professor at HKUST(Guangzhou)
Efficient LLMsMLSysHPOHPC