When Words Fall Short: Iterative Synergy Between Verbalized Reasoning and Hidden Features for LLM Confidence Estimation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnect between verbalized reporting and standalone estimators in confidence estimation for large language models, as well as their underutilization of internal representations. To this end, we propose IPoET, a framework that introduces an iterative policy-estimator co-training mechanism. By alternately optimizing and dynamically fusing verbalized reasoning trajectories with hidden-state features, IPoET challenges the conventional assumption that verbalized reports are inherently superior to estimators. In-domain and cross-domain experiments conducted on Qwen and Llama backbones demonstrate that IPoET significantly outperforms existing estimator-based and verbalization-based baselines, achieving state-of-the-art performance. These results confirm that the proposed framework successfully realizes synergistic enhancement between the two paradigms, offering a more effective approach to reliable confidence estimation in large language models.
📝 Abstract
Confidence estimation is crucial for developing trustworthy large language models (LLMs), with most methods following estimator-based or verbalization-based paradigms. While recent research increasingly focuses on improving verbalized self-reports of confidence, we challenge the prevailing view that this approach surpasses independent confidence estimators. Our empirical study shows that a dedicated confidence estimator can substantially outperform verbalized confidence, indicating that LLMs' internal representations contain richer confidence signals. Building on this finding, we propose Iterative Policy-Estimator Training (IPoET), a framework that synergizes the complementary strengths of verbalized reasoning traces and informative representations. IPoET alternates policy optimization with estimator updating, integrating estimator-derived confidence feedback into policy learning and refreshing the estimator on new policy rollouts. Experiments across diverse datasets and Qwen and Llama backbones demonstrate that, by iteratively exploiting richer hidden features and adapting to the evolving policy distribution, IPoET consistently outperforms both estimator- and verbalization-based baselines in-domain and achieves superior or comparable results across all out-of-domain metrics. For more details, refer to https://github.com/xyk829/ipoet.
Problem

Research questions and friction points this paper is trying to address.

Confidence Estimation
Large Language Models
Verbalized Reasoning
Hidden Features
Trustworthy AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

Confidence Estimation
Iterative Policy-Estimator Training
Verbalized Reasoning
Hidden Features
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yekun Xu
Yekun Xu
Unknown affiliation
A
Ante Wang
Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China
Jingyi Ren
Jingyi Ren
Tsinghua University
X
Xuanyi Chen
Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China
Weizhi Ma
Weizhi Ma
Tsinghua University
LLM and AgentsRecommendationAI for Healthcare
Y
Yang Liu
College of AI, Tsinghua University, Beijing, China; Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China; Dept. of Comp. Sci. & Tech., Institute for AI, Tsinghua University, Beijing, China