Shared Experience, Separate Learning: Companion Confidence Calibration for LLMs

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the overconfidence problem in large language models arising from the coupling of capability and confidence learning, proposing the CoCal framework. This method introduces a novel "shared experience, decoupled learning" paradigm that enables capability and confidence to be independently optimized along a shared trajectory via separate parameters. In reinforcement learning with verifiable rewards (RLVR) scenarios, CoCal trains a lightweight companion model on hidden states under verifiable feedback supervision, achieving concurrent confidence calibration without compromising task performance. Experiments demonstrate that CoCal significantly improves confidence estimation accuracy across the Qwen3 model series while preserving original task performance and exhibiting strong cross-domain generalization.
📝 Abstract
Reliable self-assessment is essential for large language models (LLMs), yet they often remain highly confident when their answers are wrong. We study \emph{concurrent confidence calibration}, where confidence is learned alongside capability improvement rather than calibrated only after training. Reinforcement learning from verifiable rewards (RLVR) provides a natural setting for this paradigm, as it continuously produces responses paired with verifiable correctness feedback. Existing concurrent methods, however, learn both capability and confidence through reinforcement learning within shared policy parameters, potentially coupling two fundamentally different learning problems. We instead propose \emph{shared experience but separate learning}: capability and confidence learn from the same trajectories, but through separate optimization mechanisms and parameters. Based on this principle, we introduce \textbf{CoCal (Companion Confidence Calibration)}, which trains a lightweight companion from rollout hidden states and verifier-derived correctness supervision while leaving task optimization unchanged. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task performance, outperforming both RL-based concurrent methods and matched post-hoc calibration. The learned companion further generalizes across domains and policy shifts, while the benefits of CoCal persist at both scales.
Problem

Research questions and friction points this paper is trying to address.

confidence calibration
large language models
self-assessment
concurrent learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Confidence Calibration
Large Language Models
Reinforcement Learning from Verifiable Rewards
Companion Network
Decoupled Optimization
🔎 Similar Papers