Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations

📅 2026-09-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过Latent Trajectory Dynamics和Action Representation Probe两种方法,利用模型内部表示来提高多轮代理设置中任务成功的预测准确性,优于表面生成和基于序列的校准基线。
📝 Abstract
As agentic systems getting adopted rapidly in safety critical applications, it is vital to measure the confidence associated with the agentic actions. In comparison to the traditional machine learning systems, agentic workflows have complex failure modes with planning, tool invocation and dynamic environment interactions. In this paper, we investigate whether model's internal representations provide stronger signals of eventual task success in multi-turn agentic setups. We introduce two complementary methods: Latent Trajectory Dynamics (LTD), which summarizes changes in residual-stream representations across an an interaction trajectory, and the Action Representation Probe (ARP), which predicts success from representations formed at action decisions. Across three interactive benchmarks (Bash, SQL, Python) and three model families (Qwen14B, Qwen7B, DeepSeek6.7B), our methods consistently outperform surface level generation and sequence-based calibration baselines providing a zero-overhead reliability monitor that requires neither prompt alterations nor multi-sample rollouts.
Problem

Research questions and friction points this paper is trying to address.

agentic systems
confidence measurement
internal representations
task success
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Trajectory Dynamics
Action Representation Probe
internal representations
multi-turn agentic setups
zero-overhead reliability monitor
🔎 Similar Papers