Proper Scoring Rules for Agentic Uncertainty Quantification

📅 2026-05-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in existing methods for evaluating uncertainty in language agents, which conflate ranking utility with probabilistic fidelity and fail to accurately capture the trajectory of success probabilities under prefix conditions. Drawing on preordered proper scoring theory, the paper introduces Trajectory Proper Scoring (TPS)—the first family of strictly proper scoring rules applicable to both complete and truncated trajectories—that rigorously incentivizes language models to emit stepwise uncertainty estimates aligned with their true success probabilities. By integrating preordered proper scoring, trajectory-level rule design, truncation handling, and projection-based approximation, TPS demonstrates markedly higher sensitivity to calibration shifts than conventional metrics across StrategyQA, Tau2-Bench, HotpotQA, and WebShop. Notably, truncation-aware approximations can substantially alter evaluation outcomes, revealing that current approaches capture only weak proxies of true calibration.
📝 Abstract
Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace $q_t = P^π(Y=1 | H_t)$. Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact $q_Z$-weighted reduced score and a tractable approximation when $q_Z$ is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Quantification
Proper Scoring Rules
Trajectory Evaluation
Calibration
Language Model Agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trajectory Proper Score
Uncertainty Quantification
Proper Scoring Rules
Prefix-conditioned Calibration
Censored Trajectories
S
Suresh Raghu
Independent Researcher
S
Satwik Pandey
Independent Researcher
S
Shashwat Pandey
Independent Researcher