Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of sparse terminal reward credit assignment and the difficulty of obtaining step-level supervision in long-horizon tool calling. To this end, we propose CITA, a framework that trains comparative reasoning models to evaluate the long-term value of candidate tool calls prior to execution. The core innovation lies in automatically generating paired supervision signals through Bayesian simulators and large language model-based semantic judgments, thereby enabling comparative learning without human annotation. Experimental results demonstrate that CITA significantly improves tool-calling F1 scores and task success rates across multiple benchmarks while achieving accurate step-level value estimation.
📝 Abstract
Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.
Problem

Research questions and friction points this paper is trying to address.

long-horizon tool use
credit assignment
step-level rewards
comparative value estimation
tool-use agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comparative Value Estimation
Long-Horizon Tool-Use Agents
Comparative Inference Model
Bayesian Tool-Graph Simulator
Step-level Rewards
🔎 Similar Papers
No similar papers found.
Y
Yu Li
School of Computer Science and Engineering, Southeast University, Nanjing, China
Z
Zheng Zhang
Shandong Jianzhu University, China
X
Xin Liu
UniSA STEM, University of South Australia, Adelaide, Australia
S
Shengtian Yang
School of Computer Science and Engineering, Southeast University, Nanjing, China
G
Guangfeng Cai
School of Computer Science and Engineering, Southeast University, Nanjing, China
Lei Feng
Lei Feng
Professor, Southeast University
Machine LearningData ScienceStatistics