Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
This study addresses the challenges of sparse terminal reward credit assignment and the difficulty of obtaining step-level supervision in long-horizon tool calling. To this end, we propose CITA, a framework that trains comparative reasoning models to evaluate the long-term value of candidate tool calls prior to execution. The core innovation lies in automatically generating paired supervision signals through Bayesian simulators and large language model-based semantic judgments, thereby enabling comparative learning without human annotation. Experimental results demonstrate that CITA significantly improves tool-calling F1 scores and task success rates across multiple benchmarks while achieving accurate step-level value estimation.