🤖 AI Summary
This study addresses the problem of cross-segment credit misassignment caused by gradient noise in reinforcement learning for tool calling. To this end, we propose SLCA, a framework that introduces a novel segment-locked credit assignment mechanism. By integrating an LLM pattern-guided simulator with hierarchical rewards, SLCA achieves decoupled advantage estimation at the segment level, effectively eliminating advantage contamination without requiring additional intermediate state sampling. Experimental results demonstrate that, when applied to a 7B-parameter model, SLCA improves performance by 1.36 and 9.15 percentage points on the BFCL and τ²-Bench benchmarks, respectively. Furthermore, it significantly accelerates convergence and reduces overall training costs.
📝 Abstract
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on $τ^2$-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.