SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of cross-segment credit misassignment caused by gradient noise in reinforcement learning for tool calling. To this end, we propose SLCA, a framework that introduces a novel segment-locked credit assignment mechanism. By integrating an LLM pattern-guided simulator with hierarchical rewards, SLCA achieves decoupled advantage estimation at the segment level, effectively eliminating advantage contamination without requiring additional intermediate state sampling. Experimental results demonstrate that, when applied to a 7B-parameter model, SLCA improves performance by 1.36 and 9.15 percentage points on the BFCL and τ²-Bench benchmarks, respectively. Furthermore, it significantly accelerates convergence and reduces overall training costs.
📝 Abstract
Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration without costly real APIs and stable training, we first construct the Schema-Guided LLM Simulator (SGLS) as foundational training infrastructure. Building on this, SLCA decouples advantage estimation at the structural segment level within a single group of rollouts, without requiring additional rollouts from intermediate states. Supported by Hierarchical Rewards (HierR), SLCA routes execution advantages to tool tokens and preference advantages to summary tokens, eliminating advantage contamination (the dominant cross-segment credit misattribution channel) within each policy update. On a 7B backbone, SLCA-GRPO accelerates convergence and outperforms standard GRPO, ToolPO, and RLTR by +2.53 pp on in-domain evaluation, +1.36 pp on the Berkeley Function-Calling Leaderboard (BFCL), and +9.15 pp on $τ^2$-Bench under the same training budgets, achieving higher accuracy with reduced tool redundancy and costs.
Problem

Research questions and friction points this paper is trying to address.

Tool-Calling Agents
Reinforcement Learning
Cross-Segment Credit Misattribution
Advantage Contamination
Output Heterogeneity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Segment-Locked Credit Assignment
Tool-Calling Reinforcement Learning
Schema-Guided LLM Simulator
Hierarchical Rewards
Cross-Segment Credit Misattribution
🔎 Similar Papers
No similar papers found.
Y
Yan Zhan
Peking University
Shaobo Liu
Shaobo Liu
Unknown affiliation
Q
Qiunan Liu
Tencent PCG QQ Team
Y
Yuanjun Shi
Tencent PCG QQ Team
S
Siqi Xu
Tencent PCG QQ Team
W
WeiYi Hou
Tencent PCG QQ Team
X
Xiang Xu
Tencent PCG QQ Team
Zekang Li
Zekang Li
Tencent PCG QQ Team
W
Weizhou Pan
Tencent PCG QQ Team
J
Jiahong Yan
Tencent PCG QQ Team