Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration

📅 2026-04-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of sparse rewards and cross-turn credit assignment in multi-turn tool-use tasks, which hinder the effectiveness of reinforcement learning (RL). The authors propose a hybrid training framework integrating MT-GRPO and GTPO, trained on a realistic customer-service user simulator grounded in large language models (LLMs). They introduce an iterative reward calibration mechanism to refine per-turn reward design, enhancing reward discriminability. Notably, this is the first study to successfully apply RL training on the Tau-Bench benchmark. The proposed GTPO mixed advantage estimation effectively mitigates the misalignment between reward discriminability and advantage direction. Experimental results show performance gains of 2.9 and 11.5 percentage points for Qwen3.5-4B and Qwen3-30B-A3B, achieving 66.7% and 69.5% success rates, respectively—surpassing GPT-4.1/GPT-4o with the smaller model and approaching Claude Sonnet 4.5 with the larger one.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPMultiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAISemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Training tool-calling agents with reinforcement learning on multi-turn tasks remains challenging due to sparse outcome rewards and difficult credit assignment across conversation turns. We present the first application of MT-GRPO (Multi-Turn Group Relative Policy Optimization) combined with GTPO (Generalized Token-level Policy Optimization) for training a tool-calling agent on realistic customer service tasks with an LLM-based user simulator. Through systematic analysis of training rollouts, we discover that naively designed dense per-turn rewards degrade performance by up to 14 percentage points due to misalignment between reward discriminativeness and advantage direction. We introduce Iterative Reward Calibration, a methodology for designing per-turn rewards using empirical discriminative analysis of rollout data, and show that our GTPO hybrid advantage formulation eliminates the advantage misalignment problem. Applied to the Tau-Bench airline benchmark, our approach improves Qwen3.5-4B from 63.8 percent to 66.7 percent (+2.9pp) and Qwen3-30B-A3B from 58.0 percent to 69.5 percent (+11.5pp) -- with the trained 4B model exceeding GPT-4.1 (49.4 percent) and GPT-4o (42.8 percent) despite being 50 times smaller, and the 30.5B MoE model approaching Claude Sonnet 4.5 (70.0 percent). To our knowledge, these are the first published RL training results on Tau-Bench. We release our code, reward calibration analysis, and training recipes.
Problem

Research questions and friction points this paper is trying to address.

tool-calling agents
multi-turn tasks
reinforcement learning
credit assignment
sparse rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Iterative Reward Calibration
Multi-Turn Reinforcement Learning
Tool-Calling Agents
GTPO
MT-GRPO
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wachiravit Modecrua
Amity Research and Application Center (ARAC)
K
Krittanon Kaewtawee
Amity Research and Application Center (ARAC)
K
Krittin Pachtrachai
Amity Research and Application Center (ARAC)
T
Touchapon Kraisingkorn
Amity Research and Application Center (ARAC)