Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the cold-start problem caused by sparse rewards in reinforcement learning for long-horizon agents by proposing the GATS method. GATS employs online policy distillation to provide token-level guidance signals and introduces a novel finding that distillation benefits depend on the performance gap between teacher and student models. Based on this insight, an adaptive weight scheduling mechanism is designed to automatically attenuate and eventually remove teacher guidance as the gap narrows, enabling a smooth transition from guided learning to autonomous exploration while supporting teacher models smaller than the student. Evaluated on benchmarks such as ALFWorld, GATS improves success rates by 4.37%–11.87% over GRPO and achieves the best average performance across all configurations.
📝 Abstract
Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at https://github.com/Ricardo-H/guide-then-let-go.
Problem

Research questions and friction points this paper is trying to address.

Sparse-Reward Reinforcement Learning
Cold-Start Problem
Long-Horizon Agents
On-Policy Distillation
Teacher-Student Performance Gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gap-Adaptive Teacher Scheduling
On-Policy Distillation
Sparse-Reward Reinforcement Learning
Cold-Start Problem
Agentic RL
🔎 Similar Papers
No similar papers found.