🤖 AI Summary
This study addresses the vulnerability of multi-turn agent online distillation to critical-step errors, which frequently induce trajectory deviation and task failure. To mitigate this issue, we propose a teacher-confidence-based adaptive intervention mechanism that employs dynamic threshold scheduling to distinguish uncertainty scenarios. Specifically, the teacher intervenes to provide corrective guidance under high uncertainty, whereas standard loss optimization—integrating supervised fine-tuning with KL divergence minimization—is executed under low uncertainty to refine the student policy. Experimental evaluations demonstrate that the proposed approach significantly outperforms existing baselines on benchmarks such as ALFWorld, achieving a 15.8% relative improvement in score on the WebShop task.
📝 Abstract
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.