Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of error accumulation in multi-turn tasks, which causes student models to deviate from the teacher's supervision scope. To mitigate this, we propose GC-OPD, a method that eliminates the need for task-specific optimization of the teacher model. Instead, it leverages graph-structured indexing to retrieve historical states and alternative execution paths, thereby enriching the context for teacher scoring. By integrating execution evidence with retrieval-augmented generation, GC-OPD achieves more precise online policy distillation. Experimental results demonstrate that GC-OPD significantly improves task success rates on benchmarks such as ScienceWorld, outperforming existing baseline methods. Overall, this work presents an effective new paradigm for alleviating exposure bias in multi-turn interactions.
📝 Abstract
On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla OPD from 24.70% to 48.78% on ScienceWorld (4B student), from 53.36% to 85.26% on ALFWorld Unseen, and from 29.10% to 37.65% on WebShop. At matched student sizes, it also achieves higher mean success than every evaluated OPD baseline using GRPO-trained teachers on ScienceWorld and ALFWorld; the strongest such ScienceWorld 4B baseline reaches 46.66%. GC-OPD requires no task-specific teacher optimization.
Problem

Research questions and friction points this paper is trying to address.

on-policy distillation
compounding errors
language agents
multi-turn tasks
teacher supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Graph-Conditioned
Agent Distillation
Execution Evidence
Off-the-Shelf Teachers
🔎 Similar Papers
No similar papers found.
X
Xiaohan Yi
Yuanbao Team, Tencent; Tsinghua University
Wen Luo
Wen Luo
Peking University
Y
Yani Huang
Yuanbao Team, Tencent
J
Junfeng Zhan
Yuanbao Team, Tencent
A
Asher Qin
Yuanbao Team, Tencent
P
Peilin Zhao
School of Artificial Intelligence, Shanghai Jiao Tong University
Xi Xiao
Xi Xiao
Oak Ridge National Laboratory | University of Alabama at Birmingham
LLM / MLLM EfficiencyImage / Video GenerationImage / Video Understanding