Stochastic Teacher Intervention for Agentic On-Policy Distillation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in multi-turn agent online distillation where early student errors accumulate, causing trajectory deviation from the teacher distribution and rendering supervision ineffective. To overcome this, we propose the STI-OPD framework, which introduces a novel stochastic intervention mechanism based on policy divergence to replace fixed thresholds, dynamically substituting student actions to balance control with exploration. Furthermore, an importance-weighted reverse KL loss is incorporated to correct sampling bias in mixed trajectories, thereby enhancing optimization reliability. Comprehensive evaluations on tool-use reasoning and long-horizon interaction benchmarks demonstrate that our approach consistently outperforms state-of-the-art baselines, validating the effectiveness of both the stochastic intervention strategy and the importance weighting module.
📝 Abstract
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Multi-turn agentic tasks
Error accumulation
Distribution drift
Token-level supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Stochastic Teacher Intervention
Agentic Tasks
KL Divergence
Importance-Weighted Reverse KL
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1