Learning from Teacher Continuations at Student States

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of covariate shift in offline distillation, fragmented supervision in online settings, and reliance on teacher probability distributions by proposing OLIVE. In this method, a student model dynamically generates prefixes that are autoregressively continued by a teacher model. Employing an asynchronous training architecture with cross-entropy loss, OLIVE directly updates the student policy without requiring access to token-level teacher probabilities, thereby enabling efficient knowledge transfer. Experimental results demonstrate that, under equivalent computational budgets, OLIVE reduces training time by 23.8% and improves performance by 13% on the ScienceWorld benchmark. Furthermore, its inference performance surpasses that of existing methods while maintaining favorable plasticity.
📝 Abstract
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
covariate shift
on-policy distillation
supervised fine-tuning
language model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Distillation
Student-State Prefix Generation
Autoregressive Teacher Continuation
Asynchronous Training
Knowledge Distillation
🔎 Similar Papers
2024-07-18arXiv.orgCitations: 0