Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitation of on-policy contextual distillation, where employing instance-level ground-truth answers as privileged information degrades out-of-distribution (OOD) generalization. To overcome this, we propose substituting such answers with generic formatted instructions targeting common errors as the teacher model’s privileged information, optimizing the student model via Kullback-Leibler divergence minimization. Notably, this work provides the first demonstration that concise, universal instructions exhibit superior transferability compared to information-dense, instance-specific answers. Extensive experiments across eight autoformalization tasks reveal that our approach improves OOD accuracy by 4 to 17 percentage points in seven settings while preserving in-domain performance, establishing a more effective paradigm for knowledge distillation in formal reasoning.
πŸ“ Abstract
On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instructions that target common student mistakes observed on the training samples, and show that such simple instructions can outperform gold as the OPCD privilege. In autoformalization tasks, using a matched formatting instruction as the privilege could outperform gold in OOD accuracy by a large margin. In 7 out of 8 experiments using ProverQA, ProofWriter, and ProntoQA as datasets, and Qwen3-Thinking and Olmo3-Thinking families as models, matched instruction privileges outperform gold in OOD by 4 to 17 points, while remaining on par with gold in-domain. Each instruction is only a few sentences (and thus contains much less information compared to all instance-specific gold) and is applied uniformly to every training sample. These results indicate that a general instruction, which applies equally to source and target domain examples, can be substantially more transferable than instance-specific gold in OPCD while maintaining in-domain performance.
Problem

Research questions and friction points this paper is trying to address.

On-Policy Context Distillation
Out-of-Distribution Generalization
Knowledge Distillation
Privileged Information
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Context Distillation
Instruction Privileges
Out-of-Distribution Generalization
Knowledge Distillation
Autoformalization
πŸ”Ž Similar Papers
No similar papers found.