Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient coverage of supervision signals caused by fixed parameters in online self-distillation by proposing a neighborhood online self-distillation framework. Methodologically, an expert pool is constructed via local parameter perturbations, decoupling anchor directions from support levels. Greedy selection combined with MaxPeak dynamic routing is then employed to achieve complementary multi-expert supervision. Furthermore, a truncated forward KL divergence is adopted to transform reference corrections into effective optimization signals within the student's state space. Evaluated on mathematical benchmarks such as AIME, the proposed approach improves the average accuracy of Qwen-series models by 1.67 to 2.75 points, demonstrating its effectiveness.
📝 Abstract
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their pool covers more such positions than the unperturbed privileged teacher. We introduce Neighborhood OPSD (N-OPSD) to turn these corrections into supervision at student-visited states. Offline, greedy selection builds a compact pool of frozen experts by rewarding filtered reference-token gains beyond the pool's current best at each position. The highest-peak expert need not provide the best training target. Online routing therefore separates the anchor direction from its level of support. MaxPeak selects the anchor token, and quantile selection chooses among experts whose top token matches it. The student learns from the chosen expert's full next-token distribution through the clipped forward-KL objective inherited from OPSD. We evaluate on AIME 2024, AIME 2025, and HMMT February 2025. Across three independent runs per method, Neighborhood OPSD improves the three-benchmark Average@12 over OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B, respectively. Student-prefix continuations support using the pool beyond the reference trajectories used for selection. Matched ablations support filtered reference-token gains as a selection criterion. Accounting for overlap within the pool and routing by state further improve student accuracy. Inference uses only the distilled student.
Problem

Research questions and friction points this paper is trying to address.

on-policy self-distillation
mathematical reasoning
parameter perturbation
supervision coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-policy self-distillation
Neighborhood perturbation
Expert pool routing
Mathematical reasoning
Forward-KL objective
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
X
Xincheng Wei
The Chinese University of Hong Kong, Shenzhen
Y
Yifan Ding
Meituan, LongCat Team
Y
Yoshua Li
Meituan, LongCat Team
Y
Yuquan Lu
Meituan, LongCat Team
Ziheng Li
Ziheng Li
Peking University
Machine LearningNatural Language Processing
Y
Yi Lu
Meituan, LongCat Team, University of Toronto
D
Dongsheng Ma
Meituan, LongCat Team, Peking University
Rongxiang Weng
Rongxiang Weng
Meituan LLM Team
Large Language ModelsComputational Linguistics
X
Xunliang Cai
Meituan, LongCat Team