Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing self-distillation methods, where fixed token selection strategies fail to adaptively identify optimal distillation positions. To overcome this, we propose BiToK-SD, a framework that formulates Top-K token selection as a bi-level optimization problem with differentiable threshold relaxation for the first time. Through strategy evolution, the student model dynamically learns optimal distillation positions during training, enabling efficient online self-distillation. Experimental results demonstrate that BiToK-SD achieves state-of-the-art average performance on mathematical reasoning benchmarks while incurring only lightweight computational overhead. This work significantly enhances both the flexibility and efficiency of knowledge distillation by replacing rigid token selection with an adaptive, end-to-end trainable mechanism.
📝 Abstract
Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token positions uniformly or select tokens using fixed heuristic criteria, assigning the same distillation strength to the selected positions rather than adaptively learning which tokens are most beneficial for distillation. To address these limitations, we propose BiToK-SD (Bilevel Top-K Token Selection for Self-Distillation), a bilevel-optimization-based token selection method that learns where distillation should be applied during on-policy self-distillation. Specifically, BiToK-SD is formulated as a bilevel optimization problem, where the lower-level problem models Top-K token selection as a differentiable threshold-based relaxation, allowing the selected positions to adapt as the student policy evolves, while the upper-level problem performs knowledge distillation on the selected positions. Experiments on mathematical reasoning benchmarks show that BiToK-SD achieves the best average performance among all compared methods while requiring only lightweight additional computation.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
self-distillation
token selection
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Bilevel Optimization
Token Selection
Large Language Models
Differentiable Relaxation
🔎 Similar Papers
2024-06-19International Conference on Computational LinguisticsCitations: 1