Disentangling Self-Distillation: Measuring and Modeling Acquisition and Retention

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the entanglement of three critical axes—generation source, teacher coupling, and KL divergence direction—in existing self-distillation methods, which has led to conflicting conclusions. We present the first work to decouple these dimensions within a unified analytical framework. Leveraging Qwen2.5 and Ministral models, we conduct 1,200 experiments combining full factorial design with controlled theoretical modeling to systematically quantify each axis’s impact. Our findings reveal an inherent trade-off between knowledge acquisition and retention. Specifically, in contradictory tasks, teacher-generated data significantly enhances acquisition, while optimal strategies vary across architectures. This work provides precise theoretical explanations for performance variations under diverse task conditions.
📝 Abstract
Self-distillation with privileged context adapts a language model from demonstrations by letting the model, once conditioned on a reference response, teach its context-free copy token by token. Our taxonomy reveals existing methods differ along three entangled axes: (i) the rollout source (student or teacher), (ii) the teacher coupling (frozen, or an exponential moving average of the student at some coupling rate) and (iii) the KL direction (reverse or forward), yet these axes are usually studied in fixed combinations and have led to conflicting conclusions. We formalize a unifying framework to encompass all self-distillation methods vs classic supervised fine-tuning: we train every combination of the three axes, on Qwen2.5-7B and Ministral-3-3B across ordinary and contradictory tasks, totaling 1,200 adaptation runs, to systematically investigate the impact of the above axes. We propose a controlled model of the same objective to explain the resulting acquisition-retention trade-offs. We find that (i) the rollout source matters mostly where the task contradicts the pretrained behavior: there teacher rollouts raise acquisition well above what student rollouts achieve, with almost no change in retention; (ii) the teacher coupling changes acquisition most, on every task: acquisition rises with the coupling rate, then falls past a task-specific rate; (iii) switching the KL direction costs retention in one model but not the other so which axis to tune first depends on the model. The controlled model reproduces the three trends.
Problem

Research questions and friction points this paper is trying to address.

Self-Distillation
Acquisition-Retention Trade-off
Language Model Adaptation
Knowledge Retention
Privileged Context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Disentangled Framework
Acquisition-Retention Trade-off
Teacher Coupling
Controlled Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Luis Zuin
Huawei Technologies Co., Ltd., Paris, France
A
Alexis Huet
Huawei Technologies Co., Ltd., Paris, France
Dario Rossi
Dario Rossi
Scientific Director at Huawei, Former Professor at Telecom ParisTech & Ecole Polytechnique
network architecturemachine learninginternet measurementscongestion_controlnetworking .
Z
Zied Ben Houidi
Huawei Technologies Co., Ltd., Paris, France