Score
Designs and implements teacher–student distillation pipelines that transfer knowledge from a privileged, physics-aware teacher into a compact student by embedding analytical physical priors or other privileged information into the distillation objective. Builds and evaluates the loss functions, model architectures, and training procedures needed to preserve predictive fidelity with limited data and computational resources so the resulting lightweight models can perform accurate, real-time inference while respecting the incorporated priors.
This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.
Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.
Conventional knowledge distillation focuses on behavioral imitation, treating the teacher model as a black box and transferring only output distributions. Method: This paper proposes “circuit distillation”—a mechanism-aware approach that transfers the teacher’s underlying computational architecture rather than superficial behavior. It aligns interpretable internal components (e.g., entity-tracking and theory-of-mind circuits) between Llama3-based teacher and student models via functional circuit correspondence matching and representation similarity loss. Contribution/Results: To our knowledge, this is the first work achieving mechanism-level distillation grounded in functional circuits. It enables transfer of complex algorithmic capabilities with minimal parameter tuning—only a small subset of student parameters is fine-tuned. The method significantly enhances model interpretability and controllability. Empirical evaluation on entity tracking and theory-of-mind tasks demonstrates superior performance over conventional distillation baselines, validating that mechanistic alignment—not just statistical mimicry—is essential for effective algorithmic capability transfer.
This work addresses the long-standing absence of a unified statistical perspective on knowledge distillation, which has frequently been perceived as an engineering heuristic. We propose a unifying framework grounded in Bayesian inference that formalizes teacher model predictions as prior information, thereby enabling principled uncertainty quantification. This framework not only bridges classical distillation methods with their extensions to large language models but also integrates seamlessly with modern generative systems. Furthermore, we provide a conceptual roadmap and identify key open problems, establishing a systematic foundation for deepening the theoretical understanding of distillation mechanisms.
In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.
This work addresses the limitation of existing single-step diffusion distillation methods, which require teacher and student models to share the same latent space, thereby hindering knowledge transfer from high-capacity teachers to lightweight students such as Stable Diffusion 1.5. The study formalizes, for the first time, the cross-latent-space distillation problem and introduces a lightweight Bridge module that maps the student’s latent representations into the teacher’s space without modifying the student backbone. This module leverages the frozen student VAE decoder as a spatial prior combined with a learnable projector, optimized jointly via latent reconstruction and attention fidelity losses. The approach supports heterogeneous architectures and varying resolutions, achieving substantial performance gains—e.g., improving the HPSv3 score of SD 1.5 from 5.4 to 9.4—while preserving single-step inference, low latency, and ecosystem compatibility.
该研究通过将在线策略蒸馏视为概率传输问题,并提出RouteOPD方法,以更精确地指导学生模型学习教师模型的知识,从而提高性能和减少背景泄漏。
This work addresses a key limitation in existing on-policy self-distillation methods, which fail to effectively leverage privileged knowledge embedded in post-hoc feedback (e.g., success/failure outcomes) from student trajectories. The authors propose PAST, a novel approach that, for the first time, utilizes complete student trajectories as privileged information to adaptively refine the teacher model. While preserving the student’s distillation prefix, PAST employs trajectory-conditioned distillation to disentangle transferable policy shifts from trajectory-specific variations and theoretically characterizes the teacher’s capacity to convey knowledge to a prefix-only student. The method integrates Forward-KL distillation, student-proximity regularization, and a distribution-preserving mechanism over correct trajectories. Evaluated on three mathematical reasoning benchmarks, PAST achieves a 5.6 percentage point improvement in Avg@12 macro-average over vanilla OPSD, with ablation studies confirming the critical roles of trajectory completeness and teacher adaptivity.
This work addresses the lack of pedagogical awareness in existing knowledge distillation methods for large language models, which often reduce knowledge transfer to a one-off data synthesis process and neglect the structured nature of learning. To remedy this, the authors propose a three-stage distillation framework—Knowledge Identification, Organization, and Adaptation—inspired by educational theory. This framework systematically integrates Bloom’s Mastery Learning theory with Vygotsky’s Zone of Proximal Development to construct a dynamic, progressive, and difficulty-controlled knowledge transfer pathway. Through an IOA architecture, the approach tailors distillation strategies to the cognitive capacity of the student model. Experiments demonstrate that student models with fewer than one-tenth the parameters of the teacher achieve 94.7% of the teacher’s performance on DollyEval and show significant improvements of 19.2% and 22.3% on MATH and HumanEval benchmarks, respectively, outperforming current state-of-the-art methods.
本文通过引入样本级逆温度更新方法,解决了TTM中教师侧温度固定的问题,利用KL散度最小化来调整温度,提高知识蒸馏效果。