Score
Designs, implements, and evaluates training pipelines and objective functions that transfer knowledge from one or more pretrained teacher models into smaller, pruned, or differently structured student models to reduce size or change behavior while preserving performance. This work covers prediction- and feature-level losses (including relational and dual-level feature alignment), multi-teacher/ensemble and heterogeneous teacher strategies, multi-source and interleaved distillation, progressive/two-stage and post-training schedules, supervised/pruned-model variants, and the design and analysis of distillation objectives and procedures.
Deploying large models—including large language models (LLMs), vision-language models (VLMs), diffusion models, and 3D/Transformer architectures—on resource-constrained edge devices remains challenging due to prohibitive computational and memory overhead. To address this, this paper presents a systematic survey of recent advances in knowledge distillation (KD). We propose the first unified taxonomy covering multimodal, generative, and foundation models, comprehensively integrating response-, feature-, and relation-based distillation, cross-modal KD, self-supervised distillation, and lightweight training paradigms. We rigorously analyze architectural adaptation challenges posed by emerging model families and construct the most comprehensive KD methodology map to date. Furthermore, we publicly release an accompanying GitHub repository with curated resources, implementations, and benchmarks. This work provides theoretical foundations, practical guidelines, and an authoritative reference for efficient large-model deployment on edge devices.
Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.
This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.
This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.
This work addresses the limitations of existing knowledge distillation methods when integrating multiple heterogeneous strategies, which often suffer from implementation complexity, rigid combinations, and catastrophic forgetting. To overcome these challenges, the authors propose a Sequential Multi-Stage Knowledge Distillation (SMSKD) framework that applies distinct distillation techniques in successive stages. Each stage leverages a frozen reference model from the previous stage to anchor learned knowledge, while a sample-level adaptive loss weighting mechanism—based on the teacher’s true class probability (TCP)—dynamically balances knowledge retention and integration. The framework flexibly accommodates arbitrary distillation strategies and numbers of stages, consistently yielding significant accuracy improvements for student models across diverse teacher–student architectures, outperforming current baselines with negligible computational overhead.
In knowledge distillation, student models typically only mimic teacher outputs (e.g., logits), failing to inherit the teacher’s internal representational capacity. Method: This paper proposes GUIDE—the first distillation method extended into parameter space—leveraging embedding-layer-guided initialization and parameter-space alignment to enable direct inheritance of the teacher’s internal representation structure, rather than output-level matching alone. GUIDE incurs no additional training or inference overhead and integrates seamlessly with standard knowledge distillation. Results: Evaluated on 400M–1B parameter language models, GUIDE reduces the teacher–student quality gap by 25%–26% using only ~20B training tokens. When applied standalone, it significantly outperforms conventional knowledge distillation, demonstrating its effectiveness and generality as a novel distillation paradigm.
This work investigates the role of post-training knowledge distillation in building efficient small language models under data-scarce or resource-constrained settings. Through systematic analysis across varying data scales and teacher model strengths, the study demonstrates that knowledge distillation significantly outperforms supervised fine-tuning in low-data regimes, though this advantage diminishes as data volume increases. To address this limitation, the authors propose a two-stage distillation strategy that combines synthetic data with human-annotated examples, consistently enhancing student model performance on domain-specific tasks. Experiments on the Tulu 3 dataset further reveal that stronger instruction-tuned teacher models can restore distillation’s superiority even at higher data scales, offering a practical and effective model compression approach for resource-limited environments.
This study challenges the conventional assumption that knowledge distillation in large language model pretraining necessarily requires a strong teacher model. It systematically investigates distillation efficacy across varying teacher-student capacity pairings—including strong-to-weak, peer-level, and weak-to-strong configurations—by jointly optimizing language modeling and distillation losses. Evaluated across diverse Transformer architectures and out-of-distribution downstream tasks, the findings reveal that even weak teachers can significantly enhance student performance when the loss weighting is appropriately calibrated. Moreover, stronger teachers do not consistently yield better distillation outcomes; excessive teacher capacity may lead to performance saturation or even degradation. The work thus uncovers the underappreciated potential of weak-teacher distillation and highlights its beneficial impact on out-of-domain generalization.
This work addresses a key limitation in existing knowledge distillation methods for large language models, where the neglect of training data sequencing and mismatched teacher-student model capacities often prevents stronger teachers from effectively enhancing student performance. To overcome this, the authors propose Curriculum Learning-guided Progressive Distillation (CLPD), a novel framework that jointly models data difficulty and teacher capability for the first time. CLPD integrates an explicit data curriculum with an implicit teacher scheduling mechanism within a unified architecture. By modularly combining curriculum learning, progressive distillation, and dynamic teacher scheduling, the framework consistently outperforms standard distillation and ablated variants across multiple reasoning benchmarks, demonstrating that the co-optimization of data ordering and teacher scheduling is crucial for efficient knowledge transfer.