Score
Designs and implements methods and training pipelines that transfer predictive and representational knowledge from one or more source models (teachers) into target models (students) to produce smaller, faster, or otherwise constrained models while preserving performance. This includes creating distillation objectives and schedules that operate on class-probabilities/logits, soft targets, saliency or attention maps, intermediate features (including dual-level or relational feature alignment), and ensemble or multi-teacher/multi-source schemes, as well as progressive, post-training, two-stage, pruned-model, and supervised distillation procedures and their evaluation.
Deploying large models—including large language models (LLMs), vision-language models (VLMs), diffusion models, and 3D/Transformer architectures—on resource-constrained edge devices remains challenging due to prohibitive computational and memory overhead. To address this, this paper presents a systematic survey of recent advances in knowledge distillation (KD). We propose the first unified taxonomy covering multimodal, generative, and foundation models, comprehensively integrating response-, feature-, and relation-based distillation, cross-modal KD, self-supervised distillation, and lightweight training paradigms. We rigorously analyze architectural adaptation challenges posed by emerging model families and construct the most comprehensive KD methodology map to date. Furthermore, we publicly release an accompanying GitHub repository with curated resources, implementations, and benchmarks. This work provides theoretical foundations, practical guidelines, and an authoritative reference for efficient large-model deployment on edge devices.
This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.
Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.
To address weak semantic preservation and degraded reasoning performance in large language model (LLM) knowledge distillation—caused by teacher-student representation mismatch—this paper proposes a novel distillation framework integrating feature alignment with hierarchical representation transfer. Methodologically, it innovatively couples fine-grained hidden-layer feature-space alignment via contrastive learning, gradient-aware dynamic scheduling of representation transfer weights, and a modular decoupled distillation mechanism—thereby overcoming the locality limitations of conventional logit- or attention-based distillation and enabling cross-depth semantic consistency modeling. Evaluated on LLaMA-2 → TinyLLaMA distillation, the student model achieves 92.3% of the teacher’s original accuracy despite an 78% reduction in parameter count, while attaining a 3.1× speedup in inference latency. These results significantly outperform existing state-of-the-art methods.
This work proposes a systematic method to convert non-neural machine learning pipelines—such as those based on random forests—into neural networks, enabling unified inference and joint optimization. Leveraging knowledge distillation, the approach treats the traditional model as a “teacher” that guides the training of a neural “student” network. The framework further integrates neural architecture search with a random forest–inspired hyperparameter selection strategy to optimize the student model. Notably, this is the first effort to employ an entire non-neural machine learning pipeline as the teacher in knowledge distillation, thereby extending the scope of this technique. Experimental evaluation across 100 OpenML tasks demonstrates that the student networks consistently replicate the performance of their teacher models, confirming the feasibility and effectiveness of the proposed conversion framework.
Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.
This work addresses the challenge of deploying model ensembles in resource-constrained settings, where their computational overhead is prohibitive despite performance gains. To this end, the authors propose an efficient knowledge distillation method that aligns representations between teacher and student models through layer- and token-level projection mappings into a high-dimensional embedding space. By integrating Low-Rank Adaptation (LoRA), the approach enables parameter-efficient fine-tuning with a lightweight alignment mechanism that supports parallel training. The trainable parameters are reduced to less than 1% of those in the teacher model. Evaluated on speech recognition tasks, the method achieves substantial reductions in word error rate (WER) and outperforms existing distillation techniques.
This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.
This work addresses the inefficiency in task-specific knowledge distillation caused by misaligned feature representations between a fine-tuned teacher model and a student model. To overcome this limitation, the authors propose a joint training framework that enforces feature alignment between teacher and student through shared low-rank adapters (LoRA), replacing conventional fine-tuning or linear probing strategies. The approach not only substantially improves student model performance but also reciprocally enhances the teacher’s accuracy, while accelerating training by a factor of two. Evaluated across multiple image classification and segmentation benchmarks, the method achieves state-of-the-art results in task-specific knowledge distillation.
This work addresses the lack of pedagogical awareness in existing knowledge distillation methods for large language models, which often reduce knowledge transfer to a one-off data synthesis process and neglect the structured nature of learning. To remedy this, the authors propose a three-stage distillation framework—Knowledge Identification, Organization, and Adaptation—inspired by educational theory. This framework systematically integrates Bloom’s Mastery Learning theory with Vygotsky’s Zone of Proximal Development to construct a dynamic, progressive, and difficulty-controlled knowledge transfer pathway. Through an IOA architecture, the approach tailors distillation strategies to the cognitive capacity of the student model. Experiments demonstrate that student models with fewer than one-tenth the parameters of the teacher achieve 94.7% of the teacher’s performance on DollyEval and show significant improvements of 19.2% and 22.3% on MATH and HumanEval benchmarks, respectively, outperforming current state-of-the-art methods.
This work addresses the limitations of existing knowledge distillation methods when integrating multiple heterogeneous strategies, which often suffer from implementation complexity, rigid combinations, and catastrophic forgetting. To overcome these challenges, the authors propose a Sequential Multi-Stage Knowledge Distillation (SMSKD) framework that applies distinct distillation techniques in successive stages. Each stage leverages a frozen reference model from the previous stage to anchor learned knowledge, while a sample-level adaptive loss weighting mechanism—based on the teacher’s true class probability (TCP)—dynamically balances knowledge retention and integration. The framework flexibly accommodates arbitrary distillation strategies and numbers of stages, consistently yielding significant accuracy improvements for student models across diverse teacher–student architectures, outperforming current baselines with negligible computational overhead.