knowledge distillation

Designs and implements methods and training pipelines that transfer predictive and representational knowledge from one or more source models (teachers) into target models (students) to produce smaller, faster, or otherwise constrained models while preserving performance. This includes creating distillation objectives and schedules that operate on class-probabilities/logits, soft targets, saliency or attention maps, intermediate features (including dual-level or relational feature alignment), and ensemble or multi-teacher/multi-source schemes, as well as progressive, post-training, two-stage, pruned-model, and supervised distillation procedures and their evaluation.

knowledgedistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-2.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$219K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Distillation Scaling Laws

Feb 12, 2025
DB
Dan Busbridge
🏛️ Apple | University of Oxford

This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.

Compare distillation and supervised learningEstimate distilled model performanceOptimize compute allocation

Relational Representation Distillation

Jul 16, 2024
NG
Nikolaos Giakoumoglou
🏛️ Imperial College London

Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.

Avoiding overly strict contrastive learning constraintsCapturing structural relationships in teacher modelsPreserving relative instance relationships effectively

Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models

Apr 18, 2025
JY
Junjie Yang
🏛️ Xiamen University | Imperial College London | University of Sussex | Purdue University | University of Liverpool | JTB Technology Corp. | Emory University | The University of Texas at Dallas | Kyoto University | AppCubic | Georgia Institute of Technology

To address weak semantic preservation and degraded reasoning performance in large language model (LLM) knowledge distillation—caused by teacher-student representation mismatch—this paper proposes a novel distillation framework integrating feature alignment with hierarchical representation transfer. Methodologically, it innovatively couples fine-grained hidden-layer feature-space alignment via contrastive learning, gradient-aware dynamic scheduling of representation transfer weights, and a modular decoupled distillation mechanism—thereby overcoming the locality limitations of conventional logit- or attention-based distillation and enabling cross-depth semantic consistency modeling. Evaluated on LLaMA-2 → TinyLLaMA distillation, the student model achieves 92.3% of the teacher’s original accuracy despite an 78% reduction in parameter count, while attaining a 3.1× speedup in inference latency. These results significantly outperform existing state-of-the-art methods.

Compressing large language models while preserving accuracyEnhancing model efficiency and accuracy through knowledge distillationOptimizing knowledge transfer using attention-based and decoupling techniques

This work proposes a systematic method to convert non-neural machine learning pipelines—such as those based on random forests—into neural networks, enabling unified inference and joint optimization. Leveraging knowledge distillation, the approach treats the traditional model as a “teacher” that guides the training of a neural “student” network. The framework further integrates neural architecture search with a random forest–inspired hyperparameter selection strategy to optimize the student model. Notably, this is the first effort to employ an entire non-neural machine learning pipeline as the teacher in knowledge distillation, thereby extending the scope of this technique. Experimental evaluation across 100 OpenML tasks demonstrates that the student networks consistently replicate the performance of their teacher models, confirming the feasibility and effectiveness of the proposed conversion framework.

knowledge distillationmachine learning pipelineneural network

Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures

Jul 22, 2024
KB
Kuluhan Binici
🏛️ National University of Singapore

Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.

Computational CostKnowledge DistillationModel Adaptability

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying model ensembles in resource-constrained settings, where their computational overhead is prohibitive despite performance gains. To this end, the authors propose an efficient knowledge distillation method that aligns representations between teacher and student models through layer- and token-level projection mappings into a high-dimensional embedding space. By integrating Low-Rank Adaptation (LoRA), the approach enables parameter-efficient fine-tuning with a lightweight alignment mechanism that supports parallel training. The trainable parameters are reduced to less than 1% of those in the teacher model. Evaluated on speech recognition tasks, the method achieves substantial reductions in word error rate (WER) and outperforms existing distillation techniques.

efficient inferenceknowledge distillationlogit distillation

This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.

feature-based methodsknowledge distillationmodel performance

This work addresses the inefficiency in task-specific knowledge distillation caused by misaligned feature representations between a fine-tuned teacher model and a student model. To overcome this limitation, the authors propose a joint training framework that enforces feature alignment between teacher and student through shared low-rank adapters (LoRA), replacing conventional fine-tuning or linear probing strategies. The approach not only substantially improves student model performance but also reciprocally enhances the teacher’s accuracy, while accelerating training by a factor of two. Evaluated across multiple image classification and segmentation benchmarks, the method achieves state-of-the-art results in task-specific knowledge distillation.

feature alignmentfoundation modelsknowledge distillation

This work addresses the lack of pedagogical awareness in existing knowledge distillation methods for large language models, which often reduce knowledge transfer to a one-off data synthesis process and neglect the structured nature of learning. To remedy this, the authors propose a three-stage distillation framework—Knowledge Identification, Organization, and Adaptation—inspired by educational theory. This framework systematically integrates Bloom’s Mastery Learning theory with Vygotsky’s Zone of Proximal Development to construct a dynamic, progressive, and difficulty-controlled knowledge transfer pathway. Through an IOA architecture, the approach tailors distillation strategies to the cognitive capacity of the student model. Experiments demonstrate that student models with fewer than one-tenth the parameters of the teacher achieve 94.7% of the teacher’s performance on DollyEval and show significant improvements of 19.2% and 22.3% on MATH and HumanEval benchmarks, respectively, outperforming current state-of-the-art methods.

knowledge distillationlanguage modelspedagogical awareness

This work addresses the limitations of existing knowledge distillation methods when integrating multiple heterogeneous strategies, which often suffer from implementation complexity, rigid combinations, and catastrophic forgetting. To overcome these challenges, the authors propose a Sequential Multi-Stage Knowledge Distillation (SMSKD) framework that applies distinct distillation techniques in successive stages. Each stage leverages a frozen reference model from the previous stage to anchor learned knowledge, while a sample-level adaptive loss weighting mechanism—based on the teacher’s true class probability (TCP)—dynamically balances knowledge retention and integration. The framework flexibly accommodates arbitrary distillation strategies and numbers of stages, consistently yielding significant accuracy improvements for student models across diverse teacher–student architectures, outperforming current baselines with negligible computational overhead.

catastrophic forgettingheterogeneous methodsknowledge distillation

Hot Scholars

XH

Xuming Hu

Assistant Professor, HKUST(GZ) / HKUST
Natural Language ProcessingLarge Language Model
DY

Dawei Yin

Senior Director, Head of Search Science at Baidu
Machine LearningWeb MiningData Mining
YH

Yao Hu

浙江大学
Machine Learning
LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
BZ

Beier Zhu

Research Scientist, Nanyang Technological University
Robust Machine Learning