model distillation

Designs, implements, and evaluates training pipelines and objective functions that transfer knowledge from one or more pretrained teacher models into smaller, pruned, or differently structured student models to reduce size or change behavior while preserving performance. This work covers prediction- and feature-level losses (including relational and dual-level feature alignment), multi-teacher/ensemble and heterogeneous teacher strategies, multi-source and interleaved distillation, progressive/two-stage and post-training schedules, supervised/pruned-model variants, and the design and analysis of distillation objectives and procedures.

modeldistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.76
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$220K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures

Jul 22, 2024
KB
Kuluhan Binici
🏛️ National University of Singapore

Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.

Computational CostKnowledge DistillationModel Adaptability

Distillation Scaling Laws

Feb 12, 2025
DB
Dan Busbridge
🏛️ Apple | University of Oxford

This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.

Compare distillation and supervised learningEstimate distilled model performanceOptimize compute allocation

This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.

feature-based methodsknowledge distillationmodel performance

This work addresses the limitations of existing knowledge distillation methods when integrating multiple heterogeneous strategies, which often suffer from implementation complexity, rigid combinations, and catastrophic forgetting. To overcome these challenges, the authors propose a Sequential Multi-Stage Knowledge Distillation (SMSKD) framework that applies distinct distillation techniques in successive stages. Each stage leverages a frozen reference model from the previous stage to anchor learned knowledge, while a sample-level adaptive loss weighting mechanism—based on the teacher’s true class probability (TCP)—dynamically balances knowledge retention and integration. The framework flexibly accommodates arbitrary distillation strategies and numbers of stages, consistently yielding significant accuracy improvements for student models across diverse teacher–student architectures, outperforming current baselines with negligible computational overhead.

catastrophic forgettingheterogeneous methodsknowledge distillation

GUIDE: Guided Initialization and Distillation of Embeddings

Oct 07, 2025
KT
Khoa Trinh
🏛️ Google Research

In knowledge distillation, student models typically only mimic teacher outputs (e.g., logits), failing to inherit the teacher’s internal representational capacity. Method: This paper proposes GUIDE—the first distillation method extended into parameter space—leveraging embedding-layer-guided initialization and parameter-space alignment to enable direct inheritance of the teacher’s internal representation structure, rather than output-level matching alone. GUIDE incurs no additional training or inference overhead and integrates seamlessly with standard knowledge distillation. Results: Evaluated on 400M–1B parameter language models, GUIDE reduces the teacher–student quality gap by 25%–26% using only ~20B training tokens. When applied standalone, it significantly outperforms conventional knowledge distillation, demonstrating its effectiveness and generality as a novel distillation paradigm.

Achieving quality improvements without training or inference overheadEnhancing student model quality beyond standard distillation limitationsReducing teacher-student performance gap through parameter space alignment

Latest Papers

What's happening recently
View more

This work investigates the role of post-training knowledge distillation in building efficient small language models under data-scarce or resource-constrained settings. Through systematic analysis across varying data scales and teacher model strengths, the study demonstrates that knowledge distillation significantly outperforms supervised fine-tuning in low-data regimes, though this advantage diminishes as data volume increases. To address this limitation, the authors propose a two-stage distillation strategy that combines synthetic data with human-annotated examples, consistently enhancing student model performance on domain-specific tasks. Experiments on the Tulu 3 dataset further reveal that stronger instruction-tuned teacher models can restore distillation’s superiority even at higher data scales, offering a practical and effective model compression approach for resource-limited environments.

Data ScarcityInstruction TuningKnowledge Distillation

This study challenges the conventional assumption that knowledge distillation in large language model pretraining necessarily requires a strong teacher model. It systematically investigates distillation efficacy across varying teacher-student capacity pairings—including strong-to-weak, peer-level, and weak-to-strong configurations—by jointly optimizing language modeling and distillation losses. Evaluated across diverse Transformer architectures and out-of-distribution downstream tasks, the findings reveal that even weak teachers can significantly enhance student performance when the loss weighting is appropriately calibrated. Moreover, stronger teachers do not consistently yield better distillation outcomes; excessive teacher capacity may lead to performance saturation or even degradation. The work thus uncovers the underappreciated potential of weak-teacher distillation and highlights its beneficial impact on out-of-domain generalization.

knowledge distillationlarge language modelsmodel generalization

This work addresses a key limitation in existing knowledge distillation methods for large language models, where the neglect of training data sequencing and mismatched teacher-student model capacities often prevents stronger teachers from effectively enhancing student performance. To overcome this, the authors propose Curriculum Learning-guided Progressive Distillation (CLPD), a novel framework that jointly models data difficulty and teacher capability for the first time. CLPD integrates an explicit data curriculum with an implicit teacher scheduling mechanism within a unified architecture. By modularly combining curriculum learning, progressive distillation, and dynamic teacher scheduling, the framework consistently outperforms standard distillation and ablated variants across multiple reasoning benchmarks, demonstrating that the co-optimization of data ordering and teacher scheduling is crucial for efficient knowledge transfer.

curriculum learningdata orderingknowledge distillation

Hot Scholars

LY

Ling Yang

Postdoc@Princeton University, PhD@Peking University
LLMDiffusion ModelsReinforcement LearningComplex Data Modeling
JX

Jinan Xu

Professor of School of Computer and Information Technology, Beijing Jiaotong University
NLPMachine TranslationLLM
JF

Junfeng Fang

National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science
XW

Xinggang Wang

Professor, Huazhong University of Science and Technology
Artificial IntelligenceComputer VisionAutonomous DrivingObject Detection
WL

Weijian Luo

Peking University
Human-preferred Generative ModelsLarge Vision-language Models