physics-guided knowledge distillation

Designs and implements teacher–student distillation pipelines that transfer knowledge from a privileged, physics-aware teacher into a compact student by embedding analytical physical priors or other privileged information into the distillation objective. Builds and evaluates the loss functions, model architectures, and training procedures needed to preserve predictive fidelity with limited data and computational resources so the resulting lightweight models can perform accurate, real-time inference while respecting the incorporated priors.

physics-guidedknowledgedistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Distillation Scaling Laws

Feb 12, 2025
DB
Dan Busbridge
🏛️ Apple | University of Oxford

This paper addresses the lack of theoretical guidance for allocating computational budget between teacher and student models in knowledge distillation. We establish, for the first time, a computational scaling law for distillation learning, quantitatively characterizing how student performance varies with teacher–student compute allocation and total budget. Methodologically, we propose a predictive distillation scaling law, derived via large-scale cross-model-size distillation experiments, computational modeling, and empirical law fitting, yielding optimal compute allocation strategies for two practical scenarios. Key contributions include: (1) uncovering how the performance crossover point—where distillation surpasses supervised pretraining—evolves with model scale; (2) identifying the critical compute threshold beyond which multi-student distillation consistently outperforms supervised training; and (3) providing a reusable, generalizable compute configuration paradigm for industrial-scale model compression, substantially reducing deployment risks in large-scale distillation.

Compare distillation and supervised learningEstimate distilled model performanceOptimize compute allocation

Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures

Jul 22, 2024
KB
Kuluhan Binici
🏛️ National University of Singapore

Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.

Computational CostKnowledge DistillationModel Adaptability

Circuit Distillation

Sep 29, 2025
SW
Somin Wadhwa
🏛️ Northeastern University

Conventional knowledge distillation focuses on behavioral imitation, treating the teacher model as a black box and transferring only output distributions. Method: This paper proposes “circuit distillation”—a mechanism-aware approach that transfers the teacher’s underlying computational architecture rather than superficial behavior. It aligns interpretable internal components (e.g., entity-tracking and theory-of-mind circuits) between Llama3-based teacher and student models via functional circuit correspondence matching and representation similarity loss. Contribution/Results: To our knowledge, this is the first work achieving mechanism-level distillation grounded in functional circuits. It enables transfer of complex algorithmic capabilities with minimal parameter tuning—only a small subset of student parameters is fine-tuned. The method significantly enhances model interpretability and controllability. Empirical evaluation on entity tracking and theory-of-mind tasks demonstrates superior performance over conventional distillation baselines, validating that mechanistic alignment—not just statistical mimicry—is essential for effective algorithmic capability transfer.

Aligning internal representations between corresponding circuit componentsDistilling computational mechanisms from teacher modelsTransferring algorithmic capabilities via interpretable internal mechanisms

This work addresses the long-standing absence of a unified statistical perspective on knowledge distillation, which has frequently been perceived as an engineering heuristic. We propose a unifying framework grounded in Bayesian inference that formalizes teacher model predictions as prior information, thereby enabling principled uncertainty quantification. This framework not only bridges classical distillation methods with their extensions to large language models but also integrates seamlessly with modern generative systems. Furthermore, we provide a conceptual roadmap and identify key open problems, establishing a systematic foundation for deepening the theoretical understanding of distillation mechanisms.

Bayesian FormulationKnowledge DistillationLarge Language Models

Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Oct 15, 2024
WX
Wenda Xu
🏛️ UC Santa Barbara | Google Cloud AI Research | CMU | Google DeepMind

In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.

Addresses teacher-student knowledge gap in distillation.Enhances student model performance across diverse tasks.Improves training data quality in knowledge distillation.

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing single-step diffusion distillation methods, which require teacher and student models to share the same latent space, thereby hindering knowledge transfer from high-capacity teachers to lightweight students such as Stable Diffusion 1.5. The study formalizes, for the first time, the cross-latent-space distillation problem and introduces a lightweight Bridge module that maps the student’s latent representations into the teacher’s space without modifying the student backbone. This module leverages the frozen student VAE decoder as a spatial prior combined with a learnable projector, optimized jointly via latent reconstruction and attention fidelity losses. The approach supports heterogeneous architectures and varying resolutions, achieving substantial performance gains—e.g., improving the HPSv3 score of SD 1.5 from 5.4 to 9.4—while preserving single-step inference, low latency, and ecosystem compatibility.

Cross-Space Distillationdiffusion modelsknowledge distillation

This work addresses a key limitation in existing on-policy self-distillation methods, which fail to effectively leverage privileged knowledge embedded in post-hoc feedback (e.g., success/failure outcomes) from student trajectories. The authors propose PAST, a novel approach that, for the first time, utilizes complete student trajectories as privileged information to adaptively refine the teacher model. While preserving the student’s distillation prefix, PAST employs trajectory-conditioned distillation to disentangle transferable policy shifts from trajectory-specific variations and theoretically characterizes the teacher’s capacity to convey knowledge to a prefix-only student. The method integrates Forward-KL distillation, student-proximity regularization, and a distribution-preserving mechanism over correct trajectories. Evaluated on three mathematical reasoning benchmarks, PAST achieves a 5.6 percentage point improvement in Avg@12 macro-average over vanilla OPSD, with ablation studies confirming the critical roles of trajectory completeness and teacher adaptivity.

on-policy self-distillationprivileged informationreasoning models

This work addresses the lack of pedagogical awareness in existing knowledge distillation methods for large language models, which often reduce knowledge transfer to a one-off data synthesis process and neglect the structured nature of learning. To remedy this, the authors propose a three-stage distillation framework—Knowledge Identification, Organization, and Adaptation—inspired by educational theory. This framework systematically integrates Bloom’s Mastery Learning theory with Vygotsky’s Zone of Proximal Development to construct a dynamic, progressive, and difficulty-controlled knowledge transfer pathway. Through an IOA architecture, the approach tailors distillation strategies to the cognitive capacity of the student model. Experiments demonstrate that student models with fewer than one-tenth the parameters of the teacher achieve 94.7% of the teacher’s performance on DollyEval and show significant improvements of 19.2% and 22.3% on MATH and HumanEval benchmarks, respectively, outperforming current state-of-the-art methods.

knowledge distillationlanguage modelspedagogical awareness

本文通过引入样本级逆温度更新方法,解决了TTM中教师侧温度固定的问题,利用KL散度最小化来调整温度,提高知识蒸馏效果。

Knowledge DistillationSample-wise Inverse-TemperatureTemperature-Adaptive

Hot Scholars

YL

Yixin Li

Stony Brook University
PET InstrumentMedical ImagingX-ray Imaging
QH

Qingyong Hu

Ph.D. of Computer Science, University of Oxford
3D VisionPhotogrammetryPoint Cloud ProcessingAutonomous Driving