distribution matching distillation

Design and implement teacher→student training procedures that transfer a model by explicitly matching probability distributions (e.g., outputs, logits, or intermediate representations) between teacher and student so the student reproduces the teacher’s predictive distribution. Work includes specifying which distributions to match, constructing distribution-matching loss terms, and building students that achieve similar fidelity with fewer inference steps or lower compute.

distributionmatchingdistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Representational Alignment Supports Effective Machine Teaching

Jun 06, 2024
IS
Ilia Sucholutsky
🏛️ Princeton University | University of Cambridge | Stevens Institute of Technology | MPI | Anthropic | NYU | MIT | University College London | UC Berkeley | The Alan Turing Institute

Existing machine learning teaching frameworks overlook representational alignment between teachers and students, prioritizing only model accuracy. Method: We propose GRADE, a representation-alignment-driven pedagogical optimization framework. Leveraging controlled machine–machine and machine–human teaching experiments, we formally define and quantify the relationship between representational alignment and teaching utility, introducing the alignment-driven teaching utility curve. We further design GRADE-Match, a cross-modal teacher–student matching algorithm that optimizes representational adaptation. Contribution/Results: Experiments demonstrate that improved representational alignment significantly enhances student task accuracy—moderated by class size and representation diversity. In simulated teaching settings, GRADE-Match achieves an average 12.3% improvement in learning outcomes. GRADE establishes a novel paradigm for interpretable, optimization-aware intelligent teaching systems grounded in representational alignment principles.

Characterize teacher expertise and student learningOptimize student-teacher matching with GRADEStudy pedagogy and representational alignment

Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Oct 15, 2024
WX
Wenda Xu
🏛️ UC Santa Barbara | Google Cloud AI Research | CMU | Google DeepMind

In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.

Addresses teacher-student knowledge gap in distillation.Enhances student model performance across diverse tasks.Improves training data quality in knowledge distillation.

Which distribution were you sampled from? Towards a more tangible conception of data

Jul 24, 2024
BH
Benedikt Holtgen
🏛️ University of Tübingen | Tübingen AI Center

This paper challenges the reliance of machine learning research in the social sciences on abstract data-generating distributions, arguing that such assumptions lack empirical grounding in finite-population settings and engender interpretability and reproducibility issues. Method: The authors advocate replacing distributional assumptions with finite-population modeling, systematically advancing five core arguments grounded in statistical foundations, philosophical epistemology, and ML empirical analysis. They reconstruct the premises of learning theory by explicitly identifying the implicit assumptions and boundary conditions underlying distributional modeling. Contribution/Results: The proposed framework enhances theoretical coherence, modeling transparency, causal traceability, and practical applicability. It provides a novel paradigm and methodological foundation for sampling design, bias correction in evaluation, and reproducibility research—thereby addressing critical limitations of conventional distribution-based approaches in social-science ML applications.

Avoid assuming data-generating probability distributions in social MLChallenge reliance on abstract distributions for fairness in algorithmsPropose alternative frameworks focusing on populations not distributions

Modelling Structured Data Learning with Restricted Boltzmann Machines in the Teacher-Student Setting

Oct 21, 2024
RT
Robin Th'eriault
🏛️ Scuola Normale Superiore di Pisa | The University of British Columbia | University of Bologna

This work investigates the learning mechanism of restricted Boltzmann machines (RBMs) for structured data within a teacher–student framework, specifically addressing whether a student RBM can recover the true latent representation from data generated by a teacher RBM exhibiting correlations among hidden variables. Using statistical-physics-based mean-field analysis and temperature-regularized inference, we systematically characterize how structural strength—quantified by the number of latent patterns and inter-pattern correlation in weight rows—affects the critical sample size required for successful learning. We find that enhanced structure drastically reduces the necessary sample complexity; in the absence of correlations, performance is independent of pattern count; and excessively low inference temperatures suppress pattern acquisition, leading to learning failure. Crucially, we establish for the first time that the student can achieve exact one-to-one or one-to-many pattern matching—surpassing the conventional two-hidden-unit limitation. These results provide the first analytically tractable generative-model foundation for the “lottery ticket hypothesis.”

Analyzing impact of pattern correlations on critical data requirementsExploring temperature effects on teacher pattern learnabilityStudying RBM learning of structured data in teacher-student setting

Enhancing Accuracy in Generative Models via Knowledge Transfer

May 27, 2024
XT
Xinyu Tian
🏛️ University of Minnesota

This study addresses the insufficient task-specific accuracy of generative models. Methodologically, it proposes a novel knowledge transfer framework that— for the first time—integrates shared embedding with distributional metrics (e.g., KL divergence) to establish a unified cross-task transfer learning mechanism. Theoretical analysis proves that structural commonality across tasks significantly enhances generative fidelity and yields the first general analytical framework for transfer learning applicable to both diffusion models and normalizing flows. Empirically, the method consistently outperforms non-transfer baselines: diffusion models achieve markedly improved generation accuracy, while normalizing flows not only demonstrate measurable performance gains but also uncover new theoretical insights—particularly regarding the distinct behavioral regimes under transfer versus non-transfer settings.

Bridging source-target tasks using shared embedding frameworkEnhancing diffusion and normalizing flow models' performanceImproving generative model accuracy via knowledge transfer

Latest Papers

What's happening recently
View more

This study investigates whether aligning training data by reasoning method or by mathematical topic better facilitates knowledge transfer during the fine-tuning of large language models for mathematics. Employing a balanced experimental design, we conduct 40 comparative fine-tuning experiments across multiple base models, validated through embedding similarity analysis and statistical significance testing. Our empirical findings reveal that cross-topic data sharing reasoning methods consistently outperforms same-topic data employing different methods, yielding average improvements exceeding 10 percentage points with confidence intervals strictly excluding zero. Based on these results, we establish reasoning method as the primary criterion for data curation. This finding challenges the conventional paradigm of organizing training data by domain classification, offering a new framework for constructing mathematical instruction datasets.

Fine-tuningLarge Language ModelsMathematical Transfer

This work challenges the common practice in knowledge distillation of naively matching a teacher model’s absolute feature representations, which overlooks the fact that such representations are only equivalent up to orthogonal transformations and isotropic scaling. From a geometric perspective, the paper proposes a new paradigm centered on representation equivalence classes: the student should instead learn class-invariant structures of the teacher’s representations—such as Gram matrices, centered kernel alignment (CKA), or principal subspaces—or leverage coordinate alignment for effective supervision. This framework unifies feature matching, relational distillation, and grafting approaches, revealing that logit-level matching is ultimately key to capability transfer. Experiments on Qwen2.5 and Llama-3.1 demonstrate that high CKA similarity alone is insufficient for performance recovery, while successful grafting hinges on boundary overlap in the training data coverage, thereby validating the proposed theory.

capability recoverygeometric invarianceknowledge distillation

This study investigates the impact of prompt data on teacher-student knowledge transfer in online policy distillation. By systematically analyzing prompt quantity, source, and selection strategies alongside parameter and functional alignment techniques, it compares cross-model configurations under reinforcement learning (RL) and supervised fine-tuning (SFT) post-training scenarios. The findings reveal that prompt utility depends primarily on the specific teacher-student pairing rather than intrinsic prompt properties. Furthermore, by distinguishing prompt efficiency from interchangeability, we demonstrate that merely substituting the teacher model can reverse the relative effectiveness of mathematics versus coding prompts. Empirical results confirm that a small number of prompts can approximate the performance achieved with large-scale prompt pools, and that random sampling outperforms targeted selection in certain settings.

data efficiencyknowledge transferon-policy distillation

This study addresses the challenge in multi-teacher knowledge distillation where base model preferences interfere with post-training signals, thereby constraining policy composition and routing performance. To overcome this, we propose the Δ-MOPD framework, which introduces a novel objective construction mechanism based on relative teacher shifts. By transferring log-probability shifts and anchoring student initialization, this approach effectively decouples base preferences from post-training increments, establishing objective construction as an independent design dimension. Experimental results demonstrate that under a three-teacher configuration, the method yields a 4.11-point improvement on mathematical tasks and an average gain of 1.95 points across five benchmarks. Furthermore, staged routing significantly reduces sequential discrepancies to 6.42 points. These findings validate the efficacy of the proposed framework for multi-teacher online distillation.

endpoint policyknowledge transferlogit shift

Hot Scholars

MG

Mohsen Ghafoorian

Sr. Staff Computer Vision Research Scientist, Qualcomm
Efficient Machine LearningComputer VisionVideo Diffusion3D Vision
AK

Adil Karjauv

Machine Learning R&D, Qualcomm
machine learning
TX

Tianfan Xue

Information Engineering Department, The Chinese University of Hong Kong
Computer VisionMachine LearningComputational Photography
MZ

Muhan Zhang

Peking University
Machine LearningGraph Neural NetworkLarge Language Models