bidirectional kl distillation

Designs and implements distillation objectives and training procedures that align a student model to a teacher by combining forward and reverse Kullback–Leibler divergence terms applied at per-token or marginal distribution levels and by matching logits and intermediate features. This work includes building per-token/marginal KL losses, spectral or subspace-aware projection and LoRA-style adaptations, optionally combining KL with MMD or other penalties, and tuning temperatures and loss weights to stabilize optimization, preserve principal modes and long-tail probabilities, and control training dynamics.

bidirectionalkldistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.

Exposure BiasImitation LearningKnowledge Distillation

Must-Read Papers

Most classic and influential ideas
View more

Although knowledge distillation is widely employed to enhance model generalization, its theoretical underpinnings remain poorly understood. This work models the teacher–student training dynamics as a coupled stochastic process and introduces a novel “distillation divergence” to quantify the discrepancy between teacher and student. Building upon this, we develop an information-theoretic framework for generalization analysis and derive upper and lower bounds on the student’s generalization error that explicitly depend on the distillation divergence. Notably, we show that the local flatness of the teacher model strictly tightens the upper bound. In the Gaussian linear setting, we further provide an interpretable decomposition of the error into bias, variance, and a rank bottleneck, offering both theoretical insights and practical principles for designing effective distillation algorithms.

distillation divergencegeneralizationinformation theory

Relational Representation Distillation

Jul 16, 2024
NG
Nikolaos Giakoumoglou
🏛️ Imperial College London

Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.

Avoiding overly strict contrastive learning constraintsCapturing structural relationships in teacher modelsPreserving relative instance relationships effectively

Robustness-Reinforced Knowledge Distillation With Correlation Distance and Network Pruning

Nov 23, 2023
SK
Seonghak Kim
🏛️ Korea Advanced Institute of Science and Technology

Existing knowledge distillation (KD) methods rely heavily on Kullback–Leibler (KL) divergence, which suffers from insufficient or biased knowledge transfer under high- or low-entropy teacher distributions; moreover, standard data augmentation can inadvertently degrade KD performance. To address these issues, we propose a robust KD framework comprising three key innovations: (i) replacing KL divergence with correlation-based distance to mitigate entropy sensitivity; (ii) integrating structured network pruning to enhance student model robustness; and (iii) identifying and mitigating the adverse interference of data augmentation in KD via a multi-stage teacher–student co-optimization strategy. Extensive experiments on CIFAR-100, FGVC-Aircraft, TinyImageNet, and ImageNet demonstrate state-of-the-art performance: student models achieve average accuracy gains of 1.2–2.7% over prior methods and exhibit significantly improved resilience to input noise and adversarial perturbations.

Addressing KL divergence limitations in knowledge transfer efficiencyImproving student model performance via robust knowledge distillationMitigating adverse effects of data augmentation in distillation

Kendall's τ Coefficient for Logits Distillation

Sep 26, 2024
YG
Yuchen Guan
🏛️ Tsinghua University

In knowledge distillation, the KL divergence loss suffers from gradient magnitudes proportional to teacher logits, leading to insufficient updates for low-probability classes and weakened inter-class relationship modeling. To address this, we propose Rank-Kendall Knowledge Distillation (RKKD), the first method to incorporate a differentiable Kendall’s τ coefficient into the distillation objective. RKKD replaces absolute logit value matching with relative ranking consistency among logit channels, establishing a temperature-free rank-order constraint. This formulation avoids optimization direction bias in soft-label matching and preserves discriminative information from small-magnitude logits, explicitly maintaining fine-grained inter-class ordinal relationships. Extensive experiments on CIFAR-100 and ImageNet demonstrate that RKKD consistently improves student accuracy across diverse teacher-student architecture pairs. Notably, it delivers stable performance gains for lightweight student models and exhibits strong generalization across datasets and model scales.

Gradient imbalance weakens inter-class information transferOptimizing KL divergence in distillation leads to sub-optimal solutionsPropose Kendall's τ ranking loss to rebalance gradients

Existing multi-teacher knowledge distillation methods lack a theoretically grounded mechanism for adaptive weight assignment, often relying on heuristic strategies. This work proposes the first operator-agnostic axiomatic framework that enables principled adaptive weighting across three granularities—tokens, tasks, and contexts—while supporting hierarchical composition and safety constraints. By leveraging axiomatic modeling, product-structure normalization, and perturbation robustness analysis, our approach decouples theoretical guarantees from specific weighting formulations, ensuring applicability to heterogeneous models and distribution-shifted scenarios. We prove the existence (and non-uniqueness) of weighting operators satisfying the proposed axioms, establish convergence and stability guarantees for the associated optimization, and provide a formal characterization of knowledge distillation under safety constraints.

adaptive weightingaxiomatic frameworkknowledge distillation

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying model ensembles in resource-constrained settings, where their computational overhead is prohibitive despite performance gains. To this end, the authors propose an efficient knowledge distillation method that aligns representations between teacher and student models through layer- and token-level projection mappings into a high-dimensional embedding space. By integrating Low-Rank Adaptation (LoRA), the approach enables parameter-efficient fine-tuning with a lightweight alignment mechanism that supports parallel training. The trainable parameters are reduced to less than 1% of those in the teacher model. Evaluated on speech recognition tasks, the method achieves substantial reductions in word error rate (WER) and outperforms existing distillation techniques.

efficient inferenceknowledge distillationlogit distillation

This work addresses a critical limitation in low-rank knowledge distillation: output-level distillation fails to explicitly align the low-rank subspaces of teacher and student models, leading to subspace misalignment and reduced compression efficiency. To resolve this, the paper introduces a spectral alignment mechanism that jointly optimizes three sources of error—subspace misalignment, coefficient mismatch, and irreducible residual—through data-weighted student subspace reference updates and a differentiable principal angle loss. The proposed method integrates LoRA adaptation, subspace projection, and data-weighted spectral decomposition. Empirical results demonstrate that it reduces subspace misalignment error from 51% to nearly zero on synthetic tasks. On six GLUE benchmarks, it outperforms the strongest spectral baseline on five tasks at rank r=4 and achieves state-of-the-art performance on SST-2 and CoLA at r=8.

knowledge distillationLoRAlow-rank adaptation

Hot Scholars

XM

Xing Ma

Meituan, NLP engineer
Dialog SystemLarge Language ModelConversation Analysis
BL

Bei Li

Meituan LLM Team
Machine TranslationDeep LearningLarge Language Models
JR

Jeffrey Regier

Assistant Professor, Department of Statistics, University of Michigan
Bayesian statisticsmachine learningbioinformaticsastronomy