Score
Designs and implements distillation objectives and training procedures that align a student model to a teacher by combining forward and reverse Kullback–Leibler divergence terms applied at per-token or marginal distribution levels and by matching logits and intermediate features. This work includes building per-token/marginal KL losses, spectral or subspace-aware projection and LoRA-style adaptations, optionally combining KL with MMD or other penalties, and tuning temperatures and loss weights to stabilize optimization, preserve principal modes and long-tail probabilities, and control training dynamics.
This study systematically investigates the feedback-to-update mechanism in on-policy distillation (OPD), where data are generated by the current policy. Framing OPD as a feedback-to-update problem, the work introduces a formula-driven categorization framework that unifies two major update pathways: distributional loss and policy gradient–style log-ratio updates. It further incorporates novel perspectives from temporal credit assignment and temporal vocabulary routing. By leveraging techniques such as KL divergence orientation, generalized advantage estimation (GAE), and counterfactual routing, the analysis reveals that OPD performance critically depends on state compatibility and support set construction. The paper establishes a comprehensive analytical framework for OPD, derives explicit bias bounds, and proposes new methods—GAE-OPD and CR-OPD—to enhance training stability, alongside actionable diagnostic tools and a practical implementation checklist.
This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.
Although knowledge distillation is widely employed to enhance model generalization, its theoretical underpinnings remain poorly understood. This work models the teacher–student training dynamics as a coupled stochastic process and introduces a novel “distillation divergence” to quantify the discrepancy between teacher and student. Building upon this, we develop an information-theoretic framework for generalization analysis and derive upper and lower bounds on the student’s generalization error that explicitly depend on the distillation divergence. Notably, we show that the local flatness of the teacher model strictly tightens the upper bound. In the Gaussian linear setting, we further provide an interpretable decomposition of the error into bias, variance, and a rank bottleneck, offering both theoretical insights and practical principles for designing effective distillation algorithms.
Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.
Existing knowledge distillation (KD) methods rely heavily on Kullback–Leibler (KL) divergence, which suffers from insufficient or biased knowledge transfer under high- or low-entropy teacher distributions; moreover, standard data augmentation can inadvertently degrade KD performance. To address these issues, we propose a robust KD framework comprising three key innovations: (i) replacing KL divergence with correlation-based distance to mitigate entropy sensitivity; (ii) integrating structured network pruning to enhance student model robustness; and (iii) identifying and mitigating the adverse interference of data augmentation in KD via a multi-stage teacher–student co-optimization strategy. Extensive experiments on CIFAR-100, FGVC-Aircraft, TinyImageNet, and ImageNet demonstrate state-of-the-art performance: student models achieve average accuracy gains of 1.2–2.7% over prior methods and exhibit significantly improved resilience to input noise and adversarial perturbations.
In knowledge distillation, the KL divergence loss suffers from gradient magnitudes proportional to teacher logits, leading to insufficient updates for low-probability classes and weakened inter-class relationship modeling. To address this, we propose Rank-Kendall Knowledge Distillation (RKKD), the first method to incorporate a differentiable Kendall’s τ coefficient into the distillation objective. RKKD replaces absolute logit value matching with relative ranking consistency among logit channels, establishing a temperature-free rank-order constraint. This formulation avoids optimization direction bias in soft-label matching and preserves discriminative information from small-magnitude logits, explicitly maintaining fine-grained inter-class ordinal relationships. Extensive experiments on CIFAR-100 and ImageNet demonstrate that RKKD consistently improves student accuracy across diverse teacher-student architecture pairs. Notably, it delivers stable performance gains for lightweight student models and exhibits strong generalization across datasets and model scales.
Existing multi-teacher knowledge distillation methods lack a theoretically grounded mechanism for adaptive weight assignment, often relying on heuristic strategies. This work proposes the first operator-agnostic axiomatic framework that enables principled adaptive weighting across three granularities—tokens, tasks, and contexts—while supporting hierarchical composition and safety constraints. By leveraging axiomatic modeling, product-structure normalization, and perturbation robustness analysis, our approach decouples theoretical guarantees from specific weighting formulations, ensuring applicability to heterogeneous models and distribution-shifted scenarios. We prove the existence (and non-uniqueness) of weighting operators satisfying the proposed axioms, establish convergence and stability guarantees for the associated optimization, and provide a formal characterization of knowledge distillation under safety constraints.
This work addresses the challenge of deploying model ensembles in resource-constrained settings, where their computational overhead is prohibitive despite performance gains. To this end, the authors propose an efficient knowledge distillation method that aligns representations between teacher and student models through layer- and token-level projection mappings into a high-dimensional embedding space. By integrating Low-Rank Adaptation (LoRA), the approach enables parameter-efficient fine-tuning with a lightweight alignment mechanism that supports parallel training. The trainable parameters are reduced to less than 1% of those in the teacher model. Evaluated on speech recognition tasks, the method achieves substantial reductions in word error rate (WER) and outperforms existing distillation techniques.
This work addresses a critical limitation in low-rank knowledge distillation: output-level distillation fails to explicitly align the low-rank subspaces of teacher and student models, leading to subspace misalignment and reduced compression efficiency. To resolve this, the paper introduces a spectral alignment mechanism that jointly optimizes three sources of error—subspace misalignment, coefficient mismatch, and irreducible residual—through data-weighted student subspace reference updates and a differentiable principal angle loss. The proposed method integrates LoRA adaptation, subspace projection, and data-weighted spectral decomposition. Empirical results demonstrate that it reduces subspace misalignment error from 51% to nearly zero on synthetic tasks. On six GLUE benchmarks, it outperforms the strongest spectral baseline on five tasks at rank r=4 and achieves state-of-the-art performance on SST-2 and CoLA at r=8.