Score
Designs and implements model distillation procedures and training objectives that combine teacher–student supervision (e.g., KLD or logits matching) with contrastive losses that explicitly separate positive examples from sampled negatives. This includes building contrastive loss augmentations, multi‑teacher and self‑supervised distillation variants, and the associated negative‑sampling and ranking‑oriented pipelines to improve precision‑focused metrics.
To address insufficient discriminability of student models and structural inconsistency between teacher and student representations in knowledge distillation, this paper proposes a joint distillation framework that synergistically optimizes discriminability and consistency. Methodologically, it introduces a contrastive learning loss to enhance inter-class separability, couples it with a distributional consistency regularizer to align latent-space structures, and incorporates learnable temperature and bias parameters to dynamically balance the dual objectives—eliminating reliance on fixed hyperparameters for adaptive optimization. Evaluated on CIFAR-100 and ImageNet, the method achieves state-of-the-art performance, with student models even surpassing teacher accuracy. Strong generalization is further validated via cross-dataset transfer on Tiny ImageNet and STL-10. The core contribution lies in the first unified formulation of discriminative modeling and structural consistency modeling within a learnable trade-off mechanism.
Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.
Existing LLM knowledge distillation methods apply uniform loss functions to all teacher-student generated data, neglecting the intrinsic alignment between loss design and heterogeneous data types—such as instructions, code, preferences, and multimodal inputs—thereby limiting performance gains. This work proposes a data-type-aware contrastive distillation framework. It introduces (i) a novel bidirectional contrastive loss that explicitly distinguishes teacher and student responses; (ii) the first dynamic coupling mechanism between loss functions and data types; and (iii) integrated techniques including hierarchical response modeling, task-adaptive weighting, and cross-modal extension. Evaluated on instruction-following and code generation, the method significantly outperforms state-of-the-art distillation approaches. It further supports preference alignment and vision-language joint distillation, enabling compact student models to retain over 92% of teacher model capability across diverse tasks.
Existing knowledge distillation methods suffer from conflicting distillation targets when teacher predictions are erroneous, and hard-label correction disrupts inter-class semantic correlations. To address this, we propose Refined Logit Distillation (RLD), the first framework introducing a dynamic logit refinement mechanism: a label-guided, differentiable logit recalibration that jointly applies confidence weighting and temperature-adaptive scaling—enabling erroneous predictions to be corrected while fully preserving inter-class semantic structure. RLD employs a hybrid loss combining cross-entropy and KL divergence, mitigating semantic distortion inherent in conventional logit-based distillation. Extensive experiments on CIFAR-100 and ImageNet demonstrate consistent improvements, with student models achieving average Top-1 accuracy gains of 1.2%–2.3% over strong baselines including KD, RKD, and VID. The implementation is publicly available.
Existing knowledge distillation (KD) methods rely heavily on Kullback–Leibler (KL) divergence, which suffers from insufficient or biased knowledge transfer under high- or low-entropy teacher distributions; moreover, standard data augmentation can inadvertently degrade KD performance. To address these issues, we propose a robust KD framework comprising three key innovations: (i) replacing KL divergence with correlation-based distance to mitigate entropy sensitivity; (ii) integrating structured network pruning to enhance student model robustness; and (iii) identifying and mitigating the adverse interference of data augmentation in KD via a multi-stage teacher–student co-optimization strategy. Extensive experiments on CIFAR-100, FGVC-Aircraft, TinyImageNet, and ImageNet demonstrate state-of-the-art performance: student models achieve average accuracy gains of 1.2–2.7% over prior methods and exhibit significantly improved resilience to input noise and adversarial perturbations.
Although knowledge distillation is widely employed to enhance model generalization, its theoretical underpinnings remain poorly understood. This work models the teacher–student training dynamics as a coupled stochastic process and introduces a novel “distillation divergence” to quantify the discrepancy between teacher and student. Building upon this, we develop an information-theoretic framework for generalization analysis and derive upper and lower bounds on the student’s generalization error that explicitly depend on the distillation divergence. Notably, we show that the local flatness of the teacher model strictly tightens the upper bound. In the Gaussian linear setting, we further provide an interpretable decomposition of the error into bias, variance, and a rank bottleneck, offering both theoretical insights and practical principles for designing effective distillation algorithms.
In knowledge distillation, student models often inherit shortcut features that teachers initially rely on but later suppress, compromising generalization and robustness. This work proposes Anti-Shortcut Distillation (ASD), a novel framework that leverages temporal information from the teacher’s training trajectory by constructing positive and negative semantic anchors from early and final checkpoints. ASD employs a push-pull mechanism to steer students away from shortcut directions, integrating a temporal contrastive loss (L<sub>tc</sub>) and a shortcut-suppression loss (L<sub>ss</sub>) based on the principal eigenvector of the second-order moment of feature displacements. Built upon InfoNCE and memory bank techniques, ASD consistently improves accuracy across 13 teacher–student model pairs and achieves the lowest mean corruption error (86.1 mCE) on CIFAR-100-C, significantly enhancing robust subspace representations in student models.
This work challenges the common practice in knowledge distillation of naively matching a teacher model’s absolute feature representations, which overlooks the fact that such representations are only equivalent up to orthogonal transformations and isotropic scaling. From a geometric perspective, the paper proposes a new paradigm centered on representation equivalence classes: the student should instead learn class-invariant structures of the teacher’s representations—such as Gram matrices, centered kernel alignment (CKA), or principal subspaces—or leverage coordinate alignment for effective supervision. This framework unifies feature matching, relational distillation, and grafting approaches, revealing that logit-level matching is ultimately key to capability transfer. Experiments on Qwen2.5 and Llama-3.1 demonstrate that high CKA similarity alone is insufficient for performance recovery, while successful grafting hinges on boundary overlap in the training data coverage, thereby validating the proposed theory.