patch-level distillation

Designs and implements teacher–student distillation methods that transfer knowledge at the patch (local-feature) level by aligning multi-level teacher and student features, including patch-wise feature maps or similarity distributions (PKT). Builds and evaluates student models that retain local patch information while enabling cheaper global-only inference and reducing inference computational cost.

patch-leveldistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.47
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$184K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of conventional knowledge distillation caused by mismatched feature distributions between teacher and student models. To mitigate this issue, the authors propose DSKD, a novel approach that integrates a lightweight diffusion model to perform denoising sampling on student features under the guidance of the teacher’s classifier. A self-distillation mechanism is then established between the original and denoised student features to enhance representation learning. Furthermore, locality-sensitive hashing (LSH) is employed to enable efficient feature alignment, effectively alleviating mapping discrepancies. Extensive experiments demonstrate that DSKD consistently outperforms existing distillation methods across multiple vision recognition tasks and model architectures, achieving superior performance and strong generalization capability.

Feature Distribution MismatchIncompatible Knowledge TransferKnowledge Distillation

Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures

Jul 22, 2024
KB
Kuluhan Binici
🏛️ National University of Singapore

Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.

Computational CostKnowledge DistillationModel Adaptability

This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.

feature-based methodsknowledge distillationmodel performance

Relational Representation Distillation

Jul 16, 2024
NG
Nikolaos Giakoumoglou
🏛️ Imperial College London

Existing knowledge distillation methods struggle to model the structured relationships among internal representations of teacher models, while mainstream contrastive learning objectives (e.g., InfoNCE) impose overly stringent instance discrimination constraints, disrupting relative semantic similarities among semantically proximal samples. To address these limitations, we propose Relational Representation Distillation (RRD). Its core innovations are: (1) a dual-temperature Softmax mechanism—employing a high temperature to emphasize dominant relational patterns and a low temperature to preserve secondary semantic similarities; and (2) a theoretically unified loss that bridges InfoNCE and KL divergence, enabling relative distribution alignment. Evaluated on multi-task transfer learning benchmarks, RRD significantly improves teacher–student representation alignment. Notably, on several downstream tasks, student models trained with RRD even surpass their teachers in performance—demonstrating both the effectiveness of structured relational modeling and its strong generalization capability.

Avoiding overly strict contrastive learning constraintsCapturing structural relationships in teacher modelsPreserving relative instance relationships effectively

Existing methods for assessing the quality of AI-generated images struggle to balance accuracy and efficiency. To address this challenge, this work proposes a multi-level transfer framework based on knowledge distillation, which employs a teacher model with hybrid local–global processing and a lightweight student model relying solely on global features. Efficient representation transfer is achieved through multi-level feature distillation and fusion. Experimental results on four AIGIQA datasets demonstrate that the student model reduces computational overhead by 67.7% while maintaining evaluation performance comparable to that of the teacher model, significantly outperforming current state-of-the-art approaches.

AI-generated image quality assessmentcomputational efficiencyimage generation

Latest Papers

What's happening recently
View more

Logit-Based Losses Limit the Effectiveness of Feature Knowledge Distillation

Nov 18, 2025
NC
Nicholas Cooper
🏛️ University of Colorado Boulder

Existing feature-based knowledge distillation (KD) methods still rely on logit-level losses (e.g., cross-entropy), hindering effective transfer of intermediate-layer feature knowledge. Method: We propose the first purely feature-driven KD framework that completely eliminates logit supervision. Instead, it trains student backbone networks via intermediate-feature alignment and geometric analysis of latent-space representations. We introduce a novel metric to quantitatively assess feature knowledge quality, enabling adaptive selection of optimal teacher layers; additionally, we design a distribution-aware alignment loss grounded in feature geometry to enhance representation consistency. Contribution/Results: Our method achieves significant improvements over state-of-the-art approaches on three image classification benchmarks, with up to 15% absolute Top-1 accuracy gain. It is the first work to empirically validate both the feasibility and superiority of high-fidelity feature knowledge transfer without any logit-level supervision.

Achieves state-of-the-art accuracy improvements across diverse neural network architecturesIntroduces knowledge quality metric to identify optimal teacher layers for transferProposes logit-free feature distillation to overcome limitations of logit-based losses

This work addresses the limitations of existing knowledge distillation methods when integrating multiple heterogeneous strategies, which often suffer from implementation complexity, rigid combinations, and catastrophic forgetting. To overcome these challenges, the authors propose a Sequential Multi-Stage Knowledge Distillation (SMSKD) framework that applies distinct distillation techniques in successive stages. Each stage leverages a frozen reference model from the previous stage to anchor learned knowledge, while a sample-level adaptive loss weighting mechanism—based on the teacher’s true class probability (TCP)—dynamically balances knowledge retention and integration. The framework flexibly accommodates arbitrary distillation strategies and numbers of stages, consistently yielding significant accuracy improvements for student models across diverse teacher–student architectures, outperforming current baselines with negligible computational overhead.

catastrophic forgettingheterogeneous methodsknowledge distillation

Although knowledge distillation is widely employed to enhance model generalization, its theoretical underpinnings remain poorly understood. This work models the teacher–student training dynamics as a coupled stochastic process and introduces a novel “distillation divergence” to quantify the discrepancy between teacher and student. Building upon this, we develop an information-theoretic framework for generalization analysis and derive upper and lower bounds on the student’s generalization error that explicitly depend on the distillation divergence. Notably, we show that the local flatness of the teacher model strictly tightens the upper bound. In the Gaussian linear setting, we further provide an interpretable decomposition of the error into bias, variance, and a rank bottleneck, offering both theoretical insights and practical principles for designing effective distillation algorithms.

distillation divergencegeneralizationinformation theory

This study addresses the limitation of traditional knowledge distillation, where blindly imitating the teacher’s probability distribution compromises the student model’s inherent discriminative capacity. To overcome this, we propose Task-Preserving Knowledge Distillation (TPKD). Grounded in a theoretical analysis of KL divergence, TPKD precisely decouples full imitation from conditional learning, facilitating superior knowledge transfer by retaining label gradients while rectifying conditional gradients. Theoretically, we demonstrate that the rectified conditional direction preserves over half of the first-order descent benefits. Experimental results show that TPKD achieves accuracies of 88.05% and 93.81% on CIFAR-100 and CLINC150, respectively, significantly outperforming standard distillation methods.

Class ProbabilitiesDiscriminationFull Imitation

Hot Scholars

JL

Jianghao Lin

Shanghai Jiao Tong University
Large Language ModelsAI AgentsRecommender Systems
RK

Rajgopal Kannan

MIS Division, Army Research Office; Electrical Engg, USC
Graph Learning and AnalyticsAccelerationOptimizationCPS
YJ

Yuhua Jiang

Tsinghua University
reinforcement learning
YY

Yong Yu

Materials Engineer
Polymer matrix compositeadhesivemodelingtest development
HS

Hina Saeeda

Post-Doctoral Researcher, at Chalmers |Gothenburg University, Sweden
Software EngineeringAgile Software EngineeringSE4AIRE4AI