spatially-stratified distillation

Designs and implements knowledge-distillation procedures that partition data by spatial regions and transfer teacher information to students with region-specific targets and loss terms. This includes creating spatial alignment algorithms and weighting schemes that enforce strong feature alignment in overlapping areas and apply discounted or sparsity-aware weights in low-density regions to regularize cross-region consistency.

spatially-stratifieddistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of conventional knowledge distillation caused by mismatched feature distributions between teacher and student models. To mitigate this issue, the authors propose DSKD, a novel approach that integrates a lightweight diffusion model to perform denoising sampling on student features under the guidance of the teacher’s classifier. A self-distillation mechanism is then established between the original and denoised student features to enhance representation learning. Furthermore, locality-sensitive hashing (LSH) is employed to enable efficient feature alignment, effectively alleviating mapping discrepancies. Extensive experiments demonstrate that DSKD consistently outperforms existing distillation methods across multiple vision recognition tasks and model architectures, achieving superior performance and strong generalization capability.

Feature Distribution MismatchIncompatible Knowledge TransferKnowledge Distillation

TAS: Distilling Arbitrary Teacher and Student via a Hybrid Assistant

Oct 16, 2024
GL
Guopeng Li
🏛️ Wuhan University | Tencent YouTu Lab

Cross-architecture knowledge distillation (CAKD) faces significant challenges in aligning features across heterogeneous models—such as CNNs, Vision Transformers (ViTs), and MLP-based architectures—due to divergent inductive biases and functional disparities in their modules. Method: We propose a novel hybrid assistant model that integrates convolutional and self-attention mechanisms, serving as an interpretable, transferable knowledge bridge between teacher and student. To enhance feature mapping robustness and discriminability, we replace conventional MSE loss with a spatially agnostic InfoNCE contrastive loss. Our framework further incorporates cross-modal feature distillation and spatial smoothing regularization within a unified CAKD paradigm. Contribution/Results: Evaluated on CIFAR-100 and ImageNet-1K, our method achieves state-of-the-art accuracy gains of up to 11.47% and 3.67%, respectively—substantially outperforming existing CAKD approaches. This work establishes a new paradigm for cooperative learning among architecturally diverse models.

Addressing feature gaps in Cross-Architecture Knowledge Distillation (CAKD)Enhancing knowledge transfer between diverse model architecturesImproving heterogeneous feature alignment via spatial-agnostic loss

Asymmetric Decision-Making in Online Knowledge Distillation:Unifying Consensus and Divergence

Mar 09, 2025
ZC
Zhaowei Chen
🏛️ JIIOV Technology | University of Southern California

Online knowledge distillation (OKD) suffers from asymmetric feature learning between teacher and student models and insufficient teacher diversity. Method: This paper proposes an asynchronous decision framework that, for the first time, reveals consensus and divergence patterns of intermediate features in foreground regions between teacher and student. It introduces an asymmetric learning objective: the student focuses on high-consensus spatial features to enhance robustness, while the teacher emphasizes low-similarity regions to preserve feature diversity. The method requires no pre-trained teacher and jointly enables consensus learning and divergence learning, naturally supporting multi-task distillation. Contribution/Results: Our approach achieves state-of-the-art performance across online and offline knowledge distillation, semantic segmentation, and diffusion model distillation. It significantly improves both feature learning efficiency and generalization capability of student models.

Enhances feature consensus learning in student models.Improves performance in online and offline knowledge distillation tasks.Promotes feature diversity in teacher models.

Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures

Jul 22, 2024
KB
Kuluhan Binici
🏛️ National University of Singapore

Conventional knowledge distillation tightly couples teacher and student architectures, resulting in poor cross-architecture generalization and prohibitive retraining costs for each new student. Method: We propose the Generalized Teacher Network (GTN), the first architecture-agnostic teacher framework that models the student pool as a weight-sharing supernet and employs a capacity-aware conditional mechanism to dynamically adapt the teacher to diverse student architectures. GTN jointly trains the teacher and students in a single, distillation-aware optimization pass. Contribution/Results: GTN eliminates the need for per-student teacher training; its overhead is amortized across the student pool. Evaluated on multi-architecture student pools, GTN consistently improves accuracy by 1.2–2.8% over baseline distillation methods, significantly enhancing deployment flexibility and computational efficiency.

Computational CostKnowledge DistillationModel Adaptability

Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Oct 15, 2024
WX
Wenda Xu
🏛️ UC Santa Barbara | Google Cloud AI Research | CMU | Google DeepMind

In knowledge distillation, supervised methods suffer from train-inference distribution mismatch, while on-policy approaches yield inaccurate teacher feedback due to low-quality student-generated samples. This paper proposes Speculative Distillation—a novel framework where the student first generates candidate token sequences, and the teacher dynamically corrects only low-confidence tokens, enabling high-fidelity knowledge transfer under inference-time distribution alignment. Its core innovation is the first online, token-level, teacher-student collaborative correction mechanism, integrating confidence-driven interleaved sampling, teacher-guided dynamic reweighting, and multi-task joint training. Evaluated across machine translation, summarization, mathematical reasoning, and instruction-following tasks, the method consistently outperforms both supervised and on-policy distillation baselines. It demonstrates robust performance gains across diverse model scales, data regimes, and initialization strategies.

Addresses teacher-student knowledge gap in distillation.Enhances student model performance across diverse tasks.Improves training data quality in knowledge distillation.

Latest Papers

What's happening recently
View more

This study investigates the inconsistent performance of feature-level knowledge distillation across diverse student architectures. Conducting controlled experiments on CIFAR-100 with a ResNet-50 teacher and a range of student models—including CustomResNet variants and MobileNetV2—under unified training settings, the work systematically evaluates feature alignment methods such as Attention Transfer and FitNets against logit-only distillation. For the first time under cross-architecture conditions, multiple feature distillation strategies are rigorously compared, revealing that logit-based knowledge distillation consistently outperforms training from scratch; Attention Transfer’s efficacy is highly dependent on student architecture; FitNets underperform logit distillation in all 15 evaluated settings; and fixed auxiliary loss coefficients induce gradient scale imbalances in student networks, highlighting the critical impact of hyperparameter sensitivity on distillation effectiveness.

feature-based methodsknowledge distillationmodel performance

This study addresses the challenge in knowledge distillation where varying data budgets shift the optimal teacher capacity, complicating effective sample selection. By analyzing relational ranking and score geometry, this work reveals for the first time why smaller teachers excel under low-data regimes and proposes the DVA framework. Employing a small teacher as a proxy, DVA models relational diversity through difficulty filtering and class-conditional volume maximization, establishing a training-free dynamic data selection mechanism that jointly optimizes difficulty matching and signal diversity. Without requiring training-dependent dynamic statistics, the proposed approach achieves performance comparable to dynamic methods while consistently outperforming static baselines, offering a novel paradigm for efficient knowledge distillation.

Data PruningData SelectionKnowledge Distillation

This study addresses the capacity gap in reasoning distillation arising from distributional discrepancies between teacher and student models, which hinders conventional methods from simultaneously achieving high supervision quality and distillation efficiency. To overcome this limitation, this work proposes TeacherGRPO, a framework that directly adapts the teacher model to the student's distribution via reinforcement learning. Building upon Group Relative Policy Optimization (GRPO), the approach introduces a curriculum-based selective alignment mechanism to focus on high-signal reasoning discrepancies. Furthermore, an importance-adaptive length regularization is designed to suppress redundancy while preserving critical reasoning steps. Extensive evaluations demonstrate that TeacherGRPO significantly outperforms existing baselines across multiple reasoning benchmarks, effectively enhancing the reasoning capabilities of smaller models. The implementation code has been made publicly available.

Capacity GapKnowledge DistillationReasoning Distillation

This study addresses the challenge that structural heterogeneity in cross-modal feature representations precludes unit-level alignment in conventional knowledge distillation. To overcome this limitation, the authors propose an abstraction mechanism based on vector-quantized codebooks, which transforms teacher features into concept-level anchors to guide student learning. By introducing a task-relevance and compatibility-guided codebook selection algorithm, the proposed method circumvents the assumption of structural consistency across feature spaces, thereby enabling concept-level distillation between heterogeneous modalities without requiring direct alignment. The effectiveness of this framework is validated across multiple classification and semantic segmentation tasks, demonstrating substantial improvements in cross-modal knowledge transfer performance under heterogeneous feature conditions.

Cross-modal knowledge distillationFeature-level alignmentStructurally heterogeneous features

This work addresses the long-standing lack of theoretical guidance in selecting the temperature parameter for knowledge distillation, a choice often made empirically or via grid search, frequently yielding suboptimal student performance. For the first time, the study systematically investigates the interaction between temperature and key training factors—such as optimizer choice and whether the teacher model is pretrained or fine-tuned—and identifies representative scenarios that critically influence optimal temperature selection. Building upon a classification distillation framework and a comprehensive cross-factor experimental design, the authors propose a context-aware temperature selection strategy. This approach moves beyond conventional fixed or heuristic tuning paradigms, offering practical guidelines tailored to diverse training configurations and consistently achieving significant improvements in student model performance.

hyperparameter tuningknowledge distillationstudent performance

Hot Scholars

YX

Yi Xu

Goertek Alpha Labs
Computer visioncomputer graphicsmachine learningaugmented reality
ML

Mingfu Liang

Meta | Northwestern University
Machine LearningComputer VisionIncremental LearningContinual Learning
XZ

Xu Zou

Z.ai
language generationreasoningworld modeling
JZ

Jiahuan Zhou

Peking University
Computer VisionMachine LearningDeep Learning
HW

Hanli Wang

Tongji University
Multimedia ComputingComputer VisionImage ProcessingMachine Learning