TAS: Distilling Arbitrary Teacher and Student via a Hybrid Assistant

📅 2024-10-16
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Cross-architecture knowledge distillation (CAKD) faces significant challenges in aligning features across heterogeneous models—such as CNNs, Vision Transformers (ViTs), and MLP-based architectures—due to divergent inductive biases and functional disparities in their modules. Method: We propose a novel hybrid assistant model that integrates convolutional and self-attention mechanisms, serving as an interpretable, transferable knowledge bridge between teacher and student. To enhance feature mapping robustness and discriminability, we replace conventional MSE loss with a spatially agnostic InfoNCE contrastive loss. Our framework further incorporates cross-modal feature distillation and spatial smoothing regularization within a unified CAKD paradigm. Contribution/Results: Evaluated on CIFAR-100 and ImageNet-1K, our method achieves state-of-the-art accuracy gains of up to 11.47% and 3.67%, respectively—substantially outperforming existing CAKD approaches. This work establishes a new paradigm for cooperative learning among architecturally diverse models.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Most knowledge distillation (KD) methodologies predominantly focus on teacher-student pairs with similar architectures, such as both being convolutional neural networks (CNNs). However, the potential and flexibility of KD can be greatly improved by expanding it to novel Cross-Architecture KD (CAKD), where the knowledge of homogeneous and heterogeneous teachers can be transferred flexibly to a given student. The primary challenge in CAKD lies in the substantial feature gaps between heterogeneous models, originating from the distinction of their inherent inductive biases and module functions. To this end, we introduce an assistant model as a bridge to facilitate smooth feature knowledge transfer between heterogeneous teachers and students. More importantly, within our proposed design principle, the assistant model combines the advantages of cross-architecture inductive biases and module functions by merging convolution and attention modules derived from both student and teacher module functions. Furthermore, we observe that heterogeneous features exhibit diverse spatial distributions in CAKD, hindering the effectiveness of conventional pixel-wise mean squared error (MSE) loss. Therefore, we leverage a spatial-agnostic InfoNCE loss to align features after spatial smoothing, thereby improving the feature alignments in CAKD. Our proposed method is evaluated across some homogeneous model pairs and arbitrary heterogeneous combinations of CNNs, ViTs, and MLPs, achieving state-of-the-art performance for distilled models with a maximum gain of 11.47% on CIFAR-100 and 3.67% on ImageNet-1K. Our code and models will be released.
Problem

Research questions and friction points this paper is trying to address.

Addressing feature gaps in Cross-Architecture Knowledge Distillation (CAKD)
Improving heterogeneous feature alignment via spatial-agnostic loss
Enhancing knowledge transfer between diverse model architectures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Assistant model bridges heterogeneous teacher-student feature gaps
Combines convolution and attention modules for fusion
Uses spatial-agnostic InfoNCE loss for feature alignment
🔎 Similar Papers
No similar papers found.
Wuhan University | Tencent YouTu Lab
G
Guopeng Li
School of Computer Science, Wuhan University
Q
Qiang Wang
Tencent YouTu Lab
K
Ke Yan
Tencent YouTu Lab
S
Shouhong Ding
Tencent YouTu Lab
Y
Yuan Gao
School of Electronic Information, Wuhan University
Gui-Song Xia
Gui-Song Xia
School of Artificial Intelligence, Wuhan University, China
Artificial IntelligenceComputer VisionPhotogrammetryRemote SensingRobotics