Swapped Logit Distillation via Bi-level Teacher Alignment

📅 2025-04-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In knowledge distillation, the teacher model’s raw logit distribution often causes erroneous alignment in student predictions with low confidence. To address this, we propose Exchange Logit Distillation (ELD), a novel paradigm that constructs a dual-teacher collaborative framework by dynamically swapping teacher logits—thereby decoupling output modeling from probability calibration. We further introduce a phased loss scheduling mechanism to mitigate the risk of misguidance from overconfident (i.e., maximum) logits. Unlike conventional single-teacher approaches that assume direct hard or soft label transfer, ELD abandons this restrictive assumption. Extensive experiments on ResNet and ViT backbones demonstrate consistent improvements across multiple image classification benchmarks, outperforming state-of-the-art distillation methods. Notably, ELD yields substantial gains in both accuracy and generalization for low-capacity student models.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Diffusion Models for VisionSearch and Optimization: Learning to Search

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Knowledge distillation (KD) compresses the network capacity by transferring knowledge from a large (teacher) network to a smaller one (student). It has been mainstream that the teacher directly transfers knowledge to the student with its original distribution, which can possibly lead to incorrect predictions. In this article, we propose a logit-based distillation via swapped logit processing, namely Swapped Logit Distillation (SLD). SLD is proposed under two assumptions: (1) the wrong prediction occurs when the prediction label confidence is not the maximum; (2) the"natural"limit of probability remains uncertain as the best value addition to the target cannot be determined. To address these issues, we propose a swapped logit processing scheme. Through this approach, we find that the swap method can be effectively extended to teacher and student outputs, transforming into two teachers. We further introduce loss scheduling to boost the performance of two teachers' alignment. Extensive experiments on image classification tasks demonstrate that SLD consistently performs best among previous state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

Addresses incorrect predictions in knowledge distillation
Proposes swapped logit processing for teacher-student alignment
Enhances performance via loss scheduling and dual teachers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Swapped logit processing for distillation
Bi-level teacher alignment method
Loss scheduling enhances performance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Stephen Ekaputra Limantoro
Department of Electrical Engineering and Computer Science, National Yang Ming Chiao Tung University, Hsinchu 300093, Taiwan.
J
Jhe-Hao Lin
Department of Electrical Engineering and Computer Science, National Yang Ming Chiao Tung University, Hsinchu 300093, Taiwan.
Chih-Yu Wang
Chih-Yu Wang
Research Center for Information Technology Innovation, Academia Sinica
CommunicationGame TheorySocial NetworkingWireless NetworkingLTE
Y
Yi-Lung Tsai
Cyberlink, New Taipei City 231, Taiwan.
Hong-Han Shuai
Hong-Han Shuai
National Yang Ming Chiao Tung University
Deep LearningData MiningMultimedia Processing
Ching-Chun Huang
Ching-Chun Huang
National Yang Ming Chiao Tung University
Computer VisionSignal ProcessingMachine Learning
Wen-Huang Cheng
Wen-Huang Cheng
Professor, IEEE Fellow, National Taiwan University
Artificial IntelligenceMultimediaComputer VisionMachine Learning