A Closer Look at Knowledge Distillation in Spiking Neural Network Training

📅 2025-11-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The fundamental discrepancy between the continuous output distribution of artificial neural networks (ANNs) and the sparse, discrete spike-based outputs of spiking neural networks (SNNs) severely limits knowledge distillation performance. To address this, we propose a novel distillation paradigm tailored for SNNs. Our method comprises two core components: (1) saliency-scaled activation map distillation, which enables semantic-aware alignment of intermediate feature representations by weighting activations according to input saliency; and (2) noise-smoothed logits distillation, which injects Gaussian noise into SNN logits to mitigate gradient instability arising from output sparsity and discreteness. Integrated within an ANN-to-SNN conversion framework, our approach consistently improves both accuracy and convergence speed across CIFAR-10, CIFAR-100, and ImageNet. Compared to state-of-the-art distillation methods, it achieves average accuracy gains of 2.1–4.7 percentage points. The implementation is publicly available.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingComputer Vision: Diffusion Models for VisionMachine Learning: Neuro-Symbolic Learning

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Spiking Neural Networks (SNNs) become popular due to excellent energy efficiency, yet facing challenges for effective model training. Recent works improve this by introducing knowledge distillation (KD) techniques, with the pre-trained artificial neural networks (ANNs) used as teachers and the target SNNs as students. This is commonly accomplished through a straightforward element-wise alignment of intermediate features and prediction logits from ANNs and SNNs, often neglecting the intrinsic differences between their architectures. Specifically, ANN's outputs exhibit a continuous distribution, whereas SNN's outputs are characterized by sparsity and discreteness. To mitigate this issue, we introduce two innovative KD strategies. Firstly, we propose the Saliency-scaled Activation Map Distillation (SAMD), which aligns the spike activation map of the student SNN with the class-aware activation map of the teacher ANN. Rather than performing KD directly on the raw %and distinct features of ANN and SNN, our SAMD directs the student to learn from saliency activation maps that exhibit greater semantic and distribution consistency. Additionally, we propose a Noise-smoothed Logits Distillation (NLD), which utilizes Gaussian noise to smooth the sparse logits of student SNN, facilitating the alignment with continuous logits from teacher ANN. Extensive experiments on multiple datasets demonstrate the effectiveness of our methods. Code is available~footnote{https://github.com/SinoLeu/CKDSNN.git}.
Problem

Research questions and friction points this paper is trying to address.

Addressing architectural differences between continuous ANNs and discrete SNNs in knowledge distillation
Aligning sparse SNN outputs with continuous ANN features through saliency-scaled activation maps
Smoothing sparse SNN logits using Gaussian noise to match continuous ANN distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Saliency-scaled Activation Map Distillation aligns spike and activation maps
Noise-smoothed Logits Distillation smooths sparse SNN outputs
Methods address architectural differences between ANNs and SNNs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xu Liu
School of Computer Science and Information Engineering, Hefei University of Technology
N
Na Xia
School of Computer Science and Information Engineering, Hefei University of Technology
J
Jinxing Zhou
Mohamed Bin Zayed University of Artificial Intelligence
J
Jingyuan Xu
School of Computer Science and Information Engineering, Hefei University of Technology
Dan Guo
Dan Guo
IEEE senior member, Professor, Hefei University of Technology
Multimedia ComputingArtificial Intelligence