🤖 AI Summary
To address the high energy consumption of foundational AI models, which hinders their deployment in resource-constrained environments, this paper proposes a dynamically activated spiking neural network (SNN) ensemble system. The method integrates knowledge distillation with ensemble learning and introduces a novel directed feature-space decoupling strategy that decomposes teacher artificial neural network (ANN) knowledge into multiple specialized student SNNs; these are adaptively activated based on input stimuli, enabling knowledge specialization and dynamic computational offloading. On CIFAR-10, the system reduces computation by 20× compared to the teacher ANN while incurring only a 2% accuracy drop; the decoupling strategy further improves accuracy by 2.4% and significantly enhances robustness under noise—outperforming the original teacher model. The core contribution is the first realization of feature-space-decoupling-driven dynamic sparse activation in SNN ensembles, simultaneously achieving high energy efficiency, competitive accuracy, and strong robustness.
📝 Abstract
While foundation AI models excel at tasks like classification and decision-making, their high energy consumption makes them unsuitable for energy-constrained applications. Inspired by the brain's efficiency, spiking neural networks (SNNs) have emerged as a viable alternative due to their event-driven nature and compatibility with neuromorphic chips. This work introduces a novel system that combines knowledge distillation and ensemble learning to bridge the performance gap between artificial neural networks (ANNs) and SNNs. A foundation AI model acts as a teacher network, guiding smaller student SNNs organized into an ensemble, called Spiking Neural Ensemble (SNE). SNE enables the disentanglement of the teacher's knowledge, allowing each student to specialize in predicting a distinct aspect of it, while processing the same input. The core innovation of SNE is the adaptive activation of a subset of SNN models of an ensemble, leveraging knowledge-distillation, enhanced with an informed-partitioning (disentanglement) of the teacher's feature space. By dynamically activating only a subset of these student SNNs, the system balances accuracy and energy efficiency, achieving substantial energy savings with minimal accuracy loss. Moreover, SNE is significantly more efficient than the teacher network, reducing computational requirements by up to 20x with only a 2% drop in accuracy on the CIFAR-10 dataset. This disentanglement procedure achieves an accuracy improvement of up to 2.4% on the CIFAR-10 dataset compared to other partitioning schemes. Finally, we comparatively analyze SNE performance under noisy conditions, demonstrating enhanced robustness compared to its ANN teacher. In summary, SNE offers a promising new direction for energy-constrained applications.