🤖 AI Summary
This study addresses the bottleneck of serial backpropagation updates in edge spiking neural networks, which impedes real-time inference and on-device continual learning. To overcome this, we propose a bidirectional spike distillation mechanism that decouples forward inference from backward learning pathways via local target alignment, establishing a parallel computing architecture. This enables, for the first time, the simultaneous execution of event-driven inference and online learning. The method entirely eliminates dense floating-point backward chains and transfers to few-shot incremental learning scenarios without data replay while supporting multimodal perception. Experimental results demonstrate that the model achieves accuracy comparable to baselines while reducing training latency to 0.72× and energy consumption to 0.36× of conventional approaches, exhibiting highly efficient on-device adaptability.
📝 Abstract
Edge intelligence requires models to sense continuously in real time and to keep adapting on-device, all under tight compute, energy, and memory budgets. Although spiking neural networks (SNNs) enable efficient event-driven inference, standard surrogate-gradient backpropagation (BP) serializes updates and blocks ongoing inference. We investigate Bidirectional Spike-Based Distillation (BSD) as an on-device learning principle that lets edge SNNs learn while inferring. BSD couples a stimulus-driven forward network with an independent target-driven reverse network and aligns their intermediate representations through local objectives. Because the two pathways have disjoint computation graphs until alignment, forward inference, reverse inference, and stage-wise updates can run concurrently. On 25 benchmarks spanning the five sensing modalities of SOUL, BSD stays within 3.8 percentage points of matched BP baselines on average, while only its forward branch is needed at deployment. The learned representations also transfer well to few-shot class-incremental learning without replaying past data. By removing the dense floating-point backward chain, BSD reduces the projected training latency to $0.72\times$ and the estimated training energy to $0.36\times$ that of BP.