🤖 AI Summary
This study addresses the misconception in spiking neural networks (SNNs) that multi-timestep execution inherently causes high latency, noting that indiscriminately increasing waiting time degrades both accuracy and speed. This work proposes Falcon, a framework challenging the assumption that more local timesteps necessarily yield higher latency. It introduces pipeline latency search to balance the speed gains from inter-layer overlapped computation against the accuracy loss from incomplete inputs. Furthermore, Falcon achieves end-to-end optimization by integrating spike-based quantization-aware training, threshold and membrane potential fine-tuning, and compute-in-memory mapping. Experimental results demonstrate accuracies of 96.31% and 83.02% on the GSCV2 and SSC datasets, respectively, with core latencies as low as 119.64 and 124.00 microseconds.
📝 Abstract
Spiking neural networks are attractive for low-power speech command recognition, yet their latency has received far less attention than their energy efficiency, and their multi-timestep execution is widely assumed to make them slower than quantized neural networks. This paper challenges the assumption that more local timesteps necessarily imply higher network latency. By overlapping computation across adjacent layers at the timestep level, SNNs may complete execution in less time than comparable bit-serial QNNs. However, this overlap relies on spikes firing on incomplete inputs, and a spike once generated cannot be withdrawn, so its error persists and reduces accuracy. Waiting for more input before firing would seem to improve accuracy at the cost of reduced overlap. Yet we find and prove that this intuition fails at some layers, where even a small increase in waiting can change spike timing and downstream computation, making the network both slower and less accurate. We therefore propose a Pipeline Delay Search method which selects each layer's delay by balancing task-level accuracy gains against added network latency. We then adapt the selected configurations through spike-based quantization-aware training and bounded tuning of firing thresholds and initial membrane potentials. Together, these steps form Falcon, a framework for Fine-grained Analysis of Latency and Controlled firing which systematically analyzes and optimizes SNN latency under a spatial analog compute-in-memory mapping with shared digital engines. We evaluate Falcon on GSCV2 and SSC, achieving competitive accuracies of 96.31 and 83.02 at modeled network-core latencies of 119.64 and 124.00us, respectively. Together, our analysis and results show that SNNs can compute more yet finish faster, and wait longer yet predict worse, highlighting why Falcon matters for both latency and accuracy.