🤖 AI Summary
This study addresses the high power consumption of audio denoising networks on edge devices and the poor hardware compatibility of spiking neural networks (SNNs). To overcome these limitations, this work proposes a hardware-friendly heterogeneous computing SNN architecture. The Spiking-FullSubNet model is optimized through quantization-aware training and activation function simplification, while a sparsity-aware flexible processing element array is designed for FPGA implementation to enhance energy efficiency. Evaluated on the PYNQ-Z1 platform, the proposed approach achieves efficient inference with a real-time factor of 0.727, reducing per-frame energy consumption to 52.9 nJ and lowering power consumption by approximately 28×. This work presents an effective hardware-software co-design solution for low-power, real-time speech enhancement at the edge.
📝 Abstract
In this work we look to neuromorphic computing to solve the power consumption problem that audio denoising neural networks face on edge devices like smartphones, wireless headphones and hearing aids. Spiking neural networks (SNNs) have the potential to solve this problem due to their high activation sparsity and low complexity, however many SOTA SNNs require hardware that supports a mixture of operations to be able to fully perform inference. To solve this problem, we convert SOTA audio denoising neural network Spiking-FullSubNet to a hardware friendly version showing that via QAT and activation function simplification we can achieve $\approx28\times$ improvement in power consumption to 52.9nJ per 32ms audio frame when calculated for custom digital hardware in a 45nm process node. We then propose a digital circuit which by means of a sparsity-aware flexible PE array can perform inference of the heterogeneous compute load of Spiking-FullSubNet, and validate this circuit on a PYNQ-Z1 FPGA achieving a real-time factor of 0.727 at 100MHz.