🤖 AI Summary
This study addresses the inherent conflict between limited computational resources and low-latency requirements for real-time multi-channel speech enhancement on low-power microcontrollers (MCUs). To this end, it proposes a mixed-precision neural beamforming scheme that employs a CNN to estimate MVDR weights alongside a lightweight Transformer for residual correction. Notably, this work pioneers the deployment of a complete multi-channel Transformer beamforming pipeline on MCU-class devices, achieving efficient inference through three-stage float32/int8/float16 mixed precision, time-division multiplexing, and voice activity detection. Experimental results demonstrate that the system attains an STOI of 97.65% and a PESQ score of 3.676, with a frame latency of merely 15 ms. Operating at only 45.9 mW, it enables over 16 hours of battery life, thereby achieving an exceptional balance between enhancement performance and energy efficiency.
📝 Abstract
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resource-constrained microcontroller units (MCUs) used in hearables. We present an optimized methodology for the real-time execution of a neural-network-based minimum variance distortionless response (MVDR) beamformer on MCUs. Using six microphones and a three-stage mixed-precision scheme (float32 MVDR, int8 CNN, float16 Transformer), the pipeline pairs a CNN that estimates the MVDR weights with a lightweight Transformer that applies a per-frame correction. By time-slicing weight estimation with beamforming, it achieves a 15~ms per-frame latency while refreshing a complete set of CNN-derived weights every 564~ms. The deployed mixed-precision pipeline attains a short-time objective intelligibility (STOI) of 97.65\%, a scale-invariant signal-to-noise ratio (SI-SNR) of 20.26~dB, and a wideband PESQ of 3.676 at an average power of 45.9~mW. A speech activity detection (SAD) module (98.5\% accuracy, 0.62~mJ per inference) bypasses the pipeline during silence; under realistic deployment conditions, the system exceeds the 16~h all-day target on a 100~mAh battery, with an estimated lifetime of up to $\sim$20~h. To our knowledge, this is the first real-time multi-channel Transformer-based neural beamforming pipeline deployed on an MCU-class device.