Adaptive Speech-to-Spike Encoding for Spiking Neural Networks

πŸ“… 2026-06-17
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the inherent mismatch between continuous speech signals and the discrete, event-driven nature of spiking neural networks (SNNs), where conventional fixed encoders struggle to produce task-optimized spike representations. The authors propose a learnable residual speech-to-spike encoder trained end-to-end with a recurrent leaky integrate-and-fire (R-LIF) SNN, directly optimizing for class separability rather than signal reconstruction. This approach achieves the first parameter-efficient, adaptive speech encoding for SNNs and provides a systematic evaluation of biologically plausible learning rulesβ€”such as Direct Feedback Alignment (DFA)β€”on audio tasks. On Google Speech Commands v2 (GSC-v2), the method attains 94.97% accuracy; a compact variant with only 35k parameters reaches 89.8%; and DFA achieves 91.5%, demonstrating its practical viability.
πŸ“ Abstract
The mismatch between continuous acoustic signals and discrete event-driven processing remains a fundamental bottleneck for neuromorphic speech processing. Current systems typically rely on fixed spike encoders, forcing downstream Spiking Neural Networks (SNNs) to compensate for non-adaptive input representations. To address this, we present a learnable residual speech-to-spike encoder jointly trained end-to-end with a Recurrent Leaky Integrate-and-Fire (R-LIF) backbone. We validate this approach on the Google Speech Commands v2 (GSC-v2) benchmark, achieving up to 94.97% accuracy. Notably, the learned encoder remains highly parameter-efficient with a compact 35k-parameter variant that reaches 89.8%, matching or exceeding prior baselines that require an order of magnitude more parameters. Our encoder-focused analysis, including linear probing and gradient-residual inspection, indicates that the encoder does not target faithful signal reconstruction but instead learns task-aligned spike representations that enhance class separability. Finally, we benchmark bio-inspired, hardware-friendly credit assignment by comparing Direct Feedback Alignment (DFA) with surrogate-gradient BPTT under identical architectures and training conditions. We find that DFA reaches 91.5% accuracy, quantifying the performance trade-off of bio-inspired learning rules for modern neuromorphic audio.
Problem

Research questions and friction points this paper is trying to address.

speech-to-spike encoding
spiking neural networks
neuromorphic speech processing
adaptive encoding
event-driven processing
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive spike encoding
spiking neural networks
end-to-end learning
bio-inspired credit assignment
speech commands classification
πŸ”Ž Similar Papers
No similar papers found.
T
Taharim Rahman Anon
PI LLC, Sapporo, Hokkaido, Japan
J
Jakaria Islam Emon
PI LLC, Sapporo, Hokkaido, Japan