🤖 AI Summary
Directly training spiking neural networks (SNNs) on static images leads to temporal collapse and hinders effective modeling of spatiotemporal dynamics, as conventional time encoding schemes—e.g., repeated frame presentation—induce rate-coding bias rather than exploiting rich temporal structure.
Method: We propose a learnable phase-shift temporal encoding mechanism that maps static images to spike trains with adaptive timing patterns. Crucially, we decouple encoding design from network optimization, revealing that convolutional layer learnability and surrogate gradient formulation—not the encoding itself—are the primary determinants of performance. Accordingly, we design a minimal, end-to-end trainable temporal encoder.
Results: Our approach significantly narrows the accuracy gap between direct and rate coding, preserves the energy efficiency of SNNs, and enhances spatiotemporal feature representation. It establishes a new paradigm for efficient, temporally expressive modeling of static images in SNNs.
📝 Abstract
Handling static images that lack inherent temporal dynamics remains a fundamental challenge for spiking neural networks (SNNs). In directly trained SNNs, static inputs are typically repeated across time steps, causing the temporal dimension to collapse into a rate like representation and preventing meaningful temporal modeling. This work revisits the reported performance gap between direct and rate based encodings and shows that it primarily stems from convolutional learnability and surrogate gradient formulations rather than the encoding schemes themselves. To illustrate this mechanism level clarification, we introduce a minimal learnable temporal encoding that adds adaptive phase shifts to induce meaningful temporal variation from static inputs.