Parallel Time-Aligned Spiking Self-Attention for Consistent Integer-Valued Training and Spike-Driven Inference

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the temporal interaction mismatch between integer training and inference in spiking self-attention by proposing a Parallel Time-aligned Spiking Self-Attention (PT-SSA) architecture. The method reconstructs virtual slices and computes attention exclusively over temporally aligned segments, thereby eliminating cross-time interference. By integrating I-LIF neurons, SFA approximation, and Triton fused kernels, PT-SSA achieves efficient parallel computation with adaptive scaling. Experimental results demonstrate that PT-SSA substantially narrows the training-inference discrepancy and enhances spike utilization. On ImageNet, it reduces the top-1 training-inference error to 0.06%, attains an accuracy of 74.53%, and yields a 2.91× improvement in training throughput.
📝 Abstract
Integer-valued leaky integrate-and-fire (I-LIF) neurons and spike firing approximation (SFA) reduce temporal training cost by representing spike trains as firing counts and normalized firing rates, respectively. However, applying spiking self-attention (SSA) directly to these compressed query, key, and value representations introduces cross-time interactions that are absent during spike-driven inference. We term this operator-level discrepancy Temporal Interaction Mismatch (TIM). We propose Parallel Time-Aligned Spiking Self-Attention (PT-SSA), which reconstructs consecutive virtual spike slices from either I-LIF counts or SFA firing rates, computes attention only between time-aligned slices in parallel, and sums the per-step outputs. To accommodate the reduced attention output scale under SFA, we further introduce Adaptive PT-SSA, which learns a positive per-block rescaling before the output SFA neuron to improve firing-level utilization. Experiments on CIFAR-10, CIFAR-100, and ImageNet-1K show that the proposed methods substantially reduce train--inference mismatch. On CIFAR-100 with I-LIF, PT-SSA reduces the mean Top-1 gap from 1.85 to 0.29 percentage points. On ImageNet-1K, Adaptive PT-SSA reduces the Top-1 gap from 27.78 to 0.06 percentage points and achieves 74.53\% spike-driven Top-1 accuracy. A Triton-fused PT-SSA training kernel retains a $2.91\times$ throughput advantage over recurrent LIF SSA.
Problem

Research questions and friction points this paper is trying to address.

Spiking Neural Networks
Spiking Self-Attention
Train-Inference Mismatch
Temporal Interaction Mismatch
Spike-Driven Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Self-Attention
Parallel Time-Aligned
Temporal Interaction Mismatch
Integer-Valued Training
Spike-Driven Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.