🤖 AI Summary
This study addresses the trade-off between short-window semantic representation and computationally expensive nonlinear attention in spiking language models by proposing Spora, a framework that jointly designs spike encoding and attention operators. Methodologically, it introduces unipolar and bipolar binary spike encoding to achieve combinatorial value representation with sign separation. Furthermore, through binary temporal weighting, threshold-triggered residual decay, cumulative shifted dot products, and integer exponential mapping, the framework enables efficient integer arithmetic while enhancing information capacity. Experimental results demonstrate that Spora achieves a GLUE score of 76.6 (with a CoLA MCC of 44.1) at a four-step window, improving to 78.2 and 47.4, respectively, at six steps. These results significantly outperform existing SpikeLM baselines, validating the effectiveness of the proposed joint design for efficient and expressive spiking language modeling.
📝 Abstract
Spiking language models face a tradeoff between representing continuous semantic features over short temporal windows and retaining costly nonlinear attention operations. We introduce Spora, which jointly designs spike encodings and attention operators. Binary temporal weights let $T$ spikes represent compositional values with up to $T$ bits of capacity, compared with $O(\log_2 T)$ bits for spike-count readout. Unipolar Binary Spiking (UBS) uses thresholds and spike-triggered residual decay to produce non-negative integer codes; Bipolar Binary Spiking (BBS) separates sign and magnitude and learns a scale for signed activations. These representations support accumulation-and-shift dot products and integer-exponent mappings in attention. With four time steps, Spora achieves 76.6 average GLUE score and 44.1 CoLA MCC, improving over SpikeLM by 1.2 and 6.2 points, respectively. Extending BBS to six steps raises these scores to 78.2 and 47.4. Conditional-decay analysis, matched-budget activation-quantization comparisons, event-workload statistics, and fixed-point evaluation further characterize the connection between encoding fidelity and computational cost.