🤖 AI Summary
This study addresses the high inference energy consumption of large language models and the challenge of activation outliers in spike-driven quantization that necessitate long firing windows. To overcome these limitations, this work proposes a short-window spike-driven quantization framework. Methodologically, it introduces the novel Logarithmic Ternary Integrate-and-Fire (LTIF) neuron, which integrates selective outlier admission, group-wise adaptive gain, and residual dynamic modeling to significantly enhance single-step information capacity while preserving shift-accumulate compatibility. Experimental evaluations on OPT and Llama model families demonstrate that the proposed framework achieves state-of-the-art accuracy while reducing linear transformation energy consumption by approximately 80% compared to baselines, thereby realizing an effective balance between predictive performance and energy efficiency.
📝 Abstract
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.