QuantaSpike: Short-Window Spike-Driven Quantization for Large Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference energy consumption of large language models and the challenge of activation outliers in spike-driven quantization that necessitate long firing windows. To overcome these limitations, this work proposes a short-window spike-driven quantization framework. Methodologically, it introduces the novel Logarithmic Ternary Integrate-and-Fire (LTIF) neuron, which integrates selective outlier admission, group-wise adaptive gain, and residual dynamic modeling to significantly enhance single-step information capacity while preserving shift-accumulate compatibility. Experimental evaluations on OPT and Llama model families demonstrate that the proposed framework achieves state-of-the-art accuracy while reducing linear transformation energy consumption by approximately 80% compared to baselines, thereby realizing an effective balance between predictive performance and energy efficiency.
📝 Abstract
Large language models (LLMs) achieve strong performance across many tasks but rely on dense multiply-accumulate (MAC) operations during inference, resulting in high energy cost. Spiking neural networks (SNNs) offer an event-driven alternative in which synaptic integration uses lightweight accumulation. However, spike-driven LLM inference remains difficult because outlier-heavy activations typically require long firing windows or auxiliary non-spiking paths. We propose QuantaSpike, a short-window spike-driven quantization framework for LLMs built around Logarithmic Ternary Integrate-and-Fire (LTIF) neurons. LTIF uses ternary events with power-of-two membrane-response quanta, improving the information represented by each firing step while retaining shift-ACC-compatible computation. QuantaSpike combines this neuron with group-adaptive gain and selective outlier admission: normal values use residual LTIF steps, whereas admitted outliers receive one additional onset spike before entering the same residual dynamics. Across OPT and Llama-2, QuantaSpike achieves state-of-the-art or competitive perplexity and zero-shot accuracy among spike-driven LLM quantization methods. It also transfers to newer dense LLMs, remaining close to the FP16 reference on Llama-3-8B and Qwen3-8B under the same four-step firing window. Analytical linear-energy projections show that QuantaSpike reduces the energy of one linear transformation by about $80.0\%$ on OPT models and $67.1\%$ on Llama-2 models relative to SpikeQuant, providing an accurate and energy-efficient spike-driven path for LLM inference.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Spiking Neural Networks
Energy Efficiency
Outlier Activations
Spike-Driven Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Neural Networks
Large Language Models
Short-Window Quantization
LTIF Neurons
Energy-Efficient Inference
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
B
Bang Hu
School of Computer Science, Fudan University, Shanghai, China
G
Guowei Zhu
School of Computer Science, Fudan University, Shanghai, China
C
Changze Lv
School of Computer Science, Fudan University, Shanghai, China
Xiaoqing Zheng
Xiaoqing Zheng
Fudan University
Natural Language Processing and Machine Learning
Fengzhe Zhang
Fengzhe Zhang
University of Cambridge
Machine Learning
F
Fan Zhang
School of Computer Science, Fudan University, Shanghai, China
W
Wei Cao
School of Computer Science, Fudan University, Shanghai, China