Massive Spikes in LLMs are Bias Vectors: Mechanistic Uncovering and Spike-Free Quantization

📅 2026-06-01
📈 Citations: 0
Influential: 0
📄 PDF

career value

211K/year
🤖 AI Summary
This work addresses the detrimental impact of activation outliers in large language models (LLMs), which severely degrade quantization performance due to their expanded dynamic range. The study reveals, for the first time, that these outliers originate from structured vector biases in attention weight matrices rather than the commonly assumed scalar offsets. By analyzing the geometric alignment of weight projections and the rotational stability under RoPE perturbations, the authors propose INSERTQUANT, a spike-free post-training quantization framework. This approach integrates template-based vector reconstruction to preserve activation fidelity, achieving state-of-the-art tensor-level quantization accuracy on LLMs. Notably, INSERTQUANT demonstrates strong generalization beyond textual modalities, successfully extending to vision Transformers and other non-language architectures.
📝 Abstract
Massive activation spikes in Large Language Models (LLMs) severely degrade quantization by stretching dynamic ranges. While prior hypotheses characterize these as high-level scalar biases, we argue that they are merely the scalar intermediates of rigid, structural vector biases in the spike-carrying tokens. We show that these tokens converge to constant vectors after normalization that drive the attention sink and value-state drain mechanisms. We geometrically substantiate this by analyzing the coordination of projection weights: $W_K$ contrastively amplifies the vector, $W_Q$ aligns semantic tokens toward it, and $W_V$ projects it into the spectral null-space. Furthermore, we reveal that the model actively preserves these structural biases against Rotary Positional Embedding (RoPE) perturbations by localizing them in "zones of rotational stability" utilizing low-frequency bands and coherent channel pairs. Leveraging this, we propose INSERTQUANT, a post-training quantization (PTQ) framework that clamps spikes and restores their function via pre-computed template vectors. This renders activations strictly spike-free, enabling robust low-bit quantization with high fidelity. INSERTQUANT achieves parity with state-of-the-art per-tensor quantization methods on LLMs and uniquely generalizes beyond text to other modalities such as ViTs.
Problem

Research questions and friction points this paper is trying to address.

activation spikes
quantization
Large Language Models
bias vectors
dynamic range
Innovation

Methods, ideas, or system contributions that make the work stand out.

spike-free quantization
structural bias vectors
rotational stability
post-training quantization
mechanistic interpretability
Y
Yung-Chin Chen
Princeton University, NJ, USA
C
Chung Peng Lee
Princeton University, NJ, USA
Z
Ze-Wei Liou
Princeton University, NJ, USA
Naveen Verma
Naveen Verma
Professor of Electrical Engineering, Princeton University
Machine-learning hardwareintelligent sensinglarge-area electronics