🤖 AI Summary
This study addresses the challenges of significant distributional discrepancies and coupled errors in joint ultra-low-bit quantization of weights, activations, and KV caches by proposing CanonQ, a unified quantization-aware training framework. Methodologically, it introduces fixed rotations and energy normalization to map tensors into canonical coordinates, enabling the frozen reuse of cross-layer Gaussian reference codebooks. Theoretically, a normalization-aware straight-through Jacobian is derived to establish an analytical relationship between quantization distortion and gradient bias. During training, source canonicalization and task adaptation are decoupled for joint network optimization to suppress error coupling. Under the W2A4KV2 configuration, CanonQ reduces perplexity by 14.28×, improves zero-shot accuracy by 57.9%, and substantially enhances code generation and mathematical reasoning capabilities.
📝 Abstract
Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.