TORQUE: Optimizing What (not) to Quantize Before and After Rotation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization of high-precision coordinate retention strategies in rotation-based quantization. We propose a method that jointly optimizes the number and positions of retained coordinates before and after rotation to minimize quantization error under a fixed bit budget. Theoretically, we prove that retaining the top-k coordinates prior to rotation minimizes the upper bound of the quantization error, thereby reducing a complex combinatorial search to the optimization of a scalar k, for which a fast parallel selection algorithm is designed. Combined with random rotation preprocessing and offline codebook optimization, our approach achieves efficient compression. Experiments demonstrate that the proposed method significantly improves the trade-off between reconstruction accuracy and storage efficiency across Gaussian modeling, nearest neighbor retrieval, KV cache compression, and activation quantization tasks.
📝 Abstract
Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top-$k$ pre-rotation retention minimizes it for each $k$. This reduces the search over coordinate subsets to an optimization over $k$, enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression.
Problem

Research questions and friction points this paper is trying to address.

Quantization
Random Rotation
Bit Budget Optimization
Coordinate Preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantization
Random Rotation
Joint Optimization
Error Bound
Top-k Retention