TORQUE: Optimizing What (not) to Quantize Before and After Rotation
This study addresses the optimization of high-precision coordinate retention strategies in rotation-based quantization. We propose a method that jointly optimizes the number and positions of retained coordinates before and after rotation to minimize quantization error under a fixed bit budget. Theoretically, we prove that retaining the top-k coordinates prior to rotation minimizes the upper bound of the quantization error, thereby reducing a complex combinatorial search to the optimization of a scalar k, for which a fast parallel selection algorithm is designed. Combined with random rotation preprocessing and offline codebook optimization, our approach achieves efficient compression. Experiments demonstrate that the proposed method significantly improves the trade-off between reconstruction accuracy and storage efficiency across Gaussian modeling, nearest neighbor retrieval, KV cache compression, and activation quantization tasks.