CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
This study addresses the limitations of existing large language model quantization methods that rely on rotation operations, which increase decoding overhead and fail to fully exploit Hessian curvature information. We establish, for the first time, the equivalence between rotation and weighted search, proposing a rotation-free lattice quantization framework. Specifically, diagonal Hessian weights are incorporated into the Viterbi search metric to achieve curvature-aware encoding, while error-feedback residuals adaptively adjust the lattice initial state to eliminate the need for rotation. Evaluated on 4–8B and 35B mixture-of-experts models, the proposed method outperforms QTIP and Proteus by 1–3 percentage points at 2-bit precision. Furthermore, by obviating inverse rotation computations, it achieves state-of-the-art decoding speed.