🤖 AI Summary
This study addresses the limitations of existing large language model quantization methods that rely on rotation operations, which increase decoding overhead and fail to fully exploit Hessian curvature information. We establish, for the first time, the equivalence between rotation and weighted search, proposing a rotation-free lattice quantization framework. Specifically, diagonal Hessian weights are incorporated into the Viterbi search metric to achieve curvature-aware encoding, while error-feedback residuals adaptively adjust the lattice initial state to eliminate the need for rotation. Evaluated on 4–8B and 35B mixture-of-experts models, the proposed method outperforms QTIP and Proteus by 1–3 percentage points at 2-bit precision. Furthermore, by obviating inverse rotation computations, it achieves state-of-the-art decoding speed.
📝 Abstract
The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks. We show that this leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read. We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block. This also explains the rotation: it removes this within-block variation, so weighting in the native basis and rotating are substitutes. On three models the weighted native search matches a full-dimension randomized Hadamard to within about one point of downstream accuracy, and weighting after the rotation gains little. Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights' amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it. At two bits CurveTQ is 1-3 points higher in mean downstream accuracy than QTIP and Proteus on three 4-8B Instruct models, even after both are given our start state, which alone lifts either baseline by 1-3 points. It also leads on a 35B mixture of experts, to our knowledge the first trellis-coded result on such a model. With no rotation to undo, our decoder is the fastest of the three at all tested batch sizes and bit widths.