ApexQuant: Data-Free Elastic Quantization by Residual Re-Isotropization

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of large language model quantization on calibration data and its significant precision degradation at low bit-widths. We propose a calibration-free elastic quantization framework that achieves residual isotropy via random rotations and progressively attenuates quantization error through recursive re-quantization of residuals. By integrating scalar, E8 lattice, and trellis-coded quantization techniques, the required number of iterations can be predicted without accessing the weights. Our approach enables multiple precisions to share a single set of quantized weight artifacts, facilitating data-free and flexible deployment. Empirically, 4-bit quantization closely approximates full-precision performance, while 2-bit quantization attains state-of-the-art results, rendering the method particularly suitable for privacy-constrained scenarios.
📝 Abstract
We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, $E_8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
Problem

Research questions and friction points this paper is trying to address.

data-free quantization
large language models
model compression
calibration-free
multi-precision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data-Free Quantization
Residual Re-Isotropization
Elastic Quantization
Calibration-Free
Large Language Models
💼 Related Jobs
No related jobs found.
A
Aksel Fristrup
Department of Computer Science, University of Copenhagen, Denmark
S
Sumit Pandey
Towards Deep Learning
Ankit Kariryaa
Ankit Kariryaa
University of Copenhagen