🤖 AI Summary
This study addresses the deployment challenges of existing 4-bit activation quantization methods, which rely on complex operators such as smoothing and SVD. To this end, we propose JustQuant, a framework that introduces the Theseus QAD method to shift quantization complexity from inference to training. By leveraging quantization-aware distillation with a multi-level progressive supervision strategy, our approach enables high-quality 4-bit activation quantization during inference using only basic low-bit operators, entirely eliminating the need for additional complex operations. Experiments demonstrate that JustQuant significantly outperforms existing PTQ and QAT methods on Diffusion Transformers (DiTs) and diffusion language models. Furthermore, it maintains full compatibility with standard hardware, effectively overcoming the practical deployment bottleneck of low-bit quantization.
📝 Abstract
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-training quantization (PTQ) and quantization-aware training (QAT) methods have made progress in 4-bit activation quantization by introducing smoothing, SVD branches, rotations, mixed precision, or advanced formats such as NVFP4. These additional operators and data types impose demanding requirements on inference engines and hardware, limiting the broad adoption of low-precision models. Can quantization be achieved using only plain low-bit operators? To answer this question, we propose JustQuant, a simple yet effective framework that moves the complexity of low-bit quantization from deployment-time operators into the training process. We first revisit model quantization from the perspective of knowledge distillation and show that a key reason existing PTQ and QAT methods fail is that they typically exploit supervision at only a single level. We then introduce Theseus QAD, a quantization-aware distillation method that progressively applies multi-level supervision, analogous to the gradual replacement process in the Ship of Theseus. Extensive experiments on DiT and diffusion large language models show two distinct regimes. For smaller models, Theseus QAD can serve as a lightweight warm-up stage that substantially improves subsequent QAT with plain operators, while naive QAD may collapse in the same setting. For larger models, Theseus QAD provides a stronger distillation training path than ordinary QAD. Across both regimes, JustQuant improves low-bit quantization quality while avoiding the complex operators required by many existing PTQ methods.