๐ค AI Summary
This study addresses the accuracy degradation in ultra-low-bitwidth quantization caused by distribution shifts between training and deployment. To mitigate this issue, we propose a two-stage framework comprising block-level Quantization-Aware Training (QAT) initialization and online policy distillation. The core innovation lies in overcoming the limitations of offline fixed completion by leveraging the student model's access states to provide reverse KullbackโLeibler divergence signals, which are jointly optimized with sampling guidance from a frozen full-precision teacher model. Experimental results demonstrate that the proposed method achieves an average score of 57.28 on Qwen3-1.7B under W3A16 quantization, outperforming the baseline by 2.90 points and establishing new state-of-the-art performance for ultra-low-bit quantization.
๐ Abstract
Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.