Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对低比特量化模型在长生成任务中的性能下降问题,提出了一种在线策略蒸馏方法(OPD),通过在量化模型实际路径上提供教师指导来改善数学和代码推理能力。
📝 Abstract
Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathematical and code reasoning substantially impaired. Long generations often degenerate into repetitive loops, exhausting the decoding budget without completing a solution. We trace this gap to quantization-amplified exposure bias: QAD trains on fixed corpus prefixes, while quantization-induced deviations compound along the model's own autoregressive trajectories. To address this mismatch, we introduce an on-policy distillation (OPD) stage that places teacher supervision where the quantized model actually goes. Starting from a QAD checkpoint, the student generates through the quantized forward path used at deployment and receives feedback from a frozen full-precision teacher on its own prefixes, combining dense token-level guidance with task-verifier rewards. Across four models at 2.79 and 1.88 effective bits, OPD raises average BF16 performance retention from 35% to 70% on MATH-500 and from 66% to 91% on HumanEval while preserving short-form performance, with reasoning gains substantially exceeding those of continued teacher-forced QAD in matched-budget comparisons. By coupling QAD's stable low-bit initialization with OPD's on-policy reasoning recovery, our framework provides a comprehensive sub-3-bit solution that preserves broad capabilities while restoring long-form reasoning.
Problem

Research questions and friction points this paper is trying to address.

Quantization
Exposure Bias
Long-Form Reasoning
On-Policy Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

on-policy distillation
quantization-aware distillation
exposure bias
low-bit reasoning
autoregressive trajectories
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuanteng Chen
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Zhongguancun Academy
Z
Zhilei Liu
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
Peisong Wang
Peisong Wang
CASIA
Deep Neural Network Acceleration and Compression
Y
Yuantian Shao
Nanjing University of Science and Technology (NJUST)
C
Chuangyi Li
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences
Weining Wang
Weining Wang
Institute of Automation, Chinese Academy of Sciences
Video UnderstandingVideo GeneratationMulti-Modal Analysis
Shuang Qiu
Shuang Qiu
City University of Hong Kong
Reinforcement LearningAgentic AILarge Language ModelsEmbodied AI
Gang Li
Gang Li
Institute of Automation, Chinese Academy of Sciences
Computer ArchitectureMachine Learning
J
Jing Liu
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Zhongguancun Academy
J
Jian Cheng
Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Zhongguancun Academy