Progressive Multi-Ancestor Bit-Depth Distillation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training instability and severe performance degradation caused by direct quantization distillation from FP32 to INT4. To overcome these challenges, we propose PMABD, a novel framework that introduces a joint supervision mechanism combining multi-ancestor teacher knowledge transfer with progressive bit-width compression. Through multi-stage collaborative distillation, this approach effectively suppresses quantization noise. Furthermore, a saturation-based adaptive stopping criterion is incorporated to substantially enhance training stability in ultra-low-bit quantization scenarios. Experimental results demonstrate that PMABD outperforms existing state-of-the-art methods on standard benchmarks such as CIFAR, achieving a 1.06% accuracy improvement for W2A2 models.
📝 Abstract
Model compression strategies are widely employed to reduce memory footprint and network complexity, particularly for devices with constrained computational, memory, and energy resources. Prior works that rely on simultaneous conversion from floating-point high-precision (FP32) to integer low-precision (INT4) representations and distillation into smaller models suffer from unstable training and drastic degradation of prediction performance. To address these limitations, we propose a unified framework, known as \textbf{P}rogressive \textbf{M}ulti-\textbf{A}ncestor \textbf{B}it-depth \textbf{D}istillation (PMABD), that progressively compresses the network while transferring knowledge through a growing pool of higher-precision ancestor teachers. PMABD generates a sequence of intermediate teachers that each learn from all higher-precision ancestors and jointly supervise the final target student. This multi-ancestor, multi-stage design stabilizes ultra-low-bit quantization by lowering quantization noise profiles across training and ensuring stable quantization. Experiments on CIFAR-10/100 with ResNet-20/32/18, and Tiny-ImageNet with MobileNetV2 show that PMABD outperforms state-of-the-art compression frameworks, results in 1.06$\%$ increase in performance of W2A2 (ResNet-18/CIFAR-100) student model. We show that a saturation-based stopping criterion contributes to improve the performance of our final student.
Problem

Research questions and friction points this paper is trying to address.

model compression
knowledge distillation
quantization
low-bit precision
training stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Progressive Distillation
Multi-Ancestor Knowledge Transfer
Bit-Depth Quantization
Model Compression
Ultra-Low-Bit
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Adil Mubashir Chaudhry
LUMS School of Science and Engineering
O
Osama Ahmad
University of Massachusetts Amherst, USA
Z
Zubair Khalid
LUMS School of Science and Engineering
Murtaza Taj
Murtaza Taj
Associate Professor of Computer Science, LUMS School of Science & Engineering
Computer ScienceGraphicsImage ProcessingComputer Vision