🤖 AI Summary
To address the capacity degradation and difficulty in dynamically tracking bottlenecks in low-bit quantized neural networks—caused by non-differentiable rounding operations—this paper models quantization as a cascade of noisy channels and proposes a progressively differentiable noise-scaled quantization framework. Methodologically, it jointly optimizes learnable bit-widths, noise scales, and clipping ranges via straight-through estimators (STE) for end-to-end training; incorporates an outlier-penalty term to precisely enforce target bit-widths; and integrates lightweight knowledge distillation to enhance training stability and gradient smoothness. The approach achieves competitive accuracy under the extremely challenging W1A1 setting—surpassing state-of-the-art methods significantly—while maintaining efficient forward inference. Key contributions include: (i) a noise-modeling-driven differentiable quantization paradigm, and (ii) a multi-variable co-constraint mechanism unifying bit-width, noise, and clipping optimization.
📝 Abstract
Quantized neural networks can be viewed as a chain of noisy channels, where rounding in each layer reduces capacity as bit-width shrinks; the floating-point (FP) checkpoint sets the maximum input rate. We track capacity dynamics as the average bit-width decreases and identify resulting quantization bottlenecks by casting fine-tuning as a smooth, constrained optimization problem. Our approach employs a fully differentiable Straight-Through Estimator (STE) with learnable bit-width, noise scale and clamp bounds, and enforces a target bit-width via an exterior-point penalty; mild metric smoothing (via distillation) stabilizes training. Despite its simplicity, the method attains competitive accuracy down to the extreme W1A1 setting while retaining the efficiency of STE.