GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks

📅 2025-08-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the capacity degradation and difficulty in dynamically tracking bottlenecks in low-bit quantized neural networks—caused by non-differentiable rounding operations—this paper models quantization as a cascade of noisy channels and proposes a progressively differentiable noise-scaled quantization framework. Methodologically, it jointly optimizes learnable bit-widths, noise scales, and clipping ranges via straight-through estimators (STE) for end-to-end training; incorporates an outlier-penalty term to precisely enforce target bit-widths; and integrates lightweight knowledge distillation to enhance training stability and gradient smoothness. The approach achieves competitive accuracy under the extremely challenging W1A1 setting—surpassing state-of-the-art methods significantly—while maintaining efficient forward inference. Key contributions include: (i) a noise-modeling-driven differentiable quantization paradigm, and (ii) a multi-variable co-constraint mechanism unifying bit-width, noise, and clipping optimization.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsUser Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendation
📝 Abstract
Quantized neural networks can be viewed as a chain of noisy channels, where rounding in each layer reduces capacity as bit-width shrinks; the floating-point (FP) checkpoint sets the maximum input rate. We track capacity dynamics as the average bit-width decreases and identify resulting quantization bottlenecks by casting fine-tuning as a smooth, constrained optimization problem. Our approach employs a fully differentiable Straight-Through Estimator (STE) with learnable bit-width, noise scale and clamp bounds, and enforces a target bit-width via an exterior-point penalty; mild metric smoothing (via distillation) stabilizes training. Despite its simplicity, the method attains competitive accuracy down to the extreme W1A1 setting while retaining the efficiency of STE.
Problem

Research questions and friction points this paper is trying to address.

Quantized neural networks face capacity reduction from rounding noise
Fine-tuning addresses quantization bottlenecks via constrained optimization
Method enables low-bit W1A1 networks while maintaining STE efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differentiable Straight-Through Estimator with learnable parameters
Exterior-point penalty enforces target bit-width constraint
Metric smoothing via distillation stabilizes training process
💼 Related Jobs
No related jobs found.
aifoundry.org | Ainekko Co.
S
Sergey Salishev
aifoundry.org, San Francisco, CA, USA
I
Ian Akhremchik
Ainekko Co., San Francisco, CA, USA