Differentiable, Bit-shifting, and Scalable Quantization without training neural network from scratch

📅 2025-10-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing neural network quantization methods face two key bottlenecks: non-differentiability causing gradient distortion, and low quantization accuracy—especially for joint weight-activation quantization—along with poor scalability to multi-bit settings. This paper proposes the first fully differentiable bit-shift-based quantization framework, replacing multiplications with bit-shift operations and enabling exact gradient propagation via a differentiable quantization function, with theoretical convergence guarantees. The method supports arbitrary *n*-bit quantization without requiring de novo training; on ResNet-18/ILSVRC-2012, it achieves state-of-the-art 8-bit performance (accuracy drop <1%) after only 15 fine-tuning epochs. During inference, high-precision multiplications are entirely eliminated, replaced solely by lightweight CPU bit-shift instructions—yielding significant reductions in both computational cost and memory footprint.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Learning & Optimization for CVSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystems
📝 Abstract
Quantization of neural networks provides benefits of inference in less compute and memory requirements. Previous work in quantization lack two important aspects which this work provides. First almost all previous work in quantization used a non-differentiable approach and for learning; the derivative is usually set manually in backpropogation which make the learning ability of algorithm questionable, our approach is not just differentiable, we also provide proof of convergence of our approach to the optimal neural network. Second previous work in shift/logrithmic quantization either have avoided activation quantization along with weight quantization or achieved less accuracy. Learning logrithmic quantize values of form $2^n$ requires the quantization function can scale to more than 1 bit quantization which is another benifit of our quantization that it provides $n$ bits quantization as well. Our approach when tested with image classification task using imagenet dataset, resnet18 and weight quantization only achieves less than 1 percent accuracy compared to full precision accuracy while taking only 15 epochs to train using shift bit quantization and achieves comparable to SOTA approaches accuracy in both weight and activation quantization using shift bit quantization in 15 training epochs with slightly higher(only higher cpu instructions) inference cost compared to 1 bit quantization(without logrithmic quantization) and not requiring any higher precision multiplication.
Problem

Research questions and friction points this paper is trying to address.

Develops differentiable quantization without manual gradient setting in backpropagation
Enables multi-bit logarithmic quantization for both weights and activations
Achieves near-full-precision accuracy with minimal training epochs and efficient inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differentiable quantization without manual derivative setting
Shift bit quantization for both weights and activations
Scalable n-bit quantization with logarithmic form values
💼 Related Jobs
No related jobs found.
Z
Zia Badar
Karachi, Pakistan