Scheduling Weight Transitions for Quantization-Aware Training

๐Ÿ“… 2024-04-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In quantization-aware training (QAT), discrete transitions of quantized weights are difficult to control, and conventional learning rate (LR) scheduling fails to jointly optimize latent weight updates and quantization-level transitions due to their tightly coupled dynamics. Method: We propose Transition Rate (TR) as the core scheduling objective and introduce a Transition-Adaptive Learning Rate (TALR) mechanismโ€”first explicitly modeling the frequency of discrete-level weight transitions as an optimization variable, thereby enabling direct control over state switching. Our approach integrates latent-variable optimization with discrete-state transition modeling, eliminating manual hyperparameter tuning while stably constraining transition intensity. Contribution/Results: On standard benchmarks, TALR significantly improves accuracy of low-bit QAT models at equivalent bit-widths, outperforming mainstream LR scheduling strategies. It enhances training controllability and convergence stability without architectural modifications or additional inference overhead.

Technology Category

Machine Learning: Quantum Machine LearningSearch and Optimization: Learning to SearchPlanning, Routing, and Scheduling: Planning/Scheduling and Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingResponsible Web: Algorithmic accountability and transparency on the web
๐Ÿ“ Abstract
Quantization-aware training (QAT) simulates a quantization process during training to lower bit-precision of weights/activations. It learns quantized weights indirectly by updating latent weights,i.e., full-precision inputs to a quantizer, using gradient-based optimizers. We claim that coupling a user-defined learning rate (LR) with these optimizers is sub-optimal for QAT. Quantized weights transit discrete levels of a quantizer, only if corresponding latent weights pass transition points, where the quantizer changes discrete states. This suggests that the changes of quantized weights are affected by both the LR for latent weights and their distributions. It is thus difficult to control the degree of changes for quantized weights by scheduling the LR manually. We conjecture that the degree of parameter changes in QAT is related to the number of quantized weights transiting discrete levels. Based on this, we introduce a transition rate (TR) scheduling technique that controls the number of transitions of quantized weights explicitly. Instead of scheduling a LR for latent weights, we schedule a target TR of quantized weights, and update the latent weights with a novel transition-adaptive LR (TALR), enabling considering the degree of changes for the quantized weights during QAT. Experimental results demonstrate the effectiveness of our approach on standard benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Sub-optimal learning rate coupling in quantization-aware training.
Difficulty in controlling quantized weight changes manually.
Introducing transition rate scheduling for better weight transition control.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transition rate scheduling for quantization-aware training
Transition-adaptive learning rate updates latent weights
Explicit control of quantized weight transitions
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Yonsei University
J
Junghyup Lee
School of Electrical and Electronic Engineering, Yonsei University, Seoul 03722, Korea
D
Dohyung Kim
School of Electrical and Electronic Engineering, Yonsei University, Seoul 03722, Korea
J
Jeimin Jeon
School of Electrical and Electronic Engineering, Yonsei University, Seoul 03722, Korea
Bumsub Ham
Bumsub Ham
Yonsei University
Computer visionImage processing