🤖 AI Summary
To address low time and energy efficiency in deep neural network training, this paper proposes “cyclic precision scheduling”—a novel paradigm that treats numerical precision as a dynamically optimizable dimension, analogous to learning rate scheduling, by periodically alternating between high and low precision to jointly optimize training. Key contributions include: (i) the first theoretical and empirical demonstration that cyclic precision promotes convergence to wider minima and reduces gradient variance; (ii) an automatic boundary-precision identification mechanism; and (iii) a unified low-precision training framework supporting both floating-point and integer quantization. Experiments across five datasets and eleven models—including image classification and language modeling tasks—show up to 1.8× speedup, 32% energy reduction, no loss in convergence accuracy, an average 0.42% decrease in generalization error, and significantly improved training stability.
📝 Abstract
Low-precision deep neural network (DNN) training has gained tremendous attention as reducing precision is one of the most effective knobs for boosting DNNs' training time/energy efficiency. In this paper, we attempt to explore low-precision training from a new perspective as inspired by recent findings in understanding DNN training: we conjecture that DNNs' precision might have a similar effect as the learning rate during DNN training, and advocate dynamic precision along the training trajectory for further boosting the time/energy efficiency of DNN training. Specifically, we propose Cyclic Precision Training (CPT) to cyclically vary the precision between two boundary values which can be identified using a simple precision range test within the first few training epochs. Extensive simulations and ablation studies on five datasets and eleven models demonstrate that CPT's effectiveness is consistent across various models/tasks (including classification and language modeling). Furthermore, through experiments and visualization we show that CPT helps to (1) converge to a wider minima with a lower generalization error and (2) reduce training variance which we believe opens up a new design knob for simultaneously improving the optimization and efficiency of DNN training. Our codes are available at: this https URL