🤖 AI Summary
This study addresses the hardware deployment challenges arising from the high arithmetic complexity of optimizers such as AdaGrad and Adam by proposing an arithmetic simplification method. Specifically, these optimizers are reformulated into elementary forms comprising only addition, subtraction, multiplication, and division by powers of two. A static step-size function S(T) is further introduced to replace all division operations with pure integer bit-shift computations. Grounded in stochastic gradient descent theory and recursive relation transformations, this work successfully constructs an arithmetically simplified static Adam variant. Rigorous theoretical proofs are provided to guarantee convergence under these simplifications. Ultimately, the proposed approach offers an efficient solution for the low-precision hardware implementation of adaptive optimizers without compromising their convergence properties.
📝 Abstract
A gradient descent method is arithmetically simple if the operations are limited to $+,-, \times$, and division $x/2^t$ with integer $t$. An arthmetically simple gradient method is easy to implement in chip design. We show how to transform AdaGrad, Adam, and AdamW into arithmetically simple. AdamW is based on the recursion $x_{t+1}=(1-\lambda\eta)x_t-\frac{\eta }{s}m_t$ and Adam is the special case of AdamW with $\lambda=0$. We transform them into a static case with $s=S(T)$, where $T$ is the number of iterations, and $S(T)$ is a fixed function. The convergence analysis is given for a static Adam, which is also arithmetically simple.