đ¤ AI Summary
This study addresses the difficulty of selecting momentum parameters for the Adam optimizer and the prohibitive cost of hyperparameter search. We propose a covariance-preserving cubic memory rule that adaptively determines a fixed β value for full-scale training. By leveraging short pilot runs and gradient probe estimation, the method jointly models numerator and denominator statistics through local directional modeling, effectively balancing sampling variance against estimation delay. Experimental results on vision-language tasks demonstrate that our approach reduces the average validation gap by 40.7%, outperforming both grid search and the best constant settings. Ultimately, this work achieves highly efficient hyperparameter adaptation at minimal computational cost.
đ Abstract
We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training. The selected $β$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$. The mean gap is also 32.3% lower than that of the best constant $β$ chosen across all eleven workloads.