Early Memory Selection for Balanced Adam

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of selecting momentum parameters for the Adam optimizer and the prohibitive cost of hyperparameter search. We propose a covariance-preserving cubic memory rule that adaptively determines a fixed β value for full-scale training. By leveraging short pilot runs and gradient probe estimation, the method jointly models numerator and denominator statistics through local directional modeling, effectively balancing sampling variance against estimation delay. Experimental results on vision-language tasks demonstrate that our approach reduces the average validation gap by 40.7%, outperforming both grid search and the best constant settings. Ultimately, this work achieves highly efficient hyperparameter adaptation at minimal computational cost.
📝 Abstract
We propose a method for choosing the shared memory parameter $β_1=β_2=β$ in Adam from a short pilot training. The selected $β$ remains fixed during the subsequent full training. A local model of Adam's normalized direction balances sampling variability against the delay introduced by averaging past gradients. This balance gives a cubic memory rule, whose two coefficients are estimated from gradient probes at a few pilot checkpoints. The estimator uses the numerator and denominator jointly, preserving their covariance. With a 200-update pilot and sixteen probe gradients at each of four checkpoints, a seed-matched retrospective evaluation on eleven vision and language workloads reduces mean relative validation gap by 40.7% and worst-quarter mean gap by 44.3% against the grid representative of shared $β=0.95$. The mean gap is also 32.3% lower than that of the best constant $β$ chosen across all eleven workloads.
Problem

Research questions and friction points this paper is trying to address.

Adam optimizer
momentum parameter selection
shared memory parameter
gradient averaging delay
sampling variability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adam optimizer
memory parameter selection
pilot training
cubic memory rule
gradient probing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alberto FernĂĄndez-HernĂĄndez
Universitat Politècnica de València, Valencia, Spain
C
Cristian PĂŠrez-Corral
Universitat Politècnica de València, Valencia, Spain
J
Jose I. Mestre
Universitat Politècnica de València, Valencia, Spain
Manuel F. Dolz
Manuel F. Dolz
Universitat Jaume I
High Performance ComputingEnergy EfficiencyParallel Programming ModelsPerformance AnalysisDeep Learning
Enrique S. Quintana-OrtĂ­
Enrique S. Quintana-OrtĂ­
Universitat Politècnica de València, Spain