🤖 AI Summary
This study addresses the high memory overhead and coordinate-fixed limitations of conventional optimizers that store momentum histories in parameter space. We propose BOM, a method that shifts optimizer states from parameter space to task space by maintaining a moving average of prediction errors at the output layer and dynamically reprojecting them through the current network in lieu of parameter-level gradient histories. This mechanism enables backpropagated output momentum and serves as a plug-in for adaptive optimizers, yielding compact state storage. Experimental results demonstrate that BOM reduces optimizer state memory by 49.7%–99.8%, accelerates training by 4.0%, and improves model performance by 1.42 points.
📝 Abstract
Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.