🤖 AI Summary
This study addresses the unclear impact of divergence selection on the performance and optimization efficacy of the Shampoo preconditioner. To this end, it constructs a unified theoretical framework based on Bregman divergences, integrating spectral analysis with GPT-2 pretraining experiments to systematically investigate how different divergences influence Kronecker approximations and finite-sample errors. The findings reveal that specific divergences can effectively compensate for the underestimation of second-order moments, offering a novel perspective for understanding Shampoo. Furthermore, this work elucidates the underlying mechanisms driving behavioral differences among various Shampoo variants, providing principled theoretical guidance for the design and refinement of second-order optimizers.
📝 Abstract
Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.