How Bregman Divergences Shape Shampoo

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear impact of divergence selection on the performance and optimization efficacy of the Shampoo preconditioner. To this end, it constructs a unified theoretical framework based on Bregman divergences, integrating spectral analysis with GPT-2 pretraining experiments to systematically investigate how different divergences influence Kronecker approximations and finite-sample errors. The findings reveal that specific divergences can effectively compensate for the underestimation of second-order moments, offering a novel perspective for understanding Shampoo. Furthermore, this work elucidates the underlying mechanisms driving behavioral differences among various Shampoo variants, providing principled theoretical guidance for the design and refinement of second-order optimizers.
📝 Abstract
Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
Problem

Research questions and friction points this paper is trying to address.

Shampoo optimizer
Bregman divergence
preconditioning
neural network optimization
Kronecker approximation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bregman divergence
Shampoo optimizer
preconditioning
Kronecker approximation
finite-sample error
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
B
Bing Liu
Zhejiang University
W
Wenjie Zhou
University of the Chinese Academy of Sciences
C
Chengcheng Zhao
Zhejiang University
Hongtao Zhang
Hongtao Zhang
Professor, Hong Kong University of Science and Technology
Operations Management
B
Boao Kong
Peking University
Felix Dangel
Felix Dangel
Postdoc at the Vector Institute, Toronto
Second-order optimizationautomatic differentiationdeep neural networkstensor networks
W
Wu Lin
University of Central Florida