🤖 AI Summary
This work addresses the limitation of conventional optimization methods that impose uniform manifold constraints across all Transformer modules, disregarding their distinct geometric preferences in weight space. To remedy this, the authors introduce the Manifold Muon optimizer during GPT-2 pretraining, proposing a module-specific geometric optimization strategy by assigning Stiefel manifold constraints to attention layers and DGram manifold constraints to MLP layers. Experimental results demonstrate that this differentiated configuration substantially enhances training stability and model performance. In contrast, uniform DGram constraints—or inappropriate manifold assignments—tend to induce singular value growth in attention weights, leading to softmax saturation and unstable optimization. The study thus reveals an asymmetric geometric preference among Transformer components and advocates for function-aware customization of optimization geometry.
📝 Abstract
Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries. We study Manifold Muon for GPT-2 pretraining and compare layer-wise assignments of Stiefel and DGram constraints across attention and MLP blocks. Our results show a clear asymmetry: constraining attention layers with Stiefel geometry while assigning DGram geometry to MLP layers gives the best performance among the tested configurations, whereas the inverted assignment and all-DGram configuration become unstable under the shared hyperparameter setting. We trace this failure to singular value growth in DGram-constrained attention weights, which can amplify attention logits and induce softmax saturation. These findings suggest that symmetry-aware and geometry-aware optimization for transformers should be module-specific rather than uniform.