Different Layers, Different Manifolds: Module-Wise Weight-Space Geometry in Transformer Optimization

📅 2026-06-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of conventional optimization methods that impose uniform manifold constraints across all Transformer modules, disregarding their distinct geometric preferences in weight space. To remedy this, the authors introduce the Manifold Muon optimizer during GPT-2 pretraining, proposing a module-specific geometric optimization strategy by assigning Stiefel manifold constraints to attention layers and DGram manifold constraints to MLP layers. Experimental results demonstrate that this differentiated configuration substantially enhances training stability and model performance. In contrast, uniform DGram constraints—or inappropriate manifold assignments—tend to induce singular value growth in attention weights, leading to softmax saturation and unstable optimization. The study thus reveals an asymmetric geometric preference among Transformer components and advocates for function-aware customization of optimization geometry.
📝 Abstract
Weight-space geometry plays a central role in neural network optimization, yet manifold constraints are often applied uniformly across all weight matrices. In this work, we ask whether different transformer modules prefer different manifold geometries. We study Manifold Muon for GPT-2 pretraining and compare layer-wise assignments of Stiefel and DGram constraints across attention and MLP blocks. Our results show a clear asymmetry: constraining attention layers with Stiefel geometry while assigning DGram geometry to MLP layers gives the best performance among the tested configurations, whereas the inverted assignment and all-DGram configuration become unstable under the shared hyperparameter setting. We trace this failure to singular value growth in DGram-constrained attention weights, which can amplify attention logits and induce softmax saturation. These findings suggest that symmetry-aware and geometry-aware optimization for transformers should be module-specific rather than uniform.
Problem

Research questions and friction points this paper is trying to address.

weight-space geometry
manifold constraints
transformer optimization
module-specific
Stiefel and DGram
Innovation

Methods, ideas, or system contributions that make the work stand out.

module-wise geometry
manifold optimization
Stiefel manifold
DGram manifold
transformer optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kirato Yoshihara
School of Engineering Science, The University of Osaka, Osaka, Japan