🤖 AI Summary
This work proposes the Soren optimizer to address the loss of relative magnitude information across gradient modalities caused by spectral flattening in the Muon optimizer. Soren reformulates spectral reshaping as a preconditioned gradient method with established convergence guarantees. By preserving singular subspaces and applying a bounded, monotonic sigmoid transformation to singular values, it smoothly compresses dominant modes. Furthermore, a finite-depth Soft Newton-Schulz polynomial is designed for efficient approximation. Experiments demonstrate that Soren achieves superior effectiveness and robustness over existing optimizers across large language model pretraining, fine-tuning, and direct preference optimization tasks.
📝 Abstract
Matrix-valued optimizers such as Muon exploit the spectral structure of neural network updates through Newton--Schulz orthogonalization, but their near-flattening of the singular spectrum discards relative magnitude information across gradient modes. We introduce \emph{Soren} (\textbf{S}pectral \textbf{O}rthogonal \textbf{Re}shapi\textbf{n}g), a matrix-valued optimizer that preserves the singular subspaces of the gradient while applying a bounded, monotone sigmoid transformation to its singular values. This smoothly compresses dominant modes without fully flattening the spectrum. We interpret Soren as a positive-definite preconditioned gradient method and establish convergence guarantees under relative smoothness and metric Polyak--{\L}ojasiewicz geometry. To avoid explicit singular value decomposition, we further develop a finite-depth Soft Newton--Schulz (SNS) polynomial realization of the sigmoid spectral map and characterize how its spectral approximation affects the induced convergence geometry. Experiments across LLM pre-training, supervised fine-tuning, and direct preference optimization demonstrate the effectiveness and robustness of Soren against established optimizers.