Cost-free Spectral Estimation for Adaptive Newton--Schulz in Matrix Optimizers

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the precision-efficiency imbalance in existing matrix optimizers, where Newton-Schulz orthogonalization relies on fixed polynomials that disregard variations in singular value spectra. We propose an adaptive orthogonalization method that leverages cost-free scalar reductions of the intra-iteration Gram matrix to estimate spectral distributions. This enables dynamic selection of optimal polynomials without additional matrix multiplications, shifting from conservative worst-case designs to real-time data-driven adaptive computation. During GPT pretraining, our approach significantly reduces validation loss, requiring fewer iterations at equivalent precision or yielding superior orthogonalization quality under identical computational budgets.
📝 Abstract
Matrix optimizers such as Muon transform each momentum matrix through an approximate orthogonalization, typically implemented by a small number of Newton-Schulz matrix multiplications. The quality and cost of this approximation depend strongly on the singular-value spectrum of its input, yet existing implementations use the same fixed polynomial routine for every layer and throughout training. We show that this uniform treatment is unnecessary: the computations in the Newton-Schulz method already reveal enough information to make the method adaptive. The Gram matrices formed inside Newton-Schulz iterations yield spectral moments through inexpensive scalar reductions, requiring no additional matrix multiplications. From these moments, we recover an estimate of the empirical singular-value distribution and use it to select a polynomial routine specialized to the current matrix. This turns Newton--Schulz orthogonalization into a spectrum-adaptive procedure that responds to differences across both layers and training time. On saved momentum matrices, spectral estimation substantially reduces orthogonalization error at a fixed iteration budget or reaches the same accuracy with fewer iterations, and in GPT pretraining up to 1B parameters it lowers the validation loss of two matrix optimizers. Our results suggest that matrix-function operations inside optimizers need not be designed for a conservative worst-case spectrum: they can cheaply measure the spectrum they are already processing and specialize computations accordingly.
Problem

Research questions and friction points this paper is trying to address.

matrix optimizer
Newton-Schulz iteration
spectral estimation
approximate orthogonalization
singular-value spectrum
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectral Estimation
Newton-Schulz Iteration
Matrix Optimizer
Adaptive Orthogonalization
Singular-Value Spectrum
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kristi Topollai
New York University
Anna Choromanska
Anna Choromanska
New York University
machine learning