🤖 AI Summary
This study addresses the challenge that the Muon optimizer relies on dense matrix operations, rendering it difficult to implement in physical neural networks. To overcome this limitation, this work proposes Physical Muon, a method that for the first time replaces Newton-Schulz iterations with physically realizable flow balance. It reformulates orthogonalization as the equilibrium state of a continuous-time dynamical system and employs stochastic probes to perform approximate computation via matrix-vector products and local writes. Experimental results demonstrate that the proposed approach achieves extremely low cross-entropy error on a 10.95M-parameter Transformer. Furthermore, circuit simulations successfully reproduce the training dynamics. Overall, this method effectively reconciles model training performance with underlying hardware compatibility.
📝 Abstract
Physical neural networks and analog in-memory computing could reduce the energy cost of neural network training. Realizing this potential, however, requires optimizers that combine effective learning with physical implementability. SGD fits local analog updates but struggles on transformers, while Adam family is unstable against analog bias. Muon offers strong training performance, but its Newton--Schulz orthogonalization relies on dense matrix-matrix products. To address this obstacle, we introduce Physical Muon, which computes the orthogonalization as the equilibrium of a continuous-time flow. Random probes approximate the flow using matrix-vector products, reciprocal reads, and local rank-1 writes. To test whether this replacement preserves training performance, we evaluate it on a 10.95M-parameter transformer. The dense flow's mean validation cross-entropy is 0.0085 above Newton--Schulz across nine seeds per method; the probe implementation is 0.0188 above the control across two seeds. Circuit simulations further reproduce the flow dynamics and yield comparable training behavior.