Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether representation learning ceases concurrently with the stagnation of training loss. Employing a teacher-student framework with the matrix Muon optimizer and extreme normalization updates, we analyze representational dynamics during loss plateaus via the average gradient outer product (AGOP). Theoretically, we demonstrate that AGOP during loss plateaus precisely recovers the teacher subspace. Empirically, we observe that under period-2 loss oscillations, weights continue to evolve while features remain aligned, an effect further validated through layer freezing and unfreezing experiments on loss dynamics. This work reveals a decoupling mechanism between representation learning and loss minimization, indicating that models can perform effective feature extraction even when the training loss has stalled.
📝 Abstract
Does representation learning stop when the training loss stops improving? We study this question for matrix Muon, whose polar-normalised updates have a step length set by the gradient's rank rather than its norm. Near the edge of stability, full-batch Muon on teacher-student problems enters approximately period-2 loss oscillations that persist for thousands of steps: the cycle-mean loss stays flat or rises, yet the weights keep moving and the learned features continue to align with the teacher subspace. For linear teacher-student learning toys, we derive explicit cycle and alignment formulas and conditional plateau and decay bounds. For a population mean-field ReLU model, we prove that, under stated dimension, initialisation and small-head conditions, the leading eigenspace of the average gradient outer product (AGOP) recovers the teacher subspace exactly during a loss plateau, before the loss later drops. In all 33 ReLU, GELU and SiLU teacher configurations we study, direction-only alignment metrics show the student AGOP aligned with, or still aligning to, the teacher subspace during the period-2 oscillations; projected head refitting on selected configurations shows that the learned directions are useful for prediction, and further measurements distinguish AGOP alignment from weight-mass concentration. In deep residual ReLU students, freezing the downstream layers while the first layer trains with full-batch exact polar updates recreates a nearly flat cycle-mean loss with improving input-AGOP alignment; freezing and unfreezing switch between this plateau and loss decrease, and the effect is sensitive to momentum and to the choice of orthogonaliser.
Problem

Research questions and friction points this paper is trying to address.

Representation Learning
Loss Minimization
Polar Updates
Loss Oscillation
Feature Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Representation Learning
Polar Updates
Muon Optimizer
Loss Oscillation
Average Gradient Outer Product (AGOP)
🔎 Similar Papers
No similar papers found.