Spectral Edge Dynamics Reveal Functional Modes of Learning

📅 2026-04-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional interpretability methods in capturing the key dynamic directions governing grokking. The authors propose a novel approach that analyzes the spectral margin of training dynamics to uncover low-dimensional functional patterns emerging in the input domain, whose structure is dictated by the algebraic symmetries of the task. This framework transcends existing paradigms rooted in local structures of parameter or feature spaces. By integrating spectral analysis, Fourier and discrete logarithm basis transformations, multi-task training, and variance concentration metrics—and cross-validating findings via head attribution and sparse autoencoders—the method achieves up to a 5.9× improvement in directional concentration across tasks such as modular addition, multiplication, subtraction, and $x^2 + y^2$. Notably, it reveals for the first time that functional patterns in multi-task settings exhibit compositional inheritance.

Technology Category

Machine Learning: Multimodal LearningNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsSearch and Optimization: Learning to Search

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalization
📝 Abstract
Training dynamics during grokking concentrate along a small number of dominant update directions -- the spectral edge -- which reliably distinguishes grokking from non-grokking regimes. We show that standard mechanistic interpretability tools (head attribution, activation probing, sparse autoencoders) fail to capture these directions: their structure is not localized in parameter or feature space. Instead, each direction induces a structured function over the input domain, revealing low-dimensional functional modes invisible to representation-level analysis. For modular addition, all leading directions collapse to a single Fourier mode. For multiplication, the same collapse appears only in the discrete-log basis, yielding a 5.9x improvement in concentration. For subtraction, the edge spans a small multi-mode family. For $x^2+y^2$, no single harmonic basis suffices, but cross-terms of additive and multiplicative features provide a 4x variance boost, consistent with the decomposition (a+b)^2 - 2ab. Multitask training amplifies this compositional structure, with the $x^2+y^2$ spectral edge inheriting the addition circuit's characteristic frequency (2.3x concentration increase). These results suggest that training discovers low-dimensional functional modes over the input domain, whose structure depends on the algebraic symmetry of the task. These results suggest that spectral edge dynamics identify low-dimensional functional subspaces governing learning, whose representation depends on the algebraic structure of the task. Simple harmonic structure emerges only when the task admits a symmetry-adapted basis; more complex tasks require richer functional descriptions.
Problem

Research questions and friction points this paper is trying to address.

grokking
spectral edge
functional modes
mechanistic interpretability
algebraic symmetry
Innovation

Methods, ideas, or system contributions that make the work stand out.

spectral edge dynamics
functional modes
grokking
algebraic symmetry
harmonic basis
🔎 Similar Papers
No similar papers found.