Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of neural networks to fall into memorization states during training, which impedes generalization and manifests as dynamical freezing in the "grokking" phenomenon. Drawing on glassy dynamics theory, the study introduces a tripartite metric framework—comprising parameter transferability, replica correlation, and fractal dimension—to empirically identify glassy characteristics in training dynamics for the first time. Inspired by exchange Monte Carlo methods, the authors propose State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that leverages stochastic exploration in parameter space to effectively disrupt memorized states. Experiments demonstrate that SAM-Swap substantially accelerates generalization and outperforms conventional regularization techniques such as weight decay and Gaussian gradient noise, thereby validating the critical role of physics-inspired parameter exchange mechanisms in overcoming generalization bottlenecks.
📝 Abstract
Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization. While previous works have attempted to interpret it through classical machine learning mechanisms like weight norm, recent research draws an analogy from statistical physics, framing grokking as a form of computational glass relaxation. This theory defines the initial memorization as a result of `fast cooling' where the training loss is reduced so quickly that a glass state is formed, followed by a `slow relaxation' towards final generalization. Although providing a unifying framework for representative grokking theories, this perspective has remained largely at the theoretical on macroscopic level without direct empirical validation on training dynamics. Here we introduce a three-component framework to directly characterize the training dynamics via parameter mobility (PM), and two representative measurements from glassy dynamics: replica correlation (RC) and fractal dimension (FD). We demonstrate that standard optimization presents clear signatures of glass dynamics and inherently traps the grokking network in a kinetic arrested memorization state with a collapsed mobility, strong history dependence, and channel-like motions. This quantitative agreement motivates us to introduce State-Aware Monte Carlo Parameter Swapping (SAM-Swap), an optimization plug-in that can accelerate generalization, inspired by swap Monte Carlo algorithm widely used in glass dynamics. Comparing SAM-Swap, weight decay, and Gaussian gradient noise, we find that accelerated generalization is consistently associated with random exploration in the parameter space, similar to diffusion in physics.
Problem

Research questions and friction points this paper is trying to address.

grokking
loss landscape
glass dynamics
memorization
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grokking
glass dynamics
parameter swapping
Monte Carlo optimization
training dynamics