🤖 AI Summary
This study addresses the instability of fully parameterized variational Monte Carlo (VMC) training and the unclear stability mechanisms underlying parameter-efficient approaches. By leveraging stochastic subspace projection and natural gradient optimization, this work quantifies the intrinsic dimensionality of the weight space in neural-network quantum wave functions to investigate the origins of stability in low-dimensional training. It reveals that small-dimensional subspaces confer essential stability rather than mere compression effects. Furthermore, the intrinsic dimension is found to increase at quantum phase transitions with a well-defined lower bound, effectively tracking the learning difficulty of quantum states, while sign representations are shown to critically influence dimensionality. Ultimately, this project achieves divergence-free, stable training, offering new perspectives for understanding the optimization dynamics of quantum states.
📝 Abstract
How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as LoRA. We measure it for variational Monte Carlo (VMC), which trains a neural network to represent the ground state of a quantum many-body system. VMC is a demanding test, because the network generates its own training samples and every gradient is noisy, and a revealing one, because the exact answer is known and every run can be scored. We train only a small latent vector that a frozen random map turns into the network's weights, with no change to the standard natural-gradient optimizer. We find that a small dimension can be misleading, while the stability it brings is real. On a magnet with a hard sign pattern, a network that cannot represent signs reaches its best energy in 8 of 28,642 directions, but only because no such network can go lower; once signs are learnable, neither the signs nor the magnitudes are cheap. The dimension rises across a quantum phase transition, so it tracks how difficult a state is at far less compute than fitting a scaling law, yet it never falls below a floor set by the random subspace itself, even where the ground state is nearly trivial. Training in the subspace, in contrast, never diverged in our experiments, whereas full-parameter training with the same settings did, and a control with matched solvers attributes the difference to the reduced dimension.