š¤ AI Summary
Neural quantum states are prone to sampling bias under finite-sample stochastic optimization, which suppresses critical low-probability configurations and leads to metastable traps that hinder convergence to the true ground state. This work proposes Annealed Gradient Descent (AGD), a method that dynamically rebalances gradient contributions from high- and low-probability configurations to alleviate the restriction of the effective variational subspace and amplify feedback signals from physically relevant configurations. AGD establishes, for the first time, a connection between support loss due to finite sampling and subspace optimization limitations, achieving markedly improved optimization stability through a lightweight mechanism. Demonstrated on molecular systems and one- and two-dimensional JāāJā models, AGD attains chemical accuracy, effectively avoids metastable traps, and enables compact neural networks to reach state-of-the-art performance.
š Abstract
Neural quantum states offer expressive representations of quantum many-body wave functions, yet their practical accuracy can be limited by stochastic optimization rather than representational capacity. Here we identify a finite-sample instability, termed subspace trapping, in which physically important configurations become strongly underestimated, remain absent from successive sampling batches and receive insufficient gradient feedback. This self-reinforcing loss of sampled support can confine optimization to an effective subspace and produce apparently stationary states above the true ground state energy. To address this problem, we introduce annealed gradient descent (AGD), a sampling-aware update with annealing factor that temporarily increases the relative contribution of sampled low-probability configurations while limiting the dominance of high-probability ones. We establish the connection between finite-sample support loss and effective subspace optimization, and then evaluate the method across molecular systems, one and two-dimensional $J_1$-$J_2$ models. Annealed gradient descent suppresses metastable trapping, preserves physically relevant configurations and enables compact neural quantum states to attain chemical accuracy and competitive state-of-the-art performance. These results establish AGD as a lightweight complement to expressive neural architectures, improved sampling strategies for scalable quantum many-body optimization.