Scalable Physics-Inspired Transformers for Spin Glasses

📅 Unknown Date
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing variational methods in modeling large-scale spin glass systems, which suffer from restricted performance and prohibitive computational costs. To overcome these challenges, the authors propose a Transformer architecture infused with physical priors, featuring an interpretable sparse attention mechanism, spin-specific positional encodings, and efficient parallel ancestral sampling via FlashAttention. This approach enables, for the first time, full probabilistic modeling of the Boltzmann distribution of large spin glasses on a single GPU. Evaluated on Sherrington–Kirkpatrick and two- and three-dimensional Edwards–Anderson models, the method achieves up to two orders of magnitude speedup over current techniques while accurately reproducing key thermodynamic quantities—such as free energy and overlap statistics—thereby overcoming a longstanding bottleneck in finite-temperature simulations using machine learning.
📝 Abstract
Efficient sampling of the Boltzmann distribution in frustrated spin glasses is central to statistical mechanics and combinatorial optimization. Despite advances in machine-learning-based approaches, two issues persist: limited understanding of why variational models fail to benefit from increased scale, unlike the monotonic scaling law of large language models; and high computational cost on large systems that negates advantages over classical sampling methods. Here, we develop a physics-inspired transformer with interpretable sparse attention and spin-tailored positional embeddings to address these challenges. By further leveraging FlashAttention for parallel ancestral sampling, it achieves up to two orders of magnitude speedup over vanilla variational autoregressive networks, enabling neural-network simulations of spin-glass systems to unprecedented sizes on a single GPU. It can resolve full probability distributions, free energies, and overlap statistics across temperatures, for Sherrington-Kirkpatrick and 2D or 3D Edwards-Anderson models, where existing machine-learning methods encounter limitations at certain temperatures. This framework thus establishes a scalable paradigm for frustrated spin-glass systems.
Problem

Research questions and friction points this paper is trying to address.

spin glasses
Boltzmann distribution
variational models
scaling law
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physics-Inspired Transformers
Sparse Attention
Spin-Glass Sampling
FlashAttention
Scalable Variational Autoregressive Networks
🔎 Similar Papers
2024-08-05arXiv.orgCitations: 0
2023-12-17Bulletin of the American Mathematical SocietyCitations: 59