Score
Design, train, and analyze models that define an energy (compatibility) function over possible configurations and use iterative optimization or learned update rules to infer low‑energy target states. Implement and evaluate inference procedures such as normalized gradient‑descent or learned iterative updates, and develop methods for learning the energy landscape and update dynamics for structured or coupled variables.
Despite growing interest in Energy-Based Models (EBMs), their theoretical relationships with mainstream generative models—including GANs, VAEs, and normalizing flows—and their formal connections to statistical mechanics (e.g., energy functions, partition functions, MCMC sampling) remain poorly unified and conceptually fragmented. Method: We propose the first cross-paradigm unification framework tailored for physicists, establishing rigorous formal mappings between EBMs and other generative paradigms through an energy-centric lens. Our approach integrates statistical physical modeling, MCMC sampling analysis, EBM optimization theory, and systematic comparative evaluation of generative mechanisms. Contribution/Results: This work bridges the conceptual gap between generative modeling and statistical mechanics, revealing fundamental commonalities and distinctions across models in terms of energy representations, sampling dynamics, and training objectives. It enhances theoretical coherence, interpretability, and principled design of EBMs—while providing a unified foundation for analyzing sampling efficiency, convergence properties, and thermodynamic analogies in deep generative modeling.
This work addresses the challenge of training thermodynamic computing hardware—driven solely by thermal noise—to perform target computations (e.g., image classification) within a fixed observation time. We propose a gradient-descent-based parameter optimization method framed as a teacher–student paradigm: a deterministic teacher network generates ideal neural activation trajectories, while a stochastic student system models tunable thermodynamic hardware (e.g., bistable units with adjustable energy barriers and coupling strengths). Physical parameters are optimized end-to-end via backpropagation through the stochastic dynamics to minimize trajectory divergence. To our knowledge, this is the first approach to apply gradient descent directly for end-to-end training of physical thermodynamic computing substrates. Experiments demonstrate robust classification performance on MNIST, with theoretical energy consumption over seven orders of magnitude lower than conventional digital implementations. The method establishes a new paradigm for ultra-low-power, brain-inspired computing grounded in nonequilibrium thermodynamics.
This work addresses the high energy consumption of conventional GPU-based deep neural network training and the limitations of existing equilibrium propagation methods, which often suffer from slow convergence and susceptibility to local minima. Inspired by Ising machine dynamics, the authors propose a novel training paradigm that replaces the traditional dissipative Hopfield relaxation process with extended-phase-space dynamics incorporating conjugate variables. This approach preserves local two-phase learning rules while altering the physical trajectory by which neuronal states approach equilibrium. The method effectively lowers energy barriers, accelerates convergence, and enhances noise robustness. Experimental results on MNIST, Fashion-MNIST, and CIFAR-10 demonstrate performance comparable to backpropagation, with significantly improved training efficiency and stability.
This work addresses the practical limitations of expensive yet powerful propagators—such as Energetic Reasoning—in constraint programming, which, despite their strong pruning capabilities, incur substantial computational overhead. To mitigate this issue, the paper proposes a hybrid framework that integrates static machine learning with dynamic search heuristics to control propagator activation. Specifically, a supervised learning approach is employed to construct a predictor function that dynamically decides whether to invoke the costly propagator during search. The authors design a set of effective instance-specific features and train a high-accuracy classification model, achieving, for the first time, seamless integration of such a predictive mechanism into a modern constraint solver. Experimental results demonstrate the feasibility of the approach and shed light on key challenges and design principles for building efficient propagator selection strategies.
Standard Bayesian optimization (BO) suffers from a “one-step bias” due to its greedy, single-step acquisition strategy, leading to premature convergence to local optima—particularly detrimental in high-dimensional, complex black-box optimization. To address this, we propose REBMBO, a reinforcement learning–enhanced BO framework that formulates BO as a Markov decision process. REBMBO jointly models local fidelity and global structure by integrating Gaussian processes (GPs) with energy-based models, and employs proximal policy optimization (PPO) to enable adaptive, multi-step lookahead balancing exploration and exploitation. Its key contribution is the first application of reinforcement learning to learn multi-step BO policies, dynamically adjusting both exploration depth and direction. Experiments on synthetic benchmarks and real-world tasks demonstrate that REBMBO consistently outperforms state-of-the-art BO methods, exhibiting strong robustness and generalization across diverse GP kernels.
This work addresses the challenge of efficiently solving NP-hard Ising and Max-Cut problems, where conventional methods struggle due to the complex, non-convex energy landscapes. The authors propose a data-driven iterative dynamical system that parameterizes spin update rules via a shared node-level multilayer perceptron and trains it using zeroth-order optimization to circumvent gradient instability associated with backpropagation. Remarkably, with an extremely low number of parameters, the learned dynamics automatically exhibit momentum-like behavior and time-varying scheduling mechanisms, substantially enhancing search efficiency. Evaluated on standard Ising and combinatorial optimization benchmarks, the method achieves solution quality and convergence speed comparable to state-of-the-art learning-based approaches and classical Ising machine heuristics.
This work addresses the disconnection between training and sampling in static scalar energy-based generative models, the absence of a unified theoretical framework, and the lack of convergence guarantees for deterministic gradient flows. The authors unify these aspects by formulating the problem as density transport in Wasserstein space, constructing a nonlinear control system with the KL divergence serving as a Lyapunov function, where training and sampling differ only in their initial conditions. Key contributions include the first Lyapunov-stability-theory-based unified generative framework, a stopping criterion for finite-step Langevin sampling, and a proof that energy superposition preserves the Gibbs invariant measure while inheriting the Lyapunov certificate. Experiments corroborate theoretical predictions, validating the stopping criterion, the invariant measure property, and the failure of deterministic gradient flows to satisfy Lyapunov convergence conditions.
This work proposes the first compiler framework enabling end-to-end mapping of general stochastic programs to thermodynamic sampling hardware. Addressing the challenge of efficiently compiling stochastic programs—expressed as directed factor graphs or parameterized random circuits—onto native energy-based model (EBM) hardware, the approach integrates context-aware pattern matching with a trajectory-level REINFORCE post-training strategy. This combination substantially reduces compilation error and enhances approximation fidelity. Empirical evaluation demonstrates the framework’s effectiveness and generality across diverse applications, including financial market simulation, ecological probabilistic modeling, Gibbs sampling for non-native EBMs, and Bayesian design of Gaussian random circuits, thereby establishing a viable pathway toward energy-efficient stochastic computing.
Training energy-based models is prone to challenges arising from non-convexity, including sensitivity to initialization, convergence to spurious local minima, and unstable gradients. This work addresses these issues by analyzing learning dynamics through the lens of effective models, combining generalized Ising formulations with Fourier expansions of the energy function and leveraging gradient flow theory. The analysis reveals two types of fixed points—data-consistent and spurious—and uncovers a hierarchical learning mechanism wherein the model preferentially captures low-order interactions first. Introducing the notion of “effective convexity,” the study explains the implicit simplicity bias in learned distributions: perturbations near data-consistent fixed points are either stable or neutral, with neutral directions preserving the effective model structure. This theoretical framework elucidates why low-order inconsistent fixed points are rarely observed in practice and provides a mechanistic understanding of training stability.
This work addresses the challenges of training large-scale energy-based neural networks on Ising machines, which are constrained by limited hardware connectivity and inefficient optimization. To overcome these limitations, the authors propose a novel approach that integrates a coherent Ising machine (CIM) with equilibrium propagation and the Adam optimizer to efficiently find the ground state of Hopfield energy networks. Notably, this is the first method to incorporate the Adam optimizer into CIM-driven energy-based model training, significantly accelerating convergence and improving solution accuracy while enabling support for deep and convolutional architectures. Experimental results demonstrate that, on both simulated and optoelectronic hardware platforms, the proposed framework achieves performance comparable to purely software-based implementations while substantially enhancing the scalability and training efficiency of physical AI systems.
本文探讨了控制理论、最优传输、概率推理、非平衡热力学和机器学习之间的联系,通过优化自由能类函数解决高维数据中的复杂结构学习问题。