Score
Designs and implements training algorithms for energy-based models that estimate parameter gradients by driving the system to equilibria and applying small perturbations (e.g., two-phase relaxation, centered equilibrium propagation, finite-difference perturbed equilibria) to produce local weight updates. Builds and analyzes the inference/relaxation dynamics, gradient-estimation procedures, and scaling or measurement-integration strategies needed to make equilibrium-propagation training converge and work in practice for larger models.
Backpropagation through time (BPTT) suffers from high computational overhead and struggles with continuous-time trajectories—especially in dynamical systems subject to periodic boundary conditions or fixed endpoint constraints. Method: This paper introduces a Lagrangian action-based equilibrium propagation framework, leveraging the principle of extremal action. Gradients are estimated directly by perturbing system trajectories and analyzing the steady-state response of conjugate variables to parameter perturbations—bypassing explicit backward-in-time computation. Contributions/Results: First, equilibrium propagation is extended to continuous-time dynamical trajectories. Second, its semiclassical limit under periodic boundaries is derived, unifying quantum and classical variational learning perspectives. Third, the framework accommodates both fixed initial/final states and dissipative dynamics. Experiments demonstrate substantial improvements over BPTT in gradient accuracy, training stability, and computational efficiency—establishing a new paradigm for physics-informed temporal modeling.
This work addresses the high energy consumption of conventional GPU-based deep neural network training and the limitations of existing equilibrium propagation methods, which often suffer from slow convergence and susceptibility to local minima. Inspired by Ising machine dynamics, the authors propose a novel training paradigm that replaces the traditional dissipative Hopfield relaxation process with extended-phase-space dynamics incorporating conjugate variables. This approach preserves local two-phase learning rules while altering the physical trajectory by which neuronal states approach equilibrium. The method effectively lowers energy barriers, accelerates convergence, and enhances noise robustness. Experimental results on MNIST, Fashion-MNIST, and CIFAR-10 demonstrate performance comparable to backpropagation, with significantly improved training efficiency and stability.
This work extends equilibrium propagation from conservative systems to non-conservative systems with non-reciprocal interactions, enabling its application to a broader class of neural architectures such as feedforward networks while preserving exact gradient computation. The key innovation lies in introducing a dynamical correction during the learning phase that scales with the non-reciprocal coupling terms, allowing inference and learning to be carried out through steady-state synchronization. By formulating a variational principle in an augmented state space, the authors rigorously derive the exact gradient for equilibrium propagation in non-conservative settings—overcoming the traditional reliance on energy-based dynamics. Experiments on MNIST demonstrate that the proposed method achieves faster convergence and superior performance compared to prior approaches.
Conventional energy-based models (EBMs) and equilibrium propagation (EP) lack principled frameworks for learning dynamic trajectories under time-varying inputs, and existing variants fail to satisfy hardware-friendly constraints—namely, forward-only computation, constant iteration count, and local implementability—while preserving variational consistency across transient dynamics. Method: We generalize EP to dynamic EBMs via the Generalized Lagrangian Equilibrium Propagation (GLEP) framework, grounded in a trajectory-level variational principle that extends the generalized Lagrangian formalism to the full system evolution—not just steady states. We rigorously analyze boundary-condition effects on gradient estimation and derive necessary and sufficient conditions for hardware compatibility. Results: We prove that Hamiltonian Echo Learning (HEL) is the unique GLEP instance satisfying all three hardware constraints. Furthermore, we establish the first formal equivalence between GLEP and HEL, yielding a biologically plausible yet engineering-practical dynamic EP learning paradigm for spiking and analog neuromorphic hardware.
This study investigates the training dynamics of coupled learning (CL) and equilibrium propagation (EP) in the continuous-time, small-perturbation limit, revealing a parameter conservation law in physically realizable systems that parallels mass conservation. By leveraging continuous-time dynamical systems analysis and perturbation theory, the work establishes— for the first time—that this conservation law holds universally across a broad range of physical settings. Furthermore, it elucidates how this constraint governs the convergence behavior of learning in linear circuits. The findings not only enhance the reliability of CL and EP training but also provide a rigorous theoretical foundation and practical guidance for achieving efficient and stable learning in neuromorphic hardware implementations.
This work addresses the instability and limited reliability of energy-based models (EBMs) in scientific data generation, which often stem from slow mixing in Markov chain Monte Carlo sampling. The authors propose a novel training algorithm based on Parallel Tempering with Trajectories (PTT), introducing PTT into the EBM framework for the first time. By leveraging the continuity of optimization paths, the method enables equilibrium sampling throughout training without additional computational overhead, facilitates accurate estimation of burn-in times, yields high-quality equilibrium samples, and supports exact log-likelihood computation. Integrated with reservoir sampling, adaptive optimization, and persistent contrastive divergence, the approach significantly outperforms existing deep generative models on discrete tabular data, demonstrating superior robustness, stability, and sample quality—particularly in small-sample and multimodal settings—while effectively mitigating overfitting.
Efficiently training high-energy-efficiency analog resistive networks for machine learning under physical hardware locality constraints remains challenging. This work proposes an analytical gradient computation framework grounded in graph theory and Kirchhoff’s laws, establishing a unified generalized equilibrium propagation model that encompasses both equilibrium propagation and coupled learning. For the first time, this approach enables exact gradient-based training without requiring duplicate network copies. The method achieves localized weight updates using only output-layer information and supports selective tuning of a subset of resistors with minimal performance degradation. Numerical simulations confirm its convergence and effectiveness, offering a novel pathway toward hardware-friendly, brain-inspired computing architectures.
This work addresses the challenge of efficiently training diffusion models on analog hardware by proposing a training method that requires no external digital accelerators while preserving a low-rank coupling structure. By integrating Symmetric Equilibrium Propagation into a bilinear energy-based model, the authors devise a local learning rule that relies solely on readout operations, reducing gradient bias from first- to second-order within a limited relaxation time. This substantially improves alignment between the update direction and the ideal gradient, enhancing training stability. Coupled with a physically grounded model of unit energy consumption, the approach achieves 10³–10⁴ times lower energy per training step compared to an equivalent GPU baseline, all while maintaining competitive model performance.
Training energy-based models is prone to challenges arising from non-convexity, including sensitivity to initialization, convergence to spurious local minima, and unstable gradients. This work addresses these issues by analyzing learning dynamics through the lens of effective models, combining generalized Ising formulations with Fourier expansions of the energy function and leveraging gradient flow theory. The analysis reveals two types of fixed points—data-consistent and spurious—and uncovers a hierarchical learning mechanism wherein the model preferentially captures low-order interactions first. Introducing the notion of “effective convexity,” the study explains the implicit simplicity bias in learned distributions: perturbations near data-consistent fixed points are either stable or neutral, with neutral directions preserving the effective model structure. This theoretical framework elucidates why low-order inconsistent fixed points are rarely observed in practice and provides a mechanistic understanding of training stability.
This work addresses the disconnection between training and sampling in static scalar energy-based generative models, the absence of a unified theoretical framework, and the lack of convergence guarantees for deterministic gradient flows. The authors unify these aspects by formulating the problem as density transport in Wasserstein space, constructing a nonlinear control system with the KL divergence serving as a Lyapunov function, where training and sampling differ only in their initial conditions. Key contributions include the first Lyapunov-stability-theory-based unified generative framework, a stopping criterion for finite-step Langevin sampling, and a proof that energy superposition preserves the Gibbs invariant measure while inheriting the Lyapunov certificate. Experiments corroborate theoretical predictions, validating the stopping criterion, the invariant measure property, and the failure of deterministic gradient flows to satisfy Lyapunov convergence conditions.