π€ AI Summary
Efficiently training high-energy-efficiency analog resistive networks for machine learning under physical hardware locality constraints remains challenging. This work proposes an analytical gradient computation framework grounded in graph theory and Kirchhoffβs laws, establishing a unified generalized equilibrium propagation model that encompasses both equilibrium propagation and coupled learning. For the first time, this approach enables exact gradient-based training without requiring duplicate network copies. The method achieves localized weight updates using only output-layer information and supports selective tuning of a subset of resistors with minimal performance degradation. Numerical simulations confirm its convergence and effectiveness, offering a novel pathway toward hardware-friendly, brain-inspired computing architectures.
π Abstract
Machine learning is a powerful method of extracting meaning from data; unfortunately, current digital hardware is extremely energy-intensive. There is interest in an alternative analog computing implementation that could match the performance of traditional machine learning while being significantly more energy-efficient. However, it remains unclear how to train such analog computing systems while adhering to locality constraints imposed by the physical (as opposed to digital) nature of these systems. Local learning algorithms such as Equilibrium Propagation and Coupled Learning have been proposed to address this issue. In this paper, we develop an algorithm to exactly calculate gradients using a graph theoretic and analytical framework for Kirchhoff's laws. We also introduce Generalized Equilibrium Propagation, a framework encompassing a broad class of Hebbian learning algorithms, including Coupled Learning and Equilibrium Propagation, and show how our algorithm compares. We demonstrate our algorithm using numerical simulations and show that we can train resistor networks without the need for a replica or readout over all resistors, only at the output layer. We also show that under the analytical gradient approach, it is possible to update only a subset of the resistance values without a strong degradation in performance.