π€ AI Summary
This work introduces, for the first time, fundamental thermodynamic limits into basic machine learning algorithms by leveraging Landauerβs principle and information thermodynamics to quantify the minimum irreversible energy dissipation incurred by floating-point implementations of simple linear regression. By constructing an entropy production model for continuous inputs, the study analyzes the thermodynamic costs associated with both exact solutions and stochastic gradient descent. Furthermore, it derives the optimal scaling law between training set size and energy consumption under a prescribed generalization error constraint. This paper establishes the first theoretical framework characterizing the trade-off between energy efficiency and generalization in linear regression, thereby providing a physical foundation for the design of energy-aware machine learning systems.
π Abstract
The construction of models from data is a significant contributor to the energetic costs of computation. Because of this, understanding how foundational thermodynamic bounds apply to modeling algorithms will be increasingly important. Here, we study the thermodynamic costs of a basic and fundamental modeling algorithm: simple linear regression. Following Landauer, we approximate the thermodynamic lower bound on irreversibly performing both exact linear regression and linear regression via stochastic gradient descent as implemented on floating-point numbers. From this, we derive energycost aware scaling laws for the optimal dataset size for training a linear regression model given a generalization error dependent demand for inference. Additionally, we discuss a method to lower bound the entropy production from the mismatch cost for algorithms with continuous input variables.