๐ค AI Summary
Existing symbolic regression methods rely heavily on heuristic model selection, regularization, and search strategies, lacking theoretical foundations and performance guarantees.
Method: This paper introduces a novel Bayesian inferenceโbased paradigm for symbolic regression, replacing heuristics with a probabilistic framework where model discovery is formulated as posterior distribution inference. It integrates information-theoretic principles (e.g., Minimum Description Length) and statistical physics concepts (e.g., variational approximation) to naturally balance model complexity and goodness-of-fit. Crucially, it emphasizes model ensembling over selecting a single optimal expression, enabling principled uncertainty quantification and theoretically grounded generalization bounds.
Results: Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art symbolic regression methods in equation discovery accuracy, physical consistency, and robustness across diverse datasets.
๐ Abstract
Symbolic regression automates the process of learning closed-form mathematical models from data. Standard approaches to symbolic regression, as well as newer deep learning approaches, rely on heuristic model selection criteria, heuristic regularization, and heuristic exploration of model space. Here, we discuss the probabilistic approach to symbolic regression, an alternative to such heuristic approaches with direct connections to information theory and statistical physics. We show how the probabilistic approach establishes model plausibility from basic considerations and explicit approximations, and how it provides guarantees of performance that heuristic approaches lack. We also discuss how the probabilistic approach compels us to consider model ensembles, as opposed to single models.