🤖 AI Summary
This study addresses the susceptibility of Transformer-based models to overfitting, limited extrapolation capabilities, and insufficient noise robustness in symbolic regression. To overcome these challenges, we propose an optimization strategy based on training set reshaping and noise-aware fine-tuning. Specifically, the method reconstructs the distribution of training formulas to enhance extrapolative generalization, while incorporating fine-tuning on noisy data to improve robustness against perturbations, thereby enabling rapid synthesis of physical equations. Experimental results demonstrate that the proposed approach generates candidate formulas within approximately ten seconds on standard benchmarks. Furthermore, it achieves significantly superior extrapolation accuracy compared to conventional search-based algorithms while maintaining higher computational efficiency. This work provides an efficient and reliable solution for deep learning-based symbolic regression.
📝 Abstract
Finding a compact formula that fits a set of input-output pairs and predicts outputs on unseen inputs is a fundamental problem in science. Symbolic regression automates the search for such formulae: search-based methods explore the space of possible formulae directly, while transformers pre-trained on synthetic data produce formulae of comparable quality substantially faster. Existing transformers, however, are prone to overfitting --- they find formulae that fit the training data well but do not extrapolate to input ranges unseen during training. We address this by shaping the set of formulae used to train a transformer, and show that the resulting formulae extrapolate substantially better. Fine-tuning the transformer on data with noise-corrupted target values further makes the synthesized formulae robust to noise in the observations. On SRBench and LLM-SRBench our transformer synthesizes a formula in about ten seconds and extrapolates better than all evaluated methods at a comparable budget. Search-based methods surpass our accuracy only when given one to three orders of magnitude more time.