🤖 AI Summary
This study addresses the limitation of aggregated fitness scores in symbolic regression, which fail to capture error distributions across inputs, by proposing a residual-aware framework to guide formula discovery. Methodologically, a residual encoder extracts error patterns to condition a large language model for generating candidate formulas. Additionally, a dual-view relational encoder is introduced to predict the post-fitting utility of correction terms, compensating for deficiencies in conventional evaluation metrics. Experimental results demonstrate that the proposed method achieves 63.57% and 38.50% in-distribution accuracy on the LLM-SRBench benchmark, significantly outperforming existing baselines and effectively enhancing numerical equation recovery performance.
📝 Abstract
Symbolic regression combines structural search with numerical fitting, but aggregate fit scores do not describe how the remaining error varies across inputs. We introduce RISR, a residual-informed method that uses these error patterns to guide formula discovery and learn which corrections are worth fitting. A residual encoder compresses aligned inputs, targets, current predictions, and residuals into continuous tokens that condition a language model to propose formulas. For subsequent refinement, a dual-view relational encoder uses additive and regularized multiplicative residuals to predict the post-fit utility of candidate corrections. We evaluate RISR on scientific tasks from the LLM-SRBench. RISR achieves 63.57% and 38.50% ID accuracy at the 1% and 0.1% pointwise relative-error tolerances, respectively. The corresponding OOD accuracies are 56.07% and 38.24%. RISR outperforms the reported baselines using the same backbone. The results show that our residual-informed approach can improve numerical equation recovery.