🤖 AI Summary
This study addresses the difficulty of large language models (LLMs) in deriving precise scientific laws from observational data without relying on external tools. To this end, it pioneers the explicit learning of symbolic regression as an intrinsic capability of LLMs. Methodologically, the model is trained via joint numerical-symbolic supervision and physics-informed constraints. Furthermore, a large-scale physics-aware benchmark, PhysSymbArena, is constructed, alongside SymbolicSGA, an iterative reasoning optimization framework driven by quantitative feedback. Experimental results demonstrate that the proposed approach significantly improves structural recovery rates across multiple benchmarks while maintaining competitive numerical fitting performance. These findings validate the feasibility of enabling LLMs to directly master symbolic regression, offering a promising paradigm for automated scientific discovery without dependence on external solvers.
📝 Abstract
Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.