Score
Designs and builds algorithms and pipelines that discover symbolic expressions or models by combining population- or evolution‑style structure search with gradient‑based optimization of continuous parameters; implements differentiable representations of symbolic forms to enable backpropagation through coefficients and hybridizes proposal generation, scoring, and accumulated search experience to refine both expression structure and numeric parameters.
SMILE框架通过结合连续优化和离散符号恢复解决符号回归中的慢收敛问题,利用数据结构分析、参数学习及符号化简化过程提高模型的准确性和鲁棒性。
Symbolic regression faces a fundamental trade-off between model interpretability and predictive accuracy. Method: This paper proposes a bi-objective optimization framework that synergistically integrates gradient descent and evolutionary computation to simultaneously minimize structural error (ensuring syntactic correctness of symbolic expressions) and behavioral error (ensuring numerical prediction fidelity). It enables end-to-end joint optimization of symbolic equivalence (structural target) and functional approximation (behavioral target), overcoming the limitation of conventional methods relying solely on symbolic matching. Key innovations include differentiable symbolic encoding, behaviorally weighted loss design, and hybrid optimization—backpropagation for neural parameter tuning and genetic operations (crossover/mutation) for symbolic expression evolution. Results: On multiple benchmark datasets, the method achieves average improvements of 23.6% in symbolic accuracy and 18.4% in mean squared error over state-of-the-art neural-symbolic regression approaches.
Existing symbolic regression methods treat explicit (y = f(x)) and implicit (F(x, y) = 0) relationships separately and neglect structural reuse across expressions, resulting in low search efficiency and poor generalization. To address this, we propose the first unified, structure-aware evolutionary framework supporting both relationship types. Our method introduces a biologically inspired, reusable motif library; employs multi-island genetic programming; incorporates implicit derivative-guided fitness evaluation; and enforces subtree semantic consistency during crossover and mutation. This enables the first joint discovery of explicit and implicit relations. On the Nguyen benchmark, our approach achieves a 5.1% accuracy improvement and uniquely solves the highly challenging Nguyen-12 problem. It attains state-of-the-art performance on the Feynman equation suite and accelerates symbolic search by up to 100× over baseline methods on the Eureqa dataset—significantly enhancing both efficiency and robustness in scientific law discovery.
This work addresses the challenge of automatically discovering symbolic mathematical models for scientific discovery. Methodologically, it formulates symbolic expressions as syntax-tree sequences and employs policy-gradient reinforcement learning to guide the search space. Crucially, it introduces the first unified framework that jointly integrates gradient-based optimization, evolutionary algorithms, and domain-specific priors—such as physical constraints—to yield interpretable, physically consistent, and high-fidelity models. The contributions are threefold: (1) a fully differentiable, end-to-end paradigm for symbolic expression generation and optimization; (2) real-time incorporation of physical constraints with guaranteed model interpretability; and (3) state-of-the-art performance across multiple benchmark tasks, demonstrating significant improvements in accuracy, physical consistency, and interpretability—thereby validating the feasibility of automated symbolic scientific discovery.
Traditional symbolic regression methods are constrained by discrete structure search, suffering from high computational costs, unstable performance, and poor scalability. This work proposes SRCO, a novel framework that, for the first time, maps symbolic expressions into a continuous embedding space via a Transformer architecture, enabling joint differentiable optimization of both expression structures and coefficients. By operating in this continuous space, SRCO supports efficient structure search through either gradient-based or sampling-based strategies, thereby transcending the limitations of conventional discrete paradigms. Extensive experiments on both synthetic and real-world datasets demonstrate that SRCO significantly outperforms existing approaches in terms of equation accuracy, robustness, and search efficiency.
This work addresses the limitation of existing large language model (LLM)-driven symbolic regression methods, which lack explicit modeling of mathematical expression structure and consequently struggle to reliably recover exact formulas. To overcome this, the authors propose FunctionEvolve, a novel framework that introduces explicit structural guidance into LLM-based symbolic regression for the first time. FunctionEvolve organizes evolutionary search using expression trees and integrates structural summarization, local tree editing, and structure-aware coefficient optimization to enable efficient, structure-transparent exploration. Evaluated on 129 synthetic benchmarks from LLM-SRBench, FunctionEvolve achieves 82.9% SA@50 and 55.8% SA@1, outperforming current baselines by factors of 4.5 and 3.6, respectively, thereby substantially enhancing the recovery of precise symbolic expressions.
Existing large language model (LLM)-based symbolic regression methods rely solely on scalar metrics—such as mean squared error—for feedback, thereby overlooking the rich structural information embedded in the data and limiting both expressive power and search efficiency. This work proposes a programmatic context-augmented LLM-based evolutionary search framework that, for the first time, integrates executable code into the LLM-driven symbolic regression process. By dynamically generating and executing code to interact with the dataset, the method actively extracts fine-grained contextual signals beyond aggregated scores to guide expression evolution. Evaluated on benchmarks including LLM-SRBench, the approach substantially outperforms strong baselines, achieving significant improvements in both accuracy and search efficiency.
Traditional approaches struggle to extract interpretable mathematical structures from numerical solutions or neural network approximations of partial differential equations (PDEs). This work proposes the Agentic Symbolic Search (ASYS) framework, which uniquely integrates agent-guided symbolic search with domain knowledge–driven inductive biases. ASYS employs differentiable symbolic programming to generate candidate analytical expressions and combines evolutionary search for structural optimization with gradient-based parameter tuning. The method overcomes limitations inherent in purely analytical, grid-based numerical, and data-driven PDE representation strategies. Evaluated on five challenging PDE problems, ASYS not only recovers known analytical solutions but also discovers novel interpretable expressions, including a two-dimensional interface formula for the Allen–Cahn equation and a nine-parameter contraction law for the Keller–Segel model—both previously unreported in the literature.
Symbolic regression suffers from a combinatorially explosive search space, and directly generating mathematical expressions with large language models (LLMs) often lacks numerical rigor. This work proposes LLM-PySR, a novel framework that repurposes the LLM as a search controller rather than a formula generator: the LLM specifies variables, operators, transformations, and expression depth to guide the PySR system in enumerating and fitting candidate expressions, which are subsequently filtered using deterministic metrics. Evaluated on 74 AI-Feynman equations and seven complex formula recovery tasks, the method achieves an optimal trade-off among accuracy, simplicity, stability, and computational cost. Notably, it successfully uncovers a compact piecewise-linear relationship between voltage shift and cycle life from real battery data.
This work explores the use of fixed-depth symbolic regression to automatically discover neural network optimizers that outperform hand-designed counterparts. The method systematically searches within an expression space composed of gradients, momentum, and adaptive terms to identify compact and efficient weight update rules. It introduces a novel mechanism for constructing nonlinear rational expressions and represents the first application of fixed-depth symbolic regression to optimizer discovery. Experimental results across 30 benchmark–architecture combinations demonstrate that the discovered update rules surpass extensively tuned state-of-the-art optimizers in 25 cases, achieving an average reduction of 44.47% in mean squared error. These findings reveal both structural commonalities and diversity among highly effective update rules.