Score
Designs and implements symbolic search systems that use large language models to guide enumeration and tree-search for compact symbolic expressions or small programs. This work specifies variables, operators and transforms; sets search depth and hard/soft constraints; prioritizes promising expression subspaces via LLM scoring or heuristics; constrains enumeration to support robust numerical fitting and pruning; and integrates sampling, solver/optimizer, and implementation-level interfaces to produce and evaluate candidate expressions.
This paper systematically reviews bottlenecks and recent advances in applying large language models (LLMs) to mathematical reasoning and optimization. It identifies key limitations: arithmetic inaccuracy, logical inconsistency, lack of theorem verifiability, and poor support for structured symbolic computation. To address these, the work proposes three innovations: (1) neuro-symbolic hybrid architectures, (2) multi-step self-correcting reasoning mechanisms, and (3) structured prompt engineering. Methodologically, it integrates chain-of-thought reasoning, tool-augmented inference, and instruction fine-tuning, and—novelty—establishes interoperable interfaces between LLMs and classical optimization frameworks, including mixed-integer programming and linear-quadratic optimal control, enabling multi-agent optimization strategies. The study rigorously delineates LLMs’ capabilities and limitations in formal mathematical tasks, significantly improving reliability in complex reasoning and mechanized proof generation. Results provide actionable pathways for engineering optimization, quantitative finance, and fundamental scientific research. (149 words)
This work addresses the challenges of using large language models (LLMs) to generate solvers for combinatorial optimization problems, where directly optimizing search strategies often introduces errors or degrades performance. The authors construct CP-SynC-XL, a benchmark comprising 100 problem types and 4,577 instances, to systematically evaluate three modeling paradigms: native Python, Python with OR-Tools, and MiniZinc with OR-Tools. Their analysis reveals a “heuristic trap”: compelling LLMs to generate optimized search logic yields only marginal speedups (1.03–1.12×) while significantly compromising correctness. Among the paradigms, Python+OR-Tools demonstrates superior performance. Based on these findings, the study proposes a “reformulate carefully, optimize sparingly” principle, advocating that LLMs should focus on formalizing variables, constraints, and objectives rather than synthesizing search strategies, thereby enhancing solver reliability.
Symbolic regression suffers from a combinatorially explosive search space, and directly generating mathematical expressions with large language models (LLMs) often lacks numerical rigor. This work proposes LLM-PySR, a novel framework that repurposes the LLM as a search controller rather than a formula generator: the LLM specifies variables, operators, transformations, and expression depth to guide the PySR system in enumerating and fitting candidate expressions, which are subsequently filtered using deterministic metrics. Evaluated on 74 AI-Feynman equations and seven complex formula recovery tasks, the method achieves an optimal trade-off among accuracy, simplicity, stability, and computational cost. Notably, it successfully uncovers a compact piecewise-linear relationship between voltage shift and cycle life from real battery data.
Existing large language model (LLM)-based symbolic regression methods rely solely on scalar metrics—such as mean squared error—for feedback, thereby overlooking the rich structural information embedded in the data and limiting both expressive power and search efficiency. This work proposes a programmatic context-augmented LLM-based evolutionary search framework that, for the first time, integrates executable code into the LLM-driven symbolic regression process. By dynamically generating and executing code to interact with the dataset, the method actively extracts fine-grained contextual signals beyond aggregated scores to guide expression evolution. Evaluated on benchmarks including LLM-SRBench, the approach substantially outperforms strong baselines, achieving significant improvements in both accuracy and search efficiency.
Existing large language models (LLMs) face significant challenges in direct symbolic execution—including low precision, high computational overhead, and strong dependence on large-scale models and high-end hardware—hindering practical deployment. To address these limitations, we propose a novel LLM-driven lightweight symbolic execution paradigm. Our approach employs path-guided task decomposition to decouple complex program analysis into fine-grained, resource-efficient subtasks. We introduce the first path-constraint generalization method based on universal code representations—rather than restricted formal languages—enabling language-agnostic constraint modeling. We further implement AutoExe, a lightweight LLM-native symbolic execution engine. Experimental results demonstrate that our method substantially improves both analysis accuracy and path-exploration scalability for small-scale LLMs running on consumer-grade hardware, matching the performance of traditional symbolic execution tools. To the best of our knowledge, this is the first work achieving highly accessible and broadly generalizable LLM-native symbolic execution.
This work addresses the inefficiency and unreliability of large language models (LLMs) in program synthesis tasks requiring extensive combinatorial search. The authors propose a novel approach that leverages a small number of LLM reasoning traces to compile, via an encoding agent, a reusable symbolic program synthesizer operating over a constrained domain-specific language (DSL), eliminating the need for LLM calls during testing. This method is the first to transform LLM reasoning traces into a zero-inference-overhead, reusable symbolic solver that functions independently while also enabling neuro-symbolic enhancement when combined with an LLM. It further supports zero-shot cross-domain transfer. On PBEBench-Hard, the approach achieves 84.7% accuracy—16.3 percentage points higher than test-time-scaled LLMs—and reaches 85.8% when augmented with an LLM while reducing token usage by 78%. In historical linguistics tasks, it attains 80.1% zero-shot accuracy.
Symbolic regression (SR) suffers from combinatorial explosion of the search space, severe overfitting, and poor model interpretability. This paper proposes a semantic-driven, LLM-augmented evolutionary framework in which large language models serve as semantic operators—guiding candidate expression generation and mutation via natural language rationales, thereby replacing traditional syntax-based, blind search. Evaluated on the FSReD benchmark, our method achieves noise-robust, high-accuracy modeling while significantly improving expression conciseness, physical interpretability, and mechanistic alignment. In a high-energy physics parameterization task, it discovers compact models with explicit physical meaning. The core innovation lies in the first deep integration of LLMs’ semantic reasoning capability into the evolutionary search loop, enabling a paradigm shift from “syntactic evolution” to “concept-driven scientific discovery.”
This work addresses the issue of structural redundancy in symbolic regression caused by multiple node labelings of expression-directed acyclic graphs (DAGs), which inflates the search space and leads to redundant fitness evaluations. To resolve this, the authors propose IsalSR, a novel framework that introduces, for the first time, a complete labeled DAG isomorphism invariant as a canonical representation. By encoding DAGs into strings via a two-level compact alphabet and generating a pruned canonical form, IsalSR uniquely normalizes all semantically equivalent expressions. This approach fundamentally eliminates structural redundancy, substantially compressing the search space and avoiding repeated evaluations. Consequently, it enhances both search efficiency and solution diversity, leading to improved overall algorithmic performance.
This work addresses the performance bottleneck in search-based program synthesis caused by the high computational cost of fine-grained abstract semantics, which, while effective at pruning incorrect programs, hinders overall efficiency. To overcome this limitation, the authors propose an offline pre-synthesis approach that first constructs a tree automaton over the input space to precisely capture the abstract semantics of the domain-specific language (DSL). This automaton enables the generation of an efficient pruning oracle that decouples abstract semantic reasoning from the online synthesis process. By doing so, the method achieves, for the first time, highly efficient synthesis under fine-grained abstract semantics. Empirical evaluations demonstrate substantial performance improvements over state-of-the-art techniques across three diverse domains: SQL query synthesis, string transformation, and matrix manipulation.