llm-guided symbolic search

Designs and implements symbolic search systems that use large language models to guide enumeration and tree-search for compact symbolic expressions or small programs. This work specifies variables, operators and transforms; sets search depth and hard/soft constraints; prioritizes promising expression subspaces via LLM scoring or heuristics; constrains enumeration to support robust numerical fitting and pruning; and integrates sampling, solver/optimizer, and implementation-level interfaces to produce and evaluate candidate expressions.

llm-guidedsymbolicsearch

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenges of using large language models (LLMs) to generate solvers for combinatorial optimization problems, where directly optimizing search strategies often introduces errors or degrades performance. The authors construct CP-SynC-XL, a benchmark comprising 100 problem types and 4,577 instances, to systematically evaluate three modeling paradigms: native Python, Python with OR-Tools, and MiniZinc with OR-Tools. Their analysis reveals a “heuristic trap”: compelling LLMs to generate optimized search logic yields only marginal speedups (1.03–1.12×) while significantly compromising correctness. Among the paradigms, Python+OR-Tools demonstrates superior performance. Based on these findings, the study proposes a “reformulate carefully, optimize sparingly” principle, advocating that LLMs should focus on formalizing variables, constraints, and objectives rather than synthesizing search strategies, thereby enhancing solver reliability.

combinatorial problemsconstraint modelingheuristic optimization

Symbolic regression suffers from a combinatorially explosive search space, and directly generating mathematical expressions with large language models (LLMs) often lacks numerical rigor. This work proposes LLM-PySR, a novel framework that repurposes the LLM as a search controller rather than a formula generator: the LLM specifies variables, operators, transformations, and expression depth to guide the PySR system in enumerating and fitting candidate expressions, which are subsequently filtered using deterministic metrics. Evaluated on 74 AI-Feynman equations and seven complex formula recovery tasks, the method achieves an optimal trade-off among accuracy, simplicity, stability, and computational cost. Notably, it successfully uncovers a compact piecewise-linear relationship between voltage shift and cycle life from real battery data.

combinatorial searchequation discoverylanguage models

Existing large language model (LLM)-based symbolic regression methods rely solely on scalar metrics—such as mean squared error—for feedback, thereby overlooking the rich structural information embedded in the data and limiting both expressive power and search efficiency. This work proposes a programmatic context-augmented LLM-based evolutionary search framework that, for the first time, integrates executable code into the LLM-driven symbolic regression process. By dynamically generating and executing code to interact with the dataset, the method actively extracts fine-grained contextual signals beyond aggregated scores to guide expression evolution. Evaluated on benchmarks including LLM-SRBench, the approach substantially outperforms strong baselines, achieving significant improvements in both accuracy and search efficiency.

Data FeedbackEvolutionary SearchLarge Language Models

Large Language Model powered Symbolic Execution

Apr 02, 2025
YL
Yihe Li
🏛️ National University of Singapore

Existing large language models (LLMs) face significant challenges in direct symbolic execution—including low precision, high computational overhead, and strong dependence on large-scale models and high-end hardware—hindering practical deployment. To address these limitations, we propose a novel LLM-driven lightweight symbolic execution paradigm. Our approach employs path-guided task decomposition to decouple complex program analysis into fine-grained, resource-efficient subtasks. We introduce the first path-constraint generalization method based on universal code representations—rather than restricted formal languages—enabling language-agnostic constraint modeling. We further implement AutoExe, a lightweight LLM-native symbolic execution engine. Experimental results demonstrate that our method substantially improves both analysis accuracy and path-exploration scalability for small-scale LLMs running on consumer-grade hardware, matching the performance of traditional symbolic execution tools. To the best of our knowledge, this is the first work achieving highly accessible and broadly generalizable LLM-native symbolic execution.

Enhancing LLM-based program analysis accuracy and scaleGeneralizing path constraints without formal language translationReducing hardware requirements for symbolic execution tasks

Latest Papers

What's happening recently
View more

This work addresses the inefficiency and unreliability of large language models (LLMs) in program synthesis tasks requiring extensive combinatorial search. The authors propose a novel approach that leverages a small number of LLM reasoning traces to compile, via an encoding agent, a reusable symbolic program synthesizer operating over a constrained domain-specific language (DSL), eliminating the need for LLM calls during testing. This method is the first to transform LLM reasoning traces into a zero-inference-overhead, reusable symbolic solver that functions independently while also enabling neuro-symbolic enhancement when combined with an LLM. It further supports zero-shot cross-domain transfer. On PBEBench-Hard, the approach achieves 84.7% accuracy—16.3 percentage points higher than test-time-scaled LLMs—and reaches 85.8% when augmented with an LLM while reducing token usage by 78%. In historical linguistics tasks, it attains 80.1% zero-shot accuracy.

combinatorial searchlarge language modelsprogram synthesis

Iterated Agent for Symbolic Regression

Oct 09, 2025
ZS
Zhuo-Yang Song
🏛️ Peking University | University of California, Berkeley | University of South China | Beijing Normal University | Beijing Computational Science Research Center

Symbolic regression (SR) suffers from combinatorial explosion of the search space, severe overfitting, and poor model interpretability. This paper proposes a semantic-driven, LLM-augmented evolutionary framework in which large language models serve as semantic operators—guiding candidate expression generation and mutation via natural language rationales, thereby replacing traditional syntax-based, blind search. Evaluated on the FSReD benchmark, our method achieves noise-robust, high-accuracy modeling while significantly improving expression conciseness, physical interpretability, and mechanistic alignment. In a high-energy physics parameterization task, it discovers compact models with explicit physical meaning. The core innovation lies in the first deep integration of LLMs’ semantic reasoning capability into the evolutionary search loop, enabling a paradigm shift from “syntactic evolution” to “concept-driven scientific discovery.”

Addressing combinatorial explosion and overfitting in symbolic regressionAutomating mathematical expression discovery from dataGenerating interpretable models using semantic-guided evolutionary search

This work addresses the issue of structural redundancy in symbolic regression caused by multiple node labelings of expression-directed acyclic graphs (DAGs), which inflates the search space and leads to redundant fitness evaluations. To resolve this, the authors propose IsalSR, a novel framework that introduces, for the first time, a complete labeled DAG isomorphism invariant as a canonical representation. By encoding DAGs into strings via a two-level compact alphabet and generating a pruned canonical form, IsalSR uniquely normalizes all semantically equivalent expressions. This approach fundamentally eliminates structural redundancy, substantially compressing the search space and avoiding repeated evaluations. Consequently, it enhances both search efficiency and solution diversity, leading to improved overall algorithmic performance.

Expression DAGNode-numbering SchemesSearch Space

This work addresses the performance bottleneck in search-based program synthesis caused by the high computational cost of fine-grained abstract semantics, which, while effective at pruning incorrect programs, hinders overall efficiency. To overcome this limitation, the authors propose an offline pre-synthesis approach that first constructs a tree automaton over the input space to precisely capture the abstract semantics of the domain-specific language (DSL). This automaton enables the generation of an efficient pruning oracle that decouples abstract semantic reasoning from the online synthesis process. By doing so, the method achieves, for the first time, highly efficient synthesis under fine-grained abstract semantics. Empirical evaluations demonstrate substantial performance improvements over state-of-the-art techniques across three diverse domains: SQL query synthesis, string transformation, and matrix manipulation.

abstract semanticsprogram synthesisscalability

Hot Scholars

ZL

Zhihao Lin

Phd Student, University of Glasgow
optimizationcontrol theoryreinforcement learningSLAM.
DD

Dong Du

Associate Professor, Nanjing University of Science and Technology
Computer Graphics3D Computer Vision
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
SH

Shibo Hao

Ph.D. student, UC San Diego
machine learninglarge language model