Score
Designs and implements methods that use large language models to generate, score, and encode prior beliefs over candidate operators or symbolic primitives and to propose candidate operators for downstream search. Builds operator-partitioning schemes and structured hypothesis outputs that bias or initialize symbolic search and parameter fitting by restricting or weighting regions of the operator search space.
This paper addresses the suboptimal synergy between large language models (LLMs) and evolutionary computation (EC). We propose the first bidirectional empowerment framework and co-evolutionary paradigm. Methodologically, we systematically integrate EC techniques—including genetic algorithms, evolution strategies, and Bayesian optimization—with LLM capabilities such as instruction tuning, chain-of-thought reasoning, and program synthesis. This enables EC to optimize LLM training, prompting, and architecture design, while LLMs enhance EC in algorithm design, hyperparameter optimization, and heuristic generation. Our contributions include: (i) formalizing cross-modal optimization pathways; (ii) surveying over 100 state-of-the-art works; (iii) empirically demonstrating that EC significantly improves LLM efficiency and robustness, whereas LLMs elevate EC’s automation level and interpretability; and (iv) revealing substantial synergistic gains on natural language understanding and complex search tasks. We further distill several key open challenges for future research.
In symbolic regression, selection operators are typically handcrafted, and existing LLM-driven evolutionary approaches suffer from code bloat and a lack of semantic guidance—hindering both interpretability and evolutionary efficiency. This paper proposes the first learning-based evolutionary framework that leverages large language models (LLMs) to automatically synthesize efficient and interpretable selection operators. We introduce two novel mechanisms: semantic-aware evaluation and code-bloat control, integrated with domain-knowledge-enhanced prompt engineering to jointly optimize operator generation for semantic validity and evolutionary efficacy. Evaluated on standard symbolic regression benchmarks, our automatically generated operators consistently outperform nine expert-designed baselines. To our knowledge, this is the first work demonstrating that LLMs can not only match but surpass human experts in algorithmic design—specifically, in crafting high-performing, semantically grounded selection operators for evolutionary symbolic regression.
Large language models (LLMs) suffer performance degradation on complex optimization tasks due to inconsistent multi-step planning. To address this, we propose MCTS-OPS, a novel neuro-symbolic framework that introduces Monte Carlo Tree Search (MCTS) into prompt sequence optimization for the first time. It formalizes prompt selection as a reward-driven sequential decision-making process, enabling dynamic exploration and iterative refinement. By tightly integrating LLM-based reasoning with symbolic search, MCTS-OPS enhances logical consistency and improves code generation quality. In network optimization benchmarks, MCTS-OPS achieves a 2–4× improvement in average reward over baseline methods, reduces result variance by 3×, and increases the success rate of finding optimal solutions on challenging instances by approximately 10%. These results demonstrate a substantial enhancement in LLMs’ capability to solve structured, multi-stage optimization problems.
Existing Alpha mining approaches suffer from either low computational efficiency (symbolic methods) or failure to preserve hierarchical structure (direct LLM application). To address these limitations, we propose the Tree-Structured Thought Evolution (TSTE) framework, which explicitly models Alpha’s inherent tree-like logical structure as an evolvable reasoning path. TSTE synergistically integrates large language models’ generative capabilities with symbolic regression principles and introduces dedicated evolutionary operators to optimize high-level reasoning architectures—without requiring full code execution. This work is the first to incorporate tree-structured reasoning into Alpha mining, substantially reducing reliance on domain expertise and computational resources. Empirical evaluation across four real-world financial market datasets demonstrates that TSTE generates superior Alpha factors in significantly less time, outperforming conventional linear evolutionary methods in both efficiency and factor performance.
This work addresses symbol regression without predefined functional bases. The proposed method introduces an image-driven, multimodal end-to-end framework: first, a vision-language model (VLM) generates initial mathematical expressions directly from function plots; second, Kolmogorov–Arnold networks (KANs) decompose multivariate regression into learnable univariate edge functions, embodying the “univariate suffices” principle; third, prompt-engineered language models guide genetic optimization and symbolic simplification to discover conditional, interpretable closed-form expressions. The approach requires no handcrafted function set, supports arbitrary prompt-based control and constraint modeling, and enables direct mapping from visual input to symbolic output. Evaluated on standard benchmarks, it significantly improves expression accuracy, generalization, and interpretability—establishing, for the first time, a multimodal symbol regression paradigm bridging images and symbolic expressions.
Large language models (LLMs) exhibit weak search capabilities in planning tasks and heavily rely on manual intervention. Method: This paper proposes AutoToS—a fully automated “Thought of Search” framework that shifts the search space from language modeling to verifiable code generation. AutoToS leverages both generic and domain-specific unit test feedback to iteratively guide LLMs in synthesizing semantically correct and logically complete successor functions and goal-detection code, integrating program semantic verification, domain-knowledge injection, and multi-scale LLM collaborative reasoning. Contribution/Results: AutoToS achieves 100% solution accuracy across all benchmark planning tasks, requiring only a minimal number of feedback iterations on average. It is compatible with open- and closed-source LLMs of diverse parameter scales. To our knowledge, this is the first work to realize end-to-end, human-free automation of planning search, establishing a novel paradigm for trustworthy LLM-based planning.
This work proposes a novel framework that integrates symbolic representations (via SymPy/SMT) with a closed-loop adaptive mechanism to address the limitations of existing reinforcement learning approaches in generating mathematical training data. Current methods often lack adaptability to the model’s evolving capabilities and offer insufficient control over problem structure. The proposed approach models problems in a symbolic space, ensuring structural controllability and verifiable solutions, while dynamically adjusting problem difficulty to match the learner’s current proficiency. By decoupling mathematical reasoning from linguistic expression, the framework enables strategy optimization through prompt-based learning within the symbolic space. Empirical results demonstrate that this method substantially enhances the mathematical problem-solving performance of small-scale open-source language models and yields training data with high diversity and precise structural control.
This work addresses the challenge of prompt optimization in test-driven code generation with large language models (LLMs). Methodologically, it introduces the first Bayesian optimization (BO) framework for prompt engineering: an auxiliary LLM maps discrete prompts into a continuous embedding space, enabling construction of a Gaussian process surrogate model; random projection and dimensionality-scaling priors are incorporated to mitigate degradation in high-dimensional modeling. Optimization leverages execution feedback from test cases to adaptively search for prompts maximizing functional correctness. Evaluated on HumanEval+, the approach significantly outperforms fixed and manually tuned prompts across multiple base LLMs, improving code generation accuracy with rapid convergence—often within few BO iterations. The core contributions are (1) establishing a differentiable prompt optimization paradigm, and (2) enhancing the stability and efficiency of BO in high-dimensional embedding spaces.
This work addresses the significant limitations of large language models (LLMs) in autonomously executing algorithms and performing complex structured reasoning. To overcome these challenges, the authors propose the LLM-DAL framework, which employs a supervised training approach based on Decompositional Algorithmic Learning. This method explicitly guides the model to decompose and internalize algorithmic reasoning steps during training. By doing so, LLM-DAL substantially enhances the model’s capability to execute and generalize on algorithmic tasks—particularly complex arithmetic functions—thereby breaking through inherent bottlenecks in structured reasoning. The framework offers a novel pathway toward improving the systematic reasoning abilities of large language models, demonstrating that explicit decomposition of algorithmic processes can lead to more robust and generalizable performance in tasks requiring precise, stepwise logic.
Symbolic regression suffers from a combinatorially explosive search space, and directly generating mathematical expressions with large language models (LLMs) often lacks numerical rigor. This work proposes LLM-PySR, a novel framework that repurposes the LLM as a search controller rather than a formula generator: the LLM specifies variables, operators, transformations, and expression depth to guide the PySR system in enumerating and fitting candidate expressions, which are subsequently filtered using deterministic metrics. Evaluated on 74 AI-Feynman equations and seven complex formula recovery tasks, the method achieves an optimal trade-off among accuracy, simplicity, stability, and computational cost. Notably, it successfully uncovers a compact piecewise-linear relationship between voltage shift and cycle life from real battery data.
LLM inference cost is increasingly becoming a critical resource bottleneck; existing optimality analyses often decouple training from inference and neglect dynamic trade-offs in strategy selection. This paper proposes Directed Stochastic Skill Search (DS3), a framework that models inference as stochastic traversal over a skill graph, establishing the first unified theoretical model for joint training–inference optimization. Leveraging a tripartite graph structure and stochastic process theory, we derive closed-form expressions for computational cost and accuracy of inference strategies—including Chain-of-Thought (CoT) and Tree-of-Thought (ToT)—and rigorously characterize emergence conditions for mechanisms such as Bag-of-Natural-proofs (BoN) and majority voting from first principles. Our theory reproduces scaling phenomena—e.g., linear accuracy growth under logarithmic compute—and yields a critical threshold under which small models can surpass large ones via efficient inference. The results provide quantitative design principles for low-cost, high-reliability LLM inference.