Improving Monte Carlo Tree Search for Symbolic Regression

📅 2025-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional Monte Carlo Tree Search (MCTS) suffers from insufficient global exploration in symbolic regression due to sequential expression construction and classical bandit policies. To address this, we propose an enhanced MCTS framework integrating reinforcement learning and evolutionary principles. Our key contributions are: (1) an Extreme Bandit strategy that improves global exploration under sparse rewards; (2) mutation- and crossover-driven state-jump actions enabling non-local transitions across the search space; and (3) a hybrid search mechanism synergizing sequential expression generation with state jumps, accompanied by theoretical performance guarantees under time constraints. Experiments on multiple benchmark datasets show that our method achieves expression recovery rates comparable to leading symbolic regression libraries. Moreover, its solutions consistently dominate the Pareto frontier in the accuracy–complexity trade-off, demonstrating significantly improved robustness and search efficiency.

Technology Category

Search and Optimization: Sampling/Simulation-based SearchMachine Learning: Neuro-Symbolic LearningReasoning under Uncertainty: Stochastic Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Symbolic regression aims to discover concise, interpretable mathematical expressions that satisfy desired objectives, such as fitting data, posing a highly combinatorial optimization problem. While genetic programming has been the dominant approach, recent efforts have explored reinforcement learning methods for improving search efficiency. Monte Carlo Tree Search (MCTS), with its ability to balance exploration and exploitation through guided search, has emerged as a promising technique for symbolic expression discovery. However, its traditional bandit strategies and sequential symbol construction often limit performance. In this work, we propose an improved MCTS framework for symbolic regression that addresses these limitations through two key innovations: (1) an extreme bandit allocation strategy tailored for identifying globally optimal expressions, with finite-time performance guarantees under polynomial reward decay assumptions; and (2) evolution-inspired state-jumping actions such as mutation and crossover, which enable non-local transitions to promising regions of the search space. These state-jumping actions also reshape the reward landscape during the search process, improving both robustness and efficiency. We conduct a thorough numerical study to the impact of these improvements and benchmark our approach against existing symbolic regression methods on a variety of datasets, including both ground-truth and black-box datasets. Our approach achieves competitive performance with state-of-the-art libraries in terms of recovery rate, attains favorable positions on the Pareto frontier of accuracy versus model complexity. Code is available at https://github.com/PKU-CMEGroup/MCTS-4-SR.
Problem

Research questions and friction points this paper is trying to address.

Enhancing MCTS for symbolic regression with innovative strategies
Addressing limitations of traditional bandit methods in expression discovery
Improving search efficiency and robustness through evolutionary-inspired actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Extreme bandit strategy for global optimization
State-jumping actions like mutation and crossover
Reward landscape reshaping for improved efficiency
💼 Related Jobs
No related jobs found.
Z
Zhengyao Huang
Center for Machine Learning Research, Peking University, Beijing, China
D
Daniel Zhengyu Huang
Beijing International Center for Mathematical Research, Center for Machine Learning Research, Peking University, Beijing, China
T
Tiannan Xiao
Huawei Technologies Ltd., Beijing, China
D
Dina Ma
Huawei Technologies Ltd., Beijing, China
Z
Zhenyu Ming
Huawei Technologies Ltd., Beijing, China
H
Hao Shi
Department of Mathematical Sciences, Tsinghua University, Beijing, China
Y
Yuanhui Wen
Huawei Technologies Ltd., Beijing, China