Score
Design and implement autotuning systems that use Monte Carlo Tree Search to autonomously explore and select high-performance kernel implementation variants; this includes building the MCTS search policy with mechanisms such as progressive widening and pruning and calibrating log-reward signals for reliable selection. Produce and evaluate optimized source-level kernel variants and analyze search trajectories, convergence, and runtime-measured performance to drive further tuning.
This work addresses the high variability and long-tail latency in Monte Carlo Tree Search (MCTS) during test-time computation expansion, which stems from inefficient search trajectories and limits the effectiveness of existing optimizations when search progress stalls. The authors propose a negative early-exit mechanism that proactively prunes unproductive trajectories and integrates an adaptive boosting strategy to dynamically reallocate the freed computational resources, thereby mitigating resource contention among parallel searches. Implemented within the vLLM inference framework, this approach significantly reduces end-to-end p99 latency and improves system throughput while preserving the accuracy of large language model inference.
Existing LLM-driven AutoML agents suffer from low code-generation diversity and suboptimal node selection due to reliance on scalar feedback. To address these issues, we propose Introspective Monte Carlo Tree Search (Introspective MCTS), a novel framework featuring: (i) a reflective node expansion mechanism that leverages parent- and sibling-node feedback—first of its kind; (ii) an LLM-based value model for proactive evaluation of the solution space prior to execution; and (iii) a hybrid reward function integrating LLM-generated scores with ground-truth performance metrics to smooth search guidance. Evaluated across diverse machine learning tasks, our framework significantly improves exploration quality and generalization capability. On mainstream open-source Agentic AutoML benchmarks, it achieves a 6% absolute performance gain. This work establishes a new, interpretable, and iterative decision-optimization paradigm for LLM-based AutoML.
This paper addresses the challenge in Monte Carlo Tree Search (MCTS) of effectively incorporating prior knowledge into the tree policy while balancing exploration efficiency and theoretical guarantees. To this end, we propose Inverse-RPO, a general framework that establishes, for the first time, a principled derivation paradigm for prior-aware UCT algorithms. Inverse-RPO systematically extends variance-aware, prior-free UCB methods (e.g., UCB-V) into variance-sensitive prior tree policies, ensuring both theoretical rigor and empirical superiority. By integrating regularized policy modeling and enhanced variance estimation, our method achieves zero-overhead integration within the mctx library. Empirical evaluation across multiple benchmark tasks demonstrates consistent and significant improvements over PUCT. The implementation is fully open-sourced, enabling straightforward reproduction and facilitating further research extensions.
Traditional Monte Carlo Tree Search (MCTS) suffers from insufficient global exploration in symbolic regression due to sequential expression construction and classical bandit policies. To address this, we propose an enhanced MCTS framework integrating reinforcement learning and evolutionary principles. Our key contributions are: (1) an Extreme Bandit strategy that improves global exploration under sparse rewards; (2) mutation- and crossover-driven state-jump actions enabling non-local transitions across the search space; and (3) a hybrid search mechanism synergizing sequential expression generation with state jumps, accompanied by theoretical performance guarantees under time constraints. Experiments on multiple benchmark datasets show that our method achieves expression recovery rates comparable to leading symbolic regression libraries. Moreover, its solutions consistently dominate the Pareto frontier in the accuracy–complexity trade-off, demonstrating significantly improved robustness and search efficiency.
Diffusion models exhibit strong generative capabilities for planning tasks but suffer from non-scalable test-time computation (TTC): performance plateaus rather than improving monotonically with increased inference budget. To address this, we propose Diffusion-MCTS—the first framework integrating Monte Carlo Tree Search (MCTS) into the diffusion paradigm. It reformulates the denoising process as a tree-structured search, incorporating value-guided node selection, conditional resampling, and backtracking from suboptimal branches to enable iterative evaluation, pruning, and refinement. This design supports dynamic exploration-exploitation trade-offs, overcoming the fundamental TTC bottleneck of conventional diffusion-based planners. Empirically, on long-horizon planning tasks, solution quality improves monotonically with computational budget—significantly outperforming diffusion baselines—and demonstrates both TTC scalability and robustness.
This work proposes Particle Monte Carlo Tree Search (PMCTS), a novel parallel framework for Monte Carlo Tree Search (MCTS) that overcomes the inherent sequential limitations of traditional MCTS in parallel environments. PMCTS introduces a particle-based sampling mechanism that effectively integrates neural network policies and value functions, enabling large-scale parallel inference while preserving theoretical guarantees for policy improvement—the first such guarantee in parallel MCTS. The method maintains rigorous theoretical foundations and demonstrates substantial performance gains over existing heuristic parallel baselines, achieving superior scalability and consistent improvements across multiple tasks.
本文提出了一种基于机器学习的LLVM IR性能排序方法,通过转移学习减少自动调优中的评估次数,提高了高性能计算系统的调优效率。
This study addresses the challenges of latent semantic errors and repetitive repair failures in code generation using open-source large language models. To this end, it proposes an execution-guided, memory-augmented Monte Carlo Tree Search (MCTS) framework. Specifically, the method employs an LLM to guide MCTS in organizing candidate programs while leveraging long-term memory retrieval to share cross-branch failure experiences, thereby preventing redundant trial-and-error. Furthermore, a reflection mechanism combined with branch-local debugging contexts is introduced to precisely distinguish failed assertions from missing evidence, enabling stable and efficient code search. Extensive evaluations on the HumanEval and MBPP benchmarks demonstrate that the proposed framework significantly outperforms direct generation methods under most configurations, while also confirming its compatibility across different compiler backends.
本文探讨了蒙特卡洛树搜索(MCTS)与每访问蒙特卡洛控制方法在轨迹生成和动作价值更新层面本质上的一致性,指出MCTS可视为以搜索语言表达的每访问蒙特卡洛控制。
Existing learning-based Monte Carlo Tree Search (MCTS) approaches for query optimization exhibit limited generalization under diverse workloads and struggle to consistently reproduce performance gains. This work identifies their out-of-distribution generalization shortcomings through a reproducibility study and proposes a novel MCTS framework that eschews end-to-end learning entirely, relying solely on the database’s built-in cost model. To enhance search efficiency and robustness, the framework incorporates a new Extreme UCT selection strategy. Evaluated on the Join Order Benchmark (JOB) and its more complex extension, JOB-Complex, the proposed method significantly outperforms learning-based MCTS optimizers such as AlphaJoin and HyperQO, and further surpasses state-of-the-art industrial query optimizers in complex join scenarios.