Score
Designs and implements Monte Carlo Tree Search systems that explore and rank a space of candidate compositions, using formal or heuristic contracts to guide node expansion, selection, and backpropagation so that promising or risky compositions are discovered earlier. Builds prioritization and budget-allocation strategies that assign limited queries or evaluations across tree branches to maximize discovery yield of high-risk or high-value composition candidates.
This work addresses the challenge of explaining decisions made by Monte Carlo Tree Search (MCTS), which are often opaque due to its asymmetric tree structure and simulation-based value estimation. To this end, the authors propose an end-to-end interpretability method that leverages a large language model (LLM) to generate evidence-based natural language explanations directly from MCTS search trajectories. The approach first classifies user intent to understand the query, then dynamically evaluates the sufficiency of evidence within the search tree and triggers targeted expansions when necessary. Explanations are synthesized using visit counts, value estimates, and risk-aware information. This study presents the first framework capable of producing adaptive, evidence-grounded explanations for probabilistic search algorithms without relying on handcrafted templates or intermediate formal representations, demonstrating that LLMs can effectively serve as end-to-end interpreters for MCTS.
To address the low search efficiency and poor interpretability of strategies in complex combinatorial games, this paper proposes a multi-level simplification model-driven enhancement of Monte Carlo Tree Search (MCTS). Our method constructs progressive game simplification models and integrates heuristic performance prediction with policy ensembling to estimate algorithmic behavior within the simplified space, thereby guiding tree expansion and policy selection in the original MCTS. The core contribution is the first establishment of a closed-loop paradigm—“simplification modeling → performance prediction → search guidance”—unifying strategy predictability and interpretability. Evaluated on multiple challenging combinatorial game benchmarks, the proposed approach significantly improves MCTS’s strategic quality (average +12.7% win rate) and convergence speed (speedup ratio up to 2.3×), demonstrating the strong guiding efficacy of simplification-based analysis for solving complex games.
To address the challenge of jointly optimizing local precision and global semantic coherence in complex, multi-step, heterogeneous information retrieval tasks, this paper proposes HG-MCTS, a large language model–based search assistant that formalizes retrieval as a knowledge-augmented, progressive information gathering process. We introduce an adaptive checklist-guided Monte Carlo Tree Search (MCTS) framework, integrating a multi-faceted sparse reward mechanism that jointly accounts for exploration, retrieval efficacy, and progress awareness. Furthermore, we incorporate dynamic subgoal modeling and cross-step knowledge caching to enhance long-horizon reasoning and reduce redundancy. Evaluated on realistic, complex retrieval benchmarks, HG-MCTS achieves significant improvements in knowledge coverage completeness and answer accuracy, while substantially reducing search path redundancy. It consistently outperforms state-of-the-art baselines across all key metrics, demonstrating superior capability in balancing fine-grained relevance with holistic semantic understanding.
Existing Monte Carlo Tree Diffusion (MCTD) methods are constrained by fixed training trajectory lengths, supporting only single-trajectory local search without global planning capability. To address this limitation, we propose Compositional MCTD—a novel framework that for the first time formalizes planning as compositional reasoning across trajectory segments, implemented via three types of combiners: online, distributed, and pre-planning. Our approach integrates diffusion models with Monte Carlo tree search, incorporating parallel exploration, plan-graph caching, and global search strategies to enable efficient, scalable sequential decision-making for long-horizon tasks. Experiments demonstrate significant improvements in success rates and trajectory coherence on long-range tasks, alongside markedly enhanced inference efficiency.
This work addresses the problem of determining whether the root value in a Monte Carlo tree search exceeds a given threshold, where internal nodes alternate between MAX and MIN operations and leaf node values correspond to the means of unknown distributions. To tackle this, the authors propose a δ-correct sequential sampling algorithm built upon the Track-and-Stop framework, featuring an innovative ratio-corrected D-Tracking strategy for arm selection. The method preserves asymptotic optimality in sample complexity while substantially reducing the actual number of samples required in practice. Furthermore, it improves computational efficiency by lowering the per-round time complexity from linear to logarithmic. Empirical evaluations demonstrate the algorithm’s dual advantages in both sample efficiency and computational speed.
Existing scientific discovery methods often conflate the quality of hypotheses with the effectiveness of experimental execution and prematurely prune search histories due to context length limitations, thereby overlooking high-quality initial hypotheses. This work proposes the ARTS framework, which for the first time decouples the evaluation of hypothesis value from execution quality by leveraging a reasoning language model to analyze execution logs and distinguish the root causes of failure, thereby guiding subsequent exploration. Additionally, ARTS introduces a test-time training mechanism that compresses knowledge from the search tree into model parameters, circumventing context window constraints. Evaluated on 22 tasks across MLGym and MLEBench, ARTS achieves a normalized score improvement of over 15.3%. A fine-tuned Qwen3-4B model under this framework matches the performance of Gemini-3 Pro and GPT-o3-reasoning at one-fifth the inference cost and demonstrates superior results on reinforcement learning tasks.
This work addresses the challenges of combinatorial explosion and long-range credit assignment in multi-hop reasoning over knowledge graphs for uncovering drug–disease mechanisms. The authors propose TESSERA, a novel framework that leverages large language models (LLMs) not as end-to-end generators but as local discriminators and state evaluators, integrating structural constraints from the knowledge graph with Monte Carlo Tree Search (MCTS) to enable neuro-symbolic, controllable reasoning. This approach generates interpretable, biologically plausible multi-step paths, demonstrating on two knowledge graphs its ability to both recover known mechanisms and identify alternative yet reasonable pathways. Ablation studies further validate the effectiveness of the dual-component LLM design.
This work addresses the limited interpretability of existing MCTS-Minimax hybrid approaches for multi-agent decision-making and the tendency of standard Monte Carlo Tree Search (MCTS) to overlook critical actions or become trapped in local tactical optima. To enhance strategic depth, the authors propose embedding shallow full-width Minimax search within the rollout phase of MCTS. Furthermore, they introduce a novel integration of process mining techniques—such as Alpha Miner and Inductive Miner—with large language models to structurally model agent behavior trajectories and generate human-readable causal and root-cause explanations. Experimental validation in a small checkers environment demonstrates the effectiveness of the approach, offering a scalable framework for explainable hybrid agents in complex strategic scenarios.
This work addresses the over-exploitation issue in Monte Carlo Tree Search (MCTS) for automated heuristic design under limited computational budgets. To mitigate this, the authors propose Clade-AHD, a novel framework that introduces, for the first time, a clade-level Bayesian belief modeling mechanism. Instead of relying on traditional node-level point estimates, Clade-AHD aggregates subtree evaluations using Beta distributions and employs Thompson sampling to guide exploration decisions. This approach effectively balances exploration and exploitation in sparse and noisy evaluation environments, substantially improving heuristic quality. Experimental results demonstrate that Clade-AHD outperforms existing methods on complex combinatorial optimization tasks while significantly reducing computational overhead.