mcts-based composition

Designs and implements Monte Carlo Tree Search procedures that assemble outputs by constructing and exploring a tree of segment or component nodes, using selection/expansion/rollout/backpropagation to generate and evaluate candidate combinations. Builds scoring and refinement mechanisms that use self-reward signals to identify best-scoring compositions and to refine selected combinations into a final assembled output.

mcts-basedcomposition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Simplifications to Guide Monte Carlo Tree Search in Combinatorial Games

Jan 13, 2025
MH
Michael Haythorpe
🏛️ Flinders University | Defence Science and Technology Group

To address the low search efficiency and poor interpretability of strategies in complex combinatorial games, this paper proposes a multi-level simplification model-driven enhancement of Monte Carlo Tree Search (MCTS). Our method constructs progressive game simplification models and integrates heuristic performance prediction with policy ensembling to estimate algorithmic behavior within the simplified space, thereby guiding tree expansion and policy selection in the original MCTS. The core contribution is the first establishment of a closed-loop paradigm—“simplification modeling → performance prediction → search guidance”—unifying strategy predictability and interpretability. Evaluated on multiple challenging combinatorial game benchmarks, the proposed approach significantly improves MCTS’s strategic quality (average +12.7% win rate) and convergence speed (speedup ratio up to 2.3×), demonstrating the strong guiding efficacy of simplification-based analysis for solving complex games.

Complex GamesEffective StrategiesImproved Search Methods

To address the lack of interpretability and susceptibility to tactical traps in Monte Carlo Tree Search (MCTS) for multi-player board games, this paper proposes Minimax-Augmented MCTS (MA-MCTS). Our method integrates a shallow minimax search into the rollout phase to enhance local robustness and—novelly—incorporates process mining techniques into MCTS policy modeling to reconstruct decision-making processes with explicit interpretability. This breaks the traditional “black-box” limitation of MCTS by generating human-readable strategy execution flowcharts, enabling real-time diagnosis and intervention. In 3v3 checkers experiments, MA-MCTS achieves a 23.6% improvement in critical-move identification accuracy and a 31.4% increase in tactical-trap avoidance rate. The main contributions are: (1) the first synergistic enhancement framework unifying MCTS and minimax for multi-player settings; and (2) a process-mining–based paradigm for interpretable MCTS policy modeling.

Combines MCTS with minimax to avoid tactical trapsExplains MCTS decision-making in board gamesUses process mining to analyze 3v3 checkers strategies

This work addresses the problem of determining whether the root value in a Monte Carlo tree search exceeds a given threshold, where internal nodes alternate between MAX and MIN operations and leaf node values correspond to the means of unknown distributions. To tackle this, the authors propose a δ-correct sequential sampling algorithm built upon the Track-and-Stop framework, featuring an innovative ratio-corrected D-Tracking strategy for arm selection. The method preserves asymptotic optimality in sample complexity while substantially reducing the actual number of samples required in practice. Furthermore, it improves computational efficiency by lowering the per-round time complexity from linear to logarithmic. Empirical evaluations demonstrate the algorithm’s dual advantages in both sample efficiency and computational speed.

Decision ThresholdMonte Carlo Tree SearchOptimal Sample Complexity

Array-Based Monte Carlo Tree Search

Aug 27, 2025
JR
James Ragan
🏛️ California Institute of Technology

Monte Carlo Tree Search (MCTS) suffers from poor scalability on modern CPUs due to branch misprediction overhead and irregular memory access patterns induced by its pointer-based tree structure, limiting simulation throughput. To address this, we propose a compact array-based MCTS implementation that replaces the explicit pointer tree with an implicit array representation—preserving the UCT algorithm’s logic while eliminating indirect memory accesses and conditional branches. This design improves cache locality and instruction-level parallelism by enabling contiguous memory layouts and predictable control flow. Experimental evaluation demonstrates a substantial increase in simulations per unit time under identical clock cycles; maximum search depth improves by up to 2.8× compared to conventional pointer-based implementations. The approach significantly enhances MCTS scalability and real-time performance in latency-sensitive decision-making tasks, without compromising algorithmic correctness or generality.

Eliminating branch prediction to improve processor efficiencyEnhancing search depth scaling in numerical simulationsOptimizing Monte Carlo Tree Search for faster decision making

RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation

Sep 15, 2024
QL
Qingyao Li
🏛️ Shanghai Jiao Tong University | Huawei Noah's Ark Lab

Existing tree-search-based code generation methods perform direct search over the raw code space, neglecting the underlying reasoning process; meanwhile, reflection-based approaches merely accumulate errors without guiding toward correct reasoning paths, degrading search quality. This paper proposes a thought-level Monte Carlo Tree Search (MCTS) framework, introducing the novel “rethink” mechanism: it transforms fine-grained code execution feedback into semantically grounded natural-language feedback, dynamically rectifies erroneous reasoning chains, and triggers node re-expansion. The method integrates LLM-based reasoning-chain modeling, execution-feedback parsing, and reflection-driven search. On HumanEval, it achieves pass@1 scores of 89.02 (+18.9) with GPT-3.5-turbo and 94.51 (+7.31) with GPT-4o-mini—substantially outperforming state-of-the-art search and feedback-augmented methods. To our knowledge, this is the first approach to establish a closed-loop pipeline spanning error identification, semantic correction, and search optimization.

Addresses reasoning process neglect in code generation tree searchImproves search quality by aligning paths with better reasoningRefines erroneous thoughts using fine-grained execution feedback

Latest Papers

What's happening recently
View more

This work addresses the challenge of explaining decisions made by Monte Carlo Tree Search (MCTS), which are often opaque due to its asymmetric tree structure and simulation-based value estimation. To this end, the authors propose an end-to-end interpretability method that leverages a large language model (LLM) to generate evidence-based natural language explanations directly from MCTS search trajectories. The approach first classifies user intent to understand the query, then dynamically evaluates the sufficiency of evidence within the search tree and triggers targeted expansions when necessary. Explanations are synthesized using visit counts, value estimates, and risk-aware information. This study presents the first framework capable of producing adaptive, evidence-grounded explanations for probabilistic search algorithms without relying on handcrafted templates or intermediate formal representations, demonstrating that LLMs can effectively serve as end-to-end interpreters for MCTS.

Decision-MakingExplainabilityMonte Carlo Tree Search

Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions

Dec 10, 2025
JX
Junlin Xiao
🏛️ The Chinese University of Hong Kong, Shenzhen | Oxford Robotics Institute | University of Oxford

To address the challenge of cross-thread statistical aggregation in root-parallel Monte Carlo tree search (MCTS) for continuous action spaces, this paper introduces Gaussian process regression (GPR) into the root-parallel MCTS framework for the first time. We propose a GPR-based value estimation method that enables reliable value extrapolation for unsampled actions and explicitly models action-space continuity. The approach preserves online planning efficiency while effectively fusing local statistics from multiple threads. Evaluated on six standard continuous control benchmarks, it significantly outperforms existing aggregation strategies—including weighted averaging and max-value aggregation—yielding substantial improvements in policy quality and planning stability, with only marginal increases in inference overhead. Our core contribution is the novel application of GPR to cross-thread value aggregation in root-parallel MCTS, establishing a new paradigm for efficient online planning in continuous action domains.

Aggregates statistics from parallel threads in continuous action spacesOutperforms existing methods across six domains with minimal time increaseUses Gaussian Process Regression for untried action value estimation

Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms

Oct 28, 2025
RS
Robin Schmöcker
🏛️ Leibniz University Hannover | University of Southern Denmark

In Monte Carlo Tree Search (MCTS), state/action abstraction often collapses multiple distinct actions into a single abstract node, leading to ambiguous action selection; conventional random tie-breaking yields suboptimal policies and degrades search efficiency. Method: We propose several novel intra-abstraction decision strategies that systematically refine action selection *within* abstract nodes, integrate them into a UCB-based MCTS framework, and synergistically combine them with imprecise abstraction techniques (e.g., pruned Optimistic Graph Abstraction). Contribution/Results: Extensive experiments across diverse benchmark environments and parameter configurations demonstrate that our strategies significantly outperform random tie-breaking baselines—accelerating convergence of abstraction-guided search, improving policy quality, and enhancing generalization stability. The approach establishes an interpretable, reusable internal decision paradigm for abstraction-augmented MCTS.

Addressing UCB tiebreak issues in pruned OGA algorithmsEvaluating alternative intra-abstraction policies for MCTSImproving MCTS sample efficiency via state abstractions

This work addresses the inefficiency of traditional Monte Carlo Tree Search (MCTS) in adversarial tabletop games with high uncertainty—such as Jaipur, Lost Cities, and Splendor—where fixed determinization strategies hinder optimal use of computational resources. To overcome this limitation, the authors propose a dual-axis dynamic resource allocation mechanism integrated within determinized MCTS: it adaptively adjusts the number of determinization trees while allocating simulations non-uniformly based on knowledge gain, thereby enabling real-time optimization of the computational budget. Experimental results demonstrate that, under identical iteration and time constraints, the proposed approach yields statistically significant improvements in win rates, confirming its effectiveness and superiority in decision-making under uncertainty.

Adversarial Board GamesDynamic Resource AllocationEnsemble Determinization MCTS

Hot Scholars

YX

Ying Xiong

Clausthal University of Technology
Petroleum geologySedimentologyGeochemistry
PW

Pengcheng Wu

Volvo Cars / KTH Royal Institute of Technology
motion planning and control of roboticsstate estimation and uncertainty quantificationsafety
AP

Alberto Pozanco

J.P. Morgan AI Research
Automated PlanningHeuristic SearchKnowledge RepresentationOptimization
WL

Weiqi Luo

School of Computer, Sun Yat-Sen Univ. Guangzhou, P.R. China
Steganography and SteganalysisMultimedia ForensicsAI Security
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning