keypoint-guided mcts

Designs and implements Monte Carlo Tree Search (MCTS) algorithms that use keypoints—salient states or landmark features—as proposals and heuristics to bias node expansion, selection, rollout, and value estimation. Work includes defining keypoint representations and matching, integrating those cues into priors and reward-guided trajectory selection, analyzing search behavior and convergence, and producing means to inspect node-to-node decision traces.

keypoint-guidedmcts

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

本文探讨了蒙特卡洛树搜索(MCTS)与每访问蒙特卡洛控制方法在轨迹生成和动作价值更新层面本质上的一致性,指出MCTS可视为以搜索语言表达的每访问蒙特卡洛控制。

Action-Value UpdatingEvery-Visit Monte Carlo ControlMonte Carlo Tree Search

Bilevel MCTS for Amortized O(1) Node Selection in Classical Planning

Aug 11, 2025
MA
Masataro Asai
🏛️ MIT-IBM Watson AI Lab | IBM Research Cambridge

In classical planning, Monte Carlo Tree Search (MCTS) node selection incurs O(log N) time overhead due to reliance on a tree-structured OPEN list, causing severe efficiency degradation with increasing search depth. To address this, we propose a two-level MCTS framework: an upper level maintains a lightweight tree structure, while the lower level performs best-first search—within a bounded expansion budget—over candidate leaf nodes, achieving amortized O(1) node selection. This work introduces, for the first time, amortized constant-time node selection in MCTS. Furthermore, tree folding is integrated to compress redundant nodes, eliminating the need for explicit maintenance of large ordered data structures. Experiments demonstrate that our approach significantly reduces node selection latency in deep-search regimes, outperforming both standard MCTS and queue-based OPEN list implementations in overall planning efficiency—while preserving solution quality and convergence guarantees.

Achieve amortized O(1) node selection in MCTSAddress large search depth bottleneck in MCTSReduce logarithmic runtime complexity in planning

To address the lack of interpretability and susceptibility to tactical traps in Monte Carlo Tree Search (MCTS) for multi-player board games, this paper proposes Minimax-Augmented MCTS (MA-MCTS). Our method integrates a shallow minimax search into the rollout phase to enhance local robustness and—novelly—incorporates process mining techniques into MCTS policy modeling to reconstruct decision-making processes with explicit interpretability. This breaks the traditional “black-box” limitation of MCTS by generating human-readable strategy execution flowcharts, enabling real-time diagnosis and intervention. In 3v3 checkers experiments, MA-MCTS achieves a 23.6% improvement in critical-move identification accuracy and a 31.4% increase in tactical-trap avoidance rate. The main contributions are: (1) the first synergistic enhancement framework unifying MCTS and minimax for multi-player settings; and (2) a process-mining–based paradigm for interpretable MCTS policy modeling.

Combines MCTS with minimax to avoid tactical trapsExplains MCTS decision-making in board gamesUses process mining to analyze 3v3 checkers strategies

Solving Stochastic Orienteering Problems with Chance Constraints Using a GNN Powered Monte Carlo Tree Search

Sep 06, 2024
MA
Marcos Abel Zuzu'arregui
🏛️ University of California, Merced

This paper addresses the Stochastic Orienteering Problem with Chance Constraints (SOP-CC): maximizing collected rewards under stochastic travel costs and a deterministic budget, while strictly bounding the probability of budget violation below a given threshold. We propose an online, anytime Monte Carlo Tree Search (MCTS) framework that, for the first time, integrates Message-Passing Graph Neural Networks (MPNNs) into the rollout phase to jointly model action utility and path failure probability—enabling efficient, robust real-time decision-making. This design accelerates search convergence and supports cross-distribution generalization. Experiments demonstrate controlled reward loss on challenging instances and strong generalization to unseen scenarios. Our approach establishes a novel paradigm for sequential chance-constrained optimization under uncertainty.

Maximizing collected reward under stochastic travel costsSolving stochastic orienteering with chance-constrained travel budgetUsing GNN-powered MCTS for online planning and execution

Latest Papers

What's happening recently
View more

This work addresses the problem of determining whether the root value in a Monte Carlo tree search exceeds a given threshold, where internal nodes alternate between MAX and MIN operations and leaf node values correspond to the means of unknown distributions. To tackle this, the authors propose a δ-correct sequential sampling algorithm built upon the Track-and-Stop framework, featuring an innovative ratio-corrected D-Tracking strategy for arm selection. The method preserves asymptotic optimality in sample complexity while substantially reducing the actual number of samples required in practice. Furthermore, it improves computational efficiency by lowering the per-round time complexity from linear to logarithmic. Empirical evaluations demonstrate the algorithm’s dual advantages in both sample efficiency and computational speed.

Decision ThresholdMonte Carlo Tree SearchOptimal Sample Complexity

This study addresses the high search costs of LLM agents in long-horizon tasks and the inability of standard MCTS to reuse decision feedback across trajectories. To this end, it proposes HyperMCTS, a training-free method that introduces hypergraph modeling to accumulate cross-trajectory normalized returns and designs a HyperUCT selection rule to aggregate overlapping evidence. By breaking prefix constraints to enable experience sharing, the approach integrates seamlessly into existing frameworks without additional training. Experiments on the DeepPlanning benchmark demonstrate that HyperMCTS improves accuracy by 2.3–7.3 percentage points, surpasses Claude Opus 4.6 with fewer API calls, and significantly enhances question-answering performance.

cross-trajectory decision accumulationLLM agentsLong-horizon tasks

This work proposes Particle Monte Carlo Tree Search (PMCTS), a novel parallel framework for Monte Carlo Tree Search (MCTS) that overcomes the inherent sequential limitations of traditional MCTS in parallel environments. PMCTS introduces a particle-based sampling mechanism that effectively integrates neural network policies and value functions, enabling large-scale parallel inference while preserving theoretical guarantees for policy improvement—the first such guarantee in parallel MCTS. The method maintains rigorous theoretical foundations and demonstrates substantial performance gains over existing heuristic parallel baselines, achieving superior scalability and consistent improvements across multiple tasks.

inference time scalingMonte Carlo Tree Searchparallelization

This work addresses the limited interpretability of existing MCTS-Minimax hybrid approaches for multi-agent decision-making and the tendency of standard Monte Carlo Tree Search (MCTS) to overlook critical actions or become trapped in local tactical optima. To enhance strategic depth, the authors propose embedding shallow full-width Minimax search within the rollout phase of MCTS. Furthermore, they introduce a novel integration of process mining techniques—such as Alpha Miner and Inductive Miner—with large language models to structurally model agent behavior trajectories and generate human-readable causal and root-cause explanations. Experimental validation in a small checkers environment demonstrates the effectiveness of the approach, offering a scalable framework for explainable hybrid agents in complex strategic scenarios.

Decision-MakingExplainabilityMCTS

This work addresses the over-exploitation issue in Monte Carlo Tree Search (MCTS) for automated heuristic design under limited computational budgets. To mitigate this, the authors propose Clade-AHD, a novel framework that introduces, for the first time, a clade-level Bayesian belief modeling mechanism. Instead of relying on traditional node-level point estimates, Clade-AHD aggregates subtree evaluations using Beta distributions and employs Thompson sampling to guide exploration decisions. This approach effectively balances exploration and exploitation in sparse and noisy evaluation environments, substantially improving heuristic quality. Experimental results demonstrate that Clade-AHD outperforms existing methods on complex combinatorial optimization tasks while significantly reducing computational overhead.

Automatic Heuristic Designcomputational budgetLarge Language Model

Hot Scholars

GS

Gokul Subramanian Ravi

Computer Science and Engineering, University of Michigan
Computer ArchitectureQuantum Computing