🤖 AI Summary
This study addresses the high search costs of LLM agents in long-horizon tasks and the inability of standard MCTS to reuse decision feedback across trajectories. To this end, it proposes HyperMCTS, a training-free method that introduces hypergraph modeling to accumulate cross-trajectory normalized returns and designs a HyperUCT selection rule to aggregate overlapping evidence. By breaking prefix constraints to enable experience sharing, the approach integrates seamlessly into existing frameworks without additional training. Experiments on the DeepPlanning benchmark demonstrate that HyperMCTS improves accuracy by 2.3–7.3 percentage points, surpasses Claude Opus 4.6 with fewer API calls, and significantly enhances question-answering performance.
📝 Abstract
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment interaction make search costly. Efficient search therefore requires effective reuse of trajectory feedback. Standard MCTS maintains prefix-specific statistics, without explicitly accumulating outcomes for decision groups that recur across different paths. To fill this gap, we propose HyperMCTS, a training-free method that augments an ordered MCTS tree with a cross-trajectory hypergraph. Hyperedges represent groups of canonical decisions and accumulate their observed returns within the current task. Our hypergraph-guided HyperUCT selection rule aggregates evidence from overlapping hyperedges into an action prior, allowing outcomes collected under one prefix to inform selection under another while preserving execution histories in the tree. On DeepPlanning, HyperMCTS improves average planning accuracy by 2.3--7.3 percentage points over the strongest baseline for each of three backbone models. It enables Qwen3.6-27B to outperform Claude Opus 4.6 (max) on Shopping Planning, while achieving higher accuracy with fewer LLM calls and output tokens than the evaluated MCTS-based baselines. SealQA experiments further demonstrate improvements in question answering.