Variance-Aware Prior-Based Tree Policies for Monte Carlo Tree Search

📅 2025-12-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge in Monte Carlo Tree Search (MCTS) of effectively incorporating prior knowledge into the tree policy while balancing exploration efficiency and theoretical guarantees. To this end, we propose Inverse-RPO, a general framework that establishes, for the first time, a principled derivation paradigm for prior-aware UCT algorithms. Inverse-RPO systematically extends variance-aware, prior-free UCB methods (e.g., UCB-V) into variance-sensitive prior tree policies, ensuring both theoretical rigor and empirical superiority. By integrating regularized policy modeling and enhanced variance estimation, our method achieves zero-overhead integration within the mctx library. Empirical evaluation across multiple benchmark tasks demonstrates consistent and significant improvements over PUCT. The implementation is fully open-sourced, enabling straightforward reproduction and facilitating further research extensions.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Stochastic OptimizationSearch and Optimization: Sampling/Simulation-based Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingEconomics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystems
📝 Abstract
Monte Carlo Tree Search (MCTS) has profoundly influenced reinforcement learning (RL) by integrating planning and learning in tasks requiring long-horizon reasoning, exemplified by the AlphaZero family of algorithms. Central to MCTS is the search strategy, governed by a tree policy based on an upper confidence bound (UCB) applied to trees (UCT). A key factor in the success of AlphaZero is the introduction of a prior term in the UCB1-based tree policy PUCT, which improves exploration efficiency and thus accelerates training. While many alternative UCBs with stronger theoretical guarantees than UCB1 exist, extending them to prior-based UCTs has been challenging, since PUCT was derived empirically rather than from first principles. Recent work retrospectively justified PUCT by framing MCTS as a regularized policy optimization (RPO) problem. Building on this perspective, we introduce Inverse-RPO, a general methodology that systematically derives prior-based UCTs from any prior-free UCB. Applying this method to the variance-aware UCB-V, we obtain two new prior-based tree policies that incorporate variance estimates into the search. Experiments indicate that these variance-aware prior-based UCTs outperform PUCT across multiple benchmarks without incurring additional computational cost. We also provide an extension of the mctx library supporting variance-aware UCTs, showing that the required code changes are minimal and intended to facilitate further research on principled prior-based UCTs. Code: github.com/Max-We/inverse-rpo.
Problem

Research questions and friction points this paper is trying to address.

Deriving prior-based UCTs from prior-free UCBs systematically
Incorporating variance estimates into tree policies for improved performance
Enhancing Monte Carlo Tree Search with variance-aware prior-based methods
Innovation

Methods, ideas, or system contributions that make the work stand out.

Inverse-RPO method derives prior-based UCTs from UCBs
Variance-aware prior-based UCTs incorporate variance estimates
Minimal code changes enable variance-aware UCTs in libraries
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Maximilian Weichart
University of Regensburg