🤖 AI Summary
This paper addresses the challenge in Monte Carlo Tree Search (MCTS) of effectively incorporating prior knowledge into the tree policy while balancing exploration efficiency and theoretical guarantees. To this end, we propose Inverse-RPO, a general framework that establishes, for the first time, a principled derivation paradigm for prior-aware UCT algorithms. Inverse-RPO systematically extends variance-aware, prior-free UCB methods (e.g., UCB-V) into variance-sensitive prior tree policies, ensuring both theoretical rigor and empirical superiority. By integrating regularized policy modeling and enhanced variance estimation, our method achieves zero-overhead integration within the mctx library. Empirical evaluation across multiple benchmark tasks demonstrates consistent and significant improvements over PUCT. The implementation is fully open-sourced, enabling straightforward reproduction and facilitating further research extensions.
📝 Abstract
Monte Carlo Tree Search (MCTS) has profoundly influenced reinforcement learning (RL) by integrating planning and learning in tasks requiring long-horizon reasoning, exemplified by the AlphaZero family of algorithms. Central to MCTS is the search strategy, governed by a tree policy based on an upper confidence bound (UCB) applied to trees (UCT). A key factor in the success of AlphaZero is the introduction of a prior term in the UCB1-based tree policy PUCT, which improves exploration efficiency and thus accelerates training. While many alternative UCBs with stronger theoretical guarantees than UCB1 exist, extending them to prior-based UCTs has been challenging, since PUCT was derived empirically rather than from first principles. Recent work retrospectively justified PUCT by framing MCTS as a regularized policy optimization (RPO) problem. Building on this perspective, we introduce Inverse-RPO, a general methodology that systematically derives prior-based UCTs from any prior-free UCB. Applying this method to the variance-aware UCB-V, we obtain two new prior-based tree policies that incorporate variance estimates into the search. Experiments indicate that these variance-aware prior-based UCTs outperform PUCT across multiple benchmarks without incurring additional computational cost. We also provide an extension of the mctx library supporting variance-aware UCTs, showing that the required code changes are minimal and intended to facilitate further research on principled prior-based UCTs. Code: github.com/Max-We/inverse-rpo.