Is Monte Carlo Tree Search Just Every-Visit Monte Carlo Control?

📅 2026-08-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了蒙特卡洛树搜索(MCTS)与每访问蒙特卡洛控制方法在轨迹生成和动作价值更新层面本质上的一致性,指出MCTS可视为以搜索语言表达的每访问蒙特卡洛控制。
📝 Abstract
Monte Carlo Tree Search (MCTS) and every-visit Monte Carlo (MC) control are usually presented as different methods. MCTS is described in the language of search (selection, expansion, simulation, and backup), whereas MC control is described in the language of reinforcement learning (trajectory sampling, return estimation, action-value updating, and policy improvement). This note argues that, at the level of trajectory generation and action-value updating, the distinction is largely terminological. The tree policy and rollout policy can be viewed as the learned and not-yet-learned parts of a single evolving policy; expansion corresponds to first visit and initialization; and backup is the ordinary every-visit Monte Carlo update. Under this interpretation, the four stages of MCTS reduce to two basic operations: trajectory sampling under the current policy and every-visit Monte Carlo updating. In this sense, MCTS is simply every-visit Monte Carlo control expressed in the language and data structure of search. The purpose of this note is expository: to make this equivalence explicit and easier to recognize.
Problem

Research questions and friction points this paper is trying to address.

Monte Carlo Tree Search
Every-Visit Monte Carlo Control
Trajectory Generation
Action-Value Updating
Innovation

Methods, ideas, or system contributions that make the work stand out.

Monte Carlo Tree Search
Every-Visit Monte Carlo Control
Policy Evolution
Trajectory Sampling
Action-Value Updating