Score
Designs and implements optimization and planning algorithms that combine Monte Carlo Tree Search with methods for handling both discrete and continuous decision variables, using simulation-based rollouts to evaluate long-horizon action or parameter sequences and identify high-value trajectories.
Variable selection heuristics in branch-and-bound (B&B) for mixed-integer linear programming (MILP) suffer from low efficiency and poor generalization. Method: We propose the first model-based reinforcement learning (MBRL) framework for B&B, which learns a dynamic environment model of the B&B search process and integrates Monte Carlo tree search (MCTS) for forward-looking, adaptive branching decisions—yielding both interpretability and sample efficiency. Contribution/Results: Our approach overcomes the dual limitations of static heuristics and model-free RL in modeling capacity and data efficiency. Evaluated on four standard MILP benchmarks, it consistently outperforms state-of-the-art RL-driven branching policies, achieving significant reductions in solving time, number of explored nodes, and optimality gap. These results validate the effectiveness and scalability of model-guided planning for combinatorial optimization.
Diffusion models exhibit strong generative capabilities for planning tasks but suffer from non-scalable test-time computation (TTC): performance plateaus rather than improving monotonically with increased inference budget. To address this, we propose Diffusion-MCTS—the first framework integrating Monte Carlo Tree Search (MCTS) into the diffusion paradigm. It reformulates the denoising process as a tree-structured search, incorporating value-guided node selection, conditional resampling, and backtracking from suboptimal branches to enable iterative evaluation, pruning, and refinement. This design supports dynamic exploration-exploitation trade-offs, overcoming the fundamental TTC bottleneck of conventional diffusion-based planners. Empirically, on long-horizon planning tasks, solution quality improves monotonically with computational budget—significantly outperforming diffusion baselines—and demonstrates both TTC scalability and robustness.
This work addresses the problem of precisely steering state distributions in nonlinear dynamical systems. Methodologically, it formulates distributional control as a discrete-time Markov decision process (MDP) with explicit state-distribution constraints, where state-feedback policies serve as the action space. A novel, differentiable, and computationally efficient distribution distance metric is introduced, and—crucially—Monte Carlo tree search (MCTS) is extended for the first time to distribution-guided control under arbitrary (including strongly nonlinear) dynamics, eliminating reliance on linearization approximations. Experiments across diverse linear and nonlinear systems demonstrate that the proposed framework achieves significantly higher fidelity in matching target distributions, exhibits robust performance, and consistently outperforms baseline methods based on linearization or moment-matching techniques.
Traditional scenario tree construction methods prioritize probabilistic fidelity but often fail to guarantee strong downstream control performance. This work proposes a novel control-performance-driven paradigm, formulating scenario allocation as an attention-based policy optimization problem. By leveraging reinforcement learning, the approach directly maximizes closed-loop control returns and generates compact branches that emphasize high-impact events under a fixed tree structure. The method incorporates an asymmetric critic to stabilize training and is evaluated against baselines such as Wasserstein reduction. In a risk-averse battery arbitrage task, it consistently achieves the highest returns across varying forecast ensemble sizes, significantly outperforming classical scenario reduction and certainty-equivalent control while demonstrating superior tail-risk management.
To address the challenge of cross-thread statistical aggregation in root-parallel Monte Carlo tree search (MCTS) for continuous action spaces, this paper introduces Gaussian process regression (GPR) into the root-parallel MCTS framework for the first time. We propose a GPR-based value estimation method that enables reliable value extrapolation for unsampled actions and explicitly models action-space continuity. The approach preserves online planning efficiency while effectively fusing local statistics from multiple threads. Evaluated on six standard continuous control benchmarks, it significantly outperforms existing aggregation strategies—including weighted averaging and max-value aggregation—yielding substantial improvements in policy quality and planning stability, with only marginal increases in inference overhead. Our core contribution is the novel application of GPR to cross-thread value aggregation in root-parallel MCTS, establishing a new paradigm for efficient online planning in continuous action domains.
This work addresses the lack of finite-time theoretical guarantees for Monte Carlo tree search (MCTS) in partially observable Markov decision processes (POMDPs) with continuous observation spaces. To this end, the authors propose Voro-POMCPOW, an algorithm that extends the UCB exploration mechanism and introduces an adaptive observation-space partitioning framework based on Voronoi cells. This approach effectively handles action-selection dependencies and non-stationarity while preserving the original observation generator and maintaining a finite branching factor. The paper provides the first finite-time theoretical analysis for MCTS in continuous POMDPs, establishing high-probability polynomial concentration bounds on root-node value estimates and finite-time bounds on partitioning error. Empirical results demonstrate that Voro-POMCPOW achieves competitive performance while offering strong theoretical guarantees and is readily extensible to continuous MDPs.
Existing Monte Carlo Tree Diffusion (MCTD) methods are constrained by fixed training trajectory lengths, supporting only single-trajectory local search without global planning capability. To address this limitation, we propose Compositional MCTD—a novel framework that for the first time formalizes planning as compositional reasoning across trajectory segments, implemented via three types of combiners: online, distributed, and pre-planning. Our approach integrates diffusion models with Monte Carlo tree search, incorporating parallel exploration, plan-graph caching, and global search strategies to enable efficient, scalable sequential decision-making for long-horizon tasks. Experiments demonstrate significant improvements in success rates and trajectory coherence on long-range tasks, alongside markedly enhanced inference efficiency.