hybrid mcts optimization

Designs and implements optimization and planning algorithms that combine Monte Carlo Tree Search with methods for handling both discrete and continuous decision variables, using simulation-based rollouts to evaluate long-horizon action or parameter sequences and identify high-value trajectories.

hybridmctsoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Planning in Branch-and-Bound: Model-Based Reinforcement Learning for Exact Combinatorial Optimization

Nov 12, 2025
PS
Paul Strang
🏛️ EDF R&D | ENSTA Paris | CNAM Paris | ISAE-SUPAERO

Variable selection heuristics in branch-and-bound (B&B) for mixed-integer linear programming (MILP) suffer from low efficiency and poor generalization. Method: We propose the first model-based reinforcement learning (MBRL) framework for B&B, which learns a dynamic environment model of the B&B search process and integrates Monte Carlo tree search (MCTS) for forward-looking, adaptive branching decisions—yielding both interpretability and sample efficiency. Contribution/Results: Our approach overcomes the dual limitations of static heuristics and model-free RL in modeling capacity and data efficiency. Evaluated on four standard MILP benchmarks, it consistently outperforms state-of-the-art RL-driven branching policies, achieving significant reductions in solving time, number of explored nodes, and optimality gap. These results validate the effectiveness and scalability of model-guided planning for combinatorial optimization.

Improves variable selection heuristics in branch-and-bound algorithmsLearns branching strategies for Mixed-Integer Linear ProgrammingUses model-based reinforcement learning for combinatorial optimization

Monte Carlo Tree Diffusion for System 2 Planning

Feb 11, 2025
JY
Jaesik Yoon
🏛️ KAIST | Mila

Diffusion models exhibit strong generative capabilities for planning tasks but suffer from non-scalable test-time computation (TTC): performance plateaus rather than improving monotonically with increased inference budget. To address this, we propose Diffusion-MCTS—the first framework integrating Monte Carlo Tree Search (MCTS) into the diffusion paradigm. It reformulates the denoising process as a tree-structured search, incorporating value-guided node selection, conditional resampling, and backtracking from suboptimal branches to enable iterative evaluation, pruning, and refinement. This design supports dynamic exploration-exploitation trade-offs, overcoming the fundamental TTC bottleneck of conventional diffusion-based planners. Empirically, on long-horizon planning tasks, solution quality improves monotonically with computational budget—significantly outperforming diffusion baselines—and demonstrates both TTC scalability and robustness.

Enhance diffusion-based planning scalabilityImprove long-horizon task solution qualityIntegrate MCTS with diffusion models

Discrete-Time Distribution Steering using Monte Carlo Tree Search

Dec 09, 2024
AE
Alexandros E. Tzikas
🏛️ Stanford University

This work addresses the problem of precisely steering state distributions in nonlinear dynamical systems. Methodologically, it formulates distributional control as a discrete-time Markov decision process (MDP) with explicit state-distribution constraints, where state-feedback policies serve as the action space. A novel, differentiable, and computationally efficient distribution distance metric is introduced, and—crucially—Monte Carlo tree search (MCTS) is extended for the first time to distribution-guided control under arbitrary (including strongly nonlinear) dynamics, eliminating reliance on linearization approximations. Experiments across diverse linear and nonlinear systems demonstrate that the proposed framework achieves significantly higher fidelity in matching target distributions, exhibits robust performance, and consistently outperforms baseline methods based on linearization or moment-matching techniques.

Applying gradient-based solutions to distribution steering and ergodic controlComputing similarity between probability distributions in controlIntroducing interpretable distance based on cumulative distribution functions

Latest Papers

What's happening recently
View more

Traditional scenario tree construction methods prioritize probabilistic fidelity but often fail to guarantee strong downstream control performance. This work proposes a novel control-performance-driven paradigm, formulating scenario allocation as an attention-based policy optimization problem. By leveraging reinforcement learning, the approach directly maximizes closed-loop control returns and generates compact branches that emphasize high-impact events under a fixed tree structure. The method incorporates an asymmetric critic to stabilize training and is evaluated against baselines such as Wasserstein reduction. In a risk-averse battery arbitrage task, it consistently achieves the highest returns across varying forecast ensemble sizes, significantly outperforming classical scenario reduction and certainty-equivalent control while demonstrating superior tail-risk management.

control performancedecision-oriented constructionscenario tree

Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions

Dec 10, 2025
JX
Junlin Xiao
🏛️ The Chinese University of Hong Kong, Shenzhen | Oxford Robotics Institute | University of Oxford

To address the challenge of cross-thread statistical aggregation in root-parallel Monte Carlo tree search (MCTS) for continuous action spaces, this paper introduces Gaussian process regression (GPR) into the root-parallel MCTS framework for the first time. We propose a GPR-based value estimation method that enables reliable value extrapolation for unsampled actions and explicitly models action-space continuity. The approach preserves online planning efficiency while effectively fusing local statistics from multiple threads. Evaluated on six standard continuous control benchmarks, it significantly outperforms existing aggregation strategies—including weighted averaging and max-value aggregation—yielding substantial improvements in policy quality and planning stability, with only marginal increases in inference overhead. Our core contribution is the novel application of GPR to cross-thread value aggregation in root-parallel MCTS, establishing a new paradigm for efficient online planning in continuous action domains.

Aggregates statistics from parallel threads in continuous action spacesOutperforms existing methods across six domains with minimal time increaseUses Gaussian Process Regression for untried action value estimation

This work addresses the lack of finite-time theoretical guarantees for Monte Carlo tree search (MCTS) in partially observable Markov decision processes (POMDPs) with continuous observation spaces. To this end, the authors propose Voro-POMCPOW, an algorithm that extends the UCB exploration mechanism and introduces an adaptive observation-space partitioning framework based on Voronoi cells. This approach effectively handles action-selection dependencies and non-stationarity while preserving the original observation generator and maintaining a finite branching factor. The paper provides the first finite-time theoretical analysis for MCTS in continuous POMDPs, establishing high-probability polynomial concentration bounds on root-node value estimates and finite-time bounds on partitioning error. Empirical results demonstrate that Voro-POMCPOW achieves competitive performance while offering strong theoretical guarantees and is readily extensible to continuous MDPs.

continuous observation spacefinite-time analysisMCTS

Compositional Monte Carlo Tree Diffusion for Extendable Planning

Oct 24, 2025
JY
Jaesik Yoon
🏛️ KAIST | SAP | NYU

Existing Monte Carlo Tree Diffusion (MCTD) methods are constrained by fixed training trajectory lengths, supporting only single-trajectory local search without global planning capability. To address this limitation, we propose Compositional MCTD—a novel framework that for the first time formalizes planning as compositional reasoning across trajectory segments, implemented via three types of combiners: online, distributed, and pre-planning. Our approach integrates diffusion models with Monte Carlo tree search, incorporating parallel exploration, plan-graph caching, and global search strategies to enable efficient, scalable sequential decision-making for long-horizon tasks. Experiments demonstrate significant improvements in success rates and trajectory coherence on long-range tasks, alongside markedly enhanced inference efficiency.

Enables globally-aware reasoning across complete plan compositionsExtends planning beyond training trajectory length limitationsReduces search complexity through parallel exploration techniques

Hot Scholars

YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
MP

Markos Papageorgiou

Technical University of Crete
Automatic ControlOptimisationTransportationWater Systems
LQ

Lianhui Qin

UC San Diego, Computer Science and Engineering
Natural Language ProcessingMachine Learning
ZZ

Zuyuan Zhang

The George Washington University
OptimizationGame TheoryDecision-making algorithms