Constrained Command-Conditioned Reinforcement Learning with Bandit Strategy Selection in Real-Time Strategy Games

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited robustness of deep reinforcement learning agents in real-time strategy games when facing out-of-distribution opponents by proposing a hierarchical architecture that decouples strategic command generation from unit-level control. At the lower level, a constrained instruction-conditioned executor handles unit control, complemented by a policy estimation mechanism grounded in in-game observations rather than opponent identification. At the upper level, a Thompson sampling-based multi-armed bandit serves as a strategist to dynamically select discrete commands. Experimental evaluations in the MicroRTS environment demonstrate that this system achieves significantly higher win rates against most strong adversaries compared to a flat Proximal Policy Optimization baseline, effectively enhancing cross-opponent generalization capabilities.
πŸ“ Abstract
Deep reinforcement learning agents reach strong performance in real-time strategy games but can be brittle against opponents outside their training distribution. Separating strategic command selection from learned unit control allows different strategies to be selected for different opponents while reusing the same execution policy. This requires an executor that can follow different commands and measurable criteria for assessing whether it does so. We introduce a constrained command-conditioned Proximal Policy Optimization (PPO) policy, the executor, for MicroRTS, a real-time strategy environment. Discrete commands specify strategic objectives and behavioral requirements for economy, army composition, military posture, and worker policy over multiple environment steps; the executor determines the unit-level actions used to fulfill them. A Thompson-sampling bandit acts as the strategist, selecting command tuples from an estimate of the opponent's strategy built from in-game observations rather than opponent identity. In a controlled comparison with a flat PPO baseline trained with the same architecture, budget, curriculum and self-play league, the strategist-executor system wins significantly more often against three of the four strongest opponents on a training map, including the two strongest held-out ones (0.55 to 0.97 and 0.01 to 0.34), with no significant difference against the others.
Problem

Research questions and friction points this paper is trying to address.

Real-Time Strategy Games
Deep Reinforcement Learning
Out-of-Distribution Generalization
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Command-Conditioned PPO
Thompson Sampling Bandit
Strategy-Executor Separation
Real-Time Strategy Games
Constrained Reinforcement Learning
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
N
Nick Leenders
NLR Royal Netherlands Aerospace Centre, Amsterdam, The Netherlands
Roy Lindelauf
Roy Lindelauf
Data Science Center of Excellence, Faculty of Military Sciences, Breda, NL
J
Joost van Oijen
NLR Royal Netherlands Aerospace Centre, Amsterdam, The Netherlands
B
Boris Λ‡Cule
Departement Intelligent Systems, Tilburg University, Tilburg, NL