🤖 AI Summary
This study addresses the challenge of suboptimal strategy selection in dynamic partially observable games, where opponents' reasoning levels are typically unknown. To overcome this, we propose a Reinforcement Learning-Based Orchestration (RLBO) framework that treats a Level-K strategy pool as schedulable resources rather than classification labels. Through Partially Observable Markov Game (POMG) modeling and training with aggregated offline and online data, an orchestrator adaptively selects strategies to maximize expected returns. This work transcends the limitations of conventional probabilistic classification-based deployment by demonstrating that strategy levels should be dynamically scheduled as resources. Pursuit-evasion experiments show that RLBO significantly outperforms classification-based methods while achieving performance comparable to end-to-end policies with substantially fewer training steps.
📝 Abstract
Level-$K$ reasoning generates a hierarchy of policies specialized to opponents with different reasoning levels. When an opponent's level is unknown, a common deployment rule estimates that level and selects the corresponding response. In a dynamic, partially-observed game, this selection is repeated, with each choice shaping subsequent states and observations. The response associated with the most likely opponent level need not maximize expected return from the current history. We formulate this deployment problem as dynamic orchestration of a fixed, pretrained policy library in a partially observable Markov game. We compare classification-based orchestrators (CBOs) trained using offline data or on-policy data aggregation with a reinforcement-learning-based orchestrator (RLBO) trained to maximize expected discounted return. In pursuit-evasion experiments, on-policy training improves classification and return, yet RLBO achieves higher return than the on-policy and offline CBOs. Given a pretrained library, RLBO also reaches performance comparable to a policy trained directly over the pursuers'action space with fewer training timesteps. These findings support treating a hierarchy of level-$K$ policies as a resource for orchestration, not a prescription for deployment.