🤖 AI Summary
This study addresses the lack of interpretability in AI strategies for complex games by proposing a novel method for generating executable programmatic policies. Inspired by human cognitive mechanisms, the approach integrates reinforcement learning with program synthesis to transform implicitly learned agent policies into executable programs based on action sequences. Experimental evaluations conducted in chess and grid-world environments demonstrate that the proposed method effectively extracts and generates efficient action sequences directly from data. Consequently, this approach significantly enhances the transparency and interpretability of AI decision-making processes. By bridging the gap between opaque policy representations and human-readable programmatic logic, this work establishes a new paradigm for game data analysis and explainable artificial intelligence.
📝 Abstract
As part of learning to play complex games, human players develop develop abstractions for concepts and strategies of gameplay consistent with game rules to improve their performance. These concepts are applied to explain other players' actions, and to inform their own actions in-game. Understanding other players' strategies is a crucial part of such improvement, but requires time and effort. In this paper, we propose a strategy similar to human cognition for training RL agents to synthesize learned strategies and policies as executable procedures based on sequences of gameplay actions. We present methods to automatically learn such programs to play chess and to solve tasks in a grid-based environment. We show that the learned strategies produce effective actions, and can be learned from gameplay data.