π€ AI Summary
This study addresses the inter-frame inference latency of world model agents in real-time games by proposing an off-policy refinement framework based on an action-conditioned world model. Using Geometry Dash as the experimental platform, a compact environment model and an initial controller are first learned via behavioral cloning (BC). Subsequently, proximal policy optimization (PPO) is executed entirely offline over the frozen world model, enabling policy iteration without online interaction. This approach achieves a 60Hz real-time control loop on consumer-grade GPUs. Experimental results demonstrate that the refined policies significantly outperform the initial BC policies in survival time across multiple levels, effectively reconciling high-fidelity modeling with real-time performance requirements.
π Abstract
World-model agents are usually evaluated in simulators that can wait for the policy; live games impose the opposite constraint, requiring capture, prediction, and action before the next frame. We present DashVMC, which learns a compact, action-conditioned world model from approximately two hours of recorded Geometry Dash gameplay. To test whether the learned dynamics are actionable, a controller is initialized by behavioural cloning (BC) and refined with Proximal Policy Optimization (PPO) entirely in frozen-model rollouts, without further interaction with the live game. Across three controller seeds, the refined policies survive longer than their BC initializations on all three official levels and a held-out community layout. At deployment, the baseline skips visual generation and sustains a 60-Hz capture-to-action loop on a consumer GPU. Action-conditioned continuations and rollout diagnostics show that the model remains useful for control despite imperfect long-horizon fidelity.