๐ค AI Summary
This work addresses the challenge of simultaneously modeling fine-grained event-level timing and long-range structural coherence in audio-driven music game level generation. To this end, we propose a multimodal sequence-to-sequence approach that conditions on both audio segments and level metadata, employing an event-based tokenized representation that explicitly encodes beat-aligned actions and their relative temporal offsetsโthereby overcoming the limitations of conventional frame-level representations. Built upon a Transformer architecture, our model jointly learns the alignment between audio and event sequences to generate levels with precise rhythmic fidelity. Experimental results demonstrate that the proposed method significantly outperforms frame-level baselines on event-level evaluation metrics and enables systematic analysis of how audio cues contribute to rhythm-aligned prediction.
๐ Abstract
Procedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. As a result, it is hard to describe event-level timing relations and longer-range structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.