Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling

๐Ÿ“… 2026-07-10
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of simultaneously modeling fine-grained event-level timing and long-range structural coherence in audio-driven music game level generation. To this end, we propose a multimodal sequence-to-sequence approach that conditions on both audio segments and level metadata, employing an event-based tokenized representation that explicitly encodes beat-aligned actions and their relative temporal offsetsโ€”thereby overcoming the limitations of conventional frame-level representations. Built upon a Transformer architecture, our model jointly learns the alignment between audio and event sequences to generate levels with precise rhythmic fidelity. Experimental results demonstrate that the proposed method significantly outperforms frame-level baselines on event-level evaluation metrics and enables systematic analysis of how audio cues contribute to rhythm-aligned prediction.
๐Ÿ“ Abstract
Procedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. As a result, it is hard to describe event-level timing relations and longer-range structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.
Problem

Research questions and friction points this paper is trying to address.

procedural generation
music-game level
event-based modeling
audio-conditioned generation
timed gameplay events
Innovation

Methods, ideas, or system contributions that make the work stand out.

event-based modeling
token sequence
music-conditioned generation
procedural level generation
beat-aligned timing
๐Ÿ”Ž Similar Papers
No similar papers found.
K
Ke Zhang
Japan Advanced Institute of Science and Technology
C
Chu-Hsuan Hsueh
Japan Advanced Institute of Science and Technology
K
Kokolo Ikeda
Japan Advanced Institute of Science and Technology