🤖 AI Summary
This study addresses the limitation of existing audio models in reasoning over production relationships among multiple audio sources by proposing a production-source-based multi-audio question-answering framework. Methodologically, this work introduces an innovative relation-first data construction strategy, wherein labels are directly derived from directory structures rather than generated by large language models, thereby enabling the automated construction of both datasets and benchmarks. Furthermore, multi-audio relational modeling is achieved through fine-tuning large audio-language models. Experimental results demonstrate that the proposed approach significantly enhances multi-audio structural reasoning performance while effectively preserving single-audio understanding capabilities.
📝 Abstract
Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.