STEMMA: Song-to-Stem Multi-Audio Reasoning for Large Audio Language Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing audio models in reasoning over production relationships among multiple audio sources by proposing a production-source-based multi-audio question-answering framework. Methodologically, this work introduces an innovative relation-first data construction strategy, wherein labels are directly derived from directory structures rather than generated by large language models, thereby enabling the automated construction of both datasets and benchmarks. Furthermore, multi-audio relational modeling is achieved through fine-tuning large audio-language models. Experimental results demonstrate that the proposed approach significantly enhances multi-audio structural reasoning performance while effectively preserving single-audio understanding capabilities.
📝 Abstract
Music understanding often requires comparing excerpts and reasoning about relationships among songs, sections, and stems. However, existing large audio-language models (LALMs) and music question-answering datasets typically operate on single recordings or compare independently sampled tracks with no known production relationship. We introduce STEMMA, a multi-audio music question-answering framework built around production provenance: whether excerpts originate from the same track or section, and which stems belong to which mixtures. Because such relations are sparse under conventional audio-first sampling, STEMMA adopts a relation-first construction strategy: it first specifies a target relation and then queries the catalog for excerpts that satisfy it and hard negatives that do not. Labels are determined directly from catalog provenance rather than generated by a language model from metadata. We build STEMMA-Bench for evaluation and a track-disjoint training set, STEMMA-Instruct. Fine-tuning two LALMs on STEMMA-Instruct improves multi-audio reasoning, with the largest gains on structural relations directly determined by the catalog, while preserving single-audio music understanding.
Problem

Research questions and friction points this paper is trying to address.

multi-audio reasoning
large audio-language models
music understanding
question-answering
production provenance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Audio Reasoning
Large Audio Language Models
Relation-First Construction
Music Question-Answering
Production Provenance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hoyeol Sohn
Graduate School of Culture Technology, KAIST, Daejeon, Republic of Korea
Wonil Kim
Wonil Kim
Neutune, Seoul, Republic of Korea
K
Keunhyoung Kim
Neutune, Seoul, Republic of Korea
S
Sangeun Kum
Neutune, Seoul, Republic of Korea
T
Taehyoung Kim
Neutune, Seoul, Republic of Korea
D
Dongjoo Moon
Neutune, Seoul, Republic of Korea
T
Theerasak Charoenchob
Neutune, Seoul, Republic of Korea
T
Teeratep Weerapang
Neutune, Seoul, Republic of Korea
Jongpil Lee
Jongpil Lee
Neutune, Seoul, Republic of Korea
Juhan Nam
Juhan Nam
KAIST
Music TechnologyMusic Information RetrievalAudio Signal ProcessingMusic Processing