Frame-Level Pansori Mode Classification with Complementary Audio Representations

πŸ“… 2026-08-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of characterizing Pansori’s musical modes (jo), which cannot be adequately defined by scale alone and instead require integration of multidimensional features such as pitch-class sets, microtonal ornamentation (sigimsae), and vocal timbre. The authors construct the first expert-annotated, frame-level dataset spanning 46 hours and covering all five traditional batang modes, employing four complementary representations: Mel-spectrograms, fundamental frequency contours, MIDI piano rolls, and embeddings from a multicultural self-supervised encoder. To prevent shortcut learning, they implement rigorous track-level and temporal-level data splits, combined with source separation and cross-modal analysis. Their model achieves robust generalization, with F1 scores on full-work held-out tests dropping only 2.1–3.6 points. This work presents the first large-scale, fine-grained mode recognition system for Pansori, demonstrating that the model captures essential modal characteristics rather than memorizing specific pieces, with qualitative results aligning closely with musicological literature and contemporary score analyses.
πŸ“ Abstract
Pansori is a traditional Korean vocal genre whose mode system (jo) is defined not by scale alone but by the entanglement of pitch collection, microtonal ornament (sigimsae), and vocal timbre. In this study, we introduce a 46-hour frame-level pansori mode annotation, expert-labeled across all five canonical batang, and evaluate four complementary input representations (mel spectrogram, F0 contour, MIDI piano roll, and a multi-cultural SSL encoder) under two split strategies designed to detect shortcut learning. Across the three well-represented modes, performance degrades by only 2.1--3.6 points of F1 when entire works are held out, indicating that the models learn mode-relevant features rather than memorizing repertoire. Per-class results further show that source separation removes the percussion cue on which changjo depends, and that generic multi-cultural pre-training fails specifically on the Ujo--Gyemyeonjo distinction. Qualitative analysis of cross-modal disagreement recovers musicologically documented phenomena and agrees with published score-based analyses of modern changjak pansori.
Problem

Research questions and friction points this paper is trying to address.

Pansori
mode classification
audio representation
frame-level annotation
jo
Innovation

Methods, ideas, or system contributions that make the work stand out.

frame-level annotation
complementary audio representations
shortcut learning detection
multi-cultural SSL encoder
pansori mode classification
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.