🤖 AI Summary
This work addresses the insufficient semantic modeling of discrete audio tokens and the performance degradation caused by unsupervised tokenization in automatic audio captioning (AAC). To overcome the semantic gap inherent in conventional unsupervised tokenization, we propose the first supervised, learnable tokenizer explicitly designed for audio event understanding. Our method leverages fine-grained audio labels as supervisory signals to explicitly model semantically meaningful audio events. Integrating label-guided training with discrete representation learning, we establish an end-to-end AAC evaluation framework on the Clotho dataset. Experimental results demonstrate that our learned tokens yield up to a 2.3-point improvement in BLEU-4 over various unsupervised tokenization baselines, significantly enhancing semantic fidelity and caption quality. This represents the first supervised tokenization approach tailored to audio event semantics in AAC, effectively bridging the semantic modeling gap in discrete audio representation learning.
📝 Abstract
Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.