Discrete Audio Representations for Automated Audio Captioning

📅 2025-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the insufficient semantic modeling of discrete audio tokens and the performance degradation caused by unsupervised tokenization in automatic audio captioning (AAC). To overcome the semantic gap inherent in conventional unsupervised tokenization, we propose the first supervised, learnable tokenizer explicitly designed for audio event understanding. Our method leverages fine-grained audio labels as supervisory signals to explicitly model semantically meaningful audio events. Integrating label-guided training with discrete representation learning, we establish an end-to-end AAC evaluation framework on the Clotho dataset. Experimental results demonstrate that our learned tokens yield up to a 2.3-point improvement in BLEU-4 over various unsupervised tokenization baselines, significantly enhancing semantic fidelity and caption quality. This represents the first supervised tokenization approach tailored to audio event semantics in AAC, effectively bridging the semantic modeling gap in discrete audio representation learning.

Technology Category

Machine Learning: Unsupervised & Self-Supervised LearningNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.Computer Vision: Diffusion Models for Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Discrete audio representations, termed audio tokens, are broadly categorized into semantic and acoustic tokens, typically generated through unsupervised tokenization of continuous audio representations. However, their applicability to automated audio captioning (AAC) remains underexplored. This paper systematically investigates the viability of audio token-driven models for AAC through comparative analyses of various tokenization methods. Our findings reveal that audio tokenization leads to performance degradation in AAC models compared to those that directly utilize continuous audio representations. To address this issue, we introduce a supervised audio tokenizer trained with an audio tagging objective. Unlike unsupervised tokenizers, which lack explicit semantic understanding, the proposed tokenizer effectively captures audio event information. Experiments conducted on the Clotho dataset demonstrate that the proposed audio tokens outperform conventional audio tokens in the AAC task.
Problem

Research questions and friction points this paper is trying to address.

Exploring audio tokens' applicability in automated audio captioning
Addressing performance degradation in token-driven AAC models
Introducing supervised tokenizer for better audio event capture
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised audio tokenizer for semantic understanding
Comparative analysis of tokenization methods in AAC
Audio tokens outperform conventional ones in AAC
🔎 Similar Papers
No similar papers found.
Jingguang Tian
Jingguang Tian
Midea AI Research Institute, Shanghai, China
speech LLMaudio understanding
Haoqin Sun
Haoqin Sun
Nankai University
Affective computingSpeech signal processingAudio understanding
X
Xinhui Hu
Hithink RoyalFlush AI Research Institute, Hangzhou, China
X
Xinkang Xu
Hithink RoyalFlush AI Research Institute, Hangzhou, China