🤖 AI Summary
AudioSet labels suffer from low accuracy and incomplete coverage, limiting downstream audio classification performance. To address this, we propose a three-stage relabeling framework leveraging general-purpose audio-language foundation models. Our method introduces a cross-modal prompt chaining mechanism that decouples audio understanding, label generation, and semantic alignment, while incorporating semantic consistency verification to ensure label quality. Fully automated and annotation-free, the framework significantly improves both label accuracy and structural coherence. Extensive evaluation on state-of-the-art models—including AST, PANNs, SSAST, and AudioMAE—demonstrates consistent average improvements of 2.1–4.7 percentage points in audio classification accuracy, with strong generalization across architectures. Our key contribution is the first systematic application of prompt chaining engineering to audio label reconstruction, establishing a scalable, high-quality paradigm for building structured audio datasets.
📝 Abstract
AudioSet is a widely used benchmark in the audio research community and has significantly advanced various audio-related tasks. However, persistent issues with label accuracy and completeness remain critical bottlenecks that limit performance in downstream applications.To address the aforementioned challenges, we propose a three-stage reannotation framework that harnesses general-purpose audio-language foundation models to systematically improve the label quality of AudioSet. The framework employs a cross-modal prompting strategy, inspired by the concept of prompt chaining, wherein prompts are sequentially composed to execute subtasks (audio comprehension, label synthesis, and semantic alignment). Leveraging this framework, we construct a high-quality, structured relabeled version of AudioSet-R. Extensive experiments conducted on representative audio classification models--including AST, PANNs, SSAST, and AudioMAE--consistently demonstrate substantial performance improvements, thereby validating the generalizability and effectiveness of the proposed approach in enhancing label reliability.The code is publicly available at: https://github.com/colaudiolab/AudioSet-R.