automatic audio annotation

Designs and builds automated audio annotation systems and pipelines that convert raw audio into structured labels (event classes, timestamps, confidence scores) and generate large-scale training annotations. These systems scale to hundreds of classes and diverse recording conditions, handling metadata, label quality control, and batch or streaming processing to produce consistent annotations for model training and evaluation.

automaticaudioannotation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the scarcity of high-quality labeled data for audio classification in domain-specific scenarios such as domestic environments. To this end, the authors propose TriA Pipeline, the first large-scale automated audio annotation framework tailored for such settings, which efficiently generates the TriA dataset comprising 2,130 hours of audio across 431 event classes. Furthermore, they introduce a prior knowledge–guided data filtering mechanism to construct a refined subset, TriA_GK. Experimental results on three household audio classification tasks demonstrate that models trained on TriA_GK achieve relative improvements of 3.97% in average accuracy and 3.35% in Macro-F1 score over baseline methods, highlighting the effectiveness of the proposed approach.

audio classificationaudio datasetsdata annotation

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation

Aug 21, 2025
YS
Yulin Sun
🏛️ National University of Defense Technology

AudioSet labels suffer from low accuracy and incomplete coverage, limiting downstream audio classification performance. To address this, we propose a three-stage relabeling framework leveraging general-purpose audio-language foundation models. Our method introduces a cross-modal prompt chaining mechanism that decouples audio understanding, label generation, and semantic alignment, while incorporating semantic consistency verification to ensure label quality. Fully automated and annotation-free, the framework significantly improves both label accuracy and structural coherence. Extensive evaluation on state-of-the-art models—including AST, PANNs, SSAST, and AudioMAE—demonstrates consistent average improvements of 2.1–4.7 percentage points in audio classification accuracy, with strong generalization across architectures. Our key contribution is the first systematic application of prompt chaining engineering to audio label reconstruction, establishing a scalable, high-quality paradigm for building structured audio datasets.

Addressing label quality bottlenecks in audio classificationEnhancing label reliability for downstream audio tasksImproving AudioSet label accuracy and completeness

Taming Data and Transformers for Scalable Audio Generation

Jun 27, 2024
MH
Moayed Haji Ali
🏛️ Rice University | Snap Inc.

To address three key bottlenecks in environmental sound generation—data scarcity, low-quality captions, and limited model scalability—we propose a data-model co-scaling paradigm. First, we construct AutoReCap-XL, the first ultra-large-scale, high-quality audio-text dataset comprising 47 million segments. Second, we design AutoCap, a high-fidelity automatic captioning model integrating Q-Former with audio metadata and incorporating synthetic caption distillation. Third, we introduce GenAu, a scalable Transformer architecture (1.25B parameters) tailored for long-duration audio generation, enhanced via contrastive learning and architectural scaling optimization. Experiments show AutoCap achieves a CIDEr score of 83.2 (+3.2% absolute improvement), while GenAu outperforms prior models significantly (FAD ↓4.7%, IS ↑11.1%, CLAP ↑13.5%). All code, models, and datasets are publicly released.

Addressing data scarcity in ambient audio generationEnhancing scalability of transformer-based audio modelsImproving caption quality for audio-text datasets

Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging

Mar 26, 2025
LT
Ludovic Tuncay
🏛️ IRIT | Université de Toulouse | CNRS | Toulouse INP

AudioSet suffers from ontology-driven label inconsistency: audio events that should be positive instances are frequently mislabeled as negative, resulting in systematic under-annotation. To address this, we propose Hierarchical Label Propagation (HLP), the first method to explicitly incorporate the audio event ontology structure into label correction—deterministically propagating positive labels upward along the ontology hierarchy. HLP is architecture-agnostic, seamlessly integrating with mainstream models including CNNs (CNN6, ConvNeXT) and Transformers (PaSST), without requiring model retraining. Experiments demonstrate that HLP increases positive label density from 1.98 to 2.39 per clip, covering 109 classes. It consistently improves mean Average Precision (mAP) on both AudioSet and FSD50K, with more pronounced gains for smaller models (e.g., +1.2% mAP for CNN6), revealing a synergistic interaction between data quality enhancement and model capacity.

Address inconsistent annotations in AudioSet audio taggingImprove model performance across architectures with HLPPropagate labels hierarchically to correct mislabeled categories

SoundCollage: Automated Discovery of New Classes in Audio Datasets

Oct 30, 2024
RC
Ryuhaerang Choi
🏛️ KAIST | Nokia Bell Labs | University of Glasgow

Unsupervised discovery of novel acoustic categories in audio datasets—without prior category definitions or manual annotations—remains a significant challenge. Method: We propose an end-to-end, reusable unsupervised audio class discovery framework integrating time-frequency decomposition preprocessing, self-supervised representation learning, and clustering-driven model feedback labeling. We further introduce “clarity”, a novel metric quantifying semantic consistency of discovered classes. Contribution/Results: The framework eliminates reliance on predefined categories or human labeling, enabling identification and structured organization of previously unknown sound types. Experiments demonstrate substantial improvements: downstream classifiers achieve +34.7% accuracy on discovered-class samples and +4.5% on an independent test set. The implementation is publicly available, establishing a new paradigm for sustainable, large-scale audio dataset mining.

AI Sound Identification AccuracyAutomatic Sound RecognitionNew Sound Discovery

Latest Papers

What's happening recently
View more

Existing evaluation methods for audio captioning struggle to accurately assess the fidelity of multimodal semantics and acoustic attributes in structured audio descriptions. This work proposes the first multi-axis evaluation framework tailored for structured audio captioning, integrating large language model (LLM)-based semantic judgments with deterministic acoustic metrics across five orthogonal dimensions: label sets, descriptive content, logical reasoning, numerical measurements, and spectral contours. The framework incorporates a controlled perturbation protocol to validate its ability to distinguish between semantic preservation and acoustic distortion. Experiments on the AudioCards dataset demonstrate that the proposed approach effectively differentiates semantically consistent paraphrases from genuine errors, significantly outperforming existing methods in both reliability and sensitivity.

audio descriptioncaption metricsevaluation framework

This work addresses the prevalent issue of class-dependent unreliable supervision in audio annotation data—manifested through label redundancy, confusion among similar classes, and weakened evidential support—which introduces bias during model training. To tackle this, the authors propose the Class-level Supervision Unreliability (CSU) framework, which explicitly models three types of non-missing-label supervision noise for the first time and learns adaptive supervision weights per class to dynamically modulate label credibility during training. Notably, CSU requires no modifications to model architecture or inference procedures. It leverages a hybrid training strategy combining real and synthetically generated audio and introduces a new benchmark, ESC-FreeGen50. Extensive experiments demonstrate that CSU significantly enhances the robustness of diverse models against various supervision noises on AudioSet and controlled settings, confirming its effectiveness and generalizability.

audio taggingclass-wise biaslabel noise

This work addresses the challenge in audio question answering where models often over-rely on textual priors or are misled by mismatched audio inputs. To mitigate this, the authors propose a diagnostic data curation strategy based on model confusion patterns. By probing model responses under normal, silent, and shuffled audio conditions, they identify and retain samples exhibiting strong audio dependence for fine-tuning. Using the Qwen3-Omni-30B-A3B-Instruct model, the approach integrates counterfactual audio testing, response normalization, and multi-model ensembling. Fine-tuning solely on the curated dataset yields a 67.27% accuracy on the official development set, significantly outperforming a local baseline of 65.90%, thereby demonstrating the method’s effectiveness in enhancing models’ reliance on genuine audio evidence.

audio dependencyaudio question answeringcounterfactual evaluation

High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.

Automatic AnnotationOut-of-Distribution GeneralizationSinging Voice Synthesis

Hot Scholars

HO

Hawau Olamide Toyin

PhD student at MBZUAI
Speech Synthesis and RecognitionStuttering SpeechNLPML
DB

Dmitry Bogdanov

Music Technology Group, Universitat Pompeu Fabra
Music Information RetrievalAudio Signal ProcessingMachine LearningSound and Music Computing
CH

Chao-Han Huck Yang

Sr. Research Scientist, NVIDIA Research
Robust Speech RecognitionLanguage ModelsPost-TrainingSequence Modeling
SC

Simon Colton

Prof. Computational Creativity at Queen Mary University of London, UK & Monash University, Australia
Computational CreativityAIGame DesignComputer Art