Score
Designs and builds automated audio annotation systems and pipelines that convert raw audio into structured labels (event classes, timestamps, confidence scores) and generate large-scale training annotations. These systems scale to hundreds of classes and diverse recording conditions, handling metadata, label quality control, and batch or streaming processing to produce consistent annotations for model training and evaluation.
This study addresses the scarcity of high-quality labeled data for audio classification in domain-specific scenarios such as domestic environments. To this end, the authors propose TriA Pipeline, the first large-scale automated audio annotation framework tailored for such settings, which efficiently generates the TriA dataset comprising 2,130 hours of audio across 431 event classes. Furthermore, they introduce a prior knowledge–guided data filtering mechanism to construct a refined subset, TriA_GK. Experimental results on three household audio classification tasks demonstrate that models trained on TriA_GK achieve relative improvements of 3.97% in average accuracy and 3.35% in Macro-F1 score over baseline methods, highlighting the effectiveness of the proposed approach.
AudioSet labels suffer from low accuracy and incomplete coverage, limiting downstream audio classification performance. To address this, we propose a three-stage relabeling framework leveraging general-purpose audio-language foundation models. Our method introduces a cross-modal prompt chaining mechanism that decouples audio understanding, label generation, and semantic alignment, while incorporating semantic consistency verification to ensure label quality. Fully automated and annotation-free, the framework significantly improves both label accuracy and structural coherence. Extensive evaluation on state-of-the-art models—including AST, PANNs, SSAST, and AudioMAE—demonstrates consistent average improvements of 2.1–4.7 percentage points in audio classification accuracy, with strong generalization across architectures. Our key contribution is the first systematic application of prompt chaining engineering to audio label reconstruction, establishing a scalable, high-quality paradigm for building structured audio datasets.
To address three key bottlenecks in environmental sound generation—data scarcity, low-quality captions, and limited model scalability—we propose a data-model co-scaling paradigm. First, we construct AutoReCap-XL, the first ultra-large-scale, high-quality audio-text dataset comprising 47 million segments. Second, we design AutoCap, a high-fidelity automatic captioning model integrating Q-Former with audio metadata and incorporating synthetic caption distillation. Third, we introduce GenAu, a scalable Transformer architecture (1.25B parameters) tailored for long-duration audio generation, enhanced via contrastive learning and architectural scaling optimization. Experiments show AutoCap achieves a CIDEr score of 83.2 (+3.2% absolute improvement), while GenAu outperforms prior models significantly (FAD ↓4.7%, IS ↑11.1%, CLAP ↑13.5%). All code, models, and datasets are publicly released.
AudioSet suffers from ontology-driven label inconsistency: audio events that should be positive instances are frequently mislabeled as negative, resulting in systematic under-annotation. To address this, we propose Hierarchical Label Propagation (HLP), the first method to explicitly incorporate the audio event ontology structure into label correction—deterministically propagating positive labels upward along the ontology hierarchy. HLP is architecture-agnostic, seamlessly integrating with mainstream models including CNNs (CNN6, ConvNeXT) and Transformers (PaSST), without requiring model retraining. Experiments demonstrate that HLP increases positive label density from 1.98 to 2.39 per clip, covering 109 classes. It consistently improves mean Average Precision (mAP) on both AudioSet and FSD50K, with more pronounced gains for smaller models (e.g., +1.2% mAP for CNN6), revealing a synergistic interaction between data quality enhancement and model capacity.
Unsupervised discovery of novel acoustic categories in audio datasets—without prior category definitions or manual annotations—remains a significant challenge. Method: We propose an end-to-end, reusable unsupervised audio class discovery framework integrating time-frequency decomposition preprocessing, self-supervised representation learning, and clustering-driven model feedback labeling. We further introduce “clarity”, a novel metric quantifying semantic consistency of discovered classes. Contribution/Results: The framework eliminates reliance on predefined categories or human labeling, enabling identification and structured organization of previously unknown sound types. Experiments demonstrate substantial improvements: downstream classifiers achieve +34.7% accuracy on discovered-class samples and +4.5% on an independent test set. The implementation is publicly available, establishing a new paradigm for sustainable, large-scale audio dataset mining.
Existing evaluation methods for audio captioning struggle to accurately assess the fidelity of multimodal semantics and acoustic attributes in structured audio descriptions. This work proposes the first multi-axis evaluation framework tailored for structured audio captioning, integrating large language model (LLM)-based semantic judgments with deterministic acoustic metrics across five orthogonal dimensions: label sets, descriptive content, logical reasoning, numerical measurements, and spectral contours. The framework incorporates a controlled perturbation protocol to validate its ability to distinguish between semantic preservation and acoustic distortion. Experiments on the AudioCards dataset demonstrate that the proposed approach effectively differentiates semantically consistent paraphrases from genuine errors, significantly outperforming existing methods in both reliability and sensitivity.
This work addresses the prevalent issue of class-dependent unreliable supervision in audio annotation data—manifested through label redundancy, confusion among similar classes, and weakened evidential support—which introduces bias during model training. To tackle this, the authors propose the Class-level Supervision Unreliability (CSU) framework, which explicitly models three types of non-missing-label supervision noise for the first time and learns adaptive supervision weights per class to dynamically modulate label credibility during training. Notably, CSU requires no modifications to model architecture or inference procedures. It leverages a hybrid training strategy combining real and synthetically generated audio and introduces a new benchmark, ESC-FreeGen50. Extensive experiments demonstrate that CSU significantly enhances the robustness of diverse models against various supervision noises on AudioSet and controlled settings, confirming its effectiveness and generalizability.
This work addresses the challenge in audio question answering where models often over-rely on textual priors or are misled by mismatched audio inputs. To mitigate this, the authors propose a diagnostic data curation strategy based on model confusion patterns. By probing model responses under normal, silent, and shuffled audio conditions, they identify and retain samples exhibiting strong audio dependence for fine-tuning. Using the Qwen3-Omni-30B-A3B-Instruct model, the approach integrates counterfactual audio testing, response normalization, and multi-model ensembling. Fine-tuning solely on the curated dataset yields a 67.27% accuracy on the official development set, significantly outperforming a local baseline of 65.90%, thereby demonstrating the method’s effectiveness in enhancing models’ reliance on genuine audio evidence.
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.