build audio classifiers

Design, build, and train machine-learning pipelines that map raw audio signals or derived acoustic and spectral features to categorical labels. Work includes preprocessing and feature extraction, selecting and training supervised classification models, applying augmentation and calibration methods, and evaluating detection accuracy and probabilistic calibration.

buildaudioclassifiers

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Standardized algorithm evaluation protocols are lacking in audio-based device health monitoring, hindering reproducible research and industrial deployment. This paper introduces the first standardized machine learning evaluation framework for this task: it systematically extracts 127 acoustic features spanning time, frequency, and time–frequency domains; rigorously benchmarks 12 classifiers under a unified cross-validation protocol and nonparametric statistical testing (e.g., Wilcoxon signed-rank test) on both synthetic and real-world industrial datasets; and proposes a novel ensemble strategy achieving 94.2% accuracy and an F1-score of 0.942 across diverse scenarios—significantly outperforming the best individual classifier (+8–15 percentage points, *p* < 0.01). The framework is open-sourced with a comprehensive benchmarking protocol, providing a reproducible, statistically validated foundation for principled algorithm selection and practical deployment.

Developing comprehensive framework for rigorous evaluation of machine learning modelsEstablishing validated benchmarking protocol for industrial monitoring solutionsStandardizing algorithm selection methodologies for audio-based equipment condition monitoring

Traditional audiogram testing in remote hearing rehabilitation relies on calibrated equipment and complex procedures, limiting accessibility. Method: This study proposes a machine learning framework to automatically classify hearing-impaired individuals into standard Bisgaard audiogram types using calibration-free Adaptive Categorical Loudness Scaling (ACALOS) data. It integrates unsupervised clustering, seven multiclass classifiers—including logistic regression—and explainable AI techniques, with PCA dimensionality reduction (first two components explaining >50% variance) applied to a large-scale ACALOS dataset (N=847). Contribution/Results: Logistic regression achieved the highest classification accuracy, demonstrating that clinical audiogram typology can be reliably predicted solely from loudness perception data. This approach eliminates hardware calibration requirements and establishes a novel, scalable paradigm for hearing assessment in resource-constrained or remote settings.

Classify audiogram types from loudness scaling data using machine learningEvaluate unsupervised, supervised, and explainable machine learning techniquesPredict standard Bisgaard audiogram types without calibration for remote settings

Audio-Language Models for Audio-Centric Tasks: A survey

Jan 25, 2025
YS
Yi Su
🏛️ Northwestern Polytechnical University

A systematic survey of Audio-Language Models (ALMs) for general-purpose audio tasks is currently lacking. Method: This paper introduces, for the first time, a six-dimensional comprehensive taxonomy—covering architectural design, pretraining paradigms, downstream adaptation strategies, benchmark datasets, evaluation protocols, and future challenges—and proposes a structured technical roadmap. Our methodology integrates multimodal representation learning, contrastive and generative pretraining, instruction tuning, multi-task collaborative optimization, and agent-based system design. Contribution/Results: The work fills a critical gap in the ALM literature by delivering the first authoritative, holistic survey; it provides researchers and practitioners with a rigorous technical reference and practical guidance, thereby significantly advancing human-like auditory modeling research and real-world applications.

Audio Language ModelsComprehensive ReviewFuture Directions

To address the challenges of subjective quality assessment and the inefficiency of purely data-driven models in speech/audio coding, this paper proposes a tightly integrated hybrid neural coding framework that synergistically combines model-driven and data-driven paradigms. Methodologically, it introduces a novel multi-level hybrid architecture that deeply couples psychoacoustic-weighted loss, customized time-frequency domain prediction (TF-Codec/MDCTNet), an LPCNet-based backbone, and a neural post-processing module, trained end-to-end via an autoencoder paradigm. The core contribution lies in systematically bridging the performance gap between classical signal modeling and end-to-end deep learning. Experimental results demonstrate that, at ultra-low bitrates of 1.6–3.2 kbps, the proposed method achieves a P.808 MOS gain of ≥0.5 over baselines, yielding subjective audio quality approaching that of wideband codecs, while increasing computational overhead by less than 15%.

Efficiency ImprovementNeural Voice and Audio CodingQuality Evaluation

Listenable Maps for Zero-Shot Audio Classifiers

May 27, 2024
FP
F. Paissan
🏛️ Fondazione Bruno Kessler | Mila | Québec AI Institute | Concordia University | Université Laval

To address the limited interpretability of zero-shot audio classifiers, this paper proposes LMAC-ZS—the first decoder-based posterior explanation method for zero-shot audio classification. LMAC-ZS explicitly reveals the model’s decision rationale in text-audio cross-modal similarity computation via *audible heatmaps*, establishing the first audible explanation paradigm for zero-shot settings. To ensure explanation fidelity, we introduce a novel loss function that strictly enforces decoder outputs to preserve the original CLAP model’s similarity scores. By integrating cross-modal similarity distillation with posterior interpretability modeling, LMAC-ZS achieves high alignment between explanations and zero-shot predictions on the CLAP benchmark. Qualitative analysis demonstrates that the generated audible maps exhibit clear semantic content and strong correlation with diverse text prompts, enabling human-perceivable, attribution-based interpretation.

Ensuring faithfulness in text-audio similarity explanationsInterpreting decisions of zero-shot audio classifiersProducing meaningful explanations for classifier decisions

Latest Papers

What's happening recently
View more

This study addresses the scarcity of high-quality labeled data for audio classification in domain-specific scenarios such as domestic environments. To this end, the authors propose TriA Pipeline, the first large-scale automated audio annotation framework tailored for such settings, which efficiently generates the TriA dataset comprising 2,130 hours of audio across 431 event classes. Furthermore, they introduce a prior knowledge–guided data filtering mechanism to construct a refined subset, TriA_GK. Experimental results on three household audio classification tasks demonstrate that models trained on TriA_GK achieve relative improvements of 3.97% in average accuracy and 3.35% in Macro-F1 score over baseline methods, highlighting the effectiveness of the proposed approach.

audio classificationaudio datasetsdata annotation

This work addresses the distortion of performance metrics in real-world audio classification evaluation under limited annotation budgets, where sampling bias often leads to inaccurate assessments. To mitigate distributional shift in the evaluation subset, the study proposes importance weighting based on density ratio estimation—a technique introduced here for the first time in audio evaluation contexts. Specifically, sample weights are computed in the feature space using three distinct approaches: kernel density estimation (KDE), logistic regression, and k-nearest neighbors (kNN). Experimental results demonstrate that this methodology substantially reduces the gap between performance estimates derived from a sparsely labeled subset and the true model performance, thereby enhancing both the accuracy and robustness of audio classification evaluation under annotation constraints.

annotation budgetaudio classificationevaluation datasets

This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.

acoustic signal preprocessingend-to-end classificationfeature-free

Existing general-purpose audio pretraining is constrained by weak, noisy, and limited-scale labels, lacking a unified strong supervision framework. This work proposes the first Unified Tag System (UTS) that integrates speech, music, and environmental sounds, and establishes a high-fidelity audio captioning pipeline to enable a new pretraining paradigm centered on high-quality, strongly supervised data. Through systematic evaluation of multiple pretraining objectives within this framework, the study demonstrates that data quality and coverage are critical to performance gains, and further reveals that different pretraining objectives substantially influence the model’s specialization capabilities across downstream tasks.

audio pre-trainingdata-centriclabel quality

This study addresses the challenges of insufficient training data and poor cross-domain generalization in underwater acoustic machine learning, primarily due to the scarcity of publicly available labeled datasets. To this end, the authors construct a new underwater audio dataset comprising over one thousand annotated recordings spanning eight categories of biological and mechanical sound sources. They propose a lightweight CNN architecture incorporating boundary-aware loss and feature alignment to effectively mitigate class imbalance and domain shift. Evaluated within a newly established cross-domain framework, the method achieves an in-domain accuracy of 96.35% and demonstrates a 42.60% improvement in zero-shot ship detection performance on the ShipsEar dataset. The work also releases the data curation pipeline and reproducible benchmarks to support future research in underwater acoustics.

acoustic classificationcross-domain generalizationdataset scarcity

Hot Scholars

SR

Soundarya Ramesh

National University of Singapore
Sensing systemsComputer Security
CG

Chitralekha Gupta

Senior Research Fellow at National University of Singapore
Music Information RetrievalAudio Signal ProcessingMachine LearningDeep Learning
EN

Emre Neftci

Institute Director, Forschungszentrum Jülich; Professor, RWTH Aachen
Neuromorphic EngineeringComputational NeuroscienceCognitive Systems and BehaviorMachine Learning
DJ

Dasaem Jeong

Sogang University
Music Information RetrievalExpressive Performance ModelingMachine Learning
AS

Akihisa Shitara

Graduate School of Library, Information and Media Studies, University of Tsukuba
AccessibilityHuman InterfaceHuman Computer Interactiond/Deaf and Hard of Hearing