acoustic feature extraction

Designs and implements signal-processing and analysis pipelines that convert raw audio into quantitative features and measurements—e.g., spectrograms, cepstral and spectral coefficients, temporal and spectro-dynamic descriptors, prosodic and loudness metrics, timbre measures, segmentation outputs, and pretrained audio embeddings—including the preprocessing steps needed to produce them. Builds and evaluates feature extraction and measurement methods, benchmarks feature types, and maps perceptual or subjective acoustic properties to objective parameters for use in downstream systems such as classification, recognition, synthesis, or encoding.

acousticfeatureextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.

acoustic signal preprocessingend-to-end classificationfeature-free

This study investigates how to select spectrogram representations that align with the architecture of downstream classifiers to enhance performance in audio and speech analysis tasks. By systematically exploring the design space of spectrograms—encompassing time–frequency resolution, temporal span, and element-wise scaling—and integrating convolutional neural networks with time–frequency analysis techniques, the work comprehensively evaluates diverse spectrogram configurations across multiple tasks. The findings reveal key principles for the co-optimization of front-end features and back-end models, delineate the suitability of various spectrogram representations for specific scenarios, and provide both theoretical grounding and practical guidance for task-oriented, efficient feature engineering and model design.

audio analysisclassifier architecturefeature representation

Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality

Dec 11, 2025
PM
Pablo M. Delgado
🏛️ Fraunhofer Institute for Integrated Circuits IIS | Ball State University | Netflix, Inc.

This study investigates how stereo processing—specifically Mid/Side versus Left/Right encoding—affects subjective audio quality perception, and evaluates the predictive accuracy of mainstream objective metrics (e.g., PEAQ, ITU-R BS.1387, DNSMOS) under spatial distortions. Leveraging the ODAQ stereo extension dataset with corresponding Mean Opinion Scores (MOS), we conduct time–frequency domain metric comparisons and statistical modeling. Our analysis quantitatively reveals, for the first time, the critical interplay between bottom-up auditory mechanisms and top-down contextual factors in stereo quality prediction. Results show that timbre-oriented metrics remain robust under simple distortions but degrade significantly under spatial distortions; current models exhibit systematic bias due to their neglect of spatial dimensions. We propose a novel three-dimensional perceptual evaluation paradigm integrating temporal, spectral, and spatial cues—providing both theoretical foundation and methodological support for next-generation audio quality metrics.

Evaluating stereo processing impact on audio qualityImproving models for timbral and spatial quality perceptionTesting objective metrics with subjective ratings dataset

This work proposes an end-to-end time-domain audio processing framework based on reservoir computing, addressing the limitations of traditional methods that rely on computationally intensive time–frequency transforms such as MFCCs and struggle to balance real-time performance, energy efficiency, and alignment with the human auditory system’s efficacy. By integrating biologically inspired auditory feature extraction with reservoir computing and replacing conventional frequency-domain transformations with lightweight convolutional operations, the proposed approach significantly reduces computational overhead while preserving discriminative feature representation. It eliminates the need for complex preprocessing and enables efficient, low-power real-time speech analysis, making it well-suited for embedded systems and voice-driven applications. This study thus establishes a highly energy-efficient and deployable paradigm for neuromorphic audio processing.

audio signal processingfeature extractionMFCC

Discrete Audio Tokens: More Than a Survey!

Jun 12, 2025
PM
Pooneh Mousavi
🏛️ Concordia University | Mila-Quebec AI Institute | The Hebrew University of Jerusalem | Carnegie Mellon University | Microsoft | Université de Toulon | Google | Apple | Laval University | University of Cambridge | University of Illinois at Urbana-Champaign | National Taiwan University | Université de Montréal

Existing discrete audio tokenization research lacks unified, cross-task and cross-domain evaluation. Method: We systematically survey and benchmark state-of-the-art methods across speech, music, and general audio domains, proposing the first comprehensive taxonomy spanning codec architecture, quantization mechanisms, training paradigms, streaming support, and application dimensions. We design a multi-objective joint optimization framework integrating reconstruction loss, semantic fidelity, and LLM alignment, unifying VQ/RVQ, GAN/MAE, and streaming token generation techniques. Contribution/Results: We conduct horizontal evaluation and controlled ablation studies across 12 standardized benchmarks, identifying critical bottlenecks. We open-source a standardized tokenizer database and core results, establishing—for the first time—the empirical trade-off boundary among reconstruction quality, inference latency, and generalization capability.

Analyze trade-offs and highlight open challengesEvaluate tokenizers on reconstruction and downstream tasksSystematic review and benchmark of discrete audio tokenizers

Latest Papers

What's happening recently
View more

Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and speech transmission index, accessible tools that combine rigorous signal processing with intuitive visualization remain scarce. This paper presents AcoustiVision Pro, an open-source web-based platform for comprehensive room impulse response (RIR) analysis. The system computes twelve distinct acoustic parameters from uploaded or dataset-sourced RIRs, provides interactive 3D visualizations of early reflections, generates frequency-dependent decay characteristics through waterfall plots, and checks compliance against international standards including ANSI S12.60 and ISO 3382. We introduce the accompanying RIRMega and RIRMega Speech datasets hosted on Hugging Face, containing thousands of simulated room impulse responses with full metadata. The platform supports real-time auralization through FFT-based convolution, exports detailed PDF reports suitable for engineering documentation, and provides CSV data export for further analysis. We describe the mathematical foundations underlying each acoustic metric, detail the system architecture, and present preliminary case studies demonstrating the platform's utility across diverse application domains including classroom acoustics, healthcare facility design, and recording studio evaluation.

acoustic characterizationinteractive visualizationopen-source platform

This study addresses the lack of integrated analysis and visualization tools for high-dimensional acoustic features in bird vocalizations. The authors propose the first open-source, end-to-end framework that combines pYIN-based fundamental frequency estimation, MFCC extraction, and PCA dimensionality reduction to construct a unified timbral space, enabling high-quality audio reconstruction via the Griffin-Lim algorithm. Innovatively, the system integrates neural audio synthesis with a dual-view 3D interactive interface built on Three.js, supporting dynamic trajectory comparison and independent playback of original and synthesized vocalizations. Experimental results demonstrate a Mel-spectral correlation coefficient exceeding 0.92 in bird song reconstruction, confirming the framework’s high-fidelity preservation of perceptual acoustic structure.

3D visualizationacoustic analysisbioacoustics

Binaspect -- A Python Library for Binaural Audio Analysis, Visualization & Feature Generation

Oct 29, 2025
DB
Dan Barry
🏛️ University College Dublin | Google LLC

This work addresses the challenge of analyzing spatial cue degradation in binaural audio under blind-source conditions. We propose an azimuth-aware modeling method that requires no prior knowledge of head-related transfer functions (HRTFs). By fusing modified interaural time difference (ITD) and interaural level difference (ILD) spectrograms and applying time-frequency adaptive clustering, we construct a robust and interpretable time–azimuth histogram (Azimuthogram), enabling multi-source azimuth separation and degradation localization. The method eliminates reliance on anatomical head models and supports degradation visualization across standard spatial audio processing pipelines—including codec-based compression (e.g., bitrate reduction), Ambisonic rendering, and vector-base amplitude panning (VBAP) localization. The resulting structured azimuth features significantly improve performance in no-reference quality prediction and spatial audio classification tasks, establishing a novel paradigm for reference-free spatial audio assessment.

Analyzes binaural audio degradation through codec and rendering processesGenerates interpretable azimuth maps for multiple sound source localizationProduces structured features for machine learning in spatial audio tasks

This study systematically investigates the encoding mechanisms and recoverability of low-level acoustic attributes—namely reverberation, loudness, spectral centroid, and relative pitch—in CLAP audio embeddings. By training linear and nonlinear probing models on frozen CLAP embeddings and conducting cross-dataset and cross-model generalization analyses alongside geometric direction consistency tests, the work reveals for the first time that reverberation, loudness, and relative pitch are approximately linearly encoded, whereas the spectral centroid requires nonlinear modeling. Moreover, the linear directions associated with these attributes remain consistent across datasets and align with their corresponding textual description embeddings. These findings demonstrate that all target attributes can be reliably recovered from CLAP embeddings, confirming the model’s strong generalization capability and cross-modal consistency among eight foundational audio models.

acoustic attributesaudio embeddingsfoundation models

LLM2Fx-Tools: Tool Calling For Music Post-Production

Dec 01, 2025
SD
Seungheon Doh
🏛️ KAIST | Sony AI | Sony Group Corporation

This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.

Generates executable audio effect sequences for music post-productionInfers effect chains from unprocessed and processed audio pairsTransfers audio effect styles from reference to new content

Hot Scholars

HY

Hung-yi Lee

National Taiwan University
deep learningspoken language understandingspeech processing
SW

Shinji Watanabe

Carnegie Mellon University
Speech recognitionSpeech processingSpeech enhancementSpeech translation
HL

Haizhou Li

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China; NUS, Singapore
Automatic Speech RecognitionSpeaker RecognitionLanguage RecognitionVoice Conversion
ZW

Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing
YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment