Score
Designs and implements signal-processing and analysis pipelines that convert raw audio into quantitative features and measurements—e.g., spectrograms, cepstral and spectral coefficients, temporal and spectro-dynamic descriptors, prosodic and loudness metrics, timbre measures, segmentation outputs, and pretrained audio embeddings—including the preprocessing steps needed to produce them. Builds and evaluates feature extraction and measurement methods, benchmarks feature types, and maps perceptual or subjective acoustic properties to objective parameters for use in downstream systems such as classification, recognition, synthesis, or encoding.
This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.
This study investigates how to select spectrogram representations that align with the architecture of downstream classifiers to enhance performance in audio and speech analysis tasks. By systematically exploring the design space of spectrograms—encompassing time–frequency resolution, temporal span, and element-wise scaling—and integrating convolutional neural networks with time–frequency analysis techniques, the work comprehensively evaluates diverse spectrogram configurations across multiple tasks. The findings reveal key principles for the co-optimization of front-end features and back-end models, delineate the suitability of various spectrogram representations for specific scenarios, and provide both theoretical grounding and practical guidance for task-oriented, efficient feature engineering and model design.
This study investigates how stereo processing—specifically Mid/Side versus Left/Right encoding—affects subjective audio quality perception, and evaluates the predictive accuracy of mainstream objective metrics (e.g., PEAQ, ITU-R BS.1387, DNSMOS) under spatial distortions. Leveraging the ODAQ stereo extension dataset with corresponding Mean Opinion Scores (MOS), we conduct time–frequency domain metric comparisons and statistical modeling. Our analysis quantitatively reveals, for the first time, the critical interplay between bottom-up auditory mechanisms and top-down contextual factors in stereo quality prediction. Results show that timbre-oriented metrics remain robust under simple distortions but degrade significantly under spatial distortions; current models exhibit systematic bias due to their neglect of spatial dimensions. We propose a novel three-dimensional perceptual evaluation paradigm integrating temporal, spectral, and spatial cues—providing both theoretical foundation and methodological support for next-generation audio quality metrics.
This work proposes an end-to-end time-domain audio processing framework based on reservoir computing, addressing the limitations of traditional methods that rely on computationally intensive time–frequency transforms such as MFCCs and struggle to balance real-time performance, energy efficiency, and alignment with the human auditory system’s efficacy. By integrating biologically inspired auditory feature extraction with reservoir computing and replacing conventional frequency-domain transformations with lightweight convolutional operations, the proposed approach significantly reduces computational overhead while preserving discriminative feature representation. It eliminates the need for complex preprocessing and enables efficient, low-power real-time speech analysis, making it well-suited for embedded systems and voice-driven applications. This study thus establishes a highly energy-efficient and deployable paradigm for neuromorphic audio processing.
Existing discrete audio tokenization research lacks unified, cross-task and cross-domain evaluation. Method: We systematically survey and benchmark state-of-the-art methods across speech, music, and general audio domains, proposing the first comprehensive taxonomy spanning codec architecture, quantization mechanisms, training paradigms, streaming support, and application dimensions. We design a multi-objective joint optimization framework integrating reconstruction loss, semantic fidelity, and LLM alignment, unifying VQ/RVQ, GAN/MAE, and streaming token generation techniques. Contribution/Results: We conduct horizontal evaluation and controlled ablation studies across 12 standardized benchmarks, identifying critical bottlenecks. We open-source a standardized tokenizer database and core results, establishing—for the first time—the empirical trade-off boundary among reconstruction quality, inference latency, and generalization capability.
Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and speech transmission index, accessible tools that combine rigorous signal processing with intuitive visualization remain scarce. This paper presents AcoustiVision Pro, an open-source web-based platform for comprehensive room impulse response (RIR) analysis. The system computes twelve distinct acoustic parameters from uploaded or dataset-sourced RIRs, provides interactive 3D visualizations of early reflections, generates frequency-dependent decay characteristics through waterfall plots, and checks compliance against international standards including ANSI S12.60 and ISO 3382. We introduce the accompanying RIRMega and RIRMega Speech datasets hosted on Hugging Face, containing thousands of simulated room impulse responses with full metadata. The platform supports real-time auralization through FFT-based convolution, exports detailed PDF reports suitable for engineering documentation, and provides CSV data export for further analysis. We describe the mathematical foundations underlying each acoustic metric, detail the system architecture, and present preliminary case studies demonstrating the platform's utility across diverse application domains including classroom acoustics, healthcare facility design, and recording studio evaluation.
This study addresses the lack of integrated analysis and visualization tools for high-dimensional acoustic features in bird vocalizations. The authors propose the first open-source, end-to-end framework that combines pYIN-based fundamental frequency estimation, MFCC extraction, and PCA dimensionality reduction to construct a unified timbral space, enabling high-quality audio reconstruction via the Griffin-Lim algorithm. Innovatively, the system integrates neural audio synthesis with a dual-view 3D interactive interface built on Three.js, supporting dynamic trajectory comparison and independent playback of original and synthesized vocalizations. Experimental results demonstrate a Mel-spectral correlation coefficient exceeding 0.92 in bird song reconstruction, confirming the framework’s high-fidelity preservation of perceptual acoustic structure.
This work addresses the challenge of analyzing spatial cue degradation in binaural audio under blind-source conditions. We propose an azimuth-aware modeling method that requires no prior knowledge of head-related transfer functions (HRTFs). By fusing modified interaural time difference (ITD) and interaural level difference (ILD) spectrograms and applying time-frequency adaptive clustering, we construct a robust and interpretable time–azimuth histogram (Azimuthogram), enabling multi-source azimuth separation and degradation localization. The method eliminates reliance on anatomical head models and supports degradation visualization across standard spatial audio processing pipelines—including codec-based compression (e.g., bitrate reduction), Ambisonic rendering, and vector-base amplitude panning (VBAP) localization. The resulting structured azimuth features significantly improve performance in no-reference quality prediction and spatial audio classification tasks, establishing a novel paradigm for reference-free spatial audio assessment.
This study systematically investigates the encoding mechanisms and recoverability of low-level acoustic attributes—namely reverberation, loudness, spectral centroid, and relative pitch—in CLAP audio embeddings. By training linear and nonlinear probing models on frozen CLAP embeddings and conducting cross-dataset and cross-model generalization analyses alongside geometric direction consistency tests, the work reveals for the first time that reverberation, loudness, and relative pitch are approximately linearly encoded, whereas the spectral centroid requires nonlinear modeling. Moreover, the linear directions associated with these attributes remain consistent across datasets and align with their corresponding textual description embeddings. These findings demonstrate that all target attributes can be reliably recovered from CLAP embeddings, confirming the model’s strong generalization capability and cross-modal consistency among eight foundational audio models.
This work addresses the challenge of enabling large language models (LLMs) to comprehend raw audio and autonomously generate executable audio effect chains (Fx-chains) for music post-production. We propose the first multimodal tool-calling framework tailored for audio effect synthesis, integrating audio representations, structured tool interfaces, chain-of-thought (CoT) planning, and autoregressive sequence modeling to achieve end-to-end mapping from input audio to effect types, ordering, and parameters. We introduce LP-Fx, a high-quality, human-annotated dataset for audio effect chaining, and pioneer the application of LLM tool-calling paradigms to audio processing. Experiments demonstrate that our system generates semantically coherent and parameter-plausible Fx-chains; successfully transfers processing characteristics in style-transfer tasks; and achieves strong interpretability and response fidelity, as validated by both human and LLM-based evaluation.