acoustic measurement

Measuring and characterizing acoustic properties (spectra, loudness, timbre, impulse responses) under varied excitations and environments, and designing experiments to quantify how signal and micro-scale parameter changes alter measurable perceptual and spectral features.

acousticmeasurement

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Sensitivity of Room Impulse Responses in Changing Acoustic Environment

Jan 02, 2025
KP
Karolina Prawda
🏛️ University of York

Dynamic acoustic environmental changes—such as variations in surface absorption, introduction of scattering objects, or human movement—significantly perturb room impulse responses (RIRs), degrading performance in echo cancellation and sound source localization. This paper proposes a novel RIR change quantification framework based on short-term coherence and sensitivity scoring, enabling online environmental change detection and classification via time-frequency similarity analysis of consecutive RIRs. To our knowledge, this is the first method capable of highly discriminative identification among three canonical acoustic disturbances: atmospheric fluctuations, absorptive surface changes, and human presence. Experimental validation in real rooms demonstrates robust sensitivity to sub-millimeter human displacements, confirming its capability to detect minute environmental variations. The framework provides an interpretable, quantifiable foundation for adaptive acoustic systems, bridging the gap between physical environmental dynamics and algorithmic responsiveness.

Acoustic Environment ChangesEcho Cancellation PerformanceSound Technology Improvement

This work introduces a novel task—material-controllable room impulse response (RIR) generation—aiming to synthesize high-fidelity acoustic responses dynamically, conditioned on user-specified material configurations (e.g., floor, wall finishes) and multimodal audio-visual observations of indoor scenes. To support this, we present Acoustic Wonderland, the first acoustic dataset enabling fine-grained material combinations and synchronized multi-view audio-visual recordings. We further propose a new audio-visual–material fusion encoder-decoder architecture that explicitly models material properties and their geometric-acoustic mapping. Experiments demonstrate substantial improvements over existing baselines and state-of-the-art methods in RIR prediction accuracy, material sensitivity, and generation diversity. Notably, our approach enables real-time, interactive editing of material parameters during inference—a capability unprecedented in prior acoustic simulation frameworks.

Develop a benchmark dataset for material-aware RIR prediction methodsGenerate acoustic profiles for indoor scenes based on material configurationsPredict Room Impulse Responses (RIRs) using audio-visual scene properties

Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality

Dec 11, 2025
PM
Pablo M. Delgado
🏛️ Fraunhofer Institute for Integrated Circuits IIS | Ball State University | Netflix, Inc.

This study investigates how stereo processing—specifically Mid/Side versus Left/Right encoding—affects subjective audio quality perception, and evaluates the predictive accuracy of mainstream objective metrics (e.g., PEAQ, ITU-R BS.1387, DNSMOS) under spatial distortions. Leveraging the ODAQ stereo extension dataset with corresponding Mean Opinion Scores (MOS), we conduct time–frequency domain metric comparisons and statistical modeling. Our analysis quantitatively reveals, for the first time, the critical interplay between bottom-up auditory mechanisms and top-down contextual factors in stereo quality prediction. Results show that timbre-oriented metrics remain robust under simple distortions but degrade significantly under spatial distortions; current models exhibit systematic bias due to their neglect of spatial dimensions. We propose a novel three-dimensional perceptual evaluation paradigm integrating temporal, spectral, and spatial cues—providing both theoretical foundation and methodological support for next-generation audio quality metrics.

Evaluating stereo processing impact on audio qualityImproving models for timbral and spatial quality perceptionTesting objective metrics with subjective ratings dataset

Non-verbal Perception of Room Acoustics using Multi Dimensional Scaling Metho

Nov 12, 2025
LB
Leonie Bohlke
🏛️ Taubert und Ruhe GmbH | Institute of Systematic Musicology, University of Hamburg | KREBS + KIEFER Ingenieure

Traditional room acoustics research relies on expert recall or rating scales to establish subjective–objective correlations, suffering from subjective bias and high cognitive load. This study proposes a novel, language-free perceptual paradigm for subjective acoustic evaluation: virtual musical stimuli are generated via binaural impulse response convolution; multidimensional scaling (MDS) is applied to uncover the latent structure of auditory spatial perception; and psychoacoustic parameters—including echo density, fractal correlation dimension, roughness, loudness, and early decay time—are employed to model and interpret perceptual dimensions. For the first time, five statistically independent and interpretable perceptual dimensions are extracted without linguistic labels or prior knowledge, all showing strong correlations (r > 0.85) with key objective acoustic metrics. This approach establishes a quantifiable perceptual foundation for concert hall design and high-fidelity virtual auditory rendering, providing a new, theory-grounded evaluation tool.

Developing alternative method to characterize subjective room acoustic perceptionsIdentifying perceptual dimensions of room acoustics using Multi Dimensional ScalingRelating subjective acoustic impressions to objective psychoacoustic parameters

This study addresses the inflated accuracy often reported by data-driven models in predicting room acoustic parameters, which stems largely from biases in evaluation protocols—particularly the overestimation of performance when test locations lack actual measurements. To rectify this, the work proposes a consistent evaluation framework that explicitly distinguishes between “location interpolation” and “prediction at truly unknown locations.” Using multi-condition measured data, it systematically evaluates three approaches: random forests, hybrid CNNs, and inverse distance weighting. Results show that high predictive performance (R² = 0.80–0.88) is achievable only when measured impulse responses at test locations are available as positional fingerprints. Under realistic generalization conditions—without any test-point data—performance drops substantially (R² = 0.09–0.57), though learning-based models still demonstrate practical advantages in predicting sound strength and reverberation time. This work underscores the dominant influence of evaluation protocols on reported metrics and establishes a more reliable benchmark for acoustic modeling.

evaluation protocolgeneralizationinput availability

Latest Papers

What's happening recently
View more

Room acoustics analysis plays a central role in architectural design, audio engineering, speech intelligibility assessment, and hearing research. Despite the availability of standardized metrics such as reverberation time, clarity, and speech transmission index, accessible tools that combine rigorous signal processing with intuitive visualization remain scarce. This paper presents AcoustiVision Pro, an open-source web-based platform for comprehensive room impulse response (RIR) analysis. The system computes twelve distinct acoustic parameters from uploaded or dataset-sourced RIRs, provides interactive 3D visualizations of early reflections, generates frequency-dependent decay characteristics through waterfall plots, and checks compliance against international standards including ANSI S12.60 and ISO 3382. We introduce the accompanying RIRMega and RIRMega Speech datasets hosted on Hugging Face, containing thousands of simulated room impulse responses with full metadata. The platform supports real-time auralization through FFT-based convolution, exports detailed PDF reports suitable for engineering documentation, and provides CSV data export for further analysis. We describe the mathematical foundations underlying each acoustic metric, detail the system architecture, and present preliminary case studies demonstrating the platform's utility across diverse application domains including classroom acoustics, healthcare facility design, and recording studio evaluation.

acoustic characterizationinteractive visualizationopen-source platform

This study addresses the lack of standardized, physically interpretable visualization and feature representation methods for distributed acoustic sensing (DAS) data, which hinders large-scale analysis efficiency. To overcome this limitation, the work introduces multispectral imaging concepts into DAS for the first time, proposing a band-decomposition-based multispectral signal representation framework. By decomposing strain-rate signals into predefined frequency bands and generating corresponding band-energy images, the method constructs a spatiotemporal-spectral feature space with clear physical meaning that is amenable to automated processing. Integrated with unsupervised clustering and a ResNet-18 convolutional neural network, the approach achieves 97.3% accuracy in detecting cetacean vocalizations, significantly enhancing both the visual interpretability and automatic recognition performance of bioacoustic signals.

data visualizationDistributed Acoustic Sensingfeature extraction

This work addresses the challenges of in situ acoustic surface admittance estimation, which is typically hindered by noise, model inaccuracies, and the restrictive assumptions of conventional methods. The authors propose a novel physics-informed neural operator framework—applied for the first time to in situ characterization of acoustic materials—that embeds the Helmholtz equation, linearized momentum equation, and Robin boundary conditions to directly learn frequency-dependent admittance spectra in an end-to-end manner from near-field acoustic pressure and particle velocity measurements. By circumventing explicit forward modeling and per-frequency inversion, the method achieves globally consistent admittance reconstruction that inherently respects physical constraints. Experimental results demonstrate its superior performance in accurately recovering both real and imaginary parts of the admittance over a broad frequency range, significantly outperforming purely data-driven approaches while exhibiting exceptional robustness to noise and sparse spatial sampling.

acoustic impedancein situ characterizationnear-field measurements

Existing evaluation metrics for generative spatial audio lack systematic investigation into their response characteristics under variations in spatial parameters such as azimuth and elevation. This work proposes the first sensitivity analysis framework that evaluates multiple metrics along continuous spatial trajectories, introducing three key criteria: responsiveness, smoothness, and symmetry. The framework is empirically applied to metrics including Fréchet Audio Distance (FAD), intensity vectors, and acoustic maps within controlled scenes of varying complexity. Results demonstrate that FAD based on directional embeddings and acoustic maps consistently excel across all three criteria, whereas intensity vectors exhibit significant performance degradation as scene complexity increases. These findings reveal substantial differences in the sensitivity of existing metrics to spatial variations, offering critical insights for the design and selection of evaluation methods in spatial audio generation.

First-Order AmbisonicsGenerative Spatial AudioMetric Evaluation

This study addresses the lack of objective and interpretable vocal biomarkers in mental health assessment by proposing a transparent, clinically interpretable speech analysis framework. It systematically integrates multidimensional perceptual features—including prosody, voice quality, semantic coherence, syntactic structure, and sarcasm—by jointly leveraging acoustic and linguistic information. Using an XGBoost model enhanced with SHAP and LIME for interpretability, the framework identifies key features such as jitter, shimmer, lexical-syntactic patterns, and affective intonation from real-world clinical data and multiple benchmark datasets. Experimental results demonstrate robust associations between vocal irregularities and symptom severity in depression, anxiety, and ADHD. Ablation studies further confirm the most discriminative feature subsets, offering reliable and explainable vocal biomarkers to support clinical evaluation.

clinical decision-supportmental health assessmentperceptual speech features

Hot Scholars

TV

Toon van Waterschoot

Professor, KU Leuven
audio processingspeech processingroom acousticsaudio engineering
SK

Shoichi Koyama

National Institute of Informatics, Japan
Acoustic Signal ProcessingAudio Signal ProcessingArray ProcessingAcoustics
SD

Simon Doclo

Professor, University of Oldenburg, Germany
Signal ProcessingSpeech and Audio ProcessingMicrophone Arrays
MB

Michael Beigl

Professor for Informatics, Karlsruhe Institute of Technology (KIT)
Ubiquitous ComputingWearable ComputingHealth & Activity Recognition using AIInternet of Things
MP

Mirco Pezzoli

Research fellow, Politecnico di Milano
audio and acoustic signal processingspatial audio recording and reproductionmachine learning for audiosound scene analysis