microphone array simulation

Synthesizing source-to-microphone impulse responses and simulating acoustic scenes for arbitrary array geometries and environments so algorithms (e.g., source separation, IVA variants) can be experimentally validated under realistic conditions and layout variations.

microphonearraysimulation

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

Oct 02, 2024
KL
Kai Li
🏛️ Tsinghua University | National Institute of Informatics | Chinese Institute for Brain Research

Current speech separation and enhancement models exhibit limited generalization under mobile-source scenarios, primarily due to insufficient diversity and realism in evaluation data—both real-world and synthetic datasets fail to adequately reflect practical acoustic conditions. To address this, we propose SonicSim: the first customizable acoustic simulation framework specifically designed for mobile sound sources. SonicSim leverages Habitat-sim for physically accurate, multi-source spatial modeling and integrates LibriSpeech, FSD50K, and FMA audio corpora with Matterport3D 3D indoor environments. Based on this framework, we construct SonicSet—a large-scale, high-fidelity benchmark dataset—and complement it with real-world counterpart recordings. Experiments demonstrate that models trained on SonicSet achieve significantly improved generalization on real mobile-source recordings compared to those trained on existing synthetic datasets, effectively narrowing the synthetic-to-real acoustic domain gap.

Lack of diverse data for speech processing in moving sound source scenarios.Need for customizable simulation tools to generate realistic synthetic data.Synthetic datasets lack acoustic realism for practical applications.

HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset

Nov 21, 2024
SS
Shivam Saini
🏛️ Leibniz University Hannover

Existing research is hindered by the scarcity of high-quality, diverse high-order Ambisonic (HOA) audio data—particularly room impulse responses (RIRs) suitable for sound source localization, reverberation modeling, and immersive sound field synthesis. To address this, we introduce the first large-scale 7th-order HOA-RIR dataset, synthesized via the image-source method (ISM) across extensive variations in room geometry, absorption materials, and transceiver configurations. We further propose a novel 64-channel spherical-harmonic-domain–optimized microphone array design, leveraging superposition to directly acquire RIRs in the spherical harmonic domain—bypassing conventional spatial coverage and order limitations. The resulting dataset achieves high spatial resolution and fidelity, substantially improving benchmark performance on sound source localization, reverberation prediction, and HOA sound field synthesis. This work establishes a foundational infrastructure for data-driven acoustic modeling.

Ambisonic audioimmersive audio productionsound realism

Existing virtual acoustic simulation methods struggle to simultaneously achieve physical accuracy for low-frequency phenomena—such as diffraction and interference—and real-time performance. This paper proposes a hybrid acoustic modeling framework based on two-dimensional finite-difference time-domain (2D FDTD) simulation, tightly integrated with Unreal Engine’s audio rendering pipeline. Scene geometry is projected to generate obstacle masks and boundary conditions; sine-swept excitation combined with deconvolution is employed to extract spatially resolved, multi-channel impulse responses. To our knowledge, this is the first end-to-end integration of a Python-based FDTD wavefield solver with a commercial game engine’s real-time audio system. The framework supports dynamic occlusion, reflection, diffraction, and interference while preserving physical fidelity. Experimental validation confirms that the computed impulse responses align closely with theoretical predictions. Results demonstrate significant improvements in spatial audio realism and immersion for VR and interactive media applications.

Capturing low-frequency wave phenomena like diffraction and reflectionIntegrating wave-based acoustic modeling into Unreal EngineSimulating accurate sound propagation in virtual environments

MB-RIRs: a Synthetic Room Impulse Response Dataset with Frequency-Dependent Absorption Coefficients

Jul 13, 2025
EG
Enric Gusó
🏛️ Universitat Pompeu Fabra | Eurecat, Centre Tecnològic de Catalunya

This study addresses the ecological validity deficit of synthetically generated room impulse responses (RIRs) in monaural speech enhancement. To this end, we propose a frequency-dependent multi-band absorption coefficient modeling approach and integrate source and microphone directivity within the image-source method framework to construct a high-fidelity multi-band RIR (MB-RIR) dataset. Our method abandons the conventional single-band absorption assumption, substantially improving the fidelity of synthetic RIRs in characterizing real-world acoustic environments. The MB-RIR dataset is publicly available under an open-source, royalty-free license. Experiments on a real-RIR test set demonstrate that DeepFilterNet3 trained on MB-RIRs achieves a 0.51 dB improvement in signal-to-distortion ratio (SDR) and an 8.9-point gain in MUSHRA subjective listening scores over the baseline, confirming the enhanced generalizability and practical utility of the proposed approach.

Enhancing monoaural speech enhancement performanceEvaluating frequency-dependent absorption coefficients in RIRsImproving ecological validity of synthetic RIR datasets

Room Impulse Response Generation Conditioned on Acoustic Parameters

Jul 16, 2025
SA
Silvia Arellano
🏛️ KTH Royal Institute of Technology | Dolby Laboratories

Existing RIR generation methods heavily rely on room geometry priors, limiting their applicability in scenarios where layout is unknown or perceptual fidelity is paramount—such as VR and audio post-production. This work introduces the first framework for RIR generation conditioned solely on perceptual acoustic parameters (e.g., reverberation time, direct-to-reverberant ratio), eliminating dependence on explicit geometric modeling. Our method innovatively unifies autoregressive Transformers (within Descript Audio Codec), MaskGIT, flow matching, and classifier-guided sampling to jointly process discrete tokens and continuous embeddings. Objective metrics and subjective MOS evaluations demonstrate state-of-the-art performance; notably, the MaskGIT variant achieves superior flexibility in unseen environments and enhanced auditory realism.

Evaluates autoregressive and non-autoregressive models for RIR generationFocuses on perceptual realism over strict physical accuracyGenerates room impulse responses using acoustic parameters not geometry

Latest Papers

What's happening recently
View more

This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.

Ambisonicsmicrophone arraysspatial audio

This work addresses the performance limitations of microphone arrays in practical applications caused by the scarcity of densely sampled room impulse responses (RIRs). To this end, it introduces the first diffusion model–based framework for RIR interpolation. The proposed method adapts and extends diffusion mechanisms from image inpainting to one-dimensional acoustic impulse response signals, enabling high-fidelity synthesis of RIRs at missing spatial locations. Experimental results on real-world RIR datasets demonstrate that the approach robustly accomplishes interpolation tasks and significantly enhances the performance of multi-microphone speech enhancement and spatial audio processing systems. These findings validate the efficacy and practical utility of diffusion models in realistic acoustic scenarios.

interpolationmicrophone array processingRoom Impulse Response

This work addresses the prohibitive computational complexity—scaling as $O(k^N)$—of conventional image-source models when synthesizing room impulse responses (RIRs) in high-dimensional spaces, which severely limits scalability. The paper introduces a novel approach by reformulating the high-dimensional image-source counting problem as a Gaussian circle lattice point enumeration task. It proposes a geometric convolution-based dimensional recurrence scheme that establishes cross-dimensional correlations and integrates frequency-dependent and reflection-weighting mechanisms. This method dramatically reduces computational complexity to $O(N k^2 \log k)$, enabling efficient RIR generation in arbitrary integer-coordinate high-dimensional spaces. The study further provides rigorous error bounds, runtime analysis, and empirical validation of statistical properties, demonstrating both theoretical soundness and practical efficacy.

computational complexityGauss circle problemhigh-dimensional

This study addresses the inflated accuracy often reported by data-driven models in predicting room acoustic parameters, which stems largely from biases in evaluation protocols—particularly the overestimation of performance when test locations lack actual measurements. To rectify this, the work proposes a consistent evaluation framework that explicitly distinguishes between “location interpolation” and “prediction at truly unknown locations.” Using multi-condition measured data, it systematically evaluates three approaches: random forests, hybrid CNNs, and inverse distance weighting. Results show that high predictive performance (R² = 0.80–0.88) is achievable only when measured impulse responses at test locations are available as positional fingerprints. Under realistic generalization conditions—without any test-point data—performance drops substantially (R² = 0.09–0.57), though learning-based models still demonstrate practical advantages in predicting sound strength and reverberation time. This work underscores the dominant influence of evaluation protocols on reported metrics and establishes a more reliable benchmark for acoustic modeling.

evaluation protocolgeneralizationinput availability

This work addresses the challenge of achieving high-quality higher-order Ambisonics (HOA) encoding with sparse and irregular microphone arrays by proposing Flow-HOA, the first generative framework to incorporate conditional flow matching into HOA encoding. The method employs a composite loss function that jointly optimizes time-domain waveform fidelity, multi-resolution spectral consistency, subband energy preservation, and spatial directivity constraints, learning a mapping from a simple prior distribution to time-invariant FIR filter coefficients. Trained solely on synthetic data, Flow-HOA generalizes effectively to real-world recordings, significantly outperforming existing baselines in both objective metrics and subjective listening tests, demonstrating superior signal fidelity, enhanced spatial accuracy, and reduced audio artifacts.

consumer audio captureHigher-Order AmbisonicsHOA encoding

Hot Scholars

TV

Tuomas Virtanen

Tampere University
machine listeningaudio signal processingaudio
AC

Andrew C. Singer

Dean, College of Engineering and Applied Sciences, Stony Brook University
Signal ProcessingCommunicationsUnderwater AcousticsAudio
ZQ

Zhong-Qiu Wang

Associate Professor, Southern University of Science and Technology
Computer AuditionSpeech SeparationMicrophone ArrayAudio Signal Processing
YM

Yoshiki Masuyama

Mitsubishi Electric Research Laboratories (MERL)
Audio Signal ProcessingSignal ProcessingMachine Learning
RM

Ryan M. Corey

University of Illinois Chicago
Signal ProcessingAudioHearing AidsArray Processing