Score
Synthesizing source-to-microphone impulse responses and simulating acoustic scenes for arbitrary array geometries and environments so algorithms (e.g., source separation, IVA variants) can be experimentally validated under realistic conditions and layout variations.
Current speech separation and enhancement models exhibit limited generalization under mobile-source scenarios, primarily due to insufficient diversity and realism in evaluation data—both real-world and synthetic datasets fail to adequately reflect practical acoustic conditions. To address this, we propose SonicSim: the first customizable acoustic simulation framework specifically designed for mobile sound sources. SonicSim leverages Habitat-sim for physically accurate, multi-source spatial modeling and integrates LibriSpeech, FSD50K, and FMA audio corpora with Matterport3D 3D indoor environments. Based on this framework, we construct SonicSet—a large-scale, high-fidelity benchmark dataset—and complement it with real-world counterpart recordings. Experiments demonstrate that models trained on SonicSet achieve significantly improved generalization on real mobile-source recordings compared to those trained on existing synthetic datasets, effectively narrowing the synthetic-to-real acoustic domain gap.
Existing research is hindered by the scarcity of high-quality, diverse high-order Ambisonic (HOA) audio data—particularly room impulse responses (RIRs) suitable for sound source localization, reverberation modeling, and immersive sound field synthesis. To address this, we introduce the first large-scale 7th-order HOA-RIR dataset, synthesized via the image-source method (ISM) across extensive variations in room geometry, absorption materials, and transceiver configurations. We further propose a novel 64-channel spherical-harmonic-domain–optimized microphone array design, leveraging superposition to directly acquire RIRs in the spherical harmonic domain—bypassing conventional spatial coverage and order limitations. The resulting dataset achieves high spatial resolution and fidelity, substantially improving benchmark performance on sound source localization, reverberation prediction, and HOA sound field synthesis. This work establishes a foundational infrastructure for data-driven acoustic modeling.
Existing virtual acoustic simulation methods struggle to simultaneously achieve physical accuracy for low-frequency phenomena—such as diffraction and interference—and real-time performance. This paper proposes a hybrid acoustic modeling framework based on two-dimensional finite-difference time-domain (2D FDTD) simulation, tightly integrated with Unreal Engine’s audio rendering pipeline. Scene geometry is projected to generate obstacle masks and boundary conditions; sine-swept excitation combined with deconvolution is employed to extract spatially resolved, multi-channel impulse responses. To our knowledge, this is the first end-to-end integration of a Python-based FDTD wavefield solver with a commercial game engine’s real-time audio system. The framework supports dynamic occlusion, reflection, diffraction, and interference while preserving physical fidelity. Experimental validation confirms that the computed impulse responses align closely with theoretical predictions. Results demonstrate significant improvements in spatial audio realism and immersion for VR and interactive media applications.
This study addresses the ecological validity deficit of synthetically generated room impulse responses (RIRs) in monaural speech enhancement. To this end, we propose a frequency-dependent multi-band absorption coefficient modeling approach and integrate source and microphone directivity within the image-source method framework to construct a high-fidelity multi-band RIR (MB-RIR) dataset. Our method abandons the conventional single-band absorption assumption, substantially improving the fidelity of synthetic RIRs in characterizing real-world acoustic environments. The MB-RIR dataset is publicly available under an open-source, royalty-free license. Experiments on a real-RIR test set demonstrate that DeepFilterNet3 trained on MB-RIRs achieves a 0.51 dB improvement in signal-to-distortion ratio (SDR) and an 8.9-point gain in MUSHRA subjective listening scores over the baseline, confirming the enhanced generalizability and practical utility of the proposed approach.
Existing RIR generation methods heavily rely on room geometry priors, limiting their applicability in scenarios where layout is unknown or perceptual fidelity is paramount—such as VR and audio post-production. This work introduces the first framework for RIR generation conditioned solely on perceptual acoustic parameters (e.g., reverberation time, direct-to-reverberant ratio), eliminating dependence on explicit geometric modeling. Our method innovatively unifies autoregressive Transformers (within Descript Audio Codec), MaskGIT, flow matching, and classifier-guided sampling to jointly process discrete tokens and continuous embeddings. Objective metrics and subjective MOS evaluations demonstrate state-of-the-art performance; notably, the MaskGIT variant achieves superior flexibility in unseen environments and enhanced auditory realism.
This work addresses the lack of a unified, high-quality transcoding method for spatial audio across diverse acquisition formats—such as Ambisonics or microphone arrays—and arbitrary playback systems. The authors propose a general parametric framework that estimates spatial metadata of primary sources and ambient sound in the time–frequency domain, constructs a spatial covariance model tailored to the target playback setup, and derives an optimal linear downmix matrix. This approach supports independent rotation between acquisition and playback geometries and, for the first time, unifies processing for both Ambisonics and raw microphone array inputs. It accommodates arbitrary array configurations, variable numbers of sources, and arbitrary angular power distributions of ambient sound. Listening tests demonstrate that the method significantly outperforms existing parametric renderers across various content types and playback configurations, with particularly notable perceptual improvements for low-order or geometrically constrained arrays.
This work addresses the performance limitations of microphone arrays in practical applications caused by the scarcity of densely sampled room impulse responses (RIRs). To this end, it introduces the first diffusion model–based framework for RIR interpolation. The proposed method adapts and extends diffusion mechanisms from image inpainting to one-dimensional acoustic impulse response signals, enabling high-fidelity synthesis of RIRs at missing spatial locations. Experimental results on real-world RIR datasets demonstrate that the approach robustly accomplishes interpolation tasks and significantly enhances the performance of multi-microphone speech enhancement and spatial audio processing systems. These findings validate the efficacy and practical utility of diffusion models in realistic acoustic scenarios.
This work addresses the prohibitive computational complexity—scaling as $O(k^N)$—of conventional image-source models when synthesizing room impulse responses (RIRs) in high-dimensional spaces, which severely limits scalability. The paper introduces a novel approach by reformulating the high-dimensional image-source counting problem as a Gaussian circle lattice point enumeration task. It proposes a geometric convolution-based dimensional recurrence scheme that establishes cross-dimensional correlations and integrates frequency-dependent and reflection-weighting mechanisms. This method dramatically reduces computational complexity to $O(N k^2 \log k)$, enabling efficient RIR generation in arbitrary integer-coordinate high-dimensional spaces. The study further provides rigorous error bounds, runtime analysis, and empirical validation of statistical properties, demonstrating both theoretical soundness and practical efficacy.
This study addresses the inflated accuracy often reported by data-driven models in predicting room acoustic parameters, which stems largely from biases in evaluation protocols—particularly the overestimation of performance when test locations lack actual measurements. To rectify this, the work proposes a consistent evaluation framework that explicitly distinguishes between “location interpolation” and “prediction at truly unknown locations.” Using multi-condition measured data, it systematically evaluates three approaches: random forests, hybrid CNNs, and inverse distance weighting. Results show that high predictive performance (R² = 0.80–0.88) is achievable only when measured impulse responses at test locations are available as positional fingerprints. Under realistic generalization conditions—without any test-point data—performance drops substantially (R² = 0.09–0.57), though learning-based models still demonstrate practical advantages in predicting sound strength and reverberation time. This work underscores the dominant influence of evaluation protocols on reported metrics and establishes a more reliable benchmark for acoustic modeling.
This work addresses the challenge of achieving high-quality higher-order Ambisonics (HOA) encoding with sparse and irregular microphone arrays by proposing Flow-HOA, the first generative framework to incorporate conditional flow matching into HOA encoding. The method employs a composite loss function that jointly optimizes time-domain waveform fidelity, multi-resolution spectral consistency, subband energy preservation, and spatial directivity constraints, learning a mapping from a simple prior distribution to time-invariant FIR filter coefficients. Trained solely on synthetic data, Flow-HOA generalizes effectively to real-world recordings, significantly outperforming existing baselines in both objective metrics and subjective listening tests, demonstrating superior signal fidelity, enhanced spatial accuracy, and reduced audio artifacts.