Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of existing speech enhancement models trained on fixed microphone arrays to devices with diverse geometric configurations. To overcome this challenge, we propose a Geometry-aware Dynamic Convolution (Geo-DConv) framework, which explicitly incorporates array geometry priors into array-agnostic speech enhancement for the first time. By leveraging a dynamic convolution mechanism that integrates microphone coordinate information directly into the network, Geo-DConv enables adaptive spatial filtering, effectively transforming fixed-array models into array-invariant systems. Extensive experiments on the RealMAN real-world recording dataset demonstrate that Geo-DConv consistently enhances performance across various array topologies when applied to two representative baseline models, significantly improving cross-topology generalization.
📝 Abstract
Multi-channel speech enhancement (SE) systems exhibit superior performance over single-channel methods but are constrained to fixed microphone array configurations. This restricts their real-world deployment across devices with diverse array geometries. While recent array-agnostic SE methods address variable microphone numbers and permutations, they largely fail to exploit explicit array geometry priors when available, missing a crucial cue for optimal spatial filtering. A Geometry-Aware Dynamic Convolution (Geo-DConv) framework is proposed, which explicitly leverages microphone coordinates to transform standard fixed-array SE models into robust array-invariant systems. Experiments are conducted on the recent real-recorded RealMAN multi-channel speech dataset. Results demonstrate that the proposed architecture enables two widely used fixed-array models to adapt to array-invariant settings, with consistent performance improvements across diverse array topologies.
Problem

Research questions and friction points this paper is trying to address.

array-invariant
speech enhancement
microphone array geometry
spatial filtering
multi-channel
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Aware Dynamic Convolution
Array-Invariant Speech Enhancement
Microphone Array Geometry
Multi-channel Speech Enhancement
Spatial Filtering
🔎 Similar Papers
No similar papers found.