🤖 AI Summary
This study addresses the performance limitations of blind speech separation caused by the restricted number of physical microphone arrays. To overcome this bottleneck, we propose a virtual microphone-based enhancement mechanism. Specifically, high signal-to-noise ratio virtual microphone signals are generated using a pretrained speech diffusion model combined with posterior sampling in the diffusion process. These synthetic signals are then incorporated to introduce a virtual microphone-enhanced multi-channel consistency constraint, thereby optimizing the training of the separation network. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in both two- and three-speaker scenarios, effectively improving blind speech separation performance.
📝 Abstract
Blind Source Separation(BSS) is a fundamental problem in signal processing, aiming to separate multiple source signals from their mixtures without prior knowledge of the sources or the mixing process. Traditional approaches, such as Independent Vector Analysis (IVA) exploits statistical independence of sources. Recently, diffusion-based approaches have emerged as a promising alternative by leveraging powerful generative priors. Among them, ArrayDPS formulates BSS problem as a posterior sampling problem, and utilizes a pretrained speech diffusion model to guide the recovery of clean source signals. A key factor behind its separation capability is the multi-channel consistency (MC) objective, which enforces the estimated source signals to reconstruct the observed microphone mixtures through the estimated acoustic transfer functions. However, the number of microphones in the array is often limited, which constrains the performance of ArrayDPS. To address this issue, we propose VM-ArrayDPS, a novel method that augments the microphone array with virtual microphones with higher-SNR, these microphones can offer extra MC constraints to enhance the separation performance. Experimental results demonstrate that VM-ArrayDPS significantly outperforms ArrayDPS on both 2-speaker and 3-speaker datasets, showcasing the effectiveness of virtual microphone augmentation in improving BSS performance. We also did ablation studies to show the influence of the number of virtual microphones and weight of the MC objective brought by virtual microphones.