🤖 AI Summary
This study addresses the challenges of poor audio quality, high latency, and insufficient robustness in online audio-visual target speaker extraction for real-world meeting scenarios. To this end, we propose a low-latency multimodal extraction framework based on personalized speech enhancement. The method inversely leverages personalized enhancement techniques by incorporating lip visual features for audio-visual multimodal fusion, and applies end-to-end joint fine-tuning to the audio-visual network to support streaming processing. Experimental results demonstrate that the two-speaker confusion rate is reduced to 1.6% with an algorithmic latency of only 20 ms. Furthermore, the Mean Opinion Score (MOS) improves by 0.57 to 0.63, effectively suppressing competing speech while achieving a favorable balance between real-time performance and perceptual audio quality.
📝 Abstract
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.