Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of poor audio quality, high latency, and insufficient robustness in online audio-visual target speaker extraction for real-world meeting scenarios. To this end, we propose a low-latency multimodal extraction framework based on personalized speech enhancement. The method inversely leverages personalized enhancement techniques by incorporating lip visual features for audio-visual multimodal fusion, and applies end-to-end joint fine-tuning to the audio-visual network to support streaming processing. Experimental results demonstrate that the two-speaker confusion rate is reduced to 1.6% with an algorithmic latency of only 20 ms. Furthermore, the Mean Opinion Score (MOS) improves by 0.57 to 0.63, effectively suppressing competing speech while achieving a favorable balance between real-time performance and perceptual audio quality.
📝 Abstract
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-Visual Personalized Voice Quality Enhancement (AV-PVQE), which approaches these requirements from the other direction. We start from a personalized speech enhancement model that reconstructs a requested voice at high quality but confuses the target in 46% of two-speaker mixtures despite clean enrollment. Adding mouth features at its speaker-conditioning input and jointly fine-tuning the visual and reconstruction networks reduces this rate to 1.6%, with no future frames and 20 ms of algorithmic delay. Compared with an online autoregressive audio-visual extractor, AV-PVQE yields separation gains on two synthetic benchmarks and larger gains on recorded meetings, and keeps its advantage on excerpts with more speakers than the fine-tuning mixtures. In personalized P.835 listening tests on two meeting corpora, it improves overall quality over this extractor by 0.57 and 0.63 MOS, with similar mean rating relative to the starting model. Preservation and rejection tests show that it keeps the target intact when no competing voice is present and suppresses competing speech when the target is absent.
Problem

Research questions and friction points this paper is trying to address.

Target-Speaker Extraction
Speech Enhancement
Audio-Visual
Low-Latency
Personalized Voice Quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Target-Speaker Extraction
Personalized Speech Enhancement
Low-Latency
Voice Quality Enhancement
Joint Fine-tuning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rayhan Rashed
Microsoft
S
Senja Filipi
University of Michigan
Ross Cutler
Ross Cutler
Microsoft
Computer VisionMachine LearningAcousticsOpticsVoIP