Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing target speaker extraction (TSE) methods overlook natural conversational turn-taking, hindering real-time tracking of the target speaker amid pauses and interference. This work establishes the first online audio-visual TSE benchmark and proposes the TS-VAP framework. TS-VAP pioneers the integration of historical, synchronous, and predictive context mechanisms, introducing a Speech-LLM-based target voice activity prediction module that fuses semantic, acoustic, and facial cues to anticipate future speaker behavior. This design transcends conventional separation channel limitations to guide low-latency streaming separation. Experimental results demonstrate that TS-VAP significantly enhances performance across multiple backbone networks, achieving nearly 1 dB improvement in realistic conversational scenarios.
📝 Abstract
In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of real conversations. We therefore introduce, to our knowledge, the first benchmark for online audio-visual TSE (AV-TSE), built from intact dyadic interactions with independent third-party interference. Observing that anticipating upcoming activity from semantic, acoustic, and facial cues benefits online AV-TSE, we propose an LLM-based target-speaker voice activity projection (TS-VAP) module. Unlike conventional VAP with separated speaker channels, it forecasts the future activity of the target and conversational partner directly from the overlapping mixture, drawing on the linguistic and conversational knowledge of a speech-LLM, and uses this prediction to guide a low-latency separator. We further combine this predictive context with historical and synchronous speaker context. Experiments show that TS-VAP consistently improves streaming extraction across multiple AV-TSE backbones, and that further combining historical, synchronous, and predictive context yields nearly 1 dB gain on real AV conversations. Project page: https://jjjjiaozi.github.io/TS-VAP/.
Problem

Research questions and friction points this paper is trying to address.

Target Speaker Extraction
Audio-Visual Speech Processing
Turn-taking
Streaming Extraction
Online Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target Speaker Extraction
Audio-Visual
Streaming
Voice Activity Projection
Large Language Model
💼 Related Jobs
No related jobs found.