Exploring a Single Autoregressive LLM for Unified Target Speech Extraction across Synchronous and Asynchronous Cues

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional target speech extraction, which typically requires separate modeling for distinct cues and lacks robustness under visual degradation. We propose TSE-Omni, a framework that leverages a single autoregressive large language model to unify the processing of synchronous and asynchronous multimodal cues. Its core innovation is a "self-registration" mechanism that enables audio-visual compensation by retrieving historical tokens during visual absence, eliminating the need for additional impairment-matching training. By integrating next-token prediction, semantic token generation, and streaming inference, TSE-Omni achieves performance comparable to existing baselines on the VoxCeleb2 and LRS3 datasets. Furthermore, it maintains high SpeechBERTScore under visual frame removal and multi-speaker interference, significantly enhancing extraction robustness in complex scenarios.
📝 Abstract
Target speech extraction (TSE) typically trains a separate extractor per cue, and visual-cue systems often need corruption-matched training to remain robust under visual frame corruption. We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text). TSE-Omni is driven by next-token prediction: each step predicts target speech semantic tokens from its own past outputs, which we term self-enrollment, forming a continuous target-speech context initialized by the enrollment cue (asynchronous audio or text, or a short visual prefix). This enables audio-visual compensation: the model uses synchronized visuals when intact and its token history when visual frames are missing. Under clean visuals, TSE-Omni matches strong discriminative and generative baselines (SpeechBERTScore 0.81 on VoxCeleb2 and 0.89 on LRS3 zero-shot) with higher DNSMOS. On the same VoxCeleb2 test set, after a 2 s clean visual start, removing the remaining visual frames leaves SpeechBERTScore at 0.81. It remains usable under sparse overlap and multi-speaker interference, and supports streaming inference. Project page: https://alexwxwu.github.io/tseomni-main/.
Problem

Research questions and friction points this paper is trying to address.

Target Speech Extraction
Unified Model
Visual Corruption Robustness
Synchronous and Asynchronous Cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive LLM
Target Speech Extraction
Self-enrollment
Audio-visual Compensation
Streaming Inference
🔎 Similar Papers
No similar papers found.