🤖 AI Summary
This work addresses the limitations of traditional target speech extraction, which typically requires separate modeling for distinct cues and lacks robustness under visual degradation. We propose TSE-Omni, a framework that leverages a single autoregressive large language model to unify the processing of synchronous and asynchronous multimodal cues. Its core innovation is a "self-registration" mechanism that enables audio-visual compensation by retrieving historical tokens during visual absence, eliminating the need for additional impairment-matching training. By integrating next-token prediction, semantic token generation, and streaming inference, TSE-Omni achieves performance comparable to existing baselines on the VoxCeleb2 and LRS3 datasets. Furthermore, it maintains high SpeechBERTScore under visual frame removal and multi-speaker interference, significantly enhancing extraction robustness in complex scenarios.
📝 Abstract
Target speech extraction (TSE) typically trains a separate extractor per cue, and visual-cue systems often need corruption-matched training to remain robust under visual frame corruption. We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text). TSE-Omni is driven by next-token prediction: each step predicts target speech semantic tokens from its own past outputs, which we term self-enrollment, forming a continuous target-speech context initialized by the enrollment cue (asynchronous audio or text, or a short visual prefix). This enables audio-visual compensation: the model uses synchronized visuals when intact and its token history when visual frames are missing. Under clean visuals, TSE-Omni matches strong discriminative and generative baselines (SpeechBERTScore 0.81 on VoxCeleb2 and 0.89 on LRS3 zero-shot) with higher DNSMOS. On the same VoxCeleb2 test set, after a 2 s clean visual start, removing the remaining visual frames leaves SpeechBERTScore at 0.81. It remains usable under sparse overlap and multi-speaker interference, and supports streaming inference. Project page: https://alexwxwu.github.io/tseomni-main/.