Unified Target-Speaker ASR with Text and Enrollment Speech Cues
This study addresses the insufficient complementarity caused by the separation of textual and enrollment speech cues in multi-speaker target speech recognition (TS-ASR). To this end, we propose a unified dual-cue TS-ASR framework that introduces a pioneering single-model unification mechanism. Built upon a Conformer architecture, the proposed method deeply fuses lexical and speaker information through a shared cross-attention module and incorporates a negative sampling strategy to enhance the supervision of cue effectiveness. Experimental results on 30,000 mixed-speech utterances demonstrate that combining five-character text prompts with enrollment speech reduces the character error rate (CER) to 8.80%, significantly outperforming the text-only (17.32%) and enrollment-only (29.06%) baselines.