Unified Target-Speaker ASR with Text and Enrollment Speech Cues

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient complementarity caused by the separation of textual and enrollment speech cues in multi-speaker target speech recognition (TS-ASR). To this end, we propose a unified dual-cue TS-ASR framework that introduces a pioneering single-model unification mechanism. Built upon a Conformer architecture, the proposed method deeply fuses lexical and speaker information through a shared cross-attention module and incorporates a negative sampling strategy to enhance the supervision of cue effectiveness. Experimental results on 30,000 mixed-speech utterances demonstrate that combining five-character text prompts with enrollment speech reduces the character error rate (CER) to 8.80%, significantly outperforming the text-only (17.32%) and enrollment-only (29.06%) baselines.
📝 Abstract
Target-speaker automatic speech recognition (TS-ASR) aims to recognize a designated speaker while suppressing interfering speech in multi-talker environments. Conventional TS-ASR typically relies on an enrollment utterance, whereas text-guided methods use known lexical content, such as a wake word, to identify the target speaker from the observed mixture. These two cues provide complementary information but are usually studied separately. We propose a Unified Dual-Cue TS-ASR framework that supports text cues, enrollment speech, or both within a single model. Text cues interact with the mixture representation to extract target-speaker information conditioned on known lexical content, while an independent enrollment utterance provides complementary speaker information. Cross-attention cue-conditioning modules are integrated into shared Conformer blocks, and negative-cue sampling provides cue-validity supervision during dual-cue training. Experiments on 30,000 two-speaker mixtures across five recording/domain conditions and four oracle text-cue lengths show that, with five-character text cues, the concatenated dual-cue method achieves 8.80% CER, compared with 17.32% for text-only and 29.06% for enrollment-only inference. It also outperforms parallel dual-cue fusion (9.49% CER) and yields lower dual-cue CER across all five evaluation subsets. These results demonstrate the benefit of jointly exploiting complementary lexical and speaker information for target-speaker ASR.
Problem

Research questions and friction points this paper is trying to address.

Target-Speaker ASR
Multi-talker Environments
Text Cues
Enrollment Speech
Dual-Cue
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target-Speaker ASR
Dual-Cue Framework
Cross-Attention Cue-Conditioning
Negative-Cue Sampling
Conformer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yuxiang Mei
Shanghai Normal University
Y
Yuchen Yan
Baidu
D
Dongxing Xu
Unisound
J
Jiaen Liang
Unisound
Yanhua Long
Yanhua Long
Professor, Shanghai Normal University
Speech signal processing