Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对真实对话中目标说话人提取的挑战,提出了一种新的损失函数来减少训练时过多静默的影响,并探讨了注册语音与目标语音不匹配的问题。
📝 Abstract
Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.
Problem

Research questions and friction points this paper is trying to address.

multi-speaker extraction
real conversational speech
silence
speaker enrolment
Innovation

Methods, ideas, or system contributions that make the work stand out.

new loss function
excess silence mitigation
real conversational speech enhancement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Robert Sutherland
School of Computer Science, University of Sheffield, Sheffield, United Kingdom
Stefan Goetze
Stefan Goetze
The University of Sheffield
Speech and Hearing
J
Jon Barker
School of Computer Science, University of Sheffield, Sheffield, United Kingdom