RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the domain shift degradation of target speech extraction (TSE) models trained on synthetic data when deployed in real-world complex acoustic environments. We propose a progressive synthetic-to-real adaptation framework that achieves efficient transfer through two-stage fine-tuning, combining target-aware supervision with a weak learning strategy. Furthermore, this work extends RemixIT from speech enhancement to TSE for the first time by designing a quality-filtering-based signal-level unsupervised fine-tuning mechanism. Evaluated on the EVAL-2 set of the SLT 2026 REAL-TSE Challenge, the proposed method reduces the total error rate (TER) by 6.53% and improves speaker similarity by 21.84%, significantly outperforming the baseline system.
📝 Abstract
Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.
Problem

Research questions and friction points this paper is trying to address.

Target Speech Extraction
Domain Gap
Synthetic-to-Real Adaptation
Real-world Conversational Scenarios
Innovation

Methods, ideas, or system contributions that make the work stand out.

Target Speech Extraction
Synthetic-to-Real Adaptation
RemixIT-TSE
Target-Aware Supervision
Self-training
🔎 Similar Papers
No similar papers found.
Y
Yu Wang
Shanghai Normal University, Shanghai, China
H
Haixin Guan
Unisound AI Technology Co., Ltd., Beijing, China
S
Shuang Wei
Shanghai Normal University, Shanghai, China
Yanhua Long
Yanhua Long
Professor, Shanghai Normal University
Speech signal processing