🤖 AI Summary
This study addresses the domain shift degradation of target speech extraction (TSE) models trained on synthetic data when deployed in real-world complex acoustic environments. We propose a progressive synthetic-to-real adaptation framework that achieves efficient transfer through two-stage fine-tuning, combining target-aware supervision with a weak learning strategy. Furthermore, this work extends RemixIT from speech enhancement to TSE for the first time by designing a quality-filtering-based signal-level unsupervised fine-tuning mechanism. Evaluated on the EVAL-2 set of the SLT 2026 REAL-TSE Challenge, the proposed method reduces the total error rate (TER) by 6.53% and improves speaker similarity by 21.84%, significantly outperforming the baseline system.
📝 Abstract
Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to extend RemixIT from speech enhancement to TSE and propose a progressive synthetic-to-real adaptation framework for real-world TSE with two fine-tuning stages. The first stage leverages region-wise speaker similarity and silence constraints within a target-aware adaptation framework to jointly optimize the model using synthetic and weakly supervised real-world data, injecting real-world traits while preserving synthetic-learned capabilities. The second stage further adapts the model using only real-world data through our RemixIT-TSE, where quality-filtered teacher pseudo targets, which guarantee reliable student training, provide signal-level supervision via SI-SNR loss. Experiments on the real conversational evaluation set (EVAL-2) of the SLT 2026 REAL-TSE Challenge, the proposed method achieves a 6.53% relative TER reduction, together with relative improvements of 21.84% in speaker similarity, 9.89% in DNSMOS-P808, and 4.10% in target-activity F1 over the source-domain baseline, demonstrating its effectiveness under unseen real-world conditions. Source code at https: //github.com/YuWang-Speech/RemixIT-TSE.