🤖 AI Summary
This work addresses the challenges of extracting the target speaker and achieving seamless transitions during within-trial shifts of auditory attention, which are hindered by neural noise and the inherent latency of electroencephalography (EEG). To overcome these limitations, the authors propose the SAGE framework, which employs a robust speech separator to generate two candidate streams and incorporates a switching-aware soft gating mechanism. This mechanism integrates EEG-guided dynamic fusion, latency-compensated alignment, and an uncertainty-driven conservative strategy to enable smooth attention switching and suppress transition artifacts. Evaluated in dynamic attention-switching scenarios, the system achieves 8.67 dB SI-SDR and 88.24% STOI, with an average switching latency reduced to 2.04 seconds, significantly enhancing the real-time performance and robustness of EEG-informed speech separation.
📝 Abstract
EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.