SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching

📅 2026-08-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of extracting the target speaker and achieving seamless transitions during within-trial shifts of auditory attention, which are hindered by neural noise and the inherent latency of electroencephalography (EEG). To overcome these limitations, the authors propose the SAGE framework, which employs a robust speech separator to generate two candidate streams and incorporates a switching-aware soft gating mechanism. This mechanism integrates EEG-guided dynamic fusion, latency-compensated alignment, and an uncertainty-driven conservative strategy to enable smooth attention switching and suppress transition artifacts. Evaluated in dynamic attention-switching scenarios, the system achieves 8.67 dB SI-SDR and 88.24% STOI, with an average switching latency reduced to 2.04 seconds, significantly enhancing the real-time performance and robustness of EEG-informed speech separation.
📝 Abstract
EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.
Problem

Research questions and friction points this paper is trying to address.

EEG-guided speech extraction
auditory attention switching
in-trial switching
target speaker extraction
neural decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

switch-aware gating
EEG-guided speech extraction
in-trial attention switching
latency-compensated alignment
uncertainty-driven strategy
X
Xuefei Wang
Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China
X
Ximin Chen
Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China
Y
Yuting Ding
Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China
C
Chunlin Li
School of Biomedical Engineering, Capital Medical University, Beijing, China
Fei Chen
Fei Chen
Professor, Southern University of Science and Technology
speech communicationspeech enhancementassistive hearing technologybrain-computer interfacebiomedical signal processing