AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding

📅 2024-10-16
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 1
📄 PDF
🤖 AI Summary
Auditory attention decoding (AAD) in multi-speaker scenarios typically relies on a two-stage paradigm—first reconstructing attended speech from EEG signals, then inferring attention—which suffers from poor cross-subject generalizability and suboptimal end-to-end optimization. Method: This paper proposes AADNet, an end-to-end deep neural network that jointly models speech representations and attention discrimination directly from raw EEG signals, eliminating the conventional two-stage pipeline. Contribution/Results: AADNet is the first model to enable unified cross-subject end-to-end learning for AAD. Evaluated over analysis windows of 1–40 seconds, it achieves a cross-subject classification accuracy of 82.7%, outperforming state-of-the-art methods (56.1%) by 26.6 percentage points. Comprehensive objective and subjective evaluations demonstrate that AADNet consistently surpasses linear reconstruction, canonical correlation analysis, and nonlinear reconstruction approaches. By substantially alleviating subject-specific dependency, AADNet markedly enhances model generalization across individuals.

Technology Category

Cognitive Modeling & Cognitive Systems: Neural Spike CodingMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Brain-Sensing and Analysis

Application Category

User Modeling, Personalization and Recommendation: User modeling for targeted and personalized online advertisingSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone continuous development, driven by its promising application in neuro-steered hearing devices. Most AAD algorithms are relying on the increase in neural entrainment to the envelope of attended speech, as compared to unattended speech, typically using a two-step approach. First, the algorithm predicts representations of the attended speech signal envelopes; second, it identifies the attended speech by finding the highest correlation between the predictions and the representations of the actual speech signals. In this study, we proposed a novel end-to-end neural network architecture, named AADNet, which combines these two stages into a direct approach to address the AAD problem. We compare the proposed network against the traditional approaches, including linear stimulus reconstruction, canonical correlation analysis, and an alternative non-linear stimulus reconstruction using two different datasets. AADNet shows a significant performance improvement for both subject-specific and subject-independent models. Notably, the average subject-independent classification accuracies from 56.1 % to 82.7 % with analysis window lengths ranging from 1 to 40 seconds, respectively, show a significantly improved ability to generalize to data from unseen subjects. These results highlight the potential of deep learning models for advancing AAD, with promising implications for future hearing aids, assistive devices, and clinical assessments.
Problem

Research questions and friction points this paper is trying to address.

Decoding attended speech in multi-talker environments using EEG signals
Improving auditory attention decoding accuracy with deep learning
Enhancing generalization of AAD models across unseen subjects
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-end deep learning for AAD
Combines two-stage approach into one
Improves subject-independent classification accuracy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Aarhus University | Amazon | KU Leuven | Leuven.AI - KU Leuven institute for AI | Kasteelpark Arenberg 10 | Department of Neurosciences | Research Group ExpORL
N
Nhan Duc Thanh Nguyen
Center for Ear-EEG, Department of Electrical and Computer Engineering, Aarhus University, 8200 Aarhus N, Denmark
H
Huy Phan
Amazon, Cambridge, MA 02142, USA
S
Simon Geirnaert
Department of Electrical Engineering (ESAT), Stadius Center for Dynamical Systems, Signal Processing and Data Analytics, KU Leuven and Leuven.AI - KU Leuven institute for AI, Kasteelpark Arenberg 10, B-3001 Leuven, Belgium; Department of Neurosciences, Research Group ExpORL, Herestraat 49 box 721, B-3000 Leuven, Belgium
Kaare Mikkelsen
Kaare Mikkelsen
Center for Ear-EEG, Department of Electrical and Computer Engineering, Aarhus University, 8200 Aarhus N, Denmark
Preben Kidmose
Preben Kidmose
Professor, Department of Electrical and Computer Engineering, Aarhus University.
Biomedical EngineeringSignal ProcessingMachine LearningEEGear-EEG.