GAANet: Global-guided Asymmetric Attention Network for Audio-Visual Speech Separation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the modeling limitations in multimodal feature fusion caused by symmetric processing and the absence of global information. To this end, we propose an asymmetric multi-scale fusion framework that enables audio and visual streams to extract interactive features at their respective optimal temporal resolutions, thereby eliminating the bottleneck of symmetric downsampling. Furthermore, a global-guided attention mechanism is introduced to compress each modality into compact tokens, leveraging high-level semantics to guide both intra- and inter-modal fusion across scales for efficient audiovisual separation. With only 3.3M parameters, the proposed model achieves state-of-the-art performance on the LRS2 and VoxCeleb2 datasets, yielding SI-SNRi scores of 16.5 dB and 14.0 dB, respectively, effectively balancing lightweight design with high precision.
📝 Abstract
Multi-scale design is crucial for efficient audio-visual speech separation, yet effectively modeling multi-scale information for audio-visual feature fusion remains challenging. We argue that the limited capacity of existing approaches primarily arises from: 1) treating features from different modalities in the same manner, and 2) overlooking the role of global features. To address these issues, we propose a Global-guided Asymmetric Attention Network (GAANet). Our model introduces two core innovations: first, an asymmetric multi-scale fusion framework that allows audio and visual streams to extract and interact with features at their respective optimal temporal resolutions, removing the need for symmetric temporal downsampling; second, a global-guided attention mechanism that compresses each modality into a compact global token with a temporal dimension of one, which then provides high-level semantic cues to guide both intra- and inter-modal fusion across scales. Experiments on LRS2 and VoxCeleb2 demonstrate that GAANet achieves state-of-the-art performance, reaching 16.5 dB SI-SNRi on LRS2 and 14.0 dB on VoxCeleb2, while maintaining a lightweight computational profile with only 3.3M parameters and 19.8G MACs. These results highlight the strong potential of asymmetric temporal modeling and global guidance for efficient and robust multimodal fusion. The source code is publicly accessible at https://github.com/redizzy/GAANet
Problem

Research questions and friction points this paper is trying to address.

Audio-Visual Speech Separation
Multi-scale Fusion
Multimodal Feature Fusion
Global Features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Audio-Visual Speech Separation
Asymmetric Multi-scale Fusion
Global-guided Attention
Multimodal Feature Fusion
Lightweight Network
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyuan Zhang
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, Anhui 230601, China
J
Jingyuan Xu
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei, Anhui 230601, China
Yiming Tang
Yiming Tang
Hefei University of Technology
Fuzzy reasoningMathematical logicMachine learningImage processingAffective computing
Liu Liu
Liu Liu
Hefei University of Technology
deep learningcomputer visionrobotics
Dan Guo
Dan Guo
IEEE senior member, Professor, Hefei University of Technology
Multimedia ComputingArtificial Intelligence