SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of sound-spatial correspondence in multi-source scenarios caused by early fusion in existing spatial audio foundation models. To this end, we propose SAIL, a framework that employs decoupled spatial audio Transformers to separately encode Mel spectrograms and inter-channel phase difference features. By integrating source-discriminative queries with a dual-stream Q-Former mechanism for alignment with large language models, SAIL effectively preserves both the acoustic-spatial structure and source-level correspondences. Experimental results demonstrate that SAIL significantly outperforms existing baseline models across multiple tasks, including multi-source sound event detection, direction-of-arrival and distance estimation, and spatial reasoning.
📝 Abstract
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Problem

Research questions and friction points this paper is trying to address.

spatial audio
large language models
multi-source scenes
acoustic-spatial correspondence
early fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Audio LLM
Disentangled Encoding
Dual-Stream Q-Former
Source-Discriminative Queries
Multi-Source Reasoning
🔎 Similar Papers
2024-02-02International Conference on Machine LearningCitations: 14
💼 Related Jobs
No related jobs found.
Z
Zhengding Luo
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
J
Jinyang Wu
Singapore Management University
H
Haozhe Ma
Tencent Hy Frontier Lab, Singapore
Y
Yanghao Zhou
Department of Computer Science and Technology, Beijing Institute of Technology, China
Woon-Seng Gan
Woon-Seng Gan
Professor of Audio Engineering and Director of Smart Nation Lab @ Nanyang Technological University,
Active Noise ControlMachine & Deep LearningSpatial AudioPerceptual Evaluation
Wenwu Wang
Wenwu Wang
Professor, University of Surrey, UK
signal processingmachine learningmachine listeningaudio/speech/audio-visualmultimodal fusion