SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of spatial awareness in existing audio large models, which struggle to localize sound sources and separate overlapping speech in complex acoustic environments. To this end, we propose SEA-LM, a framework that introduces a layout-adaptive first-order Ambisonics spatial audio encoder (FOACODER) to process wearable microphone array signals. By integrating beamforming with multimodal large language models, designing a spatiotemporal weighted cross-entropy loss function, and employing a two-stage curriculum learning strategy, SEA-LM jointly optimizes sound source localization and selective transcription tasks. Experimental results demonstrate that the proposed method significantly reduces angular error and hallucination rates across various smart glasses array configurations, while effectively improving transcription accuracy and overall system robustness.
📝 Abstract
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.
Problem

Research questions and friction points this paper is trying to address.

spatial audio understanding
sound localization
egocentric intelligence
overlapping sound sources
wearable microphone arrays
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spatial Audio Understanding
Multimodal Large Language Model
First Order Ambisonics
Wearable Microphone Arrays
Spatio-temporal Weighted Cross-Entropy Loss
🔎 Similar Papers
No similar papers found.