Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Bearings框架,通过自监督学习从一阶Ambisonics中提取声场嵌入,为声音场景提供缺失的空间表示,提高声音事件定位和检测性能。
📝 Abstract
Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.
Problem

Research questions and friction points this paper is trying to address.

self-supervised
soundfield embeddings
first-order Ambisonics
spatial representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-supervised framework
soundfield embeddings
first-order Ambisonics
masked auto-encoder
spatial information
🔎 Similar Papers
2024-09-15arXiv.orgCitations: 0