Shared Representation Learning for Reference-Guided Targeted Sound Detection

📅 2026-03-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of detecting and localizing target sounds in complex acoustic scenes by proposing a unified encoder-based shared representation learning framework. Departing from conventional conditional embedding mechanisms, the method jointly encodes reference and mixture audio signals within a shared semantic space and employs multi-task learning to simultaneously optimize detection and localization performance. By innovatively aligning audio embeddings and enabling end-to-end training, the approach significantly enhances generalization to unseen sound categories while simplifying model architecture. Evaluated on the URBAN-SED dataset, the proposed method achieves a segment-level F1 score of 83.15% and an overall accuracy of 95.17%, establishing a new state-of-the-art performance.

Technology Category

Machine Learning: Multimodal LearningSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Human listeners exhibit the remarkable ability to segregate a desired sound from complex acoustic scenes through selective auditory attention, motivating the study of Targeted Sound Detection (TSD). The task requires detecting and localizing a target sound in a mixture when a reference audio of that sound is provided. Prior approaches, rely on generating a sound-discriminative conditional embedding vector for the reference and pairing it with a mixture encoder, jointly optimized with a multi-task learning approach. In this work, we propose a unified encoder architecture that processes both the reference and mixture audio within a shared representation space, promoting stronger alignment while reducing architectural complexity. This design choice not only simplifies the overall framework but also enhances generalization to unseen classes. Following the multi-task training paradigm, our method achieves substantial improvements over prior approaches, surpassing existing methods and establishing a new state-of-the-art benchmark for targeted sound detection, with a segment-level F1 score of 83.15% and an overall accuracy of 95.17% on the URBAN-SED dataset.
Problem

Research questions and friction points this paper is trying to address.

Targeted Sound Detection
Reference-Guided
Sound Segregation
Auditory Attention
Shared Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shared Representation Learning
Targeted Sound Detection
Unified Encoder Architecture
Reference-Guided Audio Processing
Multi-task Learning
🔎 Similar Papers
No similar papers found.
S
Shubham Gupta
Speech Information and Processing Lab, Indian Institute of Technology Hyderabad, India
A
Adarsh Arigala
Speech Information and Processing Lab, Indian Institute of Technology Hyderabad, India
B
B. R. Dilleswari
Speech Information and Processing Lab, Indian Institute of Technology Hyderabad, India
S
Sri Rama Murty Kodukula
Speech Information and Processing Lab, Indian Institute of Technology Hyderabad, India