I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了音频大模型仅被动响应的问题,通过引入Interrupt和Silent Modeling方法实现主动音频辅助,提高对聋哑用户群体的实用性。
📝 Abstract
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{<interrupt>} and \texttt{<silent>}, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6\% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.
Problem

Research questions and friction points this paper is trying to address.

AudioLLMs
Proactive Audio Assistance
Wearable Applications
Innovation

Methods, ideas, or system contributions that make the work stand out.

proactive audio assistance
Interrupt and Silent Modeling (ISM)
audio stream monitoring
🔎 Similar Papers
A
Amit Kumar Singh Yadav
Meta Reality Labs, USA
R
Ritvik Shrivastava
Meta Reality Labs, USA
X
Xuan Zhang
Meta Reality Labs, USA
Seungwhan Moon
Seungwhan Moon
Facebook, Carnegie Mellon University
Dialog SystemsTransfer LearningMultimodal LearningNatural Language Processing
S
Shashank Jain
Meta Reality Labs, USA
P
Pinar Donmez
Meta Reality Labs, USA
B
Babak Damavandi
Meta Reality Labs, USA