AURA: Uncertainty-Routed Activation Editing for Acoustic Grounding in Speech Foundation Models

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AURA方法,通过编辑解码器交叉注意力头来解决声学基础模型中的听觉幻觉问题,显著降低了非语音音频的幻觉率。
📝 Abstract
Attention encoder-decoder (AED) Speech Foundation Models achieve strong ASR performance but can generate acoustically unsupported text when inputs contain no speech, weak acoustic evidence, or unreliable transcription. We propose AURA: Activation-editing with Uncertainty-Routed Adaptation, an ultra-efficient representation-editing method that freezes the pretrained model and applies sparse scale-and-shift edits to decoder cross-attention heads. AURA dynamically routes edits using cross-attention uncertainty features that capture over-concentration, diffuse attention, and abrupt frame shifts. We evaluate AURA on four datasets spanning non-speech hallucination and speech grounding stressors, including imperfect-label child speech, imperfect-label adult speech, and disfluent speech. On non-speech audio, AURA reduces hallucination rate from 89.18% to 1.94% without prior hallucination-head identification. On imperfect-label corpora, AURA approaches LoRA WER while using roughly 500x fewer trainable parameters. Sensitivity analysis and qualitative cross-attention examples are consistent with AURA's uncertainty-routed editing behavior, supporting dynamic activation editing as a practical path for grounding AED speech models.
Problem

Research questions and friction points this paper is trying to address.

Attention Encoder-Decoder
Acoustic Grounding
Speech Foundation Models
Uncertainty
Activation Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation-editing
Uncertainty-Routed Adaptation
Cross-attention Uncertainty
Sparse Edits
Acoustic Grounding
Natarajan Balaji Shankar
Natarajan Balaji Shankar
Graduate Student in Electrical and Computer Engineering, University of California Los Angeles
Automatic Speech RecognitionChildren's Speech
Z
Zilai Wang
Department of Electrical and Computer Engineering, University of California Los Angeles
Z
Zihan Wang
Department of Electrical and Computer Engineering, University of California Los Angeles
Mohan Shi
Mohan Shi
University of California, Los Angeles | Ex-USTC
Speech RecognitionSpeech LLMMulti-modal LLMDeep Learning
K
Kaiyuan Zhang
Department of Electrical and Computer Engineering, University of California Los Angeles
Abeer Alwan
Abeer Alwan
Professor of Electrical Engineering, UCLA
Speech Processing