Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in large audio-language models where irrelevant audio interferes with textual reasoning and aggregated accuracy obscures pairwise drift phenomena. To tackle this, we propose the ICAP-Gate mechanism, which leverages pairwise drift analysis and mechanism-guided intervention to precisely localize and control architecture-specific late-stage audio pathways. This approach establishes a design principle for selective modality influence control, enabling task-conditioned selective listening. Extensive evaluations across four models and two benchmarks demonstrate that our method significantly reduces both the influence rate and answer flip rate, effectively suppressing interference while preserving ASR performance. Furthermore, the inference latency of ICAP-Gate is substantially lower than that of self-consistency methods.
📝 Abstract
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
Problem

Research questions and friction points this paper is trying to address.

Large Audio-Language Models
Task-irrelevant audio interference
Paired drift
Multimodal reasoning
Modality influence control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Listening
ICAP-Gate
Paired Drift Analysis
Large Audio-Language Models
Mechanism-Guided Control
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yulin Sun
College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment
K
Kele Xu
College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment
Y
Yong Dou
College of Computer Science and Technology, National University of Defense Technology; State Key Laboratory of Complex & Critical Software Environment