🤖 AI Summary
This study addresses the issue in large audio-language models where irrelevant audio interferes with textual reasoning and aggregated accuracy obscures pairwise drift phenomena. To tackle this, we propose the ICAP-Gate mechanism, which leverages pairwise drift analysis and mechanism-guided intervention to precisely localize and control architecture-specific late-stage audio pathways. This approach establishes a design principle for selective modality influence control, enabling task-conditioned selective listening. Extensive evaluations across four models and two benchmarks demonstrate that our method significantly reduces both the influence rate and answer flip rate, effectively suppressing interference while preserving ASR performance. Furthermore, the inference latency of ICAP-Gate is substantially lower than that of self-consistency methods.
📝 Abstract
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.