🤖 AI Summary
This study addresses the challenges of tri-modal binding across text, audio, and vision, as well as audio-visual connection misalignment, encountered by large audio-visual models during multi-speaker dialogue reasoning. Specifically, this work elucidates a symbolic tri-modal binding mechanism and proposes a training-free audio-visual prompting approach based on Active Speaker Detection (ASD), coupled with a lightweight fine-tuning strategy. By leveraging symbolic representations, the proposed method achieves precise cross-modal alignment. Extensive experiments demonstrate that this approach yields significant performance improvements across four dialogue benchmarks and three general audio-visual benchmarks, effectively validating its robust multi-modal reasoning capabilities and strong generalizability.
📝 Abstract
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across four conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.