🤖 AI Summary
This study investigates how auditory semantic representations emerge within large audio models and influence downstream decision-making. Taking Qwen2.5-Omni as the subject, we employ Jacobian lensing to extract hidden state vectors corresponding to the concept “dog” and conduct causal analysis through directional interventions. This work provides the first causal evidence demonstrating that specific concept directions exist independently of input prompts in late pre-generation layers (layers 22 and 24), significantly modulating response tendencies in both classification and description tasks. By elucidating the internal mechanisms underlying semantic emergence in multimodal large language models, this research offers a novel perspective for understanding their decision-making processes.
📝 Abstract
Multimodal large language models answer audio questions, but how they represent auditory semantics and use them in decisions remains unclear, limiting our understanding of response formation. We study dog barking in Qwen2.5-Omni-7B using Jacobian lens (J-lens) readout and directional interventions. We define the dog direction as a J-lens-derived hidden-state vector associated with dog; adding or removing its component modulates dog-related information. We find this information decodable without dog/bark prompt cues or animal-identification requirements. Directional interventions change response tendencies and some final answers, with effects concentrated in late-layer states immediately before generation across species classification, vocalization classification, and sound description. The dog direction shows no comparable advantage over controls in animal/other classification. These results provide causal-intervention evidence that the dog direction affects output scores in a task-dependent manner, most consistently at L22 and L24 immediately before generation.