When Does the Concept of"Dog"Emerge in an Audio LLM?

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how auditory semantic representations emerge within large audio models and influence downstream decision-making. Taking Qwen2.5-Omni as the subject, we employ Jacobian lensing to extract hidden state vectors corresponding to the concept “dog” and conduct causal analysis through directional interventions. This work provides the first causal evidence demonstrating that specific concept directions exist independently of input prompts in late pre-generation layers (layers 22 and 24), significantly modulating response tendencies in both classification and description tasks. By elucidating the internal mechanisms underlying semantic emergence in multimodal large language models, this research offers a novel perspective for understanding their decision-making processes.
📝 Abstract
Multimodal large language models answer audio questions, but how they represent auditory semantics and use them in decisions remains unclear, limiting our understanding of response formation. We study dog barking in Qwen2.5-Omni-7B using Jacobian lens (J-lens) readout and directional interventions. We define the dog direction as a J-lens-derived hidden-state vector associated with dog; adding or removing its component modulates dog-related information. We find this information decodable without dog/bark prompt cues or animal-identification requirements. Directional interventions change response tendencies and some final answers, with effects concentrated in late-layer states immediately before generation across species classification, vocalization classification, and sound description. The dog direction shows no comparable advantage over controls in animal/other classification. These results provide causal-intervention evidence that the dog direction affects output scores in a task-dependent manner, most consistently at L22 and L24 immediately before generation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Jacobian lens
directional intervention
audio LLM
causal interpretability
hidden-state representation
🔎 Similar Papers
No similar papers found.
Z
Zhe Wang
Huazhong University of Science and Technology, Wuhan, China
Shiqi Liu
Shiqi Liu
Central China Normal University & Monash University & Hunan Normal University
Large Language ModelsAgentAI for Social ScienceEducational Data MiningLearning Analytics
R
Ruiyun Zhong
Huazhong University of Science and Technology, Wuhan, China
T
Tiechong Zhu
Huazhong University of Science and Technology, Wuhan, China
Y
Yihua Tan
Huazhong University of Science and Technology, Wuhan, China