🤖 AI Summary
This study addresses the unclear mechanisms underlying cross-layer fusion of visual and textual information in multimodal large language models. We propose an architecture-aware diagnostic framework that systematically compares concatenation-based and native multimodal architectures. Through alignment decoupling, attention entropy analysis, intrinsic dimensionality estimation, causal intervention, and vision-specific Centered Kernel Alignment (CKA), we reveal how feature spaces are reorganized under different architectural paradigms. Our findings indicate that concatenation-based models exhibit a text-dominant fusion trajectory, whereas native models achieve early-stage vision-language co-adaptation. This work elucidates the distinct multimodal fusion mechanisms inherent to these two architectural paradigms, providing a principled theoretical foundation for future model design.
📝 Abstract
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.