Architecture-Dependent Fusion Pathways in MLLMs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms underlying cross-layer fusion of visual and textual information in multimodal large language models. We propose an architecture-aware diagnostic framework that systematically compares concatenation-based and native multimodal architectures. Through alignment decoupling, attention entropy analysis, intrinsic dimensionality estimation, causal intervention, and vision-specific Centered Kernel Alignment (CKA), we reveal how feature spaces are reorganized under different architectural paradigms. Our findings indicate that concatenation-based models exhibit a text-dominant fusion trajectory, whereas native models achieve early-stage vision-language co-adaptation. This work elucidates the distinct multimodal fusion mechanisms inherent to these two architectural paradigms, providing a principled theoretical foundation for future model design.
📝 Abstract
Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary analysis, we use visual CKA to examine the Platonic Representation Hypothesis. Together, these analyses reveal two distinct fusion pathways: concatenation models follow a text-first, vision-later pathway, whereas native models exhibit earlier visual-textual co-adaptation and feature-space reorganization. This work provides a mechanistic perspective for understanding multimodal fusion and supports architecture-aware diagnostics of multimodal representations.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Multimodal Fusion
Architecture Paradigms
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Fusion Pathways
Causal Intervention
Intrinsic Dimensionality
Platonic Representation Hypothesis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hebao Zhu
Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
Dongxia Wu
Dongxia Wu
Stanford University
Uncertainty QuantificationAI for ScienceTrustworthy AISpatiotemporal Modeling