The Alignment Illusion in Multimodal Large Language Models

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prevalent misinterpretation of vision-text similarity in multimodal large language models (MLLMs) as evidence of cross-modal content interaction, which often obscures spurious alignment induced by shared language model pathways and weight anisotropy. Through controlled intervention experiments on thirteen mainstream MLLMs—incorporating Gaussian noise injection, CKA/SVCCA analysis, and principal angle cosine computation—this work proposes a novel metric termed the principal angle gap (PA gap). This metric effectively decouples weight-induced similarity from genuine multidimensional visual structure. Experimental results demonstrate that under hierarchical visual corruption, the PA gap tracks task accuracy more consistently than conventional scalar metrics. These findings confirm that internal alignment serves merely as a geometric diagnostic tool rather than a direct proxy for semantic content interaction.
📝 Abstract
Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing weight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Alignment Illusion
Cross-modal Interaction
Representation Similarity
Weight-induced Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Alignment Illusion
Multimodal Large Language Models
Principal-Angle Gap
Weight-induced Alignment
Geometric Diagnostic