🤖 AI Summary
Multimodal large language models (MLLMs) suffer from insufficient vision–language alignment, over-relying on linguistic priors while underutilizing fine-grained visual region understanding. This work presents the first systematic investigation into the internal visual comprehension mechanisms of MLLMs and proposes a novel paradigm—“visual depth enhancement” and “vision–language dynamic alignment”—to overcome the language-prior dominance bottleneck. Methodologically, we integrate attention mechanism analysis, visual feature disentanglement, cross-modal gated alignment, and token-level supervision targeting vision-dependent token prediction to strengthen visual representation learning and vision-guided language generation. Experiments demonstrate significant improvements: upstream vision-dependent token prediction accuracy increases notably, and average performance on vision-intensive tasks improves by 10 percentage points. These results validate that our paradigm effectively enhances multimodal alignment capability.
📝 Abstract
Achieving deep alignment between vision and language remains a central challenge for Multimodal Large Language Models (MLLMs). These models often fail to fully leverage visual input, defaulting to strong language priors. Our approach first provides insights into how MLLMs internally build visual understanding of image regions and then introduces techniques to amplify this capability. Specifically, we explore techniques designed both to deepen the model's understanding of visual content and to ensure that these visual insights actively guide language generation. We demonstrate the superior multimodal understanding of our resultant model through a detailed upstream analysis quantifying its ability to predict visually-dependent tokens as well as 10 pt boost on visually challenging tasks.