đ¤ AI Summary
This work addresses multi-level alignment between visual and linguistic representations in Large Vision-Language Models (LVLMs), systematically identifying three-tier semantic misalignmentsâobject-, attribute-, and relation-levelâand uncovering their root causes in data curation, model architecture, and inference dynamics.
Method: We propose the first explainability-oriented LVLM alignment analysis framework, featuring a novel three-tier misalignment taxonomy. Leveraging integrated cross-modal diagnosticsâincluding attribution analysis, feature visualization, and probing experimentsâwe combine theoretical modeling with empirical validation.
Contribution/Results: We unify and comparatively analyze two major mitigation strategiesâparameter freezing and fine-tuningâdemonstrating their respective efficacy across alignment levels. Our findings enable the development of a standardized alignment evaluation protocol, advancing both theoretical understanding and practical design principles for robust, trustworthy multimodal models.
đ Abstract
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information. However, the critical challenge of alignment between visual and linguistic representations is not fully understood. This survey presents a comprehensive examination of alignment and misalignment in LVLMs through an explainability lens. We first examine the fundamentals of alignment, exploring its representational and behavioral aspects, training methodologies, and theoretical foundations. We then analyze misalignment phenomena across three semantic levels: object, attribute, and relational misalignment. Our investigation reveals that misalignment emerges from challenges at multiple levels: the data level, the model level, and the inference level. We provide a comprehensive review of existing mitigation strategies, categorizing them into parameter-frozen and parameter-tuning approaches. Finally, we outline promising future research directions, emphasizing the need for standardized evaluation protocols and in-depth explainability studies.