🤖 AI Summary
This study addresses the tendency of vision-language models to overlook visual evidence due to an over-reliance on textual priors, resulting in prediction bias. To investigate this, we propose a four-stage diagnostic framework that decomposes the reasoning process through word completion tasks and integrates linguistic statistical modeling to systematically trace the propagation pathways of language bias and analyze the interplay between statistical priors and cross-modal coverage. Our findings reveal the dynamic evolution of linguistic priors across reasoning stages and identify the root causes of such biases. By establishing a systematic tracking methodology for language bias and elucidating the internal mechanisms underlying erroneous predictions, this work provides a theoretical foundation for enhancing the robustness of multimodal reasoning.
📝 Abstract
Despite substantial progress across downstream applications, vision-language models (VLMs) remain susceptible to language bias, often prioritizing linguistic cues over visual evidence and consequently producing incorrect predictions. Prior studies have proposed various approaches to understanding and mitigating language bias in VLMs, yet their findings often conflict due to the difficulty of tracing how language bias propagates within black-box VLMs. Building on the word completion task, we trace how language bias propagates through VLM inference by (1) proposing a diagnostic framework that decomposes the inference process into four distinct yet interdependent stages to trace the propagation of language bias; and (2) examining how two key factors underlying language bias, i.e., linguistic priors and cross-modal coverage, evolve across these stages and ultimately give rise to incorrect predictions. The linguistic prior captures the strength of statistical bias induced by the language model component of a VLM and represents the origin of language bias, whereas cross-modal coverage measures the extent to which linguistic cues cover the visual content. By decomposing inference into four stages and characterizing the interplay between linguistic priors and cross-modal coverage across these stages, we propose a systematic framework for tracing the propagation of language bias throughout the inference process; and uncover the underlying mechanism of language bias by revealing the interplay between linguistic priors and cross-modal coverage.