🤖 AI Summary
This study addresses the significant degradation in visual question answering (VQA) performance of existing vision-language models (VLMs) when confronted with ambiguous or redundant queries that violate Grice’s Cooperative Principle. To investigate this, we leverage large language models to automatically generate non-compliant questions containing ambiguous or spurious modifiers, and conduct controlled experiments comparing the processing capabilities of humans and mainstream VLMs. This work provides the first systematic quantification of VLM sensitivity to pragmatic violations and associated cognitive load disparities. Our findings demonstrate that informational violations substantially reduce VLM accuracy. Furthermore, we reveal a human–machine asymmetry in cognitive efficiency: humans exhibit lower cognitive load when processing AI-generated pragmatically violated questions, whereas VLMs suffer more severe performance degradation under such conditions.
📝 Abstract
We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.