🤖 AI Summary
This work addresses the lack of effective evaluation methods for assessing the joint perceptual and reasoning capabilities of large vision-language models (LVLMs) in open-world scenarios. To this end, it introduces real-world visual illusions into LVLM evaluation for the first time, constructing IllusionReasoning—a benchmark comprising diverse, human-annotated question-answer pairs—and proposes an open-ended evaluation paradigm that jointly probes perception and reasoning. Systematic experiments reveal that prevailing LVLMs consistently struggle with visual illusions, exhibiting insufficient reasoning abilities and a tendency to misrepresent objective reality. These findings indicate that current model performance is likely overestimated and highlight critical limitations in their capacity to reconcile perception with grounded reasoning, thereby offering new directions for future model development and refinement.
📝 Abstract
Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.