An Exam for Active Observers

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of effective evaluation of human-like active visual observation capabilities in existing vision-language models. To this end, we propose ActiveVision, a novel benchmark that establishes the first quantifiable framework for assessing active visual observation behaviors in multimodal large language models. Grounded in cognitive science principles, the benchmark comprises 17 dynamic visual tasks across three categories, all requiring iterative perception. By integrating self-generated visual programs with human evaluation, it compels models to engage in multiple rounds of interactive observation rather than producing single-pass static descriptions. Experimental results reveal that state-of-the-art models—including GPT-5.5 and Claude Fable 5—perform poorly on this benchmark, achieving at best 10.6% accuracy, starkly contrasting the human average of 96.1%, thereby exposing a fundamental deficiency in their active visual observation abilities.
📝 Abstract
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.
Problem

Research questions and friction points this paper is trying to address.

active observation
multimodal large language models
vision-language benchmark
visual perception
perception-reasoning loop
Innovation

Methods, ideas, or system contributions that make the work stand out.

active vision
multimodal large language models
visual perception
closed-loop perception
vision-language benchmark
🔎 Similar Papers
No similar papers found.