🤖 AI Summary
This study addresses the challenge of quantitatively evaluating the perceptual capabilities of AI models with respect to visual charts. Grounded in Stevens' power law, the proposed methodology employs reference images and relative magnitude estimation to systematically assess GPT-5.5’s decoding mechanisms across twelve visual variables in the absence of legends. The research elucidates the model's intrinsic perceptual processes and establishes a standardized framework for measuring AI visual interpretation capabilities. Furthermore, it facilitates interpretable comparisons between machine and human visual perception, substantially enhancing the comparability and scientific rigor of cross-model evaluations.
📝 Abstract
We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models see no legend. A model first views a reference visual representation and estimates its magnitude, then estimates the magnitude of each subsequent image of the same representation relative to that reference. Our evaluation of twelve visual variables makes how algorithmic models read visual encodings measurable, comparable with human perception, and more interpretable to humans.