🤖 AI Summary
This study addresses the challenge that current AI agents struggle to spontaneously detect misleading visual designs—such as decorative clutter, axis manipulation, and scale distortion—in charts without explicit prompting. It presents the first systematic evaluation of large language model–driven AI agents’ sensitivity to graphical integrity flaws under unguided conditions, leveraging the BeauVis and PREVis standardized scales to automatically assess their judgments of chart aesthetics and readability. The findings reveal that AI agents consistently assign high scores even when graphical integrity is compromised, indicating a pronounced tendency to overlook deceptive visual practices. This highlights a critical limitation in existing models’ capacity to perceive and evaluate the trustworthiness of data visualizations, underscoring the need for improved mechanisms to support faithful visual interpretation.
📝 Abstract
AI agents are increasingly used as low-cost proxies for early visualization evaluation. In an initial study of deliberately flawed charts, we test whether agents spontaneously penalise chart junk and misleading encodings without being prompted to look for errors. Using established scales (BeauVis and PREVis), the agent evaluated visualizations containing decorative clutter, manipulated axes, and distorted proportional cues. The ratings of aesthetic appeal and perceived readability often remained relatively high even when graphical integrity was compromised. These results suggest that un-nudged AI agent evaluation may underweight integrity-related defects unless such checks are explicitly elicited.