🤖 AI Summary
This study investigates whether state-of-the-art vision-language models possess coherent theory of mind (ToM) capabilities. Leveraging two established psychological paradigms—the Keysar director task and the Frith-Happé animated triangles task—and employing the Castelli scoring criteria alongside chain-of-thought reasoning analysis, the authors systematically evaluate ToM performance across nine prominent models. The findings reveal, for the first time, a significant dissociation across tasks: models exhibit child-like egocentric biases in the director task, while their performance on the triangles task aligns more closely with that of high-functioning autistic adults. Critically, no model achieves typical adult-level performance on both tasks. These results challenge the prevailing assumption that current models embody a unified, human-like ToM, underscoring instead the task-dependent nature and inherent limitations of their ToM capacities.
📝 Abstract
Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchmarks: the Keysar Director Task (visual perspective-taking under egocentric interference) and the Frith-Happé animated triangles scored with the Castelli rubric (intention attribution from pure motion). On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78\% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models. On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random stay near TD. No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles, and the most TD-like model on the triangles is child-like on the Director Task. We report group-level descriptions, not diagnostic labels for any model.