Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the superior performance of reasoning-based large language models on theory of mind (ToM) tasks stems from genuine mental-state understanding or from enhanced robustness to variations in prompts and task perturbations. By integrating reinforcement learning with verifiable rewards, a novel machine psychology experimental paradigm, and established ToM benchmarks, the authors systematically evaluate model behavior. The findings indicate that the performance advantage of reasoning models primarily arises from their ability to consistently produce correct answers across diverse prompts and perturbations, rather than from the acquisition of specialized ToM capabilities. These results challenge prevailing interpretations that attribute authentic theory of mind to large language models.
📝 Abstract
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements across a range of benchmarks. In this work, we examine the behavior of such reasoning models in ToM tasks using novel adaptations of machine psychological experiments together with results from established benchmarks. We observe that reasoning models consistently exhibit increased robustness to prompt variations and task perturbations. Our analysis suggests these gains come at least partly from models being more robust at reaching the correct answer under prompt and task variation. We read this as evidence for a robustness-based account rather than for a new ToM-specific ability.
Problem

Research questions and friction points this paper is trying to address.

Theory of Mind
Large Language Models
Robustness
Reasoning Models
Prompt Variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Theory of Mind
reasoning models
robustness
prompt variation
reinforcement learning
🔎 Similar Papers