MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing multimodal large language models in accurately reasoning about others’ beliefs, intentions, and knowledge states in natural multi-participant meeting scenarios, particularly when interpreting nonverbal cues, implicit attitudes, and false consensus—complex social dynamics that current benchmarks overlook. To bridge this gap, we introduce MeetingToM, the first multimodal Theory of Mind (ToM) evaluation benchmark specifically designed for real-world meetings. MeetingToM employs a hierarchical framework to assess models across three levels: individual mental state prediction, interlocutor understanding, and group consensus reasoning, with a novel focus on false consensus phenomena. By integrating multimodal signals from speech, behavior, and language into a unified evaluation protocol, our systematic assessment of state-of-the-art models reveals significant deficiencies in fusing nonverbal cues, inferring hidden attitudes, and detecting false consensus, thereby filling a critical void in modeling implicit social states and group-level dynamics.
📝 Abstract
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.
Problem

Research questions and friction points this paper is trying to address.

Theory of Mind
Multimodal LLMs
multi-party meetings
pseudo-consensus
social reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Theory of Mind
Multimodal LLMs
Multi-party Meetings
Pseudo-consensus
Social Reasoning
🔎 Similar Papers