🤖 AI Summary
This study addresses the pervasive yet underexplored challenge of implicit social context—such as emotions, intentions, and stances—in real-world interpersonal interactions, for which no systematic modeling framework currently exists. The work formally defines the task of Implicit Social Context Analysis and introduces a high-quality multimodal benchmark dataset comprising 3,108 annotated instances. To tackle this task, the authors propose a Conflict-driven Abductive Reasoning framework (CoDAR), which infers latent psychological states by modeling cognitive dissonance between verbal expressions and actual behaviors. Integrating multimodal large language models with fine-grained human annotations, CoDAR substantially advances performance on this task, though a notable gap remains compared to human-level understanding.
📝 Abstract
Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements. Such implicit social contexts are pervasive in real-world interactions, yet there remains a lack of a formal and systematic framework for studying them. In this paper, we introduce Implicit Social Context Analysis (MoCA), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance. We construct a high-quality benchmark containing 3,108 multimodal instances collected from real-world sources, with fine-grained cognitive annotations revealing who expresses what toward whom, as well as how and why it is conveyed. Using the MoCA dataset, we show that state-of-the-art multimodal large language models struggle significantly with this task because of their reliance on explicit cues and limited ability to reason over latent social contexts. To address this challenge, we propose Conflict-Driven Abductive Reasoning (CoDAR), a novel framework that models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, thereby enabling the inference of hidden mental states. Extensive experiments demonstrate that CoDAR substantially improves model performance. Nevertheless, a large gap from human reasoning remains, highlighting the fundamental difficulty of implicit social understanding.