๐ค AI Summary
Existing benchmarks for mental health video understanding predominantly rely on coarse-grained classification, which is insufficient to evaluate whether models genuinely possess deep psychological reasoning capabilities. To address this limitation, this work proposes MMHBenchโa multimodal, multi-perspective evaluation benchmark tailored for long-form videos, comprising 268 videos and 2,184 fine-grained questions that span third-person behavioral interpretation and first-person mental state inference. The benchmark introduces an innovative multi-perspective psychological understanding framework, integrating a role-simulation-based Multi-Agent Question Generation (MAQG) mechanism with expert validation to enable precise assessment of modelsโ psychological reasoning abilities. Experiments across 22 state-of-the-art multimodal large language models reveal significant performance gaps, highlighting the continued challenges in this domain.
๐ Abstract
Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.