🤖 AI Summary
This study addresses the limitations of existing evaluation methods in assessing multimodal large language models’ capacity for dynamic reasoning about character motivations in continuous visual narratives, particularly their inability to model cumulative and evolving behavioral drivers. To bridge this gap, we propose the first benchmark framework specifically designed for multimodal motivational reasoning in sequential visual storytelling. Grounded in Maslow’s hierarchy of needs and Reiss’s theory of basic desires, our framework introduces a novel dataset that integrates temporally ordered image sequences with psychologically grounded motivation labels. Experimental results demonstrate that while current state-of-the-art multimodal models exhibit reasonable performance in static recognition tasks, they struggle to maintain coherent motivational inference across narrative contexts, revealing a critical deficiency in dynamic social intelligence and establishing a foundational benchmark for future research in this domain.
📝 Abstract
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.