MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing evaluation methods in assessing multimodal large language models’ capacity for dynamic reasoning about character motivations in continuous visual narratives, particularly their inability to model cumulative and evolving behavioral drivers. To bridge this gap, we propose the first benchmark framework specifically designed for multimodal motivational reasoning in sequential visual storytelling. Grounded in Maslow’s hierarchy of needs and Reiss’s theory of basic desires, our framework introduces a novel dataset that integrates temporally ordered image sequences with psychologically grounded motivation labels. Experimental results demonstrate that while current state-of-the-art multimodal models exhibit reasonable performance in static recognition tasks, they struggle to maintain coherent motivational inference across narrative contexts, revealing a critical deficiency in dynamic social intelligence and establishing a foundational benchmark for future research in this domain.
📝 Abstract
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
sequential motivation
social intelligence
visual narratives
motivation inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal sequential reasoning
motivation inference
visual narrative understanding
social intelligence benchmarking
context accumulation
🔎 Similar Papers
No similar papers found.