🤖 AI Summary
Social videos often convey implicit pragmatic meanings—such as humor and irony—through multimodal cues and cultural context, yet current video-language models struggle to interpret such non-literal content effectively. To address this gap, this work proposes DrivelHub+, the first video-language benchmark specifically designed for implicit pragmatic understanding. It comprises 1,000 social videos annotated with human-constructed implicit narratives, emphasizing the inference of intentional deeper meanings from seemingly nonsensical surface content. Evaluated through natural language explanation generation and video–text bidirectional retrieval, DrivelHub+ moves beyond conventional explicit recognition paradigms to focus on cross-modal alignment and contextual reasoning. The benchmark provides a diagnostic tool for assessing the discrepancy between multimodal perception and pragmatic comprehension in existing models, revealing their current inability to reliably infer intended meaning from observed content.
📝 Abstract
Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.