π€ AI Summary
This work addresses the limitations of traditional video quality assessment, which emphasizes visual fidelity yet fails to capture the community engagement and resonance elicited by user-generated content (UGC). To bridge this gap, the paper introduces a novel task, CASTER, which evaluates whether UGC elicits positive community feedback through multimodal attribute analysis, and proposes the MEDEA architecture to enable human-centered quality judgment. The core innovations include the first-of-its-kind Social Chain-of-Thought mechanism that simulates diverse viewer perspectives to model βcommunity mind,β and a hybrid training strategy combining supervised fine-tuning with process-based reinforcement learning guided by a Social Alignment Reward to align reasoning pathways with human social cognition. Evaluated on the newly curated human-annotated benchmark CASTER-Bench, the proposed approach substantially outperforms existing models, yielding interpretable reasoning that closely aligns with real-world community responses.
π Abstract
Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC). In this work, we propose a paradigm shift from signal-centric metrics to human-centric resonance assessment. We introduce CASTER (Community-Aware Assessment of Social Textual Engagement and Resonance), a new task that evaluates whether a UGC item achieves positive community resonance based on its multimodal attributes rather than visual quality alone. To address this, we present MEDEA (Multimodal Engagement-Driven Evaluation Architecture), which introduces a novel Social Chain-of-Thought (Social-CoT) mechanism. Unlike traditional logical CoT, Social-CoT performs multimodal perspective-taking, instantiating diverse viewer personas to simulate collective cognitive and emotional reactions (i.e., the "community mind") before deriving a quality judgment. MEDEA is trained via a two-stage approach involving supervised fine-tuning and process-supervised reinforcement learning with Social Alignment Reward to ensure reasoning paths are grounded in authentic human social cognition. To support this task, we release CASTER-Bench, a comprehensive human-annotated benchmark covering diverse UGC categories. Experiments demonstrate that MEDEA significantly outperforms state-of-the-art baselines on CASTER-Bench while providing interpretable and empathetic reasoning paths that align with real community feedback.