🤖 AI Summary
This study addresses the issues of spurious causal correlations and shortcut learning arising from background confounding in affective video captioning. To this end, we construct an evaluation benchmark featuring dense spatiotemporal causal annotations and propose a novel framework that decouples causal triggers from background interference. The proposed framework integrates motion-guided spatiotemporal localization, interpretable sparse affective routing, and a counterfactual contrastive objective, effectively eliminating co-occurrence bias through causal intervention and multimodal representation synthesis. Experimental results demonstrate that our method achieves state-of-the-art performance across both semantic metrics and causal faithfulness, significantly enhancing the causal reliability of generated captions.
📝 Abstract
Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal reasoning or merely exploited background confounders. To address these limitations, we first construct {EVC-CauseGround}, a comprehensive benchmark with dense spatio-temporal causal annotations. Crucially, it introduces a carefully selected {Causal-Faithfulness Subset} to explicitly quantify genuine emotion-cause attribution. Second, we propose {Causal-EVC}, an emotion-grounding captioning framework, which introduces a Motion-guided Causal Spatiotemporal Localization module to precisely decouple causal triggers from background confounders. Besides, we introduce an Interpretable Sparse Emotion Routing module. By synthesizing counterfactual representations and formulating a novel counterfactual contrastive objective, we enforce the model to anchor its emotion predictions strictly on authentic causal triggers instead of confusing background. Extensive experiments show that Causal-EVC not only achieves the best performance on semantic metrics but also exhibits significant advantages in the causal-faithfulness subset, which demonstrates that our model could mine emotional cues from genuine visual causes and mitigate co-occurrence bias for interpretable multimodal emotion understanding.