EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing video generation models conflate emotional ambiance, semantic cues, and temporal dynamics within a single textual condition, hindering fine-grained affective control. This work proposes EmoWorld, a framework that achieves the first triple disentanglement within a frozen video diffusion Transformer by leveraging a pre-constructed library of affective directions and cues. During inference, it independently steers visual atmosphere, semantic affect, and temporal affective evolution without requiring model fine-tuning. The approach enables multidimensional emotion-controllable video generation and demonstrates cross-model portability. Evaluated on the Wan2.2 benchmark, EmoWorld significantly improves emotional alignment (+37%), affective cue detection (+36%), and temporal transition monotonicity (+15%), with consistent effectiveness validated across 27 emotion categories and diverse generation settings.
📝 Abstract
Emotion shapes how viewers interpret a scene, yet existing video generators entangle global atmosphere, affect-bearing semantic cues, and temporal progression within a single text condition. We present EmoWorld, a framework that decouples these factors within a frozen flow-matching video diffusion transformer (Video DiT). A one-time preparation stage extracts layer-specific affect directions and a reusable cue library from geometry-preserving neutral and emotion-edited panoramas. At inference, Visual Atmosphere Steering (VAS) injects atmosphere directions into hidden states, Semantic Affective Steering (SAS) isolates a separately scalable prompt residual for semantic cues, and Temporal Affective Steering (TAS) interpolates endpoint residual fields across denoising and video time. On Wan2.2, VAS improves target-emotion alignment by 19% while reducing a temporal-fluctuation proxy by 48%; SAS improves target-emotion alignment by 37% and increases detected affect-bearing cues by 36%; and TAS improves transition monotonicity by 15% over the strongest baseline. EmoWorld is evaluated across 27 emotion categories in text-to-video and image-to-video settings, demonstrates portability across multiple Video-DiT backbones, and supports camera-conditioned composition without updating generator parameters.
Problem

Research questions and friction points this paper is trying to address.

emotional video generation
affective control
decoupled representation
video diffusion models
controllable generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupled Affective Control
Video Diffusion Transformer
Visual Atmosphere Steering
Semantic Affective Steering
Temporal Affective Steering