WorldAlign: Decoupled 4D Reward for World-Consistent Video Generation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of simultaneously maintaining static background and dynamic subject consistency in video generation, as well as the unreliability of feedback signals. To this end, it proposes a decoupled 4D reward framework that separates static and dynamic regions via semantic-guided mask reprojection, aligning them with geometric priors and vision-language model (VLM) priors, respectively. By integrating a VLM-as-a-judge mechanism, the framework enables online post-training without requiring human preference annotations. Experimental results demonstrate that the proposed method significantly enhances spatiotemporal consistency while preserving natural motion dynamics on the Wan series models, outperforming existing approaches.
📝 Abstract
Faithful visual world simulation requires generated videos to maintain 4D world consistency, encompassing both static and dynamic consistency. Static consistency requires coherent 3D structure in static environments across viewpoints, while dynamic consistency requires plausible subject motion and consistent appearance over time. Geometry-aware post-training offers a promising way to improve world consistency. However, existing methods often rely on a static-scene assumption. Even those that accommodate dynamic scenes struggle to provide reliable static-consistency feedback, while dynamic consistency is often overlooked or inadequately assessed. To address these limitations, we introduce WorldAlign, a decoupled 4D reward framework that semantically separates static regions and dynamic subjects and provides feedback by aligning each with a world prior suited to its assumptions. For static regions, WorldAlign aligns static geometry with a geometric world prior through semantically guided masked reprojection, enabling more reliable static-consistency evaluation; an auxiliary camera-motion reward discourages nearly static solutions. For dynamic subjects, WorldAlign uses a strong vision-language model (VLM) as a dynamic world prior and constructs a VLM-as-a-judge reward based on sample-specific checklists that assess dynamicity, physical plausibility, shape, and texture consistency. This decoupled design enables more effective online post-training without requiring human preference annotations. Across two pretrained image-to-video generators, Wan2.1 and Wan2.2, WorldAlign jointly improves static and dynamic consistency over existing methods without suppressing overall or subject motion. These results support decoupled world-prior alignment for more faithful visual world simulation. Project page: https://worldalign.github.io/.
Problem

Research questions and friction points this paper is trying to address.

video generation
4D world consistency
static consistency
dynamic consistency
visual world simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupled 4D Reward
World-Consistent Video Generation
VLM-as-a-Judge
Masked Reprojection
Online Post-Training