SymRegFlow: Symmetry-Regularized Flow Matching for Video World Models

πŸ“… 2026-10-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the reliance of video world models on expensive paired data and their lack of novel-view supervision under continuous camera poses. To this end, we propose a symmetric regularized flow matching framework. The method generates noisy anchors by geometrically warping source views, introducing masked dual-anchor supervision and cross-anchor denoising consistency. By integrating affine Gaussian surrogate modeling with geometric reprojection, it achieves multi-view consistent video generation without ground-truth novel-view supervision. We further provide theoretical proof that this regularization recovers the clean-reference optimal solution. Experiments on the nuScenes dataset demonstrate that our approach reduces FrΓ©chet Video Distance (FVD) by over 31%, achieving state-of-the-art performance in FVD, FVMD, and FID metrics while exhibiting superior instance preservation capabilities.
πŸ“ Abstract
Flow-matching-based multi-view world models generate realistic videos, but are commonly restricted to fixed camera rigs. Extending them to continuously varying camera poses requires paired pose--video observations with dense pose coverage, which are costly to acquire. We introduce \emph{SymRegFlow}, a symmetry-regularized flow-matching framework for multi-view-consistent video generation across continuous viewpoints without ground-truth novel-view RGB supervision. For each target pose, SymRegFlow geometrically warps source views into noisy anchors and combines masked dual-anchor supervision with cross-anchor denoising-output consistency to mitigate anchor-specific errors. Under an affine Gaussian surrogate, we prove that suitable consistency regularization recovers the clean-reference optimum at fixed noise levels, strictly outperforming single- and merged-anchor baselines. Experiments on Cosmos-Drive-Dreams and nuScenes demonstrate high-quality, multi-view-consistent autonomous-driving video generation: on nuScenes, SymRegFlow achieves the lowest FVD and FVMD among the evaluated baselines, reducing FVD by over 31\% relative to the best baseline, and source-conditioned inference also attains the best FID and instance preservation.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Flow Matching
Video World Models
Symmetry Regularization
Multi-view Consistency
Novel View Synthesis
πŸ”Ž Similar Papers
No similar papers found.