Score
Designs and builds diffusion-based video generators conditioned on geometric proxies or anchors (e.g., 3D/4D proxy geometry and camera anchors) that produce temporally coherent frames consistent with the underlying geometry and camera motion. Also analyzes and enforces viewpoint and motion consistency across arbitrary camera trajectories to enable geometry-aware video synthesis aligned to the proxy.
This work addresses the challenge of simultaneously achieving high visual fidelity, natural motion dynamics, and strong temporal consistency in video diffusion models for complex 3D scenes. We propose a high-quality 3D-scene video synthesis method that requires no paired 3D–2D data. Methodologically, we introduce an image-video diffusion co-generation framework: (i) a sparse appearance-guided sampling image diffusion model generates key anchor views; (ii) a flow-aware camera-controlled and geometry-structured video diffusion model performs high-fidelity, temporally consistent intermediate-frame interpolation. Experiments demonstrate that our approach produces stylized, detail-rich, and motion-coherent 3D-scene videos across diverse complex scenarios, significantly outperforming existing baselines. Our method establishes a novel paradigm for rapid, 3D-model-free content generation, advancing the state of the art in diffusion-based 3D video synthesis.
Video generation models commonly suffer from geometric inconsistency, motion instability, and visual artifacts, limiting 3D scene realism. To address this, we propose a preference optimization framework grounded in epipolar geometry constraints—requiring no end-to-end differentiability—that embeds classical multi-view geometric priors into modern video diffusion models (e.g., Latent Diffusion Transformers), enhanced with rectified flow techniques. The method is trained on static scenes yet generalizes effectively to dynamic content. Our key contribution is the first use of pairwise epipolar constraints as stable, interpretable optimization signals, bridging the 3D consistency gap inherent in purely data-driven approaches. Experiments demonstrate significant improvements in spatial geometric fidelity and camera trajectory stability, while preserving high visual quality and substantially enhancing the 3D authenticity of generated videos.
Existing video generation methods predominantly operate in the 2D pixel space without explicit 3D structural constraints, leading to geometric temporal inconsistency, physically implausible motion, and structural artifacts. To address this, we propose a latent-space geometry-aware video generation framework built upon latent diffusion models. Our method integrates a frame-wise depth prediction module and introduces a multi-view geometric loss that aligns predicted depth maps across frames within a shared 3D coordinate system—enabling joint optimization of appearance synthesis and 3D structural modeling. Leveraging a diffusion Transformer architecture, we unify a depth prediction network with an image-level latent encoder and impose latent-space depth regularization. Extensive experiments demonstrate that our approach significantly improves geometric consistency, temporal stability, and physical plausibility of generated videos across multiple benchmarks, outperforming current state-of-the-art methods.
Video diffusion models trained solely on raw videos struggle to learn geometry-aware 3D dynamic structures, resulting in spatially inconsistent generations. To address this, we propose Geometry Forcing—a novel framework that explicitly injects geometric constraints into the diffusion process for the first time. Specifically, it leverages features from a pre-trained geometric foundation model to enforce dual alignment—angular and scale alignment—on intermediate latent representations of the diffusion model. This is achieved via cosine similarity-based directional matching and unnormalized geometric feature regression, enabling effective fusion of geometric and video representations. Geometry Forcing significantly improves generation quality and 3D consistency under varying camera viewpoints and motion conditions. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, outperforming existing methods in both visual fidelity and structural coherence.
Existing long-video generation methods struggle to simultaneously ensure pixel-level geometric consistency and efficient camera motion modeling, especially under severe occlusions and complex trajectories. This paper proposes a coupled 3D-structural anchoring and 2D generative refinement framework: Gaussian splatting constructs an online-updated geometric cache serving as a 3D-consistency anchor; differentiable geometric warping explicitly reprojects historical frames, while a spatiotemporal diffusion model (ST-Diff) inpaints occluded regions. We introduce the first spatiotemporally adaptive noise scheduling scheme—applying full noise to void regions to trigger generation, while preserving structural information in warped regions and injecting partial noise for texture refinement. Our method achieves state-of-the-art geometric consistency and visual fidelity under complex camera motions and dynamic occlusions, significantly enhancing structural robustness and texture quality in long-range video synthesis.
This work addresses the challenge of maintaining temporally consistent 3D geometry in large-scale video diffusion models, which often exhibit structural drift and implausible motion under viewpoint changes. To tackle this issue, the authors propose VideoWeave, an implicit geometry-guided latent-space post-training framework that encodes geometric features into implicit geometric latent variables and jointly models them with video latents in a shared denoising space, thereby imposing flexible geometric constraints on the generation distribution. By avoiding explicit geometric reconstruction, VideoWeave effectively mitigates error propagation from upstream components. The method is supported by GeoVid-80K, a newly curated paired dataset comprising 80,000 samples. Experiments demonstrate that VideoWeave significantly enhances geometric consistency in both text-to-video and image-to-video generation while preserving high visual fidelity.
SpatialCrafter通过引入全局3D代理和两阶段框架,解决了基于视频扩散模型的图像到场景生成中的随机幻觉、长期漂移及3D一致性不佳的问题。
This work addresses the challenges of geometric inconsistency and motion discontinuity in image-to-video generation under sparse camera pose conditions. The authors propose a diffusion-based approach that injects geometric priors from a pretrained video-to-3D model into the generative process via knowledge distillation during training. To enhance temporal coherence and spatial fidelity, the method incorporates keyframe trajectory cycle consistency and cross-frame depth structural constraints, guided by a three-stage coarse-to-fine curriculum learning strategy. Notably, the 3D priors are utilized only during training, eliminating the need for additional computation at test time. Experiments demonstrate that the proposed method significantly improves geometric consistency and motion smoothness across various sparse pose configurations, enabling high-quality video synthesis without requiring dense input trajectories.
Existing video diffusion models struggle to maintain geometric consistency when precisely controlling camera poses, as directly injecting camera parameters often leads to structural distortions. This work proposes a novel approach that encodes camera motion into temporally consistent stochastic representations, implicitly embedding pose information into the noise space through a geometry-guided reprojection-based optical flow and a noise-space warping mechanism. This design decouples motion from appearance while preserving the Gaussian prior inherent in diffusion models, ensuring geometrically consistent noise propagation under camera transformations. Experimental results demonstrate that the proposed method significantly improves both visual quality and camera trajectory fidelity in generated videos, outperforming current state-of-the-art approaches.
This work addresses the challenge of spatial control in video generation arising from viewpoint changes and camera motion by proposing a geometry-aware diffusion Transformer architecture. By integrating projected positional encoding and a depth-aware disambiguation mechanism, the method effectively fuses 3D depth information with 2D reprojection. It further introduces structured context tokens and geometry-guided cross-attention to enable precise spatial manipulation directly within the native latent space. The proposed approach significantly enhances controllability for viewpoint-dependent editing tasks, supporting camera trajectory redirection, novel view synthesis, and geometry-consistent video editing, all while preserving the strong generative priors of the underlying foundation model.