geometry-conditioned video diffusion

Designs and builds diffusion-based video generators conditioned on geometric proxies or anchors (e.g., 3D/4D proxy geometry and camera anchors) that produce temporally coherent frames consistent with the underlying geometry and camera motion. Also analyzes and enforces viewpoint and motion consistency across arbitrary camera trajectories to enable geometry-aware video synthesis aligned to the proxy.

geometry-conditionedvideodiffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of simultaneously achieving high visual fidelity, natural motion dynamics, and strong temporal consistency in video diffusion models for complex 3D scenes. We propose a high-quality 3D-scene video synthesis method that requires no paired 3D–2D data. Methodologically, we introduce an image-video diffusion co-generation framework: (i) a sparse appearance-guided sampling image diffusion model generates key anchor views; (ii) a flow-aware camera-controlled and geometry-structured video diffusion model performs high-fidelity, temporally consistent intermediate-frame interpolation. Experiments demonstrate that our approach produces stylized, detail-rich, and motion-coherent 3D-scene videos across diverse complex scenarios, significantly outperforming existing baselines. Our method establishes a novel paradigm for rapid, 3D-model-free content generation, advancing the state of the art in diffusion-based 3D video synthesis.

Addressing video diffusion models' limitations in handling complex scene fidelityCreating style-consistent videos without requiring paired 3D-image datasetsGenerating high-quality 3D scene videos from coarse geometry and camera trajectories

Epipolar Geometry Improves Video Generation Models

Oct 24, 2025
OK
Orest Kupyn
🏛️ University of Oxford | Google | TU Munich

Video generation models commonly suffer from geometric inconsistency, motion instability, and visual artifacts, limiting 3D scene realism. To address this, we propose a preference optimization framework grounded in epipolar geometry constraints—requiring no end-to-end differentiability—that embeds classical multi-view geometric priors into modern video diffusion models (e.g., Latent Diffusion Transformers), enhanced with rectified flow techniques. The method is trained on static scenes yet generalizes effectively to dynamic content. Our key contribution is the first use of pairwise epipolar constraints as stable, interpretable optimization signals, bridging the 3D consistency gap inherent in purely data-driven approaches. Experiments demonstrate significant improvements in spatial geometric fidelity and camera trajectory stability, while preserving high visual quality and substantially enhancing the 3D authenticity of generated videos.

Addressing geometric inconsistencies and unstable motion in video generation modelsImproving 3D scene consistency through epipolar geometry constraintsReducing visual artifacts while maintaining high visual quality output

GeoVideo: Introducing Geometric Regularization into Video Generation Model

Dec 03, 2025
YB
Yunpeng Bai
🏛️ The University of Texas at Austin | DAMO Academy, Alibaba Group

Existing video generation methods predominantly operate in the 2D pixel space without explicit 3D structural constraints, leading to geometric temporal inconsistency, physically implausible motion, and structural artifacts. To address this, we propose a latent-space geometry-aware video generation framework built upon latent diffusion models. Our method integrates a frame-wise depth prediction module and introduces a multi-view geometric loss that aligns predicted depth maps across frames within a shared 3D coordinate system—enabling joint optimization of appearance synthesis and 3D structural modeling. Leveraging a diffusion Transformer architecture, we unify a depth prediction network with an image-level latent encoder and impose latent-space depth regularization. Extensive experiments demonstrate that our approach significantly improves geometric consistency, temporal stability, and physical plausibility of generated videos across multiple benchmarks, outperforming current state-of-the-art methods.

Addresses temporal inconsistency in video generation geometryImproves spatio-temporal coherence and physical plausibility in videosIntroduces geometric regularization for 3D structure modeling

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Jul 10, 2025
HW
Haoyu Wu
🏛️ Microsoft Research | Tsinghua University

Video diffusion models trained solely on raw videos struggle to learn geometry-aware 3D dynamic structures, resulting in spatially inconsistent generations. To address this, we propose Geometry Forcing—a novel framework that explicitly injects geometric constraints into the diffusion process for the first time. Specifically, it leverages features from a pre-trained geometric foundation model to enforce dual alignment—angular and scale alignment—on intermediate latent representations of the diffusion model. This is achieved via cosine similarity-based directional matching and unnormalized geometric feature regression, enabling effective fusion of geometric and video representations. Geometry Forcing significantly improves generation quality and 3D consistency under varying camera viewpoints and motion conditions. Extensive experiments demonstrate state-of-the-art performance across multiple benchmarks, outperforming existing methods in both visual fidelity and structural coherence.

Aligning diffusion models with 3D features improves world consistencyEnhancing video generation with geometry-aware directional and scale alignmentVideo diffusion models lack 3D geometric-aware representations

WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion

Dec 22, 2025
HK
Hanyang Kong
🏛️ National University of Singapore | The Hong Kong Polytechnic University

Existing long-video generation methods struggle to simultaneously ensure pixel-level geometric consistency and efficient camera motion modeling, especially under severe occlusions and complex trajectories. This paper proposes a coupled 3D-structural anchoring and 2D generative refinement framework: Gaussian splatting constructs an online-updated geometric cache serving as a 3D-consistency anchor; differentiable geometric warping explicitly reprojects historical frames, while a spatiotemporal diffusion model (ST-Diff) inpaints occluded regions. We introduce the first spatiotemporally adaptive noise scheduling scheme—applying full noise to void regions to trigger generation, while preserving structural information in warped regions and injecting partial noise for texture refinement. Our method achieves state-of-the-art geometric consistency and visual fidelity under complex camera motions and dynamic occlusions, significantly enhancing structural robustness and texture quality in long-range video synthesis.

Addressing occlusion artifacts in warped content using adaptive refinementBridging 3D geometric constraints with 2D generative model capabilitiesGenerating long-range geometrically consistent 3D videos from novel views

Latest Papers

What's happening recently
View more

This work addresses the challenge of maintaining temporally consistent 3D geometry in large-scale video diffusion models, which often exhibit structural drift and implausible motion under viewpoint changes. To tackle this issue, the authors propose VideoWeave, an implicit geometry-guided latent-space post-training framework that encodes geometric features into implicit geometric latent variables and jointly models them with video latents in a shared denoising space, thereby imposing flexible geometric constraints on the generation distribution. By avoiding explicit geometric reconstruction, VideoWeave effectively mitigates error propagation from upstream components. The method is supported by GeoVid-80K, a newly curated paired dataset comprising 80,000 samples. Experiments demonstrate that VideoWeave significantly enhances geometric consistency in both text-to-video and image-to-video generation while preserving high visual fidelity.

3D structuregeometric consistencygeometric drift

This work addresses the challenges of geometric inconsistency and motion discontinuity in image-to-video generation under sparse camera pose conditions. The authors propose a diffusion-based approach that injects geometric priors from a pretrained video-to-3D model into the generative process via knowledge distillation during training. To enhance temporal coherence and spatial fidelity, the method incorporates keyframe trajectory cycle consistency and cross-frame depth structural constraints, guided by a three-stage coarse-to-fine curriculum learning strategy. Notably, the 3D priors are utilized only during training, eliminating the need for additional computation at test time. Experiments demonstrate that the proposed method significantly improves geometric consistency and motion smoothness across various sparse pose configurations, enabling high-quality video synthesis without requiring dense input trajectories.

3D geometry priorsimage-to-video generationmotion consistency

Existing video diffusion models struggle to maintain geometric consistency when precisely controlling camera poses, as directly injecting camera parameters often leads to structural distortions. This work proposes a novel approach that encodes camera motion into temporally consistent stochastic representations, implicitly embedding pose information into the noise space through a geometry-guided reprojection-based optical flow and a noise-space warping mechanism. This design decouples motion from appearance while preserving the Gaussian prior inherent in diffusion models, ensuring geometrically consistent noise propagation under camera transformations. Experimental results demonstrate that the proposed method significantly improves both visual quality and camera trajectory fidelity in generated videos, outperforming current state-of-the-art approaches.

camera pose controlgeometric consistencystructural distortion

This work addresses the challenge of spatial control in video generation arising from viewpoint changes and camera motion by proposing a geometry-aware diffusion Transformer architecture. By integrating projected positional encoding and a depth-aware disambiguation mechanism, the method effectively fuses 3D depth information with 2D reprojection. It further introduces structured context tokens and geometry-guided cross-attention to enable precise spatial manipulation directly within the native latent space. The proposed approach significantly enhances controllability for viewpoint-dependent editing tasks, supporting camera trajectory redirection, novel view synthesis, and geometry-consistent video editing, all while preserving the strong generative priors of the underlying foundation model.

camera motionscene geometryspatial controllability

Hot Scholars

SK

Seungryong Kim

Associate Professor, KAIST
Computer VisionMachine Learning
ZC

Zeyu Cai

Institute of Heavy Ion Physics, Peking University
AI for SciencePlasma PhysicsAI AgentsNumber Theory
YX

Yuliang Xiu

Westlake University | Max Planck Institute for Intelligent Systems
Computer GraphicsComputer VisionDigital Humans