geometry-anchored video generation

Designs and builds video synthesis systems that generate temporally coherent frame sequences conditioned on explicit geometric cues—such as depth maps, camera poses, 3D scene representations, or projected object poses—to perform image-to-video transformation, geometry-aware compositing, and long-horizon rollout generation. Works on modeling and enforcing geometric consistency across frames to maintain pixel- and scene-aligned pose projections, preserve object–scene relations and interaction trajectories, and reduce trajectory drift over long horizons.

geometry-anchoredvideogeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the trade-off between high computational cost in dense multi-view video novel view synthesis and quality degradation (e.g., flickering, geometric distortion, and spatiotemporal inconsistency) under sparse input conditions, this paper proposes an efficient online 3D video synthesis method. Our approach introduces two key innovations: (1) a globally geometry-constrained real-time rendering paradigm, leveraging TSDF-based spatial modeling to enforce cross-view and inter-frame geometric consistency; and (2) a temporal color-difference masking guided progressive depth map optimization scheme, jointly enhanced by a pre-trained multi-view fusion network to improve photometric fidelity. The method retains online inference capability (≈30 FPS) while significantly suppressing flickering artifacts. It achieves state-of-the-art synthesis quality across multiple benchmarks, demonstrating superior geometric stability and visual coherence under sparse input settings.

Achieving efficient high-quality novel-view video synthesisEnhancing view and temporal consistency using geometry guidanceReducing multi-view and temporal inconsistencies in synthesis

3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation

Oct 16, 2025
JL
JoungBin Lee
🏛️ KAIST AI | Sony AI | ETH Zürich | Sony Group Corporation

This work addresses three key challenges in long-video generation: weak scene consistency, imprecise camera control, and erroneous persistence of dynamic elements across temporal boundaries. To this end, we propose a video generation framework supporting arbitrary-length input sequences. Our method introduces (1) a 3D scene memory mechanism that jointly leverages dynamic SLAM and adaptive dynamic masking to explicitly decouple static geometry from dynamic content; and (2) dual spatiotemporal conditioning, which fuses spatiotemporal features from adjacent frames and incorporates static scene geometry via geometric projection—thereby ensuring long-term spatial coherence and controllable free-viewpoint rendering. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in scene consistency, camera motion accuracy, and visual quality, while maintaining computational efficiency and realistic motion dynamics.

Generating consistent long videos with precise camera controlMaintaining spatial coherence across arbitrary-length video generationSeparating static geometry from dynamic elements in videos

Existing image-to-video generation methods suffer from limitations in camera trajectory control, temporal consistency, and geometric completeness. This work proposes an end-to-end framework based on dynamic 3D Gaussian splatting, which— for the first time—employs dynamic 3D Gaussian representations for single-image-driven video synthesis. The method jointly models camera motion and object dynamics within a single forward pass. By leveraging an explicit 3D scene representation, a motion sampling mechanism conditioned on a single input image, and differentiable rendering guided by prescribed camera trajectories, it achieves efficient, controllable, and temporally coherent video generation. Experiments on KITTI, Waymo, RealEstate10K, and DL3DV-10K demonstrate that the proposed approach significantly outperforms existing methods in both video quality and inference efficiency.

3D consistencycamera-controlled video generationgeometric integrity

Existing methods for scene-consistent video generation often suffer from error accumulation due to reliance on external memory, non-differentiable operations, or decoupled multi-model architectures, leading to degraded consistency. This work proposes a “Geometry-as-Context” framework that integrates geometric information as dynamic context within an autoregressive video generation process, enabling end-to-end training through alternating estimation of current-view geometry and rendering of novel views. Key innovations include a camera-gated attention mechanism to enhance pose awareness, interleaved training of geometry and RGB sequences, and stochastic dropping of geometric context during training to support pure RGB inference. Experiments demonstrate that the proposed method significantly outperforms existing approaches under both unidirectional and round-trip camera trajectories, achieving substantial improvements in scene consistency and camera control accuracy.

3D reconstructioncamera trajectorygeometry context

GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation

Nov 28, 2025
YW
Yuhao Wan
🏛️ Nankai University | ByteDance Inc. | Renmin University of China

Existing single-image-to-3D scene generation methods suffer from geometric distortions and texture blurriness, primarily due to inherent limitations in monocular depth estimation. To address this, we propose Video2Scene—a novel framework that leverages a video diffusion model to synthesize multi-view frames, extracts globally consistent geometric features from them, and reconstructs 3D scenes using a predefined camera trajectory. Our key contributions are: (1) a geometric alignment loss enforcing structural consistency across multi-frame depth maps under camera motion constraints; and (2) a lightweight geometric adaptation module enhancing cross-frame geometric feature transferability and utilization. Evaluated on ScanNet and Matterport3D, Video2Scene significantly outperforms state-of-the-art methods, achieving substantial improvements in PSNR, LPIPS, and Chamfer Distance—quantifying both visual fidelity and geometric accuracy. Qualitative results further confirm that reconstructed scenes exhibit high geometric precision and photorealistic texture quality.

Addresses geometric distortions in image-to-3D scene generationEnhances geometric consistency using multi-frame geometry featuresImproves fidelity of 3D scenes from single images

Latest Papers

What's happening recently
View more

This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.

3D object manipulationgeometric editingsecondary effects

This work addresses the challenges of geometric inconsistency and motion discontinuity in image-to-video generation under sparse camera pose conditions. The authors propose a diffusion-based approach that injects geometric priors from a pretrained video-to-3D model into the generative process via knowledge distillation during training. To enhance temporal coherence and spatial fidelity, the method incorporates keyframe trajectory cycle consistency and cross-frame depth structural constraints, guided by a three-stage coarse-to-fine curriculum learning strategy. Notably, the 3D priors are utilized only during training, eliminating the need for additional computation at test time. Experiments demonstrate that the proposed method significantly improves geometric consistency and motion smoothness across various sparse pose configurations, enabling high-quality video synthesis without requiring dense input trajectories.

3D geometry priorsimage-to-video generationmotion consistency

Existing video re-rendering methods rely on synthetic data for supervised fine-tuning, which often fails to accurately adhere to physical scale and camera trajectories in real-world scenes, limiting their generalization. This work introduces reinforcement learning into camera-controlled video re-rendering for the first time, proposing an unpaired training framework built upon a pre-trained video generation model that jointly optimizes using real videos and synthetic camera trajectories. A geometry-aware reward mechanism is designed to explicitly enforce 3D scale consistency and physically plausible camera motion. Experiments demonstrate that the proposed approach significantly outperforms existing supervised learning baselines in both camera control accuracy and visual fidelity, without requiring synchronized multi-view real video data.

camera trajectory accuracycamera-controlled video generationphysical scale adherence

Existing video generation methods struggle to achieve precise metric-level control under out-of-distribution camera trajectories, and static 3D caches often fail after viewpoint changes. To address these limitations, this work proposes GeoStream, a framework that enables accurate camera control through autoregressive streaming generation. GeoStream introduces a self-refreshing 3D cache mechanism that dynamically updates geometric representations online and leverages frames generated by the model itself to construct on-policy geometric conditions for distillation training. This design aligns the training and inference distributions, effectively mitigating autoregressive drift and geometric feedback errors. Extensive quantitative and qualitative evaluations demonstrate that GeoStream significantly outperforms existing approaches in camera controllability.

3D cacheautoregressive streamingcamera control

This work addresses geometric inconsistencies in text-to-video generation—such as object deformation, texture drift, and non-rigid background motion—by introducing a geometric consistency reward mechanism that explicitly optimizes temporal geometric structure during reinforcement fine-tuning of diffusion models. For the first time, geometric consistency is formulated as a directly optimizable objective without modifying the model’s latent space, making the approach applicable to complex dynamic scenes involving both camera and object motion. By integrating optical flow, depth-pose estimation, and feature correspondence techniques, the method effectively disentangles rigid background from dynamic object regions and evaluates their consistency separately. Experiments demonstrate that this approach substantially reduces temporal geometric artifacts while preserving high visual fidelity, outperforming strong existing baselines.

camera motiongeometric consistencyobject deformation

Hot Scholars

SZ

Shuai Zhao

Postdoctoral, Nanyang Technological University
LLMsModel SecurityBackdoor Attack
ZL

Zhihao Liang

South China University of Technology
Computer Vision and Pattern RecognitionMachine Learning
EC

Erik Cambria

Professor @ NTU CCDS & Visiting @ MIT Media Lab
Neurosymbolic AIMultimodal InteractionNLPAffective Computing
TM

Tobias Meisen

Bergische Universität Wuppertal, previously RWTH Aachen University
Industrial AIDeep LearningDeep Reinforcement LearningSemantic Technologies