image-plane trajectory planning

Designs, implements, and evaluates algorithms and representations that compute, represent, and use trajectories defined in image/pixel coordinates (including panoramic and pixel-space paths) to plan, guide, or condition image synthesis, view consistency, or camera-motion signals. This work produces sensor-agnostic 2D-to-2D trajectory outputs and conditioning mechanisms that are decoupled from camera intrinsics to support trajectory-guided generation, cross-dataset training, and unified training/inference workflows.

image-planetrajectoryplanning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Camera Trajectory Generation: A Comprehensive Survey of Methods, Metrics, and Future Directions

Jun 01, 2025
ZD
Zahra Dehghanian
🏛️ Sharif University of Technology

The field of camera trajectory generation suffers from fragmented knowledge and inconsistent evaluation criteria, lacking a systematic survey. Method: This paper establishes the first unified knowledge framework for the domain, clarifying foundational definitions, a comprehensive technical taxonomy—including rule-based, numerical optimization, supervised/reinforcement learning, and neural motion modeling approaches—and a multidimensional evaluation system. It introduces the first holistic method classification scheme and cross-benchmark evaluation guidelines, integrating mainstream datasets and human-in-the-loop assessment protocols. Contribution/Results: The study identifies critical gaps—insufficient adaptability, low computational efficiency, and limited creative expressivity—and proposes future directions toward efficient modeling, dynamic environment adaptation, and semantics-controllable generation. The framework provides a reusable theoretical foundation and practical paradigm for computer graphics, VR, robotic navigation, and intelligent cinematic production.

Addressing limitations and future opportunities in the fieldLack of systematic survey on camera trajectory generation methodsNeed for comprehensive review of models, metrics, and datasets

Existing image-to-video generation methods suffer from limitations in camera trajectory control, temporal consistency, and geometric completeness. This work proposes an end-to-end framework based on dynamic 3D Gaussian splatting, which— for the first time—employs dynamic 3D Gaussian representations for single-image-driven video synthesis. The method jointly models camera motion and object dynamics within a single forward pass. By leveraging an explicit 3D scene representation, a motion sampling mechanism conditioned on a single input image, and differentiable rendering guided by prescribed camera trajectories, it achieves efficient, controllable, and temporally coherent video generation. Experiments on KITTI, Waymo, RealEstate10K, and DL3DV-10K demonstrate that the proposed approach significantly outperforms existing methods in both video quality and inference efficiency.

3D consistencycamera-controlled video generationgeometric integrity

Existing video generation methods rely on separate modules to model camera motion, object translation, and local deformation—limiting holistic motion control. To address this, we propose the first unified trajectory-driven motion control framework. Our approach represents multi-granularity motion—including global viewpoint changes and fine-grained local deformations—as editable keypoint trajectories, which are injected into the latent space of pre-trained image-to-video diffusion models (e.g., SVD, PixArt-Video) via a lightweight motion injector. This enables end-to-end, plug-and-play motion conditioning without modifying the backbone architecture. Crucially, our method preserves temporal coherence and semantic alignment while simultaneously controlling both macro-scale camera motions and micro-scale deformations. Extensive experiments demonstrate state-of-the-art performance across motion sketching, dynamic viewpoint synthesis, and precise motion editing—achieving superior controllability, visual fidelity, and cross-model generalization.

Integrates camera movement, object translation, and local motionProjects user-defined trajectories into latent space for controlUnified motion control in video generation using trajectory inputs

Existing video re-rendering methods struggle to simultaneously preserve appearance fidelity, ensure dynamic consistency, and enable precise camera control under novel viewpoints, while lacking explicit spatiotemporal pixel correspondences. To address these limitations, this work proposes a video diffusion Transformer conditioned on paired 3D point trajectories, achieving four-dimensional consistent and camera-controllable generation. The key innovations include a data pipeline that extracts one-to-one trajectory correspondences from multi-view videos and a dual-view trajectory conditioner that integrates geometric operations with temporal aggregation. Evaluated on a benchmark of 400 videos encompassing both static and dynamic scenes, the proposed method substantially outperforms existing approaches, reducing rotation errors by 30–65% and translation errors by 61–72%.

4D consistencycamera-controlled generationspatiotemporal correspondence

Latest Papers

What's happening recently
View more

This work addresses the challenge of flexibly supporting multimodal camera motion control in video generation. To this end, the authors propose a modality-agnostic framework that maps video, pose, and text inputs into a unified motion embedding space, enabling consistent and precise viewpoint manipulation. Key contributions include the construction of a Motion Triplet Dataset, the introduction of a geometry-driven motion representation based on camera extrinsics, and the design of a motion consistency objective in the latent space. The proposed method not only unifies multimodal inputs under a single processing pipeline but also enables novel capabilities such as motion sequence composition and cross-modal interpolation. Experiments demonstrate that the approach generates high-quality videos across all three modalities, accurately adhering to target camera trajectories and validating its effectiveness and generalization.

camera motion controlheterogeneous inputsmodality-agnostic

Existing methods for 4D scene synthesis from monocular videos are constrained by view interpolation, often failing to ensure cross-view consistency, particularly in out-of-view regions. This work proposes PanoGaussian, a unified framework that integrates panoramic trajectory guidance with an explicit dynamic Gaussian representation. By incorporating 3D dynamic physical priors into the modeling process, the method effectively mitigates scale and deformation artifacts in unseen regions. PanoGaussian enables camera-conditioned video generation, producing high-quality, temporally coherent 4D dynamic scenes even under large viewpoint changes, significantly outperforming current state-of-the-art approaches.

4D scene synthesiscross-view consistencydynamic content modeling

This work proposes a fully parallelizable, pixel-level distributed visual odometry and depth estimation algorithm that overcomes the inefficiencies of traditional approaches, which rely on transmitting redundant and noisy raw pixel data and struggle with on-sensor deployment. The method introduces, for the first time, an on-chip sensor architecture based on Gaussian Belief Propagation (GBP), where pixels exchange photometric observations and surface normal priors in parallel to reach consensus on camera motion. A keyframe-like anchoring mechanism is incorporated to effectively constrain inter-frame baselines and preserve geometric consistency. Experimental results demonstrate that the proposed approach achieves efficient and robust on-chip visual odometry and depth estimation on real-world datasets.

depth estimationon-sensor processingpixel-distributed computing

Traditional visual navigation struggles to balance global geometric consistency with topological generalization, limiting its performance in complex environments. This work proposes a novel map representation based on pixel-level relative 3D connectivity, which constructs a pixel correspondence graph in a relative coordinate frame from image sequences and generates a “WayPixel Costmap” for planning and control. By preserving high-fidelity geometric information without requiring global geometric consistency, the approach overcomes the limitations of conventional topological graphs and dense reconstructions. Experimental results demonstrate that the method significantly outperforms image-level and object-level representations across four simulated tasks and real-world scenarios, validating its accuracy and practicality for visual navigation.

3D map representationgeometric consistencypixel-relative connectivity

Hot Scholars

KZ

Kaipeng Zhang

Shanghai AI Laboratory
LLMMultimodal LLMsAIGC
YY

Yuyang Yin

Beijing Jiaotong University
Computer VisionAIGC
YW

Yunchao Wei

Professor, Beijing Jiaotong University, UTS, UIUC, NUS
Computer VisionMachine Learning
FL

Fangfu Liu

Tsinghua University
Computer Vision3D VisionMachine Learning