OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of simultaneously satisfying geometric constraints and object-aware composition in panorama- and text-driven camera trajectory generation by proposing OmniCam, an autoregressive model. The method introduces geometry-grounded pose token learning with explicit 3D object anchors and decoupled geometric-semantic conditioning streams. By integrating a panoramic point cloud encoder, hybrid absolute-rotation and relative-translation tokenization, and quaternion sign consistency handling, it achieves precise trajectory planning. Additionally, a large-scale dataset, OmniCaT, is constructed. Experiments demonstrate that OmniCam reduces trajectory error by 28%–47% and collision rate by 65.8%, significantly outperforming existing baselines. The approach has been successfully applied to video generation and robotic active perception.
📝 Abstract
Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.
Problem

Research questions and friction points this paper is trying to address.

camera trajectory generation
panoramic image
language-conditioned
scene geometry
target-aware framing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Camera Trajectory Generation
Pose Token Learning
Panoramic Point-Cloud Encoder
Autoregressive Model
Geometry-Grounded Conditioning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.