TKCAM: Text and Keyframe to Camera Trajectory Generation

๐Ÿ“… 2026-10-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge of generating high-quality, controllable camera motions for AI filmmaking by proposing a generative masked modeling framework for camera trajectory synthesis. Methodologically, it employs a residual vector quantizer to discretize motion features and constructs a two-stage masked Transformer with multimodal attention modules. The approach innovatively introduces a sparse visual keyframe conditioning mechanism combined with free-text prompts, enabling the generation of coherent and smooth trajectories under local temporal guidance. Experimental results demonstrate that the proposed method surpasses existing state-of-the-art baselines across FID, text-matching, and retrieval metrics. Furthermore, this work releases the RealEstate10K-Cap dataset along with a comprehensive evaluation benchmark to facilitate future research in controllable camera motion generation.
๐Ÿ“ Abstract
Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. We represent camera dynamics using a 12-dimensional kinematic feature comprising position, velocity, and a continuous rotation representation and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for multimodal conditioning. A central feature of our framework is sparse visual keyframe conditioning: users can provide free-form text prompts together with RGB observations at selected timestamps, which provide temporally localized visual guidance for generating coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM surpasses recent state-of-the-art baselines on Frรฉchet distance (FID), text-motion matching scores, and retrieval metrics (R@K), while additional analyses evaluate temporal smoothness and cross-domain generalization. Code is available at https://github.com/linearalgebrayhz/TKCAM.
Problem

Research questions and friction points this paper is trying to address.

Camera Trajectory Generation
Text-conditioned Motion Synthesis
Keyframe Conditioning
Controllable Camera Motion
Cross-domain Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Masked Transformer
Residual Vector Quantizer
Sparse Keyframe Conditioning
Camera Trajectory Generation
Multimodal Conditioning